Aerial visual localization
The use of synthetic panoramic images generated from a 3D model formed by multiple data sources addresses the inefficiencies of manual image collection and computational intensity in existing aerial visual localization methods, enabling efficient and effective localization of aerial vehicles.
Patent Information
- Application Number
- PCT/SG2025/050359
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-29
- Filing Date
- 2025-05-29
- Publication Date
- 2025-12-04
AI Technical Summary
Existing aerial visual localization methods require a vast amount of manually collected images and intensive computational processing, leading to inefficiencies and difficulties in practical implementation.
A method and system for training a visual localization model using synthetic panoramic images generated from a 3D model formed by multiple data sources, reducing the need for manual image collection and computational intensity.
Enables efficient and effective aerial visual localization by predicting the location of an aerial vehicle using computationally efficient synthetic data, enhancing practical implementation.
Smart Images

Figure SG2025050359_04122025_PF_FP_ABST
Abstract
Description
AERIAL VISUAL LOCALIZATIONCROSS-REFERENCE TO RELATED APPLICATION
[0001] This application claims the benefit of priority of Singapore Patent Application No. 10202401520U filed on 29 May 2024, the content of which being hereby incorporated by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] The present invention generally relates to aerial visual localization (or aerial imagebased localization), including a method of training a visual localization model for aerial visual localization and a system thereof, as well as a method of performing aerial visual localization of an aerial vehicle during flight and an aerial vehicle configured to perform the method of aerial visual localization during flight, and more particularly, with respect to absolute visual localization.BACKGROUND
[0003] With the increasing popularity of aerial vehicles, whether manned or unmanned, significant research has been devoted to developing a more robust navigational suite. Existing aerial vehicles have primarily relied on a combination of an Inertial Measurement Unit (IMU) and the Global Navigation Satellite System (GNSS) to establish their location in three- dimensional (3D) space. However, there is a growing need for a self-sufficient absolute navigational system, given the vulnerability of satellite-based and external-based localization methods to a range of interferences, including natural obstacles such as weather and multi- pathing, as well as artificial disruptions such as jamming and spoofing. External-based localization methods refer to localization methods which rely on signals and / or infrastructure outside of the vehicle for performing localization. Therefore, in the realm of aerial vehicle navigation, the sole reliance on satellite-based and / or external-based localization methods presents vulnerabilities to various interferences. This drives the need for a self-sufficient (or self-reliant) absolute navigational system, including an absolute localization system, either as a primary or sole localization system or as a secondary or backup localization system.
[0004] Visual or image-based localization methods can be divided into relative visual localization (RVL), where the pose difference between two successive frames is estimated, and absolute visual localization (AVL), where the pose in the global frame is directly determinedor estimated. In this regard, as explained above, absolute visual localization is of particular interest for enabling aerial vehicle navigation to avoid sole reliance on satellite-based and / or external-based localization methods which are vulnerable to various interferences, thus enhancing reliability and / or providing an alternative approach in aerial visual localization.
[0005] There are various existing absolute visual localization methods. However, the main challenges or problems associated with various existing absolute visual localization methods include the vast amount of images required to be collected manually (e.g., a very large amount reference images is required to be manually collected) and / or the intensive computational processing required to perform localization (e.g., for analyzing and comparing an image collected by a vehicle with the large amount of reference images for estimating the location of the vehicle), which result in difficulties and inefficiencies in practical implementation, especially for aerial visual localization.
[0006] A need therefore exists to provide methods, as well as systems thereof, for aerial visual localization (or aerial image-based localization), and more particularly, for absolute visual localization, that seek to overcome, or at least ameliorate, one or more deficiencies in conventional localization methods, and more particularly, with improved efficiency and effectiveness in aerial visual localization, thus enhancing practical implementation for aerial visual localization. It is against this background that the present invention has been developed.SUMMARY
[0007] According to a first aspect of the present invention, there is provided a method of training a visual localization model for aerial visual localization, the method comprising: obtaining a 3D model of an environment, the 3D model formed based on a plurality of different data sources; generating synthetic panoramic images based on the 3D model of the environment; and training the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
[0008] According to a second aspect of the present invention, there is provided a system for training a visual localization model for aerial visual localization, the system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to:obtain a 3D model of an environment, the 3D model formed based on a plurality of different data sources; generate synthetic panoramic images based on the 3D model of the environment; and train the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
[0009] According to a third aspect of the present invention, there is provided a method of performing aerial visual localization of an aerial vehicle during flight, the method comprising: obtaining, from an image sensor of the aerial vehicle, a panoramic image of an environment captured by the image sensor during flight; and predicting, using a visual localization model trained according to the above-mentioned method according to the first aspect of the present invention, a location of the aerial vehicle based on the panoramic image of the environment captured.
[0010] According to a fourth aspect of the present invention, there is provided an aerial vehicle comprising: a flight controller operable to control a flight of the aerial vehicle; an image sensor configured to capture a panoramic image of an environment during flight; at least one memory; and at least one processor communicatively coupled to the at least one memory and the image sensor and configured to perform aerial visual localization of the aerial vehicle during flight comprising: obtaining, from the image sensor, a panoramic image of an environment captured by the image sensor during flight; and predicting, using a visual localization model trained according to the above- mentioned method of training a visual localization model according to the first aspect of the present invention, a location of the aerial vehicle based on the panoramic image of the environment captured, wherein the flight controller is operable to control a flight of the aerial vehicle based on the predicted location of the aerial vehicle.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Embodiments of the present invention will be better understood and readily apparent to one of ordinary skill in the art from the following written description, by way of example only, and in conjunction with the drawings, in which:FIG. 1 depicts a schematic diagram of a method of training a visual localization model for aerial visual localization of an aerial vehicle, according to various embodiments of the present invention;FIG. 2 depicts a schematic block diagram of a system for training a visual localization model for aerial visual localization of an aerial vehicle, according to various embodiments of the present invention;FIG. 3 depicts a schematic diagram of a method of performing aerial visual localization of an aerial vehicle during flight, according to various embodiments of the present invention;FIG. 4 depicts a schematic block diagram of a system for performing aerial visual localization of an aerial vehicle during flight, according to various embodiments of the present invention;FIG. 5 depicts a schematic drawing of an aerial vehicle, according to various embodiments of the present invention;FIG. 6 depicts a schematic diagram of a method of training a visual localization model for aerial visual localization of an aerial vehicle, according to various example embodiments of the present invention;FIG. 7A shows an example GPS logger used to establish a correspondence between the captured image and the location ground truth for ground truth data collection;FIG. 7B shows an example drone mounted with a 360-dcgrcc camera and the GPS logger;FIG. 7C shows an example panoramic 360-degree image captured by the drone in 2: 1 equirectangular format for inference to predict the location of the drone;FIGs. 8A to 8C show overviews of example photogrammetric 3D models built for the different environments, according to various example embodiments of the present invention;FIG. 9A illustrates an example 3D model (3D landscape) with example virtual roads generated based on OpenStreetMap (OSM) data;FIG. 9B illustrates the example 3D model filled-in with virtual structures (e.g., buildings), as well as visual textures thereof, from an authoritative virtual environment data source, according to various example embodiments of the present invention;FIGs. 10A to 10D show example synthetic panoramic images generated in Unity, presented in 2: 1 equirectangular format, with example domain randomizations for environmental conditions, according to various example embodiments of the present invention;FIG. 11 shows a table (Table I) presenting a summary of datasets of different environments, according to various example embodiments of the present invention;FIGs. 12A and 12B illustrate a comparison of an example synthetic panoramic image before and after being put through the trained CycleGAN network;FIG. 13 shows a table (Table II) presenting a summary of the median error reported for the localization performance of each environment (configuration);FIGs. 14A to 14D show XY plots (top view) of the predicted locations for each of the four environments on the corresponding ground truth locations provided by their respective test sets, according to various example embodiments of the present invention;FIGs. 15A to 15D show histogram breakdowns of the Euclidean errors with the ground truth altitude for the four environments;FIG. 16 shows sample synthetic panoramic images from different variants of a dataset, according to various example embodiments of the present invention;FIG. 17 shows a Table (Table III) presenting a summary of the median error reported for the localization performance for each dataset variant;FIG. 18 depicts a setup diagram for a closed-loop control of an example UAV, according to various example embodiments of the present invention;FIG. 19 depicts a schematic block diagram of the closed-loop control of a flight controller of the UAV, according to various example embodiments of the present invention;FIG. 20 shows the setpoints together with the filtered XYZ position data estimated by the PX4 EKF2 filter (extended Kalman filter) using visual localization data, according to various example embodiments of the present invention; andFIG. 21 shows the raw unfiltered XYZ position as output from the localization network, inferred on the image stream from the onboard 360-degree camera, according to various example embodiments of the present invention.DETAILED DESCRIPTION room Various embodiments of the present invention provide aerial visual localization (or aerial image-based localization), including a method of training a visual localization model for aerial visual localization and a system thereof, as well as a method of performing aerial visuallocalization of an aerial vehicle during flight and an aerial vehicle including a processor configured to perform the above-mentioned method of aerial visual localization of the aerial vehicle during flight, and more particularly, with respect to absolute visual localization.
[0013] As described in the background, in the realm of aerial vehicle navigation, the sole reliance on satellite -based and / or external-based localization methods for aerial visual localization presents vulnerabilities to various interferences. In this regard, absolute visual localization is of particular interest for enabling aerial vehicle navigation to avoid sole reliance on satellite-based and / or external-based localization methods which axe vulnerable to various interferences, thus enhancing reliability and / or providing an alternative approach in aerial visual localization. This drives the need for a self-sufficient absolute navigational system, including an absolute localization system, either as a primary or sole localization system or as a secondary or backup localization system. However, existing absolute visual localization methods suffer from a number of key challenges or problems including the vast amount of images required to be collected manually (e.g., a very large amount reference images is required to be manually collected) and / or the intensive computational processing required to perform localization (e.g., for analyzing and comparing an image collected by a vehicle with the large amount of reference images for estimating the location of the vehicle), which result in difficulties and inefficiencies in practical implementation, especially for aerial visual localization. In this regard, various embodiments of the present invention provide methods, as well as systems thereof, for aerial visual localization, and more particularly, for absolute visual localization, that seek to overcome, or at least ameliorate, one or more deficiencies in conventional localization methods, and more particularly, with improved efficiency and effectiveness in aerial visual localization, thus enhancing practical implementation for aerial visual localization.
[0014] It will be appreciated by a person skilled in the art that the aerial visual localization of an aerial vehicle disclosed herein according to various embodiments of the present invention is not limited to any particular type of aerial vehicle, and may be applied to, or implemented in, any type of aerial vehicle, whether manned or unmanned, as desired or as appropriate, as long as localization of the aerial vehicle is desired or required to be performed.
[0015] FIG. 1 depicts a schematic diagram of a method 100 of training a visual localization model (Al or machine learning model) for aerial visual localization (or aerial image-based localization), according to various embodiments of the present invention. The method 100 comprising: obtaining (at 106) a three-dimensional (3D) model of an environment (e.g., outdoor environment), the 3D model formed based on a plurality of different data sources; generating(at 108) synthetic panoramic images based on the 3D model of the environment; and training (at 110) the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location (or pose) of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
[0016] The method 100 of training a visual localization model for aerial visual localization according to various embodiments of the present invention advantageously enables aerial visual localization to be implemented and performed with improved efficiency and effectiveness. In par ticular, according to the method 100, a 3D model of an environment is obtained and synthetic panoramic images arc then generated based on the 3D model of the environment obtained. In this regard, different 3D models may be obtained for different environments, and synthetic panoramic images may be generated from each 3D model. This significantly reduces the amount of reference images required to be collected manually, or eliminates the need to manually collect reference images altogether, since synthetic panoramic images may be digitally or automatically generated based on the 3D model of the environment, for example, along with domain randomization and / or adaptation. In addition, the 3D model is formed based on a plurality of different data sources. This advantageously enables 3D models of various environments to be built reliably and accurately, as well as with versatility, based on which synthetic panoramic images (training data) can be generated for training the visual localization model. For example, with respect to versatility, data (from data sources) of various forms can be utilized to generate a 3D model of an environment based on which synthetic panoramic images (training data) can be generated for training the visual localization model and thus, data source is not restricted or necessary to be in the form of panoramic images. Furthermore, by training the visual localization model to predict a location of the aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle, the location of the aerial vehicle can be predicted or estimated using the trained visual localization model in a computationally efficient manner, for example, in contrast with the intensive computational processes associated with various conventional visual localization methods that require analyzing and comparing an image collected by a vehicle with a very large amount of reference images for estimating the location of the vehicle. Therefore, the method 100 of training a visual localization model for aerial visual localization according to various embodiments of the present invention advantageously enables aerial visual localization to be implemented and performed with improved efficiency and effectiveness, thus enhancing practical implementation for aerial visual localization. Accordingly, in various embodiments, the aerialvisual localization is absolute visual localization. These advantages or technical effects, and / or other advantages or technical effects, will become more apparent to a person skilled in the art as the method 100 of training a visual localization model for aerial visual localization and the corresponding system thereof, as well as a method of performing aerial visual localization of an aerial vehicle during flight and an aerial vehicle configured to perform the method of aerial visual localization during flight, are described in more detail according to various embodiments and example embodiments of the present invention.
[0017] In various embodiments, the plurality of different data sources comprises two or more of a photogrammetry data source, a virtual environment data source and a street map data source. In various embodiments, the photogrammetry data source comprises images of one or more environments which may be utilized to build one or more 3D models of the one or more environments based on photogrammetric reconstruction. In this regard, for example, the images in the photogrammetry data source may be manually or self-collected (e.g., using a camera- equipped drone) and may be normal images. In various embodiments, the photogrammetry data source may comprise one or more 3D models built based on photogrammetric reconstruction for one or more environments, which may be referred to as photogrammetry 3D models. As described above, for each of the one or more environments, synthetic panoramic images (e.g., 360-degree synthetic panoramic images) may be generated based on the 3D model of the environment. In various embodiments, the virtual environment data source comprises one or more virtual representations (e.g., a 3D model or digital twin) of one or more environments, respectively, the virtual representation of each environment comprising virtual features, such as digital or virtual components of various structures in the environment, as well as visual textures thereof, such as but not limited to, buildings, bridges and so on (any structures which may exist in an environment). In various embodiments, the street map data source comprises virtual features of one or more environments relating to a street map thereof (street map or geographic data), including virtual roads, virtual terrains and so on, as well as visual textures thereof. For example, the street map or geographic data of an environment may include virtual representations of various features on the ground of the environment, such as roads and terrains.
[0018] In various embodiments, the virtual environment data source is an authoritative virtual environment data source. That is, the virtual environment data source may include one or more virtual representations (e.g., digital twin) of one or more environments, respectively, provided by one or more authorities, and such virtual representations may thus be referred to as official virtual representations. For example, the street map data source is an open street map(OSM) data source, which is a free and open map database hosted by the OpenStreetMap Foundation. rooi9] In various embodiments, the 3D model of the environment is formed based on photogrammetric reconstruction using images from the photogrammetry data source, the 3D model having augmented therein virtual features (e.g., virtual structures, such as buildings) (e.g., as well as visual textures thereof) from the virtual environment data source and / or virtual features (e.g., virtual roads) from the street map data source. Alternatively, the 3D model of the environment is formed based on a 3D model of the environment from the virtual environment data source, the 3D model having augmented therein virtual features (e.g., virtual roads) from the street map data source and / or virtual features (e.g., virtual structures, such as buildings which may be missing from the 3D model from the virtual environment data source) (e.g., as well as visual textures thereof) from the photogrammetry data source (e.g., the virtual features may be derived or obtained from the photogrammetry 3D model of the environment stored in the photogrammetry data source). Still alternatively, the 3D model of the environment may be formed based on street map data from the street map data source, the 3D model (e.g., a 3D landscape of the environment with virtual roads) having augmented therein virtual features (e.g., virtual structures, such as buildings) (e.g., as well as visual textures thereof) from the virtual environment data source and / or virtual features (e.g., virtual structures, such as buildings) from the photogrammetry data source.
[0020] In various embodiments, the above-mentioned generating (at 108) synthetic panoramic images comprises performing domain randomization to generate synthetic panoramic images with domain randomization. For example, example parameters which may be randomized include skybox, sun angle, lighting and so on.
[0021] In various embodiments, the above-mentioned generating (at 108) synthetic panoramic images further comprises performing domain adaptation on the synthetic panoramic images to enhance reality of the synthetic panoramic images.
[0022] In various embodiments, the visual localization model comprises a feature extractor network (e.g., an encoder) configured to extract features from a synthetic panoramic image inputted thereto and a regression network (e.g., a fully connected regression network) configured to output a predicted location for the synthetic panoramic image based on the extracted features.
[0023] In various embodiments, the synthetic panoramic images generated based on the 3D model of the environment are 360-degree synthetic panoramic images. In this regard, utilizing360-degree synthetic panoramic images has been found to improve the training of the visual localization model (as the amount of visual features of the environment utilized for training the visual localization model is maximized (360-degree view of the environment)), as well as improve the performance of the trained visual localization model in predicting the location of the aerial vehicle (as the amount of visual features of the environment utilized by the trained visual localization model for inferring the location of the aerial vehicle is maximized (360- degree view of the environment)).
[0024] FIG. 2 depicts a schematic block diagram of a system 200 for training a visual localization model for aerial visual localization, according to various embodiments of the present invention, corresponding to the above-mentioned method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention. The system 200 comprises: at least one memory 202; and at least one processor 204 communicatively coupled to the at least one memory 202 and configured to perform the method 100 of training a visual localization model for aerial visual localization of an aerial vehicle according to various embodiments of the present invention. Accordingly, the at least one processor 204 is configured to: obtain a 3D model of an environment, the 3D model formed based on a plurality of different data sources; generate synthetic panoramic images based on the 3D model of the environment; and train the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
[0025] It will be appreciated by a person skilled in the ait that the at least one processor 204 may be configured to perform various functions or operations through sct(s) of instructions (e.g., software modules) executable by the at least one processor 204 to perform various functions or operations. Accordingly, as shown in FIG. 2, the system 200 may comprise: a 3D model obtaining module (or a 3D model obtaining circuit) 206 configured to obtain a 3D model of an environment, the 3D model formed based on a plurality of different data sources; a synthetic panoramic image generating model (or a synthetic panoramic image generating circuit) 208 configured to generate synthetic panoramic images based on the 3D model of the environment; and a visual localization model training module (a visual localization model training circuit) 210 configured to train the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of the aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
[0026] It will be appreciated by a person skilled in the art that the above-mentioned modules are not necessarily separate modules, and two or more modules may be realized by or implemented as one functional module (e.g., a circuit or a software program) as desired or as appropriate without deviating from the scope of the present invention. For example, two or more of the 3D model obtaining module 206, the synthetic panoramic image generating module 208 and the visual localization model training module 210 may be realized (e.g., compiled together) as one executable software program (e.g., software application), which for example may be stored in the at least one memory 202 and executable by the at least one processor 204 to perform the corresponding functions or operations as described herein according to various embodiments of the present invention.
[0027] In various embodiments, the system 200 for training a visual localization model corresponds to the method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1, therefore, various operations, functions or steps configured to be performed by the least one processor 204 may correspond to various operations, functions or steps of the method 100 of training a visual localization model described hereinbefore according to various embodiments, and thus need not be repeated with respect to the system 200 for training a visual localization model for clarity and conciseness. In other words, various embodiments described herein in context of methods (e.g., the method 100 of training a visual localization model) are analogously valid for the corresponding systems or devices (e.g., the system 200 for training a visual localization model), and vice versa. For example, in various embodiments, the at least one memory 202 may have stored therein the 3D model obtaining module 206, the synthetic panoramic image generating module 208 and / or the visual localization model training module 210, which respectively correspond to various operations, functions or steps of the method 100 of training a visual localization model as described hereinbefore according to various embodiments, which are executable by the at least one processor 204 to perform the corresponding operations, functions or steps as described herein.
[0028] FIG. 3 depicts a schematic diagram of a method 300 of performing aerial visual localization of an aerial vehicle during flight, according to various embodiments of the present invention. The method 300 comprising: obtaining (at 306), from an image sensor of the aerial vehicle, a panoramic image of an environment captured by the image sensor during flight; and predicting (at 308), using a visual localization model trained according to the above-mentioned method 100 of training a visual localization model as described hereinbefore with reference toFIG. 1 according to various embodiments of the present invention, a location (or pose) of the aerial vehicle based on the panoramic image of the environment captured.
[0029] It will be appreciated by a person skilled in the art that the method 300 of performing aerial visual localization of an aerial vehicle during flight according to various embodiments of the present invention may be applied to, or implemented in, any type of aerial vehicle, whether manned or unmanned, as desired or as appropriate, as long as localization of the aerial vehicle is desired or required to be performed, and more particularly, visual localization or absolute visual localization.
[0030] Accordingly, the method 300 of performing aerial visual localization of an aerial vehicle during flight according to various embodiments of the present invention advantageously uses the trained visual localization model (trained according to the above-mentioned method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention) to predict the location (or pose) of the aerial vehicle based on the panoramic image captured. In particular, the location of the aerial vehicle is predicted or estimated using the trained visual localization model in a computationally efficient manner, for example, in contrast with the intensive computational process associated with various conventional visual localization methods that requires analyzing and comparing an image collected by a vehicle with a very large amount of reference images for estimating the location of the vehicle. Therefore, utilizing the trained visual localization model advantageously enables aerial visual localization of an aerial vehicle be implemented and performed with improved efficiency and effectiveness, thus enhancing practical implementation for aerial visual localization. Accordingly, in various embodiments, the aerial visual localization is absolute visual localization.
[0031] In various embodiments, the image sensor is a 360-degree camera configured to capture a 360-degree panoramic image of the environment.
[0032] FIG. 4 depicts a schematic block diagram of a system 400 for performing aerial visual localization of an aerial vehicle during flight, according to various embodiments of the present invention, corresponding to the above-mentioned method 300 of performing aerial visual localization as described hereinbefore with reference to FIG. 3 according to various embodiments of the present invention. The system 400 comprises: at least one memory 402; and at least one processor 404 communicatively coupled to the at least one memory 402 and configured to perform the method 300 of performing aerial visual localization of an aerial vehicle according to various embodiments of the present invention. Accordingly, the at leastone processor 204 is configured to: obtain, from an image sensor of the aerial vehicle, a panoramic image of an environment captured by the image sensor during flight; and predict, using a visual localization model trained according to the above-mentioned method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention, a location of the aerial vehicle based on the panoramic image of the environment captured. In various embodiments, the system 400 for performing aerial visual localization of an aerial vehicle may be implemented or installed in the aerial vehicle either as a primary or sole localization system or as a secondary or backup localization system.
[0033] It will be appreciated by a person skilled in the art that the at least one processor 404 may be configured to perform various functions or operations through sct(s) of instructions (e.g., software modules) executable by the at least one processor 404 to perform various functions or operations. Accordingly, as shown in FIG. 4, the system 400 may comprise: a panoramic image obtaining module (or a panoramic image obtaining circuit) 406 configured to obtain, from an image sensor of the aerial vehicle, a panoramic image of an environment captured by the image sensor during flight; and a location prediction module (or a panoramic image obtaining circuit) 408 configured to predict, using a visual localization model trained according to the above-mentioned method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention, a location of the aerial vehicle based on the panoramic image of the environment captured.
[0034] It will be appreciated by a person skilled in the art that the above-mentioned modules arc not necessarily separate modules, and two or more modules may be realized by or implemented as one functional module (e.g., a circuit or a software program) as desired or as appropriate without deviating from the scope of the present invention. For example, the panoramic image obtaining module 406 and the location prediction module 408 may be realized (e.g., compiled together) as one executable software program (e.g., software application), which for example may be stored in the at least one memory 402 and executable by the at least one processor 404 to perform the corresponding functions or operations as described herein according to various embodiments of the present invention.
[0035] In various embodiments, the system 400 for performing aerial visual localization corresponds to the method 300 of performing aerial visual localization as described hereinbefore with reference to FIG. 1 , therefore, various operations, functions or stepsconfigured to be performed by the least one processor 404 may correspond to various operations, functions or steps of the method 300 of performing aerial visual localization described hereinbefore according to various embodiments, and thus need not be repeated with respect to the system 400 for performing aerial visual localization for clarity and conciseness. In other words, as explained hereinbefore, various embodiments described herein in context of methods (e.g., the method 300 of performing aerial visual localization) are analogously valid for the corresponding systems or devices (e.g., the system 400 for performing aerial visual localization), and vice versa. For example, in various embodiments, the at least one memory 402 may have stored therein the panoramic image obtaining module 406 and / or the location prediction module 408, which respectively correspond to various operations, functions or steps of the method 300 of performing aerial visual localization as described hereinbefore according to various embodiments, which are executable by the at least one processor 404 to perform the corresponding operations, functions or steps as described herein.
[0036] A computing system, a controller, a microcontroller or any other system providing a processing capability may be provided according to various embodiments in the present invention. Such a system may be taken to include one or more processors and one or more computer- readable storage mediums. For example, the system 200 for training a visual localization model described hereinbefore may include at least one processor 204 and at least one computer-readable storage medium (or memory) 202 which are for example used in various processing carried out therein as described herein. A memory or computer -readable storage medium used in various embodiments may be a volatile memory, for example a DRAM (Dynamic Random Access Memory) or a non-volatile memory, for example a PROM (Programmable Read Only Memory), an EPROM (Erasable PROM), EEPROM (Electrically Erasable PROM), or a flash memory, e.g., a floating gate memory, a charge trapping memory, an MRAM (Magnetoresistive Random Access Memory) or a PCRAM (Phase Change Random Access Memory). Furthermore, it will be appreciated by a person skilled in the art that the system 200 may be implemented by a high-performance computer system, or a network of high- performance computer systems, known in the art for performing training, especially when a large-scale training is performed. For example, the high-performance computer system may include multiple processors including GPUs (graphics processing units) and CPUs (central processing units) optimized for advanced computing tasks, such as execution of machine learning algorithms. For example, the high-performance computer system may comprise anarray of GPUs (Graphics Processing Units) dedicated to handling parallel processing tasks to facilitate the rapid execution of machine learning algorithms.
[0037] In various embodiments, a “circuit” may be understood as any kind of a logic implementing entity, which may be special purpose circuitry or a processor executing software stored in a memory, firmware, or any combination thereof Thus, in an embodiment, a “circuit” may be a hard-wired logic circuit or a programmable logic circuit such as a programmable processor, e.g., a microprocessor (e.g., a Complex Instruction Set Computer (CISC) processor or a Reduced Instruction Set Computer (RISC) processor). A “circuit” may also be a processor executing software, e.g., any kind of computer program, e.g., a computer program using a virtual machine code, e.g., lava. Any other kind of implementation of various functions or operations may also be understood as a “circuit” in accordance with various other embodiments. Similarly, a “module” may be a portion of a system according to various embodiments in the present invention and may encompass a “circuit” as above, or may be understood to be any kind of a logic-implementing entity therefrom.
[0038] Some portions of the present disclosure may be explicitly or implicitly presented in terms of algorithms and functional or symbolic representations of operations on data within a computer memory. These algorithmic descriptions and functional or symbolic representations are the means used by those skilled in the data processing arts to convey most effectively the substance of their work to others skilled in the art. An algorithm may be, and generally, conceived to be a self-consistent sequence of steps leading to a deshed result.
[0039] The present specification also discloses a system (e.g., which may also be embodied as one or more devices or apparatuses), such as the system 200 for training a visual localization model, for performing various operations, functions or steps of various methods described herein. Such a system may be specially constructed for the required purposes or may comprise a general purpose computer system selectively activated or reconfigured by a computer program stored in the computer system. In general, various algorithms that may be presented herein are not limited to being implemented or executed by any particular computer system. Alternatively, the construction of more specialized computer system (e.g., a high-performance computer system as described hereinbefore) to perform various operations, functions or steps of various methods described herein may be provided as desired or as appropriate without going beyond the scope of the present invention.
[0040] In addition, the present specification also at least implicitly discloses computer program(s) or software / functional module(s), in that it would be apparent to a person skilled inthe art that various operations, functions or steps of various methods described herein may be put into effect by computer code. The computer program(s) is not intended to be limited to any particular programming language and implementation thereof, and it will be appreciated by a person skilled in the art that a variety of programming languages and coding thereof may be used to implement the computer program(s). Moreover, the computer program(s) is not intended to be limited to any particular control flow as there are a variety of programming languages which can use different control flows. It will be appreciated by a person skilled in the ait that a computer program may be stored on any computer-readable storage medium (non- transitory computer-readable storage medium), such as but not limited to, a magnetic disk, an optical disk or a memory chip. For example, a computer program stored on a computer-readable storage medium may be loaded and executed on a computer system to implement various operations, functions or steps of various methods described herein according to various embodiments of the present invention.
[0041] Accordingly, in various embodiments, there is provided a computer program product, embodied in one or more computer-readable storage mediums (non-transitory computer-readable storage medium), comprising instructions (e.g., the 3D model obtaining module 206, the synthetic panoramic image generating module 208 and / or the visual localization model training module 210) executable by one or more computer processors to perform the method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention. Accordingly, various computer programs or software modules described herein may be stored in a computer program product receivable by a system therein, such as the system 200 as shown in FIG. 2, for execution by at least one processor 204 of the system 200 to perform various operations, functions or steps of various methods described herein according to various embodiments of the present invention.
[0042] In various embodiments, there is also provided a computer program product, embodied in one or more computer-readable storage mediums (non-transitory computer- readable storage medium), comprising instructions (e.g., the panoramic image obtaining module 406 and / or the location prediction module 408) executable by one or more computer processors 404 to perform the method 300 of performing aerial visual localization as described hereinbefore with reference to FIG. 3 according to various embodiments of the present invention.
[0043] It will be appreciated by a person skilled in the art that various modules described herein (e.g., the 3D model obtaining module 206, the synthetic panoramic image generating module 208 and / or the visual localization model training module 210, as well as the panoramic image obtaining module 406 and / or the location prediction module 408) may be software module(s) realized by computer program(s) or set(s) of instructions executable by a computer processor to perform various functions or operations. Various modules described herein (e.g., the 3D model obtaining module 206, the synthetic panoramic image generating module 208 and / or the visual localization model training module 210), together with the at least one processor 204 and the at least one memory 202, may also be implemented as hardware module(s) being functional hardware unit(s) designed to perform various functions or operations. Similarly, various modules described herein (e.g., the panoramic image obtaining module 406 and / or the location prediction module 408), together with the at least one processor 404 and the at least one memory 402, may also be implemented as hardware module(s) being functional hardware unit(s) designed to perform various functions or operations. More particularly, in the hardware sense, a module is a functional hardware unit designed for use with other components or modules. For example, a module may be implemented using discrete electronic components, or it may form a portion of an entire electronic circuit such as an Application Specific Integrated Circuit (ASIC) or a Field Programmable Gate Array (FPGA). Numerous other possibilities exist. It will also be appreciated by a person skilled in the art that a combination of hardware and software modules may be implemented. Furthermore, various operations, functions or steps of various methods described herein may be performed in parallel rather than sequentially as desired or as appropriate (e.g., as long as it does not render the mcthod(s) inoperable or unsatisfactory for its intended purpose).
[0044] FIG. 5 depicts a schematic drawing of an aerial vehicle 500 according to various embodiments of the present invention. The aerial vehicle 500 comprising: a flight controller 502 operable to control a flight of the aerial vehicle 500; an image sensor 504 (e.g., mounted on or integrally built in the aerial vehicle 500) configured to capture a panoramic image of an environment during flight; at least one memory 402; and at least one processor 404 (corresponding to the above-mentioned system 400 of perform aerial visual localization) communicatively coupled to the at least one memory 402 and the image sensor 504 and configured to perform the above-mentioned method 300 of aerial visual localization of the aerial vehicle 500 during flight comprising: obtaining, from the image sensor 504, a panoramic image of an environment captured by the image sensor 504 during flight; and predicting, usinga visual localization model trained according to the above-mentioned method 100 of training a visual localization model as described hereinbefore with reference to FIG. 1 according to various embodiments of the present invention, a location of the aerial vehicle 500 based on the panoramic image of the environment captured. In this regard, the flight controller 502 is operable to control a flight of the aerial vehicle 500 based on the predicted location of the aerial vehicle 500. In various embodiments, the least one processor 404 may be configured to perform the aerial visual localization of the aerial vehicle 500 either as a primary or sole localization system or as a secondary or backup localization system (e.g., as a backup to satellite-based and external-based localization methods).
[0045] As explained hereinbefore, it will be appreciated by a person skilled in the art that the present invention is not limited to any particular type of aerial vehicle, whether manned or unmanned, as long as localization of the aerial vehicle is desired or required to be performed. Therefore, it will be appreciated by a person skilled in the art that the aerial vehicle 500 shown in FIG. 5 is for illustration purpose only and does not limit the aerial vehicle 500 to any particular type of aerial vehicle or any specific structural configuration thereof as long as the aerial vehicle is capable of flight navigation. For example, in the case of an unmanned aerial vehicle (UAV) (which may also be referred to as an aerial robot, a rotorcraft or a drone, such as a multirotor drone), the aerial visual localization may be implemented as part of a navigation system (comprising the flight controller) of the UAV to navigate the environment autonomously (i.e., autonomous UAV navigation). In addition, operations and flight controls of various types of aerial vehicles are well known in the art and thus need not be described herein for clarity and conciseness.
[0046] It will be appreciated by a person skilled in the art that the terminology used herein is for the purpose of describing various embodiments only and is not intended to be limiting of the present invention. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0047] Any reference to an element or a feature herein using a designation such as “first”, “second” and so forth does not limit the quantity or order of such elements or features, unless stated or the context requires otherwise. For example, such designations may be used herein asa convenient way of distinguishing between two or more elements or instances of an element. Thus, a reference to first and second elements does not necessarily mean that only two elements can be employed, or that the first element must precede the second element, unless stated or the context requires otherwise. In addition, a phrase referring to “at least one of’ a list of items refers to any single item therein or any combination of two or more items therein.
[0048] In order that the present invention may be readily understood and put into practical effect, various example embodiments of the present invention will be described hereinafter by way of examples only and not limitations. It will be appreciated by a person skilled in the ail that the present invention may, however, be embodied in various different forms or configurations and should not be construed as limited to the example embodiments set forth hereinafter. Rather, these example embodiments arc provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present invention to those skilled in the art.
[0049] Various example embodiments provide a method of performing aerial visual localization using panoramic images (preferably 360-degree panoramic images), driven by a visual localization network / model (e.g., a deep convolutional neural network (DCNN)) trained for predicting / estimating a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle. Utilizing panoramic imagery, and more particularly, 360-degree panoramic images, offers the advantage of encompassing visual features or information from wider or all angles to predict the location of the aerial vehicle. Synthetic data (synthetic panoramic images) generated from multiple different data sources, such as photogrammetry data. Open Street Map (OSM) data, and official / authoritative virtual environment data (e.g., official 3D model or 3D building data of an environment), arc used to train the visual localization network. In this regard, to address a technical problem that using deep learning-based methods usually require a rather dense distribution of training data, various example embodiments advantageously generate synthetic data (synthetic panoramic images) (e.g., which is controllable and versatile) for training the visual localization model instead of relying purely or mostly on real-world image datasets. In various example embodiments, domain adaptation (e.g., using CycleGAN (disclosed in Zhu et al., “Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks,” November 2017”, http: / / arxiv.org / abs / 1703.10593)) is also used to bridge the Sim2Real (simulation-to-reality) gap and enhance model performance. In this regard, various example embodiments note that depending on the type and quality of the synthetic data, there may be a Sim2Real gap whichmay be bridged based on domain adaptation to enhance the trained visual localization network for deployment in the actual environment. In this regard, generative Al for domain adaptation / style transfer such as Pix2Pix (disclosed in Isola, et al., “Image-to-Image Translation with Conditional Adversarial Networks,” November 2016, http: / / arxiv.org / abs / 1611.07004) for paired images and CycleGAN for unpaired images can be leveraged. In various example embodiments, virtual or visual features (e.g., virtual buildings) may be generated based on OpenStreetMap (OSM) data and utilized. OSM is a publicly available map of the world, containing crowd-sourced data of roads / streets, buildings, and other features. In experiments conducted, for example, utilizing OSM features arc shown to improve localization performance (median Euclidean error) by at least 13%, and a further 20% with cycleGAN dataset augmentation (domain adaptation). Closed loop control for aerial vehicle navigation is also achieved using the trained visual localization model, for example, enabling a quadrotor prototype to hover within a l m circle in experiments conducted.
[0050] In various example embodiments, a corresponding system for performing aerial visual localization of an aerial vehicle may be implemented or installed in an aerial vehicle either as a primary or sole localization system or as a secondary or backup localization system, and may thus provide a self-reliant visual navigation system (including an absolute localization system) for aerial vehicles leveraging 360-degree panoramic images. The sole or backup localization system is independent of external references for localization (e.g., GPS) and for example, may be applied to urban ah mobility (UAM) operations where aerial vehicles fly in close proximity to built-up areas.
[0051] FIG. 6 depicts a schematic diagram of a method 600 of training a visual localization model 640 for aerial visual localization, according to various example embodiments of the present invention. The method 600 comprises obtaining a 3D model of an environment. In this regard, as illustrated in FIG. 6, the 3D model is formed based on a plurality of different data sources, such as a photogrammetry data source 612, an authoritative virtual environment data source 614 and OpenStreetMap (OSM) data source 616. As an illustrative example, the authoritative virtual environment data source 614 may comprise one or more official virtual representations (official 3D models) of one or more environments provided by the Singapore Land Authority (SLA) and referred to as Virtual Singapore.
[0052] The method 600 further comprises generating synthetic panoramic images based on the 3D model of the environment. For example, different 3D models may be obtained for different environments, and synthetic panoramic images may be generated from each 3D model.As will be described later below in further detail according to various example embodiments of the present invention, the 3D model of the environment may be formed based on photogrammetric reconstruction using images from the photogrammetry data source 612. In this regard, the 3D model may have augmented therein virtual features (e.g., virtual structures, such as buildings) from the authoritative virtual environment data source 614 and / or virtual features (e.g., virtual roads) from the OSM data source 616. The 3D model of the environment may instead be formed based on a 3D model of the environment from the authoritative virtual environment data source 614, the 3D model having augmented therein virtual features (virtual roads) from the OSM data source 616 and / or virtual features (e.g., virtual structures, such as buildings which may be missing from the 3D model from the authoritative virtual environment data source 614) from the photogrammetry data source 612 (e.g., the virtual features may be derived or obtained from the photogrammetry 3D model of the environment stored in the photogrammetry data source 612). Still alternatively, the 3D model of the environment may be formed based on OSM data from the OSM data source 616, the 3D model (e.g., a 3D landscape of the environment with virtual roads) having augmented therein virtual features (e.g., virtual structures, such as buildings) from the authoritative virtual environment data source 614 and / or virtual features (e.g., virtual structures, such as buildings) from the photogrammetry data source 612. Accordingly, in various example embodiments, multiple data sources are utilized to build a 3D model (e.g., using the Unity software) to generate synthetic panoramic images for training the visual localization model 640.
[0053] As shown in FIG. 6, in various example embodiments, generating synthetic panoramic images may comprise performing domain randomization to generate synthetic panoramic images with domain randomization. For example, Unity Perception (disclosed in Borkman et al., “Unity Perception: Generate Synthetic Data for Computer Vision,” November 2021, http: / / arxiv.org / abs / 2107.04259) may be employed for performing domain randomization when generating synthetic panoramic images. In various embodiments, generating synthetic panoramic images may further comprise performing domain adaptation on the synthetic panoramic images to enhance reality of the synthetic panoramic images. For example, CycleGan may be employed for performing domain adaptation on the synthetic panoramic images. Accordingly, in various embodiments, multiple image augmentations may be applied prior to the synthetic panoramic images being utilized by the visual localization model 640 for training to facilitate the visual localization model 640 to generalize better to real data.
[0054] As shown in FIG. 6, the method 600 further comprises training the visual localization model 640 based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle. In various example embodiments, the visual localization model 640 comprises a feature extractor network 644 (e.g., an encoder) configured to extract features from a synthetic panoramic image 642 inputted thereto and a regression network 646 (e.g., a fully connected regression network) configured to output a predicted location for the synthetic panoramic image 642 based on the extracted features. It will be appreciated by a person skilled in the art that the present invention is not limited to the example network architecture of the visual localization model 640 shown in FIG. 6, which is an illustrative example, and any artificial intelligence (Al) or machine learning model having any network architecture or configuration capable of being trained based on images to predict a location based on an image may be utihzed as desired or as appropriate without going beyond the scope of the present invention.
[0055] In various example embodiments, the method 600 may further comprises performing train-time data augmentation by employing a number of image augmentation techniques during train-time to make the localization network generalize better, such as but not limited to RandomRain, HSV (Hue, Saturation and Value), Horizontal Shift, Coarse Dropout and Elastic Transform.
[0056] Accordingly, the method 600 of training a visual localization model for aerial visual localization according to various embodiments of the present invention particularly uses synthetic panoramic images (e.g., only synthetic panoramic images) for training, and utilizes multiple data sources including non-panoramic images. In this regard, the capability for the trained visual localization model 640 to be deployed on an actual aerial vehicle (e.g., drone) for predicting or estimating the location of the aerial vehicle will be discussed and demonstrated later below according to various example embodiments of the present invention. Therefore, various example embodiments advantageously demonstrate the building of a panoramic visual localization pipeline using data from multiple sources, and training a visual localization model using synthetic panoramic images generated based on a 3D model of the environment.DATA GENERATION
[0057] In various example embodiments, multiple data sources (e.g., a photogrammetry data source 612, an authoritative virtual environment data source 614 and an OSM data source616) are utilized to build the synthetic data generation pipeline (which may also be referred to as synthetic panoramic image pipeline). Using these data sources 612, 614, 616, a visually accurate 3D model of an environment is built that is then used to generate synthetic panoramic images. An advantage of this approach is that data (from data sources) of various forms may be utilized to generate a 3D model of an environment based on which synthetic panoramic images (training data) can be generated for training a visual localization model 640 and thus, data source is not restricted or necessary to be in the form of panoramic images.Data Sources
[0058] 1) Self-Collected Photo grammetry / Image Data Source 612: Using a camera- equipped drone, both normal and panoramic images can be acquired. For example, the normal images may be utilized for photogrammetric reconstruction to build a 3D model where synthetic panoramic images can be generated. For illustration purpose only and without limitation, FIGs. 7A and 7B show an example drone setup used for data collection. In particular, FIG. 7A shows an example RTK GPS logger used to establish a correspondence between the captured image and the location ground truth for ground truth data collection and FIG. 7B shows an example drone mounted with a 360-degree camera (e.g., Insta360 Sphere camera) and the RTK GPS logger. FIG. 7C shows an example panoramic 360-degree image in 2:1 equirectangular format captured by the drone (from Kallang Riverside in Singapore) which may be inputted to the trained visual localization model 640 for inference to predict the location of the drone based on the panoramic 360-degree image captured.
[0059] In various example embodiments, a photogrammetric 3D model may be generated using any photogrammetric 3D model building method (e.g., algorithm or software) as desired or as appropriate, such as but not limited to, the Reality Capture software. For illustration purpose, FIGs. 8A to 8C show overviews of example photogrammetric 3D models built for the different environments, namely, Aerial Arena at Singapore University of Technology and Design (SUTD), Kallang Riverside and Tuas South, respectively, which are all in Singapore.
[0060] 2) Authoritative Data Source 614: With the increasing push toward smart cities, authorities such as the Singapore Land Authority (SLA) possess or provide 3D models (e.g., digital twin data) of various regions such as high fidelity 3D models and textures of the building facades.
[0061] 3) OpenStreetMap Data Source 616: In various example embodiments, to generate a 3D model of an environment, OSM data associated with the environment may be convertedinto a 3D landscape of the environment, such as but not limited to, utilizing CityGen3D, a plugin within Unity that converts OSM data into a 3D landscape. In this regard, in various example embodiments, the 3D landscape may be generated with virtual features, such as virtual terrains and virtual roads, as well as certain virtual structures (e.g., buildings), but may lack still lack virtual features (i.e., incomplete visual data in the 3D model generated), such as additional virtual structures (e.g., buildings). To address the incomplete visual data in the 3D model generated, virtual features (e.g., virtual structures, such as buildings), as well as visual textures thereof (e.g., building facade data), from the authoritative virtual environment data source 614 and / or virtual features (e.g., virtual structures, such as buildings), as well as visual textures thereof, from the photogrammetry data source 612 may be added to the 3D landscape. As an example illustration, FIG. 9A illustrates an example 3D model (3D landscape) with example virtual roads generated using CityGen3D based on OSM data but still lack virtual features, such as virtual structures (e.g., buildings). FIG. 9B illustrates the example 3D model filled-in with (augmented with) virtual structures (e.g., buildings), as well as visual textures thereof, from the authoritative virtual environment data source 614 (e.g., SLA building model and facade data). The 3D model shown in FIG. 9B also shows certain virtual buildings (buildings shown with plain facade) originally generated using CityGen3D based on OSM building information.Rendering of Visual Images
[0062] For the synthetic data generation pipeline, the synthetic panoramic images may be generated based on a 3D model of the environment using any synthetic panoramic image generating method (e.g., algorithm or software) as desired or as appropriate, such as but not limited to, the Unity software. For example, the Unity software may be utilized to take advantage of its real-time rendering. Furthermore, for example, the Unity Perception module may be utilized to perform domain randomization to generate synthetic panoramic images with domain randomization of the environment, such as the skybox, sun angle, lighting and so on. This domain randomization is built into the synthetic dataset. Data points may also be randomized following a uniform distribution within a specified bounding box. Furthermore, data points that collide with buildings may be removed from the datasets. For illustration purpose, FIGs. 10A to 10D show example synthetic panoramic images generated in Unity (at Kallang Riverside), presented in 2: 1 equirectangular format. Example domain randomizations performed for environmental conditions are also illustrated such as randomization of lighting, skies and sun angle.
[0063] In various example embodiments, the 3D model generated for an environment may be used to generate an entire training set of synthetic panoramic images (e.g., 360-degree synthetic panoramic images) fortraining the visual localization model 640. For example, virtual cameras may be placed within the 3D model in multiple (e.g., randomized) locations to obtain synthetic panoramic images from different viewpoints and angles. The virtual cameras may be configured to apply domain randomization during the image capturing process by applying and adjusting various domain randomization parameters (e.g., different lighting and background). For example, the virtual cameras may be configured to adjust (e.g., randomize) settings or values of domain randomization parameters prior to taking each image in the 3D model. In various example embodiments, a training set of synthetic panoramic images is generated for a corresponding environment for training a visual localization model for the corresponding environment.Environments
[0064] As illustrative examples, several environments in different areas of Singapore were selected for testing. Each of them possesses a unique set of features that serves as a testbed for the visual localization model 640 (e.g., an absolute visual localization (AVL) model) according to various example embodiments of the present invention.
[0065] 1) Aerial Arena at Singapore University of Technology and Design (SUTD): TheAerial Arena is an enclosed and netted semi-outdoor testing area located within the university compound, which is approximately 1000 m2.
[0066] 2) Kallang Riverside Park: A small open field located near the Singapore NationalStadium, bounded by two condominium complexes, which is approximately 47,000 m2.
[0067] 3) Tuas South Ave 16: An industrial estate with low-lying buildings and a large field with relatively low amount of features, which is approximately 51,000 m2.
[0068] 4) 3D Virtual Singapore Dataset - Ang Mo Kio (AMK): Courtesy of the SingaporeLand Authority (SLA), a high fidelity 3D model of a region of Singapore (Ang Mo Kio region) was obtained, covering approximately a 1 km x 1 km area. Using this 3D model, synthetic panoramic image data can be generated. Due to confidentiality reasons, only the building facade texture data was provided, with no terrain and contour data. In experiments conducted, a smaller area of about 150,000 m2, along Ang Mo Kio Street 31, was tested.Dataset Format
[0069] The ground truth location for each environment is projected to a flat XYZ format in metres from the lat-lon-altitude (LLA) format of each image. Each environment has its own unique reference point to be used as the origin of the projection. A summary of the datasets used is presented in Table 1 shown in FIG. 1 1 .LOCALIZATION METHODNetwork Architecture
[0070] As described hereinbefore with reference to FIG. 6, in various example embodiments, the visual localization model 640 comprises a feature extractor network 644 (e.g., an encoder, such as Xccption) configured to extract features from a synthetic panoramic image 642 inputted thereto followed by a regression network 646 (e.g., a fully connected regression network) configured to output a predicted location for the synthetic panoramic image 642 based on the extracted features. For example, a dropout layer of 0.5 is added before each dense layer (dimension of 4096) and the final output layer is of dimension 3. For example, ReLU activation is used for the dense layers and linear activation for the output. Accordingly, an input image 642 (synthetic panoramic image) is fed to an encoder 644 to be represented as a feature vector, where the fully connected regressors 646 output the position estimate for the input image 642.Training
[0071] As an illustrative example, the framework employed for building and training the visual localization model 640 is Keras (Tensorflow). For example, a batch size of 48 and a learning rate of 0.000075 is used. For example, mean absolute error (MAE) is used as the loss function for training. Furthermore, train-time data augmentation is performed by employing a number of image augmentation techniques during train-time to make the network generalize better, such as RandomRain, HSV (Hue, Saturation and Value), Horizontal Shift (with Wraparound), Coarse Dropout and Elastic Transform. Train-time data augmentation may be performed using any image augmentation method (e.g., algorithm or software) as desired or as appropriate, such as but not limited to, the Albumentations library (Buslaev et al., “Albumentations: Fast and Flexible Image Augmentations,” Information, volume 11, no. 2, page 125, February 2020). For illustration purpose, in experiments conducted, the Albumentations library was used to perform train-time data augmentation. In addition, in various example embodiments, a method may be employed during train time for automaticallylearning certain image augmentation, such as but not limited to Auto-Augment (disclosed in Cubuk et al., “AutoAugment: Learning Augmentation Strategies from Data,” 20191EEE / CVE Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 2019, pages 113-123). For example, AutoAugment transfer may be utilized which uses learned policies from the Reduced-CIFAR 10 dataset instead of learning from scratch. For example, the operations included in the AutoAugment policy are: [Color, Equalize, ShearY, Brightness, Sharpness, AutoContrast, Rotate, Posterize, Contrast, Invert, Solarize], These are mostly colorbased transformations, and in various example embodiments, translation augmentation is removed.Data Augmentation using CycleGAN
[0072] Compared to the other 3 environments where photogrammetric reconstruction was used, the Ang Mo Kio (AMK) dataset provided by the SLA is visually quite different compared to the actual location. In order to bridge this Sim2Real gap, various example embodiments perform domain adaptation on the synthetic panoramic images (training dataset) generated based on the official 3D model of AMK provided by the SLA to enhance reality of the synthetic panoramic images. In various example embodiments, domain adaptation is applied to all of the synthetic panoramic images in the training dataset generated for a particular environment. For example, various example embodiments introduce a style transfer data augmentation technique leveraging Generative Adversarial Networks (GANs). CycleGAN utilizes unpaired images to synthesize a style transfer relationship between the original and target images and vice versa. In this regaid, CycleGAN is utilized to enrich dataset diversity by capturing visual features and color palettes not present in the original synthetic dataset. In various example embodiments, the CycleGAN is trained on images from the same environment as the environment which the training dataset of synthetic panoramic images is generated for. Accordingly, the domain adaptation may be applied to the training dataset prior to being utilized on the visual localization model 640 so as to enhance the training dataset to enhance reality (e.g., visually match) to the environment. Furthermore, for training the CycleGAN network, the training images do not need to have a ground truth label for them to be used. Training data for training the CycleGAN was collected via a combination of aerial and ground videos. For example, the CycleGAN is trained on additional aerial and ground videos (exported as images by extracting the individual frames.) which do not include a ground truth and does not necessarily have to be in 360-degree image format. As an illustrative example and in experiments conducted, Keras was used to train anddeploy the CycleGAN network. Training was run for 300 epochs and a learning rate of 0.000075 was used. The output resolution was set to 384 X 384 pixels. FIGs. 12A and 12B illustrate a comparison of an example synthetic panoramic image before and after being put through the trained CycleGAN network. Images are re-scaled from 2:1 equirectangular format to 1 : 1 aspect ratio to fit the network requirements.RESULTS AND DISCUSSIONPerformance on Independent Test Set
[0073] Experimental results of the four different environments mentioned above, namely, Aerial Arena at SUTD, Kallang Riverside Park, Tuas South Ave 16 and 3D Virtual Singapore Dataset - Ang Mo Kio (AMK) Street 31 , will now be discussed. An independent test set is reserved for each environment and contains only real images (ground truth). Images in this test set were not used during the training process for both the localization and cycleGAN networks. Median Euclidean Distance error is used as the performance metric for comparison and is calculated as: Median. The visual localization performance using the trained visual localization model 640 for each environment is presented in Table II shown in FIG. 13, which presents a summary of the median error reported for each environment.
[0074] FIGs. 14A to 14D show XY plots (top view) of the predicted locations for each of the four environments on the corresponding ground truth locations provided by their respective test sets. In FIGs. 14A to 14D, each line connects the predicted point to the corresponding ground truth. Displayed on the right of each XY plot arc the images with the top 3 highest error. From the images with the highest error presented in FIGs. 14A to 14D, it can be observed that most of the images are near the ground level. This may be due to limited visibility of the surrounding environment, and also features that are present at that proximity to the ground might not match the features that the trained visual localization network 640 is looking for.
[0075] Histogram breakdowns of the Euclidean errors with the ground truth altitude for the four environments are illustrated in FIGs. 15A to 15D, respectively. As seen from the histograms, the localization performance for most of the environments drastically decreases at lower altitudes near the ground. For most of the outdoor environments, the localization starts to be more stable at around 10-15 m above ground, where more of the surrounding features start to be visible and make up most of the image, enabling the network to localize better. Thisdiscrepancy may have manifested less in the smaller environment of AerialArena at SUTD as the surrounding features are much closer to the inference points.
[0076] For the AMK test set, the aerial images include only a smaller area due to flight permit restrictions. The Ang Mo Kio (AMK) test set was supplemented with additional ground data to span the area trained for the visual localization. The aerial data is concentrated between 300-400 m (x-coordinates) and 520-620 m (y-coordinates), which is depicted by the zoomedin view shown in FIG. 14D. As can be seen in the FIG. 14D, localization failure may happen when too many of the important features are obstructed, in this case by trees and other roadside objects.Experiments on Ang Mo Kio (AMK) environment
[0077] Several variants of the AMK dataset were created, with different combinations of OSM generated entities / features added to study their effects on visual localization performance, as shown in FIG. 16. Each dataset variant was trained with the same amount of iterations. In the first or basic dataset variant, synthetic panoramic images were generated using only the data provided by the SLA authoritative source with randomized ground color added by domain randomization. The second dataset variant include synthetic panoramic images generated using the data provided by the SLA authoritative source and OSM generated roads. The third dataset variant includes synthetic panoramic images generated further using OSM generated roads and OSM generated buildings compared to the third dataset variant with default textures included by CityGen3D. The fourth dataset variant includes synthetic panoramic images generated further with randomized color of the building facade by domain randomization compared to the third dataset variant. The fifth dataset variant includes synthetic panoramic images generated further with domain adaptation compared to the fourth dataset variant by putting the synthetic panoramic images through the cycleGAN domain adaptation, which are combined with unaugmented synthetic panoramic images. In this regard, the fifth dataset variant is combined with the un-augmented images to provide the localization model from overfitting to the artifacts of the cycleGAN augmentation process. In the third to fifth dataset variants, in the respective 3D model, OSM generated buildings are used to fill up spaces where no SLA data is provided. For illustration purpose, sample synthetic panoramic images from each dataset variant of the AMK dataset is presented in FIG. 16, where from the third to fifth dataset variants, the additional OSM generated buildings which are missing in the original / raw SLA dataset can be seen. Images are re-scaled to 1 : 1 ratio for standardization. For example, generated road texturescan be observed in the sample synthetic panoramic images of the second dataset variant shown in the 2nd column of FIG. 16. In this example, missing buildings not present in the raw SLA dataset are generated via CityGen3D and OSM data (3rdcolumn of FIG. 16 onwards).
[0078] As can be seen from the results for the different variants of the AMK dataset in Table TIT shown in FIG. 17, adding additional generated features reduced the visual localization error, compared to the basic dataset variant with only the building facade added. Using the median Euclidean error as a metric, adding the generated OSM roads helped to improve the error by around 13%, and including the OSM generated buildings with random building facade color improved the error by a further 5%. Interestingly, using the default texturing scheme built into the CityGen3D program made the performance worse. This may be due to the localization network learning the wrong textural features on the building facade instead of the building outlines. When put through the cycleGAN domain adaptation process, the performance greatly improved by a further 20% compared to the fourth dataset variant with randomized building colors. It can also be seen that the biggest improvement difference occurs between the basic dataset variant (pure building-only dataset) and the second dataset variant with the OSM generated roads. From this observation, it can be inferred that in a dense urban area, the trained localization network not only looks for visual features on the building facades, but also the road layouts to determine its position.Closed-Loop Control
[0079] In experiments conducted, for illustration purpose, the capability of the trained visual localization network 640 was tested to provide closed-loop positional reference to an example actual UAV 1800 (quadrotor). FIG. 18 depicts a setup diagram for closed-loop control of the UAV 1800 and FIG. 19 depicts a schematic block diagram of the closed-loop control of a flight controller 1806 of the UAV 1800. The UAV 1800 was outfitted with a 360-degree camera 1802 (e.g., Ricoh Theta X 360 camera), which features onboard stabilization and outputs panoramic images in equirectangular format, and is fed to an on-board computer 1804 (e.g., corresponding to the processor 404 of the aerial vehicle 500 described hereinbefore with reference to FIG. 5 according to various embodiments of the present invention) for processing. The predicted location is then fed to the UAV 1800 having the flight controller 1806 running the PX4 firmware. For example, the on-board computer 1804 may communicate the predicted position to the flight controller 1806 via MavLink. As shown in FIG. 19, the flight controller 1806 may be configured with position control 1902, velocity control 1904, angle control 1906and angular rate control 1908 based on a position / location error (or position / location difference) between the predicted position fed to the flight controller 1806 and a position setpoint. rooso] In experiments conducted, a simple hover was demonstrated utilizing the onboard 360-dgrcc camera for localization. The onboard network inference time is about 700ms. FIG. 20 shows the setpoints together with the filtered XYZ position data estimated by the PX4 EKF2 filter (extended Kalman filter) using visual localization data. FIG. 21 shows the raw unfiltcrcd XYZ position as output from the visual localization network 640, inferred on the image stream from the onboard 360-degree camera. Onboard GPS is disabled. As seen from FIG. 21, the UAV 1800 manages to successfully hover about a single point using the trained visual localization network 640 as the only absolute positional reference system. Together with the EKF, the UAV 1800 manages to stay within a 1 meter circle. A step change in the Y-desired position is input at around the 800 second mark. As can be seen in the graph, the UAV 1800 successfully adjusts its position to the new setpoint. This experiment was conducted in the Aerial Arena at SUTD environment.
[0081] Accordingly, the feasibility and capability of utilizing synthetic panoramic images generated from multiple data sources in training a visual localization model 640 (e.g., neural network absolute visual localization (AVL) model) according to various example embodiments of the present invention have been successfully demonstrated. The aerial visual localization method is successfully tested and validated on multiple environments. For example, different levels of generated features using OSM data and their effects on localization performance were also studied. For example, the model performance was shown to be enhanced by at least 13% (using Median Euclidean Error) via domain adaptation (cycleGAN). Closed-loop control on a quadrotor was also demonstrated using the trained visual localization model 640.
[0082] While embodiments of the invention have been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the scope of the invention as defined by the appended claims. The scope of the invention is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced.
Claims
CLAIMS1. A method of training a visual localization model for aerial visual localization, the method comprising: obtaining a 3D model of an environment, the 3D model formed based on a plurality of different data sources; generating synthetic panoramic images based on the 3D model of the environment; and training the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
2. The method according to claim 1, wherein the plurality of different data sources comprises two or more of a photogrammetry data source, a virtual environment data source and a street map data source.
3. The method according to claim 2, wherein the virtual environment data source is an authoritative virtual environment data source and / or the street map data source is an open street map data source.
4. The method according to claim 3, wherein the 3D model of the environment is formed based on photogrammetric reconstruction using images from the photogrammetry data source, the 3D model having augmented therein virtual features from the virtual environment data source and / or virtual features from the street map data source, the 3D model of the environment is formed based on a 3D model of the environment from the virtual environment data source, the 3D model having augmented therein virtual features from the street map data source, and / or virtual features from the photogrammetry data source, or the 3D model of the environment is formed based on street map data from the street map data source, the 3D model having augmented therein virtual features from the virtual environment data source and / or virtual features from the photogrammetry data source.
5. The method according to any one of claims 1 to 4, wherein said generating synthetic panoramic images comprises performing domain randomization to generate synthetic panoramic images with domain randomization.
6. The method according to any one of claims 1 to 5, wherein said generating synthetic panoramic images further comprises performing domain adaptation on the synthetic panoramic images to enhance reality of the synthetic panoramic images.
7. The method according to any one of claims 1 to 6, wherein the synthetic panoramic images generated based on the 3D model of the environment are 360-degree synthetic panoramic images.
8. The method according to any one of claims 1 to 7. wherein the visual localization model comprises a feature extractor network configured to extract features from a synthetic panoramic image inputted thereto and a regression network configured to output a predicted location for the synthetic panoramic image based on the extracted features.
9. A system for training a visual localization model for aerial visual localization, the system comprising: at least one memory; and at least one processor communicatively coupled to the at least one memory and configured to: obtain a 3D model of an environment, the 3D model formed based on a plurality of different data sources; generate synthetic panoramic images based on the 3D model of the environment; and train the visual localization model based on the synthetic panoramic images to obtain the visual localization model trained to predict a location of an aerial vehicle based on a panoramic image of the environment captured by the aerial vehicle.
10. The system according to claim 9, wherein the plurality of different data sources comprises two or more of a photogrammetry data source, a virtual environment data source and a street map data source.
11. The system according to claim 10, wherein the virtual environment data source is an authoritative virtual environment data source and / or the street map data source is an open street map data source.
12. The system according to claim 1 1 , wherein the 3D model of the environment is formed based on photogrammetric reconstruction using images from the photogrammetry data source, the 3D model having augmented therein virtual features from the virtual environment data source and / or virtual features from the street map data source, the 3D model of the environment is formed based on a 3D model of the environment from the virtual environment data source, the 3D model having augmented therein virtual features from the street map data source, and / or virtual features from the photogrammetry data source, or the 3D model of the environment is formed based on street map data from the street map data source, the 3D model having augmented therein virtual features from the virtual environment data source and / or virtual features from the photogrammetry data source.
13. The system according to any one of claims 9 to 12, wherein said generate synthetic panoramic images comprises performing domain randomization to generate synthetic panoramic images with domain randomization.
14. The system according to any one of claims 9 to 13, wherein said generate synthetic panoramic images further comprises performing domain adaptation on the synthetic panoramic images to enhance reality of the synthetic panoramic images.1 . The system according to any one of claims 9 to 14, wherein the synthetic panoramic images generated based on the 3D model of the environment are 360-degree synthetic panoramic images.
16. The system according to any one of claims 9 to 15, wherein the visual localization model comprises a feature extractor network configured to extract features from a synthetic panoramic image inputted thereto and a regression network configured to output a predicted location for the synthetic panoramic image based on the extracted features.
17. A method of performing aerial visual localization of an aerial vehicle during flight, the method comprising: obtaining, from an image sensor of the aerial vehicle, a panoramic image of an environment captured by the image sensor during flight; and predicting, using a visual localization model trained according to the method according to any one of claims 1 to 8, a location of the aerial vehicle based on the panoramic image of the environment captured.
18. The method according to claim 17, wherein the image sensor is a 360-degree camera configured to capture a 360-dcgrcc panoramic image of the environment.
19. An aerial vehicle comprising: a flight controller operable to control a flight of the aerial vehicle; an image sensor configured to capture a panoramic image of an environment during flight; at least one memory; and at least one processor communicatively coupled to the at least one memory and the image sensor and configured to perform aerial visual localization of the aerial vehicle during flight comprising: obtaining, from the image sensor, a panoramic image of an environment captured by the image sensor during flight; and predicting, using a visual localization model trained according to the method according to any one of claims 1 to 8, a location of the aerial vehicle based on the panoramic image of the environment captured, wherein the flight controller is operable to control a flight of the aerial vehicle based on the predicted location of the aerial vehicle.
20. The aerial vehicle according to claim 19. wherein the image sensor is a 360-degree camera configured to capture a 360-degree panoramic image of the environment.
Citation Information
Patent Citations
Method and system for visual localization
US20220148219A1