Systems and methods for automatically generating a training image set for an environment

By creating two-dimensional rendered images from georeference models and generating relevant tags and link data, the problem of training image set generation in the prior art is solved and dependent on operator skills is achieved, and high-speed and accurate image set generation is achieved.

CN112504271BActive Publication Date: 2025-06-13THE BOEING CO
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010579014.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-16
Filing Date
2020-06-23
Publication Date
2025-06-13
Estimated Expiration
2040-06-23

AI Technical Summary

Technical Problem

The prior art requires manual semantic segmentation and mask application when generating image sets for machine vision system training, which is time-consuming and dependent on operator skills, especially in the case of large data sets.

Method used

By using the processor to receive the physical coordinate set and retrieve the environment model data from the georeference model, create a two-dimensional rendered image, generate link data to associate the rendered image with the tag and native image, and ultimately store the training set.

Benefits of technology

High-speed automatic generation of semantic segmented images is realized, reducing dependence on human operators, and improving the generation efficiency and accuracy of training image sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112504271B_ABST
    Figure CN112504271B_ABST
Patent Text Reader

Abstract

The present invention relates to a system and method for automatically generating a training image set for an environment. A computer-implemented method for generating a training set of images and labels for a native environment, the method comprising: receiving a set of physical coordinates, retrieving environmental model data corresponding to a georeferenced model of the environment, and creating a plurality of two-dimensional (2D) rendered images, each image corresponding to a view of one of the set of physical coordinates. The 2D rendered images include one or more environmental features. The method further comprises generating link data that associates each 2D rendered image with (i) a label for the one or more environmental features included and (ii) a corresponding native image. Additionally, the method further comprises storing the training set, the training set including the 2D rendered images, the labels, the corresponding native images, and the link data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to systems and methods for automatically generating a training image set for an environment. Background Art

[0002] At least some known machine vision systems are trained to navigate an environment detected by an image sensor (e.g., an environment detected by a camera mounted on a machine). For example, at least some known unmanned aerial vehicles (“UAVs”) utilize a trained machine vision system to autonomously navigate an environment related to various mission objectives of the UAV. As another example, at least some known autonomous motor vehicles utilize a trained machine vision system to navigate an environment related to various objectives of autonomous driving and / or autonomously pursuing the autonomous motor vehicle. Such machine vision systems are typically trained using a suitable machine learning algorithm applied to a training image set.

[0003] Such a training image set typically includes labels and metadata to facilitate machine learning. For example, training images can be semantically segmented to identify at least one feature of interest in the environment depicted in the training image. Semantic segmentation can include masks (e.g., preselected colors superimposed on respective environmental features in each training image) to train the applied machine learning algorithm to associate detected contours with the correct environmental features in the environment. The training images can also include other labels (e.g., the name of the environmental feature in the image) and metadata (e.g., a description of the viewpoint from which the image was captured, the distance from the environmental feature (e.g., if the object is a runway marker, the label may identify the distance from the marker), etc.). Known methods for generating such a training image set are subject to some limitations. For example, an operator typically manually inputs semantic segmentation into the training images and applies a mask of the appropriate color to each environmental feature in the original image. This process is more time-consuming than desired and is dependent on the skills of the operator. In addition, a large data set of such training images can be on the order of thousands of images, which can make manual segmentation impractical. Summary of the Invention

[0004] This document describes a method for generating a training set of images and labels for a native environment. The method is implemented on a computing system including at least one processor communicating with at least one storage device. The method includes using the at least one processor to receive a plurality of sets of physical coordinates and retrieve environmental model data corresponding to a georeferenced model of the environment from the at least one storage device. The environmental model data defines a plurality of environmental features. The method further includes using the at least one processor to create a plurality of two-dimensional (2D) rendered images based on the environmental model data. Each 2D rendered image corresponds to a view of one of the plurality of sets of physical coordinates. The plurality of 2D rendered images includes one or more environmental features. The method also includes using the at least one processor to generate link data that associates each 2D rendered image with (i) labels for the one or more environmental features included and (ii) corresponding native images. Additionally, the method includes using the at least one processor to store the training set, which includes the 2D rendered images, the labels, the corresponding native images, and the link data.

[0005] This document describes a computing system for generating a training set of images and labels for a native environment. The computing system includes at least one processor communicating with at least one storage device. The at least one processor is configured to receive a plurality of sets of physical coordinates and retrieve environmental model data corresponding to a georeferenced model of the environment from the at least one storage device. The environmental model data defines a plurality of environmental features. The at least one processor is further configured to create a plurality of two-dimensional (2D) rendered images based on the environmental model data. Each 2D rendered image corresponds to a view of one of the plurality of sets of physical coordinates. The plurality of 2D rendered images includes one or more environmental features. The at least one processor is also configured to generate link data that associates each 2D rendered image with (i) labels for the one or more environmental features included and (ii) corresponding native images. Additionally, the at least one processor is configured to store the training set, which includes the 2D rendered images, the labels, the corresponding native images, and the link data.

[0006] This document describes a non - transitory computer - readable storage medium having computer - executable instructions implemented thereon for generating a training set of images and labels of an environment. When executed by at least one processor in communication with at least one storage device, the computer - executable instructions cause the at least one processor to receive a plurality of sets of physical coordinates and retrieve environment model data corresponding to a georeferenced model of the environment from the at least one storage device. The environment model data defines a plurality of environmental features. The computer - executable instructions further cause the at least one processor to create a plurality of two - dimensional (2D) rendered images based on the environment model data. Each 2D rendered image corresponds to a view of one of the plurality of sets of physical coordinates. The plurality of 2D rendered images includes one or more environmental features. The computer - executable instructions further cause the at least one processor to generate link data that associates each 2D rendered image with (i) labels for the one or more included environmental features and (ii) a corresponding native image. Additionally, the computer - executable instructions cause the at least one processor to store the training set, which includes the 2D rendered images, the labels, the corresponding native images, and the link data.

[0007] There are various improvements to the features related to the above - mentioned aspects. Other features can also be incorporated into the above - mentioned aspects. These improvements and additional features can exist alone or in any combination. For example, the various features discussed below with respect to any illustrated example can be incorporated into any of the above - mentioned aspects alone or in any combination. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a schematic diagram of an image of an exemplary native environment viewed from a first advantageous position.

[0009] Figure 2A is as viewed from a first advantageous position Figure 1 of an exemplary semantic segmentation image of the native environment

[0010] Figure 2B is Figure 2A a detail view of.

[0011] Figure 3A is an exemplary data acquisition and processing framework for generating a training set of images and labels for a native environment (e.g., Figure 1 the native environment of).

[0012] Figure 3B is Figure 3A a continuation of the data acquisition and processing framework.

[0013] Figure 3C is an exemplary data acquisition and processing framework for a native environment (such as Figure 1Schematic block diagram of an exemplary computing system for generating a training set of images and labels in a native environment

[0014] Figure 4 is for Figure 1 Examples of labels, metadata, and link data in a training set for the environment shown Figures 3A - 3C generated by the computing system shown

[0015] Figure 5A is an example of a baseline 2D rendered image that can be created by the computing system shown Figures 3A to 3C in

[0016] Figure 5B is an example of a 2D rendered image that can be created by the computing system shown Figures 3A to 3C in

[0017] Figure 6A with simulated environmental changes and a background added to it, and is an example of a physical test pattern that can be viewed by a camera used in a machine vision system

[0018] Figure 6B is Figure 6A an example of an acquired image of the physical test pattern shown

[0019] Figure 7A is a flowchart of an exemplary method for generating a training set of images and labels for a native environment (such as the native environment shown Figure 3C in Figure 1 using a computing system such as the one shown

[0020] Figure 7B is Figure 7A a continuation of the flowchart

[0021] Figure 7C is Figure 7A and Figure 7B a continuation of the flowchart

[0022] Although specific features of various examples may be shown in some figures and not in others, this is for convenience only. Any feature of any figure can be referenced and / or claimed in combination with any feature of any other figure.

[0023] Unless otherwise noted, the figures provided herein are intended to illustrate the features of examples of the present disclosure. It is believed that these features can be applied to a variety of systems including one or more examples of the present disclosure. As such, the figures are not meant to include all conventional features known to those of ordinary skill in the art that are required to practice the examples disclosed herein. Detailed Description

[0024] Examples of computer-implemented methods for generating a training set of images and labels for a native environment as described herein include creating a plurality of two-dimensional (2D) rendered images from views of a georeferenced model of the native environment. A georeferenced model is broadly defined as a model of a native environment that associates the internal coordinate system of the model to a geographic coordinate system in the physical world. For example, for a particular airport environment that includes static physical environment features (such as runways, runway markings, other aircraft-traversable areas, and airport signage, each located at specific geographic coordinates in the physical world), the georeferenced model of the environment includes corresponding virtual runways, virtual runway markings, virtual other aircraft-traversable areas, and virtual airport signage, each defined by internal model coordinates that are associated to the geographic coordinates of the corresponding physical feature. Based on an input set of spatial coordinates of a selected viewpoint (such as the geographic "physical" location coordinates and physical orientation of the viewpoint), a simulated or "rendered" perspective of the virtual environment can be obtained from the georeferenced model using a suitable rendering algorithm (such as ray tracing).

[0025] Since detailed georeferenced models have been developed for many airports, the systems and methods disclosed herein are particularly useful for (but not limited to) airport environments. They are also particularly useful for (but not limited to) regulated or controlled environments (again, such as airports), as it can be expected that the nature and location of the environmental features will not change significantly over time.

[0026] Examples also include generating link data that associates each 2D rendered image with (i) labels for one or more included environmental features and (ii) corresponding native images, and storing the 2D rendered images, labels, corresponding native images, and link data in a training set. Examples of creating 2D rendered images include: using environmental model data to detect at least one environmental feature in a corresponding view; and rendering a plurality of pixels that define the detected environmental feature in the 2D rendered image for each detected environmental feature. Examples of creating labels include associating a label corresponding to each detected environmental feature in the 2D rendered image with each 2D rendered image.

[0027] In particular, since the 2D rendered images are cleanly generated from the georeferenced model and there are no uncontrolled or unnecessary elements in the images, the pixels representing environmental features in each 2D rendered image can be precisely identified by the computing system, and the computing system can apply appropriate algorithms to automatically construct or "fill in" the semantic segmentation for the pixels of each environmental feature with little or no intervention input from a human operator. Thus, the systems and methods of the present disclosure replace the manual and subjective judgments required by prior art methods for semantic segmentation of images with high-speed, automatically generated semantic segmentation images that are objectively accurate on a pixel-by-pixel basis because each semantic segmentation is precisely based on the pixels of the environmental features in the 2D rendered image.

[0028] In some examples, the set of physical coordinates used to generate the 2D rendered image defines a path through the environment. For example, the set of physical coordinates can be obtained by recording the coordinates and orientation of a vehicle traveling along the path, such as by using an on-board global positioning satellite (GPS) system, an inertial measurement unit (IMU), and / or other on-board geolocation systems of the vehicle. Thus, a training image set can be easily created for typical scenarios encountered by a self-guiding vehicle, such as the standard ways an aircraft approaches the runways of an airport or the standard ground paths a baggage transport vehicle takes to the various gates of an airport. In some such examples, the vehicle used to "capture" the path coordinates also carries sensors (such as cameras), and the images from the sensors are tagged with the physical coordinates of the vehicle or associated with the physical coordinates of the vehicle along the path by matching timestamps with the on-board GPS system. Thus, each 2D rendered image (e.g., semantic segmentation image) automatically generated at each set of physical coordinates can be associated with the native or "real" camera image captured at that set of physical coordinates, and these camera images can be used as the native images of the training set.

[0029] Unless otherwise specified, the terms "first", "second", etc. are used herein only as labels and are not intended to impose an order, position, or hierarchical requirement on the items referred to by these terms. Moreover, for example, referring to a "second" item does not require or preclude the existence of, for example, a "first" or lower-numbered item or a "third" or higher-numbered item.

[0030] Figure 1 is an exemplary schematic diagram of an image of a native environment 100 as viewed from a first advantageous position. The environment 100 includes a plurality of static physical environmental features (collectively referred to as environmental features 110), each environmental feature being associated with corresponding geographical coordinates in the physical world. In this example, the image of the native environment 100 also includes a plurality of objects 102 that are not included in the georeferenced model. For example, the objects 102 are temporary or dynamic physical objects that can be found at different locations within the native environment 100 at different times.

[0031] In this example, the native environment 100 is an airport, and the environmental features 110 include permanent or semi-permanent features that are typically present in an airport environment. For example, the environmental features 110 include a runway 120 and a centerline 122 of the runway 120. Although only one runway 120 is shown from the advantageous position for obtaining an image in Figure 1 , it should be understood that the native environment 100 may include any appropriate number of runways 120. The environmental features 110 also include a plurality of taxiways 130 and position markers 132. For example, the position markers 132 are surface markers at locations such as runway holding positions, taxiway intersections, and / or taxiway / runway intersections. The environmental features 110 also include aprons 140 and buildings 150, such as hangars or terminals. Additionally, the environmental features 110 include a plurality of signs 160, such as runway / taxiway position and / or direction signs. Although the environmental features 110 listed above are typical representatives of an airport environment, they are not unique or required for the native environment 100.

[0032] Although aspects of the present disclosure have been described in terms of an airport environment for illustrative purposes, in an alternative embodiment, the native environment 100 is any suitable environment that includes environmental features 110 that can be characterized as static physical environmental features.

[0033] Figure 3A and Figure 3B are schematic diagrams of an exemplary data acquisition and processing framework for a training set for generating images and labels for the native environment 100. Figure 3C is a schematic block diagram of an exemplary computing system 300 that is used to generate a training set of images and labels for the native environment 100 that can be used to implement the Figure 3A and Figure 3B framework. In particular, the computing system 300 includes at least one processor 302 that is configured to generate a 2D rendered image 340 based on a georeferenced model of the native environment 100.

[0034] Starting from Figure 3C , the at least one processor 302 can be configured to perform one or more operations described herein by programming the at least one processor 302. For example, the at least one processor 302 is programmed to execute a model data manipulation module 320, an image processing module 322, a data linking module 324, and / or other suitable modules that perform the steps described below.

[0035] In this example, computing system 300 includes at least one storage device 304 operatively coupled to at least one processor 302, and programs the at least one processor 302 by encoding operations as one or more computer-executable instructions 306 and providing the executable instructions 306 in the at least one storage device 304. In some examples, the computer-executable instructions are provided as a computer program product by implementing the instructions on a non-transitory computer-readable storage medium. The at least one processor 302 includes, but is not limited to, for example, a graphics card processor, another type of microprocessor, a microcontroller, or other equivalent processing means capable of executing commands of computer-readable data or programs for executing a model data manipulation module 320, an image processing module 322, a data link module 324, and / or other suitable modules as described below. In some examples, the at least one processor 366 includes, for example, but is not limited to, a plurality of processing units coupled in a multi-core configuration. In certain examples, the at least one processor 302 includes a graphics card processor programmed to execute the image processing module 322 and a general-purpose microprocessor programmed to execute the model data manipulation module 320, the data link module 324, and / or other suitable modules.

[0036] In this example, the at least one storage device 304 includes one or more means for enabling storage and retrieval of information such as executable instructions and / or other data. The at least one storage device 304 includes one or more computer-readable media, such as, for example, but not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), solid state disk, hard disk, read-only memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), and / or non-volatile RAM (NVRAM) memory. The above memory types are merely examples and are thus not limited to the types of memory that can be used as the at least one storage device 304. The at least one storage device 304 is configured to store, but is not limited to, application source code, application object code, portions of source code of interest, portions of object code of interest, configuration data, execution events, and / or any other type of data.

[0037] In this example, computing system 300 includes a display device 372 coupled to the at least one processor 302. The display device 372 presents information, such as a user interface, to an operator of the computing system 300. In some examples, the display device 372 includes a display adapter (not shown) that is coupled to a display device (not shown), such as a cathode ray tube (CRT), a liquid crystal display (LCD), an organic light emitting diode (OLED) display, and / or an "electronic ink" display. In some examples, the display device 372 includes one or more display devices.

[0038] In this example, computing system 300 includes a user input interface 370. The user input interface 370 is coupled to the at least one processor 302 and receives input from an operator of the computing system 300. The user input interface 370 includes, for example, a keyboard, a pointing device, a mouse, a stylus, a touch-sensitive panel (such as, but not limited to, a touchpad or a touch screen), and / or an audio input interface (such as, but not limited to, a microphone). A single component, such as a touch screen, can act as both the display device 372 and the user input interface 370.

[0039] In some examples, computing system 300 includes a communication interface 374. The communication interface 374 is coupled to the at least one processor 302 and is configured to couple to communicate with one or more remote devices (such as, but not limited to, a network server) and perform input and output operations for such devices. For example, the communication interface 374 includes (but is not limited to) a wired network adapter, a wireless network adapter, a mobile telecommunications adapter, a serial communication adapter, and / or a parallel communication adapter. The communication interface 374 receives data from and / or transmits data to one or more remote devices for use by the at least one processor 302 and / or storage in the at least one storage device 304.

[0040] In other examples, the computing system 300 is implemented in any suitable manner that enables the computing system 300 to perform the steps described herein.

[0041] Still referring Figure 1 to, the at least one processor 302 can access environment model data 308 corresponding to a georeferenced model of environment 100. The environment model data 308 can be compiled using, for example, aerial maps, semantic maps, feature maps, and / or other suitable spatial data or metadata about environment 100. In this example, the environment model data 308 is stored in the at least one storage device 304, and the at least one processor 302 is programmed to retrieve the environment model data 308 from the at least one storage device 304.

[0042] The environmental model data 308 includes data associated with each environmental feature 110. More specifically, the environmental model data 308 includes, for example, a unique identifier 312 for each environmental feature 110 and the type 314 of each environmental feature 110. The environmental model data 308 also includes the spatial extent 316 of each environmental feature 110 within the environment 100. In particular, the spatial extent 316 of each environmental feature 110 stored in the environmental model data 308 is associated with the geographical coordinates of the environmental feature 110. In the example, since the object 102 is not a static physical environmental feature of the environment 100, the object 102 is not represented in the environmental model data 308.

[0043] As described above, in the example, the environment 100 is an airport, and the environmental model data 308 includes unique identifiers 213 and types 314 for each runway 120, the centerline 122 of the runway 120, taxiways 130, position markers 132, aprons 140, buildings 150, and signs 160. More specifically, each individual environmental feature 110 has a unique identifier 312 that is different from the unique identifier 312 of each other environmental feature 110 included in the georeference model. Environmental features 110 of similar categories (e.g., runways 120, signs 160) share the same type 314. The classification described for this example is non-limiting. For example, the sign 160 can be further classified into types 314 such as runway signs, apron signs, etc. Alternatively, for at least some environmental features 110, the type 314 is not included in the environmental model data 308.

[0044] The spatial extent 316 of each environmental feature 110 can be determined based on the semantic map data for the reference viewpoint, the feature map data for the reference viewpoint, and / or other metadata within the environmental model data 308. Additionally or alternatively, the spatial extent 316 is stored in a predefined data structure by using values. Different data structures can be defined corresponding to the type 314 of the environmental feature 110. For example, the spatial extent 316 is defined for some environmental features 110 using boundary coordinates. Alternatively, the spatial extent 316 is defined and / or stored in any suitable manner such that the computer system 300 can function as described herein.

[0045] The at least one processor 302 is also programmed to receive a plurality of sets of physical coordinates 332. In this example, the sets of physical coordinates 332 are stored in the at least one storage device 304, and the at least one processor 302 is programmed to retrieve the sets of physical coordinates 332 from the at least one storage device 304. Each set of physical coordinates 332 defines a favorable location in physical space from which the environment 100 can be viewed. For example, each set of physical coordinates 332 includes a location (e.g., latitude, longitude, altitude) and a direction of view from that location (e.g., heading, angle of attack, roll). Overall, the plurality of sets of physical coordinates 332 represent all of the favorable points for the images in the training set 350 to be used to train the machine vision system 362 (e.g., the spatial relationship of the viewer in six degrees of freedom relative to a geospatial coordinate system). More specifically, the training set 350 includes native images 356 and corresponding 2D rendered images 340, and the training algorithm 360 is programmed to use the corresponding pairs of native images 356 and 2D rendered images 340 to train the machine vision system 362 to identify environmental features 110.

[0046] In some examples, the sets of physical coordinates 332 define a path through the environment 100. For example, the sets of physical coordinates 332 are a series of points through which a UAV or other aircraft may approach and land on the runway 120 and / or taxi on the taxiway 130 and apron 140 towards the building 150. As another example, the sets of physical coordinates 332 are a series of points through which an autonomous baggage transport vehicle (not shown) may travel between various gates of the building 150 and the apron 140. Alternatively, the sets of physical coordinates 332 are not related to a path through the environment 100.

[0047] In some examples, the at least one processor 302 receives a plurality of camera images 330 each associated with one of a set of physical coordinates 332. For example, a test vehicle 380 equipped with a camera 382 traverses a path defined by the set of physical coordinates 332, and a recording device 384 records the camera images 330 captured along the path. It should be understood that the terms “camera image” and “camera” broadly refer to any type of image that can be obtained by any type of image capture device, and are not limited to images captured in visible light or images captured via a lens camera. In certain examples, the at least one processor 302 is programmed to receive the camera images 330 via a communication interface 374 and store the camera images 330 in the at least one storage device 304. Additionally, in some examples, the at least one processor 302 is programmed to include the camera images 330 as native images 356 in a training set 350. Alternatively, the at least one processor 302 is not programmed to receive the camera images 330, and / or does not include the camera images 330 as native images 356 in the training set 350.

[0048] In certain examples, each camera image 330 includes a corresponding geocoordinate tag, and the at least one processor 302 is programmed to receive the set of physical coordinates 332 by extracting the geocoordinate tag from the camera image 330. For example, the test vehicle also includes an on-board geolocation system 386 (e.g., a GPS receiver and / or an inertial measurement unit (IMU)), and as the test vehicle 380 captures each camera image 330 along the path, the corresponding set of physical coordinates 332 is captured from the on-board geolocation system 386 and embedded as a geocoordinate tag into the captured camera image 330. Alternatively, the camera images 330 and the set of physical coordinates 332 are recorded separately by the camera 382 and the on-board geolocation system 386, respectively, along with timestamps, and the respective timestamps are synchronized to associate each camera image 330 with the correct set of physical coordinates 332. Alternatively, the at least one processor 302 is programmed to receive the set of physical coordinates 332 in any suitable manner, such as listing numerical coordinates and orientation values in a text file, for example. In some examples, the at least one storage device 304 also stores the displacement and orientation of each camera 382 relative to the on-board geolocation system 386, and the at least one processor 302 is programmed to adjust the geocoordinate tag for each camera 382 based on the displacement and orientation of the respective camera 382 to obtain a more accurate set of physical coordinates 332.

[0049] As described above, systems in the prior art for creating a training set would require a human operator to manually identify environmental features 110 in camera images 330 and perform semantic segmentation on each environmental feature 110 in each camera image 330 to complete the training image set, which would be very time-consuming and would result in a subjective and not strictly precise pixel-by-pixel fitting of the environmental features 110. By creating 2D rendered images 340 from environmental model data 308 and automatically performing semantic segmentation on the 2D rendered images 340 to create semantic segmentation images 352, the computing system 300 provides an advantage over such prior art systems.

[0050] In this example, each 2D rendered image 340 corresponds to a view of one of a plurality of sets of physical coordinates 332. For example, the at least one processor 302 is programmed to apply a suitable rendering algorithm to the environmental model data 308 to detect each environmental feature 110 that appears in the view defined by a given set of physical coordinates 332 and render a plurality of pixels 342 for each detected environmental feature 110, where these pixels 342 define the detected environmental feature 110 in the resulting 2D rendered image 340. For example, the spatial extent 316 enables the at least one processor 302 to determine whether a corresponding environmental feature 110 appears within a bounding box or region of interest (ROI) associated with the view of the environment 100 defined by the specified set of physical coordinates 332. The algorithm corresponds the view defined by the set of physical coordinates 332 with the spatial extent 316 of each detected environmental feature 110 in the ROI to render a plurality of rendered pixels 342. Suitable rendering algorithms (such as, but not limited to, ray tracing algorithms) are known and need not be discussed in depth for the purposes of this disclosure. One such ray tracing algorithm is provided in the Unity Pro product sold by Unity Technologies ApS of San Francisco, California. In this example, each 2D rendered image 340 is stored as a Portable Network Graphics (PNG) image file in at least one storage device 304. Alternatively, each 2D rendered image 340 is stored in any suitable format such that the training set 350 can function as described herein.

[0051] In this example, the at least one processor 302 is programmed to create a 2D rendered image 340 that includes a semantic segmentation image 352. More specifically, to create each semantic segmentation image 352, the at least one processor 302 is programmed to apply a visualization mode that renders a plurality of pixels 342 corresponding to each detected environmental feature 110 in a corresponding semantic color. In some examples, the at least one processor 302 is programmed in a semantic segmentation visualization mode to associate each type 314 of environmental feature 110 with a preselected color, and for each detected environmental feature 110 in the 2D rendered image 340, to render the pixels 342 with the preselected color associated with the type 314 of the respective environmental feature 110. Thus, for example, all runways 120 can be rendered with the same bright red color to create the semantic segmentation image 352. In some examples, "background" pixels that do not correspond to detected environmental features 110 are rendered in a neutral background color to enhance the contrast with the pixels 342, or alternatively are rendered with a naturalized RGB background color palette. For example, a color key for each type 314 is included in the metadata 358 of the training set 350, or alternatively the training algorithm 360 is programmed in some other way to associate the respective preselected colors with the corresponding types 314 of environmental features 110. Alternatively, the at least one processor 302 automatically and precisely determines the pixels 342 corresponding to each of the one or more detected environmental features 110 during the rendering of the 2D rendered image 340, and automatically and precisely colors those pixels 342 to create the semantic segmentation image 352, such that the computing system 300 can generate a semantic segmentation image 352 with precise pixel-level accuracy in a high-speed process without the need for manual learning or manual manipulation of the native image 356.

[0052] In this example, the at least one processor 302 is programmed to selectively create 2D rendered images 340 in a variety of visualization modes. For example, in addition to the visualization mode for creating the semantic segmentation image 352 described above, the at least one processor 302 is programmed to apply an RGB visualization mode to create additional 2D rendered images 340 for each set of physical coordinates 332. For example, the plurality of pixels 342 corresponding to each detected environmental feature 110 are rendered with a naturalized red-green-blue (RGB) feature color palette, e.g., associated with the corresponding type 314 or unique identifier 312 of the detected environmental feature 110, and the background is further rendered with a naturalized RGB background color palette to create a 2D composite image 344 that approximates the physical appearance of the corresponding native image 356. In some such embodiments, the semantic segmentation images 352 can be conceptualized as corresponding to the underlying 2D composite image 344, but with feature-type-based semantic colors superimposed on each environmental feature 110. As another example, the at least one processor 302 is programmed to apply a depth map visualization mode to create additional 2D rendered images 340 for each set of physical coordinates 332. In the depth map visualization mode, the plurality of pixels 342 corresponding to each detected environmental feature 110 are rendered with a color scale corresponding to the physical distance of the pixels 342 from the position coordinates of the set of physical coordinates 332 to create a 2D depth map 346. The 2D composite (RGB) images 344 and 2D depth maps 346 can be used in the training set 350 to improve the performance of the training algorithm 360. In some examples, the at least one processor 302 is programmed to apply additional or alternative visualization modes to create additional 2D rendered images 340 for each set of physical coordinates 332.

[0053] Turning now to Figure 2A , a schematic diagram of an exemplary semantic segmentation image 352 of the environment 100 as viewed from a first vantage point such as Figure 1 is presented. Figure 2B is Figure 2A a detail view of. One or more environmental features 110 detected by the rendering algorithm based on the environmental feature data 310 include a runway 120, a centerline 122 of the runway 120, a taxiway 130, position markers 132, an apron 140, a building 150, and a sign 160 (as Figure 1As shown. The at least one processor 302 renders the pixels 342 corresponding to the runway 120 in a first color 220, the pixels 342 corresponding to the centerline 122 in a second color 222, the pixels 342 corresponding to the taxiway 130 in a third color 230, the pixels 342 corresponding to the position marker 132 in a fourth color 232, the pixels 342 corresponding to the apron 140 in a fifth color 240, the pixels 342 corresponding to the building 150 in a sixth color 250, and the pixels 342 corresponding to the sign 160 in a seventh color 260.

[0054] Figure 4 Are examples of labels 404 and 452, metadata 358, and link data 354 generated by the computing system 300 for the training set 350 of the environment 100. The process of generating the training set 350 also includes using the at least one processor 302 to generate link data 354 that associates each 2D rendered image 340 with a corresponding native image 356. In this example, the at least one processor 302 generates the link data 354 as a data structure 400 that includes a plurality of records 401. For example, the data structure 400 is a comma-separated values (CSV) file or a table in a database. Each record 401 includes a first pointer 402 that points to at least one of the 2D rendered images 340 and a second pointer 403 that points to the corresponding native image 356. The training algorithm 360 is configured to parse the training set 350 based on the information in the data structure 400 for the corresponding 2D rendered image 340 and native image 356 pair. It should be understood that the term "pointer" as used herein is not limited to a variable that stores an address in a computer memory, but rather more broadly refers to any information element that identifies a location (e.g., file path, storage location) accessible by an object associated with the pointer (e.g., 2D rendered image 340, native image 356).

[0055] In this example, the first pointer 402 is implemented as the file path and file name of an image file stored in the at least one storage device 304 and storing the 2D rendered image 340, and the second pointer 403 is implemented using time metadata 440 corresponding to the time when the corresponding native image 356 was captured as a camera image 330. For example, a timestamp is stored with each native image 356 (e.g., as metadata in the image file). The training algorithm 360 parses the timestamp stored with each native image 356, finds the corresponding timestamp 442 in the time metadata 440 of one of the records 401, and follows the first pointer 402 in the identified record 401 to find the 2D rendered image 340 corresponding to the native image 356.

[0056] Alternatively, the set of physical coordinates 332 is used as the second pointer 403 and is used to match each native image 356 with the corresponding record 401 in a manner similar to the above-described timestamp 442. For example, the set of physical coordinates 332 used to generate the 2D rendered image 340 is stored in the corresponding record 401 and is matched with the set of physical coordinates captured and stored with each native image 356 (e.g., as metadata in an image file). Alternatively, the second pointer 403 is implemented as a file name and a path to the stored native image 356. Alternatively, each record 401 includes a first pointer 402 and a second pointer 403 implemented in any suitable manner such that the training set 350 can function as described herein.

[0057] Alternatively, each record 401 associates the 2D rendered image 340 with the corresponding native image 356 in any suitable manner such that the training set 350 can function as described herein.

[0058] In this example, the training set 350 includes a visualization mode label 404 for each 2D rendered image 340. For example, the data structure 400 includes a label 404 for the 2D rendered image 340 associated with each record 401. The visualization mode label 404 identifies the visualization mode used to create the image, such as "SEM" for a semantic segmentation image 352, "RGB" (i.e., red-green-blue) for a 2D composite (RGB) image 344, and "DEP" for a depth map 346. Alternatively, the visualization mode label 404 is not included in the training set 350. For example, the training set 350 includes 2D rendered images 340 of a single visualization mode.

[0059] In this example, the training set 350 further includes a feature label 452 for each environmental feature 110 detected in the 2D rendered image 340. For example, the feature label 452 is a text string based on the unique identifier 312 and / or type 314 of the detected environmental feature 110. Although only one feature label 452 is shown in each record 401 of Figure 4 it should be understood that any number of feature labels 452 can be included in each record 401 based on the number of environmental features 110 detected in the corresponding 2D rendered image 340. Additionally or alternatively, the training set 350 includes any suitable additional or alternative labels such that the training set 350 can function as described herein.

[0060] In some examples, each record 401 also includes metadata 358. In Figure 4In the example, in the case where the physical coordinate set 332 does not yet exist as the second pointer 403, the metadata 358 includes the physical coordinate set 332. In this example, the physical coordinate set 332 is represented as latitude 432, longitude 434, altitude 436, and heading 438. Additional orientation variables in the physical coordinate set 332 include, for example, angle of attack and roll angle (not shown). Alternatively, each record 401 is associated with the physical coordinate set 332 in any suitable manner such that the training set 350 can function as described herein. In some embodiments, the physical coordinate set 332 is represented by additional and / or alternative data fields in any suitable coordinate system (e.g., polar coordinates or WGS84 GPS) relative to a georeference model.

[0061] In this example, the metadata 358 also includes a sensor index 406 for the corresponding native image 356. For example, the test vehicle 380 includes a plurality of cameras 382, and the camera among the plurality of cameras 382 corresponding to the second pointer 403 in the record 401 and associated with the native image 356 is identified by the sensor index 406. In some examples, as described above, the at least one storage device 304 stores the displacement and orientation of each camera 382 relative to the on-board geolocation system 386, and the at least one processor 302 retrieves the stored displacement and orientation based on the sensor index 406 to adjust the physical coordinate set 332 of the corresponding camera 382. Alternatively, the sensor index 406 is not included in the metadata 358.

[0062] In this example, the metadata 358 also includes time metadata 440. For example, the time metadata 440 includes a relative transit time 444 along the path calculated from a timestamp 442. In the case where the timestamp 442 does not yet exist as the second pointer 403, the metadata 358 also includes the timestamp 442. Alternatively, the metadata 358 does not include the time metadata 440.

[0063] In some examples, the metadata 358 also includes spatial relationship metadata 454 associated with at least some of the feature tags 452. More specifically, the spatial relationship metadata 454 defines the spatial relationship between the physical coordinate set 332 and the detected environmental feature 110 corresponding to the feature tag 452. For example, the training algorithm 360 is configured to train the machine vision system 362 to recognize the distance to certain types 314 of environmental features 110, and the spatial relationship metadata 454 is used for this purpose in the training algorithm 360. In some examples, as described above, the at least one processor 302 is programmed to use the spatial relationship metadata 454 in creating the 2D depth map 346.

[0064] In this example, the spatial relationship metadata 454 is implemented as a distance. More specifically, the at least one processor 302 is programmed to calculate, for each 2D rendered image 340, a straight-line distance from the corresponding set of physical coordinates 332 to each detected environmental feature 110, based on the environmental model data 308. Alternatively, the spatial relationship metadata 454 includes any suitable set of parameters, such as relative (x, y, z) coordinates.

[0065] In this example, the at least one processor 302 is programmed to generate additional link data 450 that associates the spatial relationship metadata 454 with the corresponding 2D rendered image 340, and store the spatial relationship metadata 454 and the additional link data 450 as part of a training set 350 in the at least one storage device 304. For example, the additional link data 450 is implemented by including each feature tag 452 and the corresponding spatial relationship metadata 454 in a record 401 of a data structure 400 corresponding to the 2D rendered image 340. Alternatively, the spatial relationship metadata 454 and / or the additional link data 450 are stored as part of the training set 350 in any suitable manner such that the training set 350 can function as described herein. For example, the additional link data 450 and the spatial relationship metadata 454 are stored in the metadata of an image file that stores the corresponding 2D rendered image 340.

[0066] It should be understood that in some embodiments, the link data 354 includes additional and / or alternative fields from Figure 4 those shown.

[0067] Figure 5A is an example of a baseline 2D rendered image 500 created from one of the environmental model data 308 and the set of physical coordinates 332 as described above. Figure 5B is an example of a 2D rendered image 340 corresponding to the baseline 2D rendered image 500 and having simulated environmental changes 502 and a background 504 added thereto.

[0068] In this example, the baseline 2D rendered image 500 is one of the 2D rendered images 340 generated from environmental model data 308 using, for example, a ray tracing algorithm and including a default background image appearance and / or default variable environmental effects (such as lighting effects for weather, time of day). In some cases, the native image 356 and / or the images seen by the machine vision system 362 through one or more of its cameras in the field of view may include various backgrounds, weather, and / or lighting for time of day. In certain examples, this may result in a mismatch between the native image 356 on the one hand and the 2D rendered images 340 in the training set 350 and the images seen by the machine vision system 362 through one or more of its cameras in the field of view, which may reduce the effectiveness of the training set 350.

[0069] In some examples, the at least one processor 302 is also programmed to create a 2D rendered image 340 having a plurality of simulated environmental variations 502 and / or backgrounds 504 to account for such diversity in real-world images. In some such embodiments, the at least one processor 302 is programmed to apply a plurality of modifications to the environmental model data 308, each modification corresponding to a different environmental variation 502 and / or a different background 504. For example, the environmental model data 308 is modified to include different positions of the sun, resulting in a rendering of the 2D rendered image 340 with environmental variations 502 and backgrounds 504 that include lighting effects representative of different times of day. For another example, the environmental model data 308 is modified to include a 3-D distribution of water droplets corresponding to a selected cloud, fog, or precipitation distribution in the environment 100, and the at least one processor 302 is programmed to associate appropriate light scattering / light diffraction characteristics with the water droplets, resulting in a rendering of the 2D rendered image 340 with environmental variations 502 and backgrounds 504 that include weather-induced visibility effects and / or cloud formation backgrounds.

[0070] Additionally or alternatively, the at least one processor 302 is programmed to apply a ray tracing algorithm only with a default background image aspect and / or default variable environmental effects to produce a baseline 2D rendered image 500, and to directly apply 2D modifications to the baseline 2D rendered image 500 to create an additional 2D rendered image 340 having a plurality of simulated environmental variations 502 and / or backgrounds 504. For example, the at least one processor 302 is programmed to overlay each baseline 2D rendered image onto a standing 2D image representative of a plurality of different backgrounds 504 to create an additional 2D rendered image 340 having, for example, different cloud formation backgrounds. In one embodiment, the at least one processor 302 identifies portions of the baseline 2D rendered image 500 corresponding to the sky, marks those portions to facilitate deletion of pixels of the background 504 when the baseline 2D rendered image 500 is overlaid on the background 504 (e.g., the at least one processor 302 treats these portions as a "green screen"), and overlays each modified baseline 2D rendered image 500 on one or more backgrounds 504 of weather and / or time of day to create one or more additional 2D rendered images 340 from each baseline 2D rendered image 500. Figure 5B Shown overlaid on a "cloudy day" background 504 Figure 5A is the baseline 2D rendered image 500, which includes one of a runway 120, a centerline 122, and a marker 160.

[0071] For another example, the at least one processor 302 is programmed to apply a 2D lighting effect algorithm to a baseline 2D rendered image 500 to create an additional 2D rendered image 340 having an environmental variation 502 corresponding to a particular time of day or other ambient light conditions in the environment 100. In Figure 5B the example shown, a "lens flare" environmental variation 502 is added to Figure 5A the shown baseline 2D rendered image 500, and it is created by a lighting modification algorithm that propagates from the upper right corner of the baseline 2D rendered image 500. For another example, a raindrop algorithm simulates raindrops 506 on a camera lens by applying a local fisheye lens optical distortion to one or more random locations on the baseline 2D rendered image 500 to create a corresponding additional 2D rendered image 340. In some examples, similar 2D modifications to the baseline 2D rendered image 500 are used to create an additional 2D rendered image 340 having an environmental variation 502 representing a reduction in visibility caused by clouds or fog.

[0072] Accordingly, the computing system 300 enables the generation of a set of training images of the environment under various environmental conditions without waiting for or relying on changes in the time of day or weather conditions.

[0073] Additionally or alternatively, the at least one processor 302 is programmed in any suitable manner such that the training set 350 can function as described herein to create 2D rendered images 340 having environmental variations 502 and / or different backgrounds 504, or is not programmed to include environmental variations 502 and / or backgrounds 504.

[0074] Similarly, in some examples, the at least one processor 302 is further programmed to apply simulated intrinsic sensor effects in the 2D rendered image 340. Figure 6A is an example of a physical test pattern 600 that can be viewed by a camera used by the machine vision system 362. Figure 6BThis is an example of the acquired image 650 of the physical test pattern 600 obtained by the camera. The physical test pattern 600 is a checkerboard pattern defined by horizontal lines 602 and vertical lines 604. However, due to the inherent sensor effects of the camera, the acquired image 650 of the physical test pattern 600 is distorted such that the horizontal lines 602 become curved horizontal lines 652 and the vertical lines 604 become curved vertical lines 654. Without further modification, the 2D rendered image 340 generated from the environmental model data 308 using, for example, a ray tracing algorithm does not include the bending distortion caused by the inherent sensor effects, as exemplified in the acquired image 650. In some examples, this can lead to a mismatch between the native image 356 and, on the one hand, the 2D rendered image 340 in the training set 350 and the image seen by the machine vision system 362 through one or more of its cameras in the field of view, which can reduce the effectiveness of the training set 350.

[0075] In this example, the at least one processor 302 is programmed to take into account such inherent sensor effects when creating the 2D rendered image 340. In other words, the at least one processor 302 is programmed to deliberately distort the non-distorted rendered image. More specifically, the at least one processor 302 is programmed to apply simulated inherent sensor effects in the 2D rendered image 340. For example, initially, a 2D rendered image 340 is created from the environmental model data 308 at a view corresponding to the set of physical coordinates 332 using, for example, a suitable ray tracing algorithm as described above, and then an inherent sensor effect mapping algorithm is applied to the initial output of the ray tracing algorithm to complete the 2D rendered image 340.

[0076] For example, one such inherent sensor effect mapping algorithm is to map the x and y coordinates of each initial 2D rendered image to xd and yd coordinates according to the following formula to generate the corresponding 2D rendered image 340:

[0077] xd = x(1 + k1 r 2 + k2 r 4 ); and

[0078] yd = y(1 + k1 r 2 + k2 r 4 );

[0079] where r = the radius from the center of the initial 2D rendered image to the point (x, y).

[0080] For example, the factors k1 and k2 of a particular camera are determined by comparing the acquired image 650 captured by the camera with the physical test pattern 600. For a camera with a fish-eye lens, a suitable extended mapping can also be used to determine and apply another factor k3. Alternatively, the at least one processor 302 is programmed to take into account such inherent sensor effects when creating the 2D rendered image 340 in any suitable manner such that the training set 350 can function as described herein, or is not programmed to include the inherent sensor effects.

[0081] Additionally or alternatively, the at least one processor 302 is programmed to apply any suitable additional processing when creating the 2D rendered image 340 and / or when processing the native image 356. For example, at least some known examples of the training algorithm 360 perform better on a training image set 350 having a relatively low image resolution. The at least one processor 302 can be programmed to reduce the image resolution of the camera image 330 before storing the camera image 330 as the native image 356 and create a 2D rendered image 340 having a corresponding reduced image resolution. For another example, at least some known examples of the training algorithm 360 perform better on a training set 350 that does not include a large range of unsegmented background images. The at least one processor 302 can be programmed to crop the camera image 330 and / or the 2D rendered image 340 before storage.

[0082] Figure 7A is a flow chart of an exemplary method 700 for generating an image and labels for a native environment (such as native environment 100) for a training set (such as training set 350). As described above, method 700 is implemented on a computing system that includes at least one processor in communication with at least one storage device (e.g., computing system 300 that includes at least one processor 302 in communication with at least one storage device 304). In this example, the steps of method 700 are implemented by the at least one processor 302. Figure 7B and Figure 7C is Figure 7A a continuation of the flow chart of.

[0083] Still referring Figures 1 to 6B to, in this example, method 700 includes receiving 702 a plurality of sets of physical coordinates 332. In some examples, the step of receiving 702 the set of physical coordinates 332 includes receiving 704 the camera image 330 captured during a physical traversal path and extracting 706 the geographical coordinate labels from the camera image 330 to obtain the set of physical coordinates 332. In certain examples, the step of receiving 702 the set of physical coordinates 332 includes receiving 708 a list of digital coordinate values.

[0084] In this example, method 700 further includes retrieving 710 environmental model data 308 corresponding to a georeferenced model of environment 100. The environmental model data 308 defines a plurality of environmental features 110.

[0085] In this example, method 700 further includes creating 712 a 2D rendered image 340 from the environmental model data 308. Each 2D rendered image 340 corresponds to a view of one of a plurality of sets of physical coordinates 332. The plurality of 2D rendered images 340 includes one or more environmental features 110. In some examples, the step of creating 712 the 2D rendered image 340 includes applying 714 a plurality of modifications to the environmental model data 308, each modification corresponding to a different environmental change 502 and / or a different background 504. Additionally or alternatively, the step of creating 712 the 2D rendered image 340 further includes determining 716 a visualization mode for each 2D rendered image 340, and rendering, for each of one or more environmental features 110, a plurality of pixels defining the environmental feature in a color corresponding to the determined visualization mode (e.g., semantic segmentation mode, RGB mode, or depth mode). Additionally or alternatively, in some examples, the step of creating 712 the 2D rendered image 340 includes applying 720 a simulated intrinsic sensor effect in the 2D rendered image 340. In some examples, as an alternative to or in addition to step 714, the step of creating 712 the 2D rendered image 340 includes directly applying 722 a 2D modification to a baseline 2D rendered image 500 to create an additional 2D rendered image 340 having a plurality of simulated environmental changes 502 and / or backgrounds 504.

[0086] In some examples, method 700 further includes creating 724 at least one of labels and metadata. For example, the step of creating 724 at least one of labels and metadata includes assigning 726 a visualization mode label 404 to each 2D rendered image 340. For another example, the step of creating 724 at least one of labels and metadata includes generating 728 a corresponding feature label 452 for each 2D rendered image 340 in which an environmental feature 110 appears, for each of one or more environmental features 110. For another example, the step of creating 724 at least one of labels and metadata includes calculating 730 metadata 358 based on the environmental model data 308, the metadata 358 including spatial relationship metadata 454 from the corresponding set of physical coordinates 332 to at least one detected environmental feature 110.

[0087] In this example, method 700 further includes generating 736 link data 354 that associates each of the 2D rendered images 340 with (i) tags of one or more included environmental features 110 and (ii) corresponding native images 356. In some examples, the step of generating 736 link data 354 includes generating 738 link data 354 as a data structure 400 that includes a plurality of records 401, and each record 401 includes a first pointer 402 that points to at least one of the 2D rendered images 340 and a second pointer 403 that points to the corresponding native image 356. In certain examples, each camera image 330 is associated with a timestamp 442 corresponding to the relative time 444 traversed along the path, and the step of generating 736 link data 354 includes generating 740 link data 354 as a data structure 400 that includes a plurality of records 401, and each record 401 includes (i) a pointer 402 that points to at least one of the 2D rendered images 340 and (ii) time metadata 440 that includes at least one of the timestamp 442 and the relative time 444 associated with the corresponding native image 356.

[0088] In certain examples, method 700 further includes, for each 2D rendered image 340, generating 744 additional link data 450 that associates metadata 358 with the 2D rendered image 340, and storing 746 the metadata 358 and the additional link data 450 as part of a training set 350. In some examples, the step of storing 746 the metadata 358 and the additional link data 450 includes including 748, in at least one record 401 of the data structure 400 of the link data 354, functional tags 452 for each of one or more environmental features 110 and spatial relationship metadata 454 from each of the one or more environmental features 110 to a corresponding set of physical coordinates 332.

[0089] In this example, method 700 further includes storing 750 the training set 350, which includes the 2D rendered images 340, tags such as visualization mode tags 404 and / or feature tags 452, corresponding native images 356, and link data 354. In some examples, the training set 350 is subsequently transmitted to a training algorithm 360 to train a machine vision system 362 to navigate the environment 100.

[0090] The examples of the computer-implemented methods and systems described above for generating a training set for a native environment utilize 2D rendered images created from views of a georeferenced model of the environment. These examples include rendering pixels that define each detected environmental feature in each 2D rendered image according to a preselected color scheme (such as semantic segmentation, red-green-blue natural scheme, or depth map), and generating link data that associates the 2D rendered images with corresponding native images. The examples also include storing a training set that includes the 2D rendered images, native images, labels, and link data. In some examples, camera images captured along a path through the environment are used as the native images in the training image set, and a set of physical coordinates is extracted from the physical location and orientation of the camera at each image capture. In some examples, at least one of external sensor effects, intrinsic sensor effects, and varying background imagery is added to the 2D rendered images to create a more robust training set.

[0091] Exemplary technical effects of the methods, systems, and devices described herein include at least one of the following: (a) high-speed automatic generation of semantic segmentation images for a training set; (b) generation of objectively accurate semantic segmentation images on a per-pixel basis; (c) generation of a training image set of an environment under various environmental conditions without waiting for or relying on physical scene adjustments; and (d) simulation of various external and / or intrinsic sensor effects in computer-generated 2D rendered images without the need for any physical cameras and / or physical scene adjustments.

[0092] In addition, the present disclosure also includes embodiments according to the following clauses:

[0093] 1. A method (700) for generating a training set (350) of images and labels for a native environment (100), the method being implemented on a computing system (300) that includes at least one processor (302) in communication with at least one storage device (304), the method comprising using the at least one processor to:

[0094] Receive a plurality of sets of physical coordinates (332);

[0095] Retrieve, from the at least one storage device, environmental model data (308) corresponding to a georeferenced model of the native environment, the environmental model data defining a plurality of environmental features (110);

[0096] Create a plurality of two-dimensional (2D) rendered images (340) based on the environmental model data, each 2D rendered image corresponding to a view of one of the plurality of sets of physical coordinates, the plurality of 2D rendered images including one or more of the plurality of environmental features;

[0097] Generate link data (354) that associates each of the 2D rendered images with (i) a label (452) for one or more included environmental features and (ii) a corresponding native image (356); and

[0098] Store the training set in the at least one storage device, the training set including the 2D rendered images, the labels, the corresponding native images, and the link data.

[0099] 2. The method (700) according to clause 1, the method further comprising using the at least one processor (302) to perform the following steps: for each environmental feature (110) among the one or more environmental features, generate a corresponding label (452) for each 2D rendered image (340) in which the environmental feature appears.

[0100] 3. The method (700) according to any one of clauses 1 to 2, wherein the plurality of sets of physical coordinates (332) define a path through the native environment (100), the method further comprising using the at least one processor (302) to perform the following steps: receive a plurality of camera images (330) recorded during physically traversing the path, wherein each of the plurality of camera images is associated with one of the plurality of sets of physical coordinates.

[0101] 4. The method (700) according to clause 3, wherein each of the plurality of camera images (330) includes a corresponding geocoordinate label, and wherein the method further comprises using the at least one processor (302) to perform the following steps: extract the geocoordinate labels from the plurality of camera images to receive the plurality of sets of physical coordinates (332).

[0102] 5. The method (700) according to clause 3, wherein each of the plurality of camera images (330) is associated with a timestamp (442) corresponding to a relative traversal time along the path, and wherein the method further comprises using the at least one processor (302) to perform the following steps: generate the link data (354) as a data structure (400) including a plurality of records (401), each record including (i) a pointer (402) to at least one of the plurality of 2D rendered images (340) and (ii) time metadata (440), the time metadata (440) including at least one of the timestamp and the relative traversal time associated with the corresponding native image (356).

[0103] 6. The method (700) according to any one of clauses 1 to 5, the method further comprising using the at least one processor (302) to perform the following steps: generating the link data (354) as a data structure (400) including a plurality of records (401), each record including a first pointer (402) pointing to at least one of the plurality of 2D rendered images and a second pointer (403) pointing to the corresponding native image (356).

[0104] 7. The method (700) according to any one of clauses 1 to 6, the method further comprising using the at least one processor (302) to perform the following steps for each 2D rendered image (340):

[0105] Based on the environmental model data (308), calculating metadata (358, 454), the metadata (358, 454) including a spatial relationship from a corresponding set of physical coordinates (332) to at least one of the one or more environmental features (110);

[0106] Generating additional link data (354) associating the metadata with the 2D rendered image; and

[0107] Storing the metadata and the additional link data as part of the training set (350) in the at least one storage device (304).

[0108] 8. The method (700) according to clause 7, wherein the link data (354) includes a data structure (400), the data structure (400) including a plurality of records (401), each record including a first pointer (402) pointing to at least one of the plurality of 2D rendered images (340) and a second pointer (403) pointing to the corresponding native image (356), wherein the method further comprises using the at least one processor (302) to perform the following steps: storing the metadata (358, 454) and the additional link data (354) by including, in at least one of the plurality of records, a label (452) for each of the one or more environmental features (110) and a spatial relationship from each of the one or more environmental features to the corresponding set of physical coordinates (332).

[0109] 9. The method (700) according to any one of clauses 1 to 8, wherein the step of using the at least one processor (302) to create the plurality of 2D rendered images (340) includes:

[0110] Determining a visualization mode for each of the 2D rendered images; and

[0111] For each of the one or more environmental features (110), render a plurality of pixels (342) defining the environmental feature in a color corresponding to the determined visualization pattern.

[0112] 10. The method (700) according to any one of clauses 1 to 9, wherein the step of creating the plurality of 2D rendered images (340) using the at least one processor (302) includes: applying simulated intrinsic sensor effects.

[0113] 11. A computing system (300) for generating a training set (350) of images and labels for a native environment (100), the computing system including at least one processor (302) in communication with at least one storage device (304), wherein the at least one processor is configured to:

[0114] Receive a plurality of sets of physical coordinates (332);

[0115] Retrieve environmental model data (308) corresponding to a georeferenced model of the native environment from the at least one storage device, the environmental model data defining a plurality of environmental features (110);

[0116] Create a plurality of two-dimensional 2D rendered images (340) according to the environmental model data, each 2D rendered image corresponding to a view of one of the plurality of sets of physical coordinates, the plurality of 2D rendered images including one or more of the plurality of environmental features;

[0117] Generate link data (354), the link data (354) associating each 2D rendered image with (i) a label (452) for one or more of the included environmental features and (ii) a corresponding native image (356); and

[0118] Store the training set in the at least one storage device, the training set including the 2D rendered images, the labels, the corresponding native images, and the link data.

[0119] 12. The computing system (300) according to clause 11, wherein the at least one processor (302) is further configured to generate a corresponding label (452) for each 2D rendered image (340) in which one of the one or more environmental features (110) appears.

[0120] 13. The computing system (300) according to any one of clauses 11 to 12, wherein the plurality of sets of physical coordinates (332) define a path through the native environment (100), and wherein the at least one processor (302) is further configured to receive a plurality of camera images (330) recorded during the physical traversal of the path, wherein each of the plurality of camera images is associated with one of the plurality of sets of physical coordinates.

[0121] 14. The computing system (300) according to clause 13, wherein each of the plurality of camera images (330) includes a corresponding geographic coordinate tag, and wherein the at least one processor (302) is further configured to receive the plurality of sets of physical coordinates (332) by extracting the geographic coordinate tags from the camera images.

[0122] 15. The computing system (300) according to clause 13, wherein each of the plurality of camera images (330) is associated with a timestamp (442) corresponding to the relative traversal time along the path, and wherein the at least one processor (302) is further configured to generate the link data (354) as a data structure (400) including a plurality of records (401), each record including (i) a pointer (402) to at least one of the plurality of 2D rendered images (340) and (ii) time metadata (440), the time metadata (440) including at least one of the timestamp and the relative traversal time associated with the corresponding native image (356).

[0123] 16. The computing system (300) according to any one of clauses 11 to 15, wherein the at least one processor (302) is further configured to generate the link data (354) as a data structure (400) including a plurality of records (401), each record including a first pointer (402) to at least one of the plurality of 2D rendered images and a second pointer (403) to the corresponding native image (356).

[0124] 17. The computing system (300) according to any one of clauses 11 to 16, wherein the at least one processor (302) is further configured for each 2D rendered image (340):

[0125] Based on the environmental model data (308), calculate metadata (358, 454), the metadata (358, 454) including a spatial relationship from the corresponding set of physical coordinates (332) to at least one of the one or more environmental features (110);

[0126] Generate additional link data (354) that associates the metadata with the 2D rendered images; and

[0127] Store the metadata and the additional link data as part of the training set (350) in the at least one storage device (304).

[0128] 18. The computing system (300) according to clause 17, wherein the link data (354) includes a data structure (400), the data structure (400) includes a plurality of records (401), each record includes a first pointer (402) pointing to at least one of the plurality of 2D rendered images (340) and a second pointer (403) pointing to the corresponding native image (356), and wherein the at least one processor (302) is further configured to store the metadata (358, 454) and the additional link data by including, in at least one of the plurality of records, a label (452) for each of the one or more environmental features (110) and a spatial relationship from each of the one or more environmental features to the corresponding set of physical coordinates (332).

[0129] 19. The computing system (300) according to any one of clauses 11 to 18, wherein the at least one processor (302) is further configured to create the plurality of 2D rendered images (340) by:

[0130] Determine a visualization mode for each of the 2D rendered images; and

[0131] For each of the one or more environmental features (110), render a plurality of pixels (342) defining the environmental feature in a color corresponding to the determined visualization mode.

[0132] 20. A non-transitory computer-readable storage medium having computer-executable instructions for generating a training set (350) of images and labels for a native environment (100) implemented thereon, wherein when executed by at least one processor (302) in communication with at least one storage device (304), the computer-executable instructions cause the at least one processor to:

[0133] Receive a plurality of sets of physical coordinates (332);

[0134] Retrieve environmental model data (308) corresponding to a georeferenced model of the native environment from the at least one storage device, the environmental model data defining a plurality of environmental features (110);

[0135] Create a plurality of two-dimensional (2D) rendered images (340) based on the environmental model data, each of the 2D rendered images corresponding to a view of one of the plurality of physical coordinate sets, the plurality of 2D rendered images including one or more environmental features of the plurality of environmental features;

[0136] Generate link data (354) that associates each of the 2D rendered images with (i) a label (452) for one or more of the included environmental features and (ii) a corresponding native image (356); and

[0137] Store the training set in the at least one storage device, the training set including the 2D rendered images, the labels, the corresponding native images, and the link data.

[0138] The systems and methods described herein are not limited to the specific examples described herein. Rather, the components of the systems and / or the steps of the methods can be used independently and separately from the other components and / or steps described herein.

[0139] As used herein, an element or step recited in the singular and preceded by "a" or "an" should be understood as not excluding a plurality of elements or steps, unless such exclusion is explicitly stated. Additionally, a reference to "an example" or "example" of the present disclosure is not intended to be construed as excluding the existence of additional examples that also include the recited features.

[0140] This written description uses examples to disclose various examples, including the best mode, to enable those skilled in the art to practice those examples, including making and using any device or system and performing any incorporated method. The scope of patentable subject matter is defined by the claims and may include other examples that occur to those skilled in the art. If such other examples have structural elements that are not different from the literal language of the claims, or if they include equivalent structural elements that are not substantially different from the literal language of the claims, then such other examples are intended to be included within the scope of the claims.

Claims

1. A method (700) for generating a training set (350) of images and labels for a native environment (100), the method being implemented on a computing system (300) that includes at least one processor (302) in communication with at least one storage device (304), the method comprising using the at least one processor to perform the following steps: Receiving a plurality of sets of physical coordinates (332); Retrieving, from the at least one storage device, environmental model data (308) corresponding to a georeferenced model of the native environment, the environmental model data defining a plurality of environmental features (110); Creating a plurality of two-dimensional rendered images (340) based on the environmental model data, the creating step comprising: using the environmental model data to detect one or more of the plurality of environmental features in a view of one of the plurality of sets of physical coordinates, each of the plurality of two-dimensional rendered images corresponding to the view of one of the plurality of sets of physical coordinates, the plurality of two-dimensional rendered images including the detected one or more of the plurality of environmental features; Generating link data (354) that associates each of the plurality of two-dimensional rendered images with (i) a label (452) and (ii) a corresponding native image (356), the label marking each of the one or more environmental features with an identifier and / or type of each of the one or more environmental features, and wherein the native image is a camera image corresponding to a view of one of the plurality of sets of physical coordinates; and Storing the training set in the at least one storage device, the training set including the plurality of two-dimensional rendered images, the labels, the corresponding native images, and the link data.

2. The method (700) according to claim 1, the method further comprising using the at least one processor (302) to perform the following steps: for each of the one or more environmental features (110), generating a corresponding label (452) for each two-dimensional rendered image (340) in which the environmental feature appears.

3. The method (700) according to claim 1 or 2, wherein, the plurality of sets of physical coordinates (332) define a path through the native environment (100), the method further comprising using the at least one processor (302) to perform the following steps: receiving a plurality of camera images (330) recorded during physically traversing the path, wherein each of the plurality of camera images is associated with one of the plurality of sets of physical coordinates.

4. The method (700) according to claim 3, wherein, Each of the plurality of camera images (330) includes a corresponding geographic coordinate tag, and wherein the method further comprises performing, by the at least one processor (302), the steps of: extracting the geographic coordinate tags from the plurality of camera images to receive the plurality of physical coordinate sets (332).

5. The method (700) according to claim 3, wherein, each of the plurality of camera images (330) is associated with a timestamp (442) corresponding to a relative traversal time along the path, and wherein the method further comprises performing, by the at least one processor (302), the steps of: generating the link data (354) as a data structure (400) including a plurality of records (401), each record including (i) a pointer (402) to at least one of the plurality of two-dimensional rendered images (340) and (ii) time metadata (440), the time metadata (440) including at least one of the timestamp and the relative traversal time associated with the corresponding native image (356).

6. The method (700) according to claim 1 or 2, the method further comprising performing, by the at least one processor (302), the steps of: generating the link data (354) as a data structure (400) including a plurality of records (401), each record including a first pointer (402) to at least one of the plurality of two-dimensional rendered images and a second pointer (403) to the corresponding native image (356).

7. The method (700) according to claim 1 or 2, the method further comprising performing, by the at least one processor (302), for each two-dimensional rendered image (340), the steps of: calculating metadata (358, 454) based on the environmental model data (308), the metadata (358, 454) including a spatial relationship from the corresponding physical coordinate set (332) to at least one of the one or more environmental features (110); generating additional link data (354) associating the metadata with the two-dimensional rendered image; and storing the metadata and the additional link data as part of the training set (350) in the at least one storage device (304).

8. The method (700) according to claim 7, wherein, The linked data (354) includes a data structure (400), the data structure (400) includes a plurality of records (401), each record includes a first pointer (402) pointing to at least one of the plurality of two-dimensional rendered images (340) and a second pointer (403) pointing to the corresponding native image (356), wherein the method further includes using the at least one processor (302) to perform the following steps: storing the metadata (358, 454) and the additional linked data (354) by including, in at least one of the plurality of records, a label (452) for each of the one or more environmental features (110) and the spatial relationship from each of the one or more environmental features to the corresponding set of physical coordinates (332).

9. The method (700) according to claim 1 or 2, wherein, the step of using the at least one processor (302) to create the plurality of two-dimensional rendered images (340) includes: determining a visualization mode for each of the plurality of two-dimensional rendered images; and rendering, for each of the one or more environmental features (110), a plurality of pixels (342) defining the environmental feature in a color corresponding to the determined visualization mode.

10. The method (700) according to claim 1 or 2, wherein, the step of using the at least one processor (302) to create the plurality of two-dimensional rendered images (340) includes: applying a simulated intrinsic sensor effect.

11. A computing system (300) for generating a training set (350) of images and labels for a native environment (100), the computing system includes at least one processor (302) communicating with at least one storage device (304), wherein, the at least one processor is configured to: receive a plurality of sets of physical coordinates (332); retrieve, from the at least one storage device, environmental model data (308) corresponding to a georeferenced model of the native environment, the environmental model data defining a plurality of environmental features (110); create a plurality of two-dimensional rendered images (340) based on the environmental model data, the creating including: using the environmental model data to detect that one or more of the plurality of environmental features appear in a view of one of the plurality of sets of physical coordinates, each of the plurality of two-dimensional rendered images corresponding to the view of one of the plurality of sets of physical coordinates, the plurality of two-dimensional rendered images including the detected one or more of the plurality of environmental features; Generate link data (354) that associates each of the plurality of two-dimensional rendered images with (i) a label (452) and (ii) a corresponding native image (356), where the label tags each of the one or more environmental features with an identifier and / or type of each of the one or more environmental features, and wherein the native image is a camera image corresponding to a view of one of the plurality of physical coordinate sets; and Store the training set in the at least one storage device, the training set including the plurality of two-dimensional rendered images, the labels, the corresponding native images, and the link data.

12. The computing system (300) according to claim 11, wherein, The at least one processor (302) is further configured to generate a corresponding label (452) for each two-dimensional rendered image (340) in which one of the one or more environmental features (110) appears; The plurality of physical coordinate sets (332) define a path through the native environment (100), and wherein the at least one processor (302) is further configured to receive a plurality of camera images (330) recorded during physical traversal of the path, wherein each of the plurality of camera images is associated with one of the plurality of physical coordinate sets; Each of the plurality of camera images (330) includes a corresponding geographic coordinate tag, and wherein the at least one processor (302) is further configured to receive the plurality of physical coordinate sets (332) by extracting the geographic coordinate tags from the plurality of camera images; and Each of the plurality of camera images (330) is associated with a timestamp (442) corresponding to a relative traversal time along the path, and wherein the at least one processor (302) is further configured to generate the link data (354) as a data structure (400) including a plurality of records (401), each record including (i) a pointer (402) to at least one of the plurality of two-dimensional rendered images (340) and (ii) time metadata (440) including at least one of the timestamp and the relative traversal time associated with the corresponding native image (356).

13. The computing system (300) according to claim 11 or 12, wherein, The at least one processor (302) is further configured to generate the link data (354) as a data structure (400) including a plurality of records (401), each record including a first pointer (402) to at least one of the plurality of two-dimensional rendered images (340) and a second pointer (403) to the corresponding native image (356).

14. The computing system (300) according to claim 11 or 12, wherein, the at least one processor (302) is further configured for each two-dimensional rendered image (340): calculate metadata (358, 454) based on the environmental model data (308), the metadata (358, 454) including a spatial relationship from the set of physical coordinates (332) to at least one environmental feature among the one or more environmental features (110); generate additional link data (354) associating the metadata with the two-dimensional rendered image; and store the metadata and the additional link data as part of the training set (350) in the at least one storage device (304); wherein the link data (354) includes a data structure (400), the data structure (400) including a plurality of records (401), each record including a first pointer (402) pointing to at least one of the plurality of two-dimensional rendered images (340) and a second pointer (403) pointing to the corresponding native image (356), and wherein the at least one processor (302) is further configured to store the metadata (358, 454) and the additional link data by including, in at least one of the plurality of records, a label (452) for each of the one or more environmental features (110) and the spatial relationship from each of the one or more environmental features to the corresponding set of physical coordinates (332).

15. The computing system (300) according to claim 11 or 12, wherein, the at least one processor (302) is further configured to create the plurality of two-dimensional rendered images (340) by: determining a visualization mode for each of the plurality of two-dimensional rendered images; and rendering a plurality of pixels (342) defining the environmental feature in a color corresponding to the determined visualization mode for each of the one or more environmental features (110).

Citation Information

Patent Citations

  • Method, device and equipment used for marking map

    CN108694882A

  • Employing three-dimensional (3D) data predicted from two-dimensional (2D) images using neural networks for 3D modeling applications and other applications

    US20190026956A1

  • Techniques for training machine learning

    US20200302241A1