Systems and methods for refining synthetic data using auxiliary inputs via generative adversarial networks

By introducing auxiliary input into the generative adversarial network and refining the synthetic data, the problem of insufficient realisticity of synthetic data is solved, and higher quality training data is achieved, which is suitable for fields such as computer vision and autonomous driving.

CN109472365BActive Publication Date: 2025-06-06FORD GLOBAL TECH LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201811035739.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-09-08
Filing Date
2018-09-06
Publication Date
2025-06-06
Estimated Expiration
2038-09-06

AI Technical Summary

Technical Problem

The process of annotating and labeling image training data in the prior art is cumbersome and expensive, especially when using synthetic data, which lacks sufficient realisticity and is difficult to directly use in training machine learning models.

Method used

By using Generative Adversarial Networks (GANs) to refine synthetic data, using auxiliary inputs such as semantic maps, depth maps, and object edges, providing tips on increasing the realisticity of synthetic data, thereby generating more realistic refined synthetic data while preserving annotation metadata and tagged metadata.

Benefits of technology

The realism of synthetic data is improved, making it more suitable for training machine learning models, especially in computer vision and autonomous driving-related applications, and the use value of synthetic data in these fields is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN109472365B_ABST
    Figure CN109472365B_ABST
Patent Text Reader

Abstract

The present invention extends to methods, systems, and computer program products for refining synthetic data using a generative adversarial network (GAN) using auxiliary inputs. The refined synthetic data can be rendered more realistically than the original synthetic data. The refined synthetic data also retains annotation metadata and labeling metadata used to train machine learning models. The GAN can be extended to use the auxiliary channel as an input to the refiner network, thereby providing hints for increasing the realism of the synthetic data. The refinement of the synthetic data enhances the use of the synthetic data in additional applications.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] 1. Field of the Invention

[0002] The present invention generally relates to the field of formulating realistic training data for training machine learning models, and more specifically to refining synthetic data through generative adversarial networks using auxiliary inputs.

[0003] 2. Related technologies

[0004] The process of annotating and labeling relevant portions of image training data (e.g., still images or videos) for use in training machine learning models can be tedious, time-consuming, and expensive. To reduce these annotation and labeling burdens, synthetic data (e.g., virtual images generated by a game or other graphics engine) can be used. Annotating synthetic data is more straightforward because annotation is a direct byproduct of generating the synthetic data. Summary of the invention

[0005] The present invention extends to methods, systems, and computer program products for refining synthetic data using auxiliary inputs via a generative adversarial network.

[0006] Aspects of the invention include using a generative adversarial network ("GAN") to refine synthetic data. The refined synthetic data can be rendered more realistically than the original synthetic data. The refined synthetic data also retains annotation metadata and label metadata used to train machine learning models. The GAN can be extended to use an auxiliary channel as an input to a refiner network to provide hints about increasing the realism of the synthetic data. The refinement of the synthetic data enhances the use of the synthetic data in additional applications.

[0007] In one aspect, a GAN is used to refine a synthetic (or virtual) image (e.g., an image generated by a game engine) into a more realistic refined synthetic (or virtual) image. The more realistic refined synthetic image retains the annotation metadata and labeling metadata of the synthetic image used to train the machine learning model. Auxiliary inputs are provided to the refiner network to provide hints about how the more realistic refined synthetic image looks. The auxiliary inputs can help apply the correct texture to different areas of the synthetic image. The auxiliary inputs can include semantic maps (e.g., to help with image segmentation), depth maps, edges between objects, etc. The refinement of the synthetic image enhances the use of the synthetic image for solving problems in computer vision, including applications related to autonomous driving, such as image segmentation, identifying drivable paths, object tracking, and object three-dimensional (3D) pose estimation.

[0008] The semantic map, depth map, and object edges ensure that the correct texture is applied to different areas of the composite image. For example, the semantic map can segment the composite image into multiple areas and identify the content of each area, such as leaves from a tree or the side of a green building. The depth map can distinguish how each image area in the composite image looks, such as different levels of detail / texture based on the distance of the object from the camera. Object edges can define the transition between different objects in the composite image.

[0009] Thus, aspects of the present invention include an image processing system that refines a synthetic (or virtual) image to improve the appearance of the synthetic (or virtual) image and provides a higher quality (e.g., more realistic) synthetic (or virtual) image for training a machine learning model. When the training of the machine learning model is complete, the machine learning model can be used with autonomous vehicles and driver assistance vehicles to accurately process and identify objects within images captured by vehicle cameras and sensors.

[0010] Generative adversarial networks (GANs) can use machine learning to train two networks, a discriminator network and a generator network, that essentially compete with each other (i.e., compete against each other). The discriminator network is trained to distinguish between real data instances (e.g., real images) and synthetic data instances (e.g., virtual images) and classify data instances as real or synthetic. The generator network is trained to produce synthetic data instances that the discriminator network classifies as real data instances. When the discriminator network cannot evaluate whether any data instance is synthetic or real, a strategic equilibrium is reached. It may be that the generator network never directly observes real data instances. Instead, the generator network receives information about real data instances as seen indirectly through the parameters of the discriminator network.

[0011] In one aspect, the discriminator network distinguishes between real images and synthetic (or virtual) images and classifies the images as real or synthetic (or virtual). In this regard, the generator network is trained to produce synthetic (or virtual) images. GAN can be extended to include a refiner network (which may or may not replace the generator network). The refiner network observes the synthetic (or virtual) image and generates variants of the synthetic (or virtual) image. The variants of the synthetic (or virtual) image are intended to exhibit characteristics of increased similarity to the real image while retaining annotation metadata and tag metadata. The refiner network attempts to refine the synthetic (or virtual) image so that the discriminator network classifies the refined synthetic (or virtual) image as a real image. The refiner network also attempts to maintain the similarity (e.g., normalized characteristics) between the input synthetic (or virtual) image and the refined synthetic (or virtual) image.

[0012] The refiner network can be extended to receive additional information that can be generated as part of the synthesis process. For example, the refiner network can receive one or more of the following: semantic maps (e.g., to help with image segmentation), depth maps, edges between objects, etc. In one aspect, the refiner network receives as input an auxiliary image that encodes a pixel-level semantic segmentation of a synthesized (or virtual) image. In another aspect, the refiner network receives as input an auxiliary image that encodes a depth map of the contents of the synthesized (or virtual) image. In yet another aspect, the refiner network can receive an auxiliary image that encodes the edges between objects in the synthesized (or virtual) image.

[0013] For example, a synthetic (or virtual) image may include leaves from a tree. Semantic segmentation may indicate that the portion of the synthetic (or virtual) image that includes leaves is actually leaves (and not, for example, the side of a green building). A depth map may be used to differentiate how leaves look depending on the distance from the camera. Edges may be used to differentiate between different objects in a synthetic (or virtual) image.

[0014] Auxiliary data can be extracted from a dataset of real images used during training of the discriminator network. Extracting auxiliary data from a dataset of real images can include using a sensor, such as a LiDAR, that is synchronized with the camera data stream. For auxiliary data representing semantic segmentation, segmentation can be performed manually or by a semantic segmentation model. The GAN can then be formulated as a conditional GAN, where the discriminator network is conditioned on the provided auxiliary data.

[0015] Therefore, GANs can utilize auxiliary data streams (such as semantic maps and depth maps) to help ensure that the correct textures are correctly applied to different areas of the synthetic (or virtual) image. GANs can generate refined synthetic (or virtual) images with increased realism while preserving annotations and / or labels for training additional models (e.g., computer vision, autonomous driving, etc.). BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The specific features, aspects and advantages of the present invention will be better understood with reference to the following description and accompanying drawings, in which:

[0017] Figure 1 An exemplary block diagram of a computing device is shown.

[0018] Figure 2 An exemplary generative adversarial network that facilitates using auxiliary inputs to refine synthetic data is shown.

[0019] Figure 3 A flowchart illustrating an exemplary method for refining synthetic data using auxiliary inputs through a generative adversarial network.

[0020] Figure 4 An exemplary data flow for refining synthetic data using auxiliary inputs through a generative adversarial network is shown. DETAILED DESCRIPTION

[0021] Figure 1 An exemplary block diagram of a computing device 100 is shown. The computing device 100 can be used to execute various programs (such as those discussed herein). The computing device 100 can be used as a server, a client, or any other computing entity. The computing device 100 can perform various communication and data transfer functions as described herein, and can execute one or more applications (such as the applications described herein). The computing device 100 can be any of a variety of computing devices, such as a mobile phone or other mobile device, a desktop computer, a notebook computer, a server computer, a handheld computer, a tablet computer, etc.

[0022] The computing device 100 includes one or more processors 102, one or more memory devices 104, one or more interfaces 106, one or more mass storage devices 108, one or more input / output (I / O) devices 110, and a display device 130, all coupled to a bus 112. The processor 102 includes one or more processors or controllers that execute instructions stored in the memory device 104 and / or the mass storage device 108. The processor 102 may also include various types of computer storage media, such as cache memory.

[0023] Memory device 104 includes various computer storage media, such as volatile memory (e.g., random access memory (RAM) 114) and / or non-volatile memory (e.g., read-only memory (ROM) 116). Memory device 104 may also include a rewritable ROM, such as flash memory.

[0024] The mass storage device 108 includes various computer storage media, such as magnetic tapes, magnetic disks, optical disks, solid-state memory (e.g., flash memory), etc. Figure 1 As depicted, the particular mass storage device is a hard drive 124. Various drives may also be included in the mass storage device 108 to enable reading from and / or writing to various computer-readable media. The mass storage device 108 includes removable media 126 and / or non-removable media.

[0025] I / O devices 110 include various devices that allow data and / or other information to be input into or retrieved from computing device 100. Exemplary I / O devices 110 include cursor control devices, keyboards, keypads, bar code scanners, microphones, monitors or other display devices, speakers, printers, network interface cards, modems, cameras, lenses, radars, CCD or other image capture devices, and the like.

[0026] Display device 130 includes any type of device capable of displaying information to one or more users of computing device 100. Examples of display device 130 include a monitor, a display terminal, a video projection device, and the like.

[0027] Interfaces 106 include various interfaces that allow computing device 100 to interact with other systems, devices or computing environments, as well as people. Exemplary interfaces 106 may include any number of different network interfaces 120, such as interfaces with: personal area networks (PANs), local area networks (LANs), wide area networks (WANs), wireless networks (e.g., near field communication (NFC), Bluetooth, Wi-Fi, etc. networks), and the Internet. Other interfaces include user interfaces 118 and peripheral device interfaces 122.

[0028] The bus 112 allows the processor 102, memory device 104, interface 106, mass storage device 108, and I / O device 110 to communicate with each other and with other devices or components coupled to the bus 112. The bus 112 represents one or more of several types of bus structures, such as a system bus, a PCI bus, an IEEE 1394 bus, a USB bus, etc.

[0029] Figure 2 An exemplary generative adversarial network (GAN) 200 that facilitates using auxiliary input to refine synthetic data is shown. The generative adversarial network (GAN) 200 can be implemented using components of the computing device 100.

[0030] As depicted, GAN 200 includes a generator 201, a refiner 202, and a discriminator 203. Generator 201 can generate and output a virtual image including synthetic image data and annotations. The synthetic image data can represent an image of a road scene. The annotation annotates the synthetic image data with ground truth data of the road scene. The annotation can be a byproduct of generating the synthetic image data. In one aspect, generator 201 is a game engine.

[0031] However, synthetic image data may lack sufficient realism, especially for higher resolution images and / or images containing more complex objects. Human observers can usually distinguish real images from virtual images generated by game engines.

[0032] In this way, the refiner 202 can access the virtual image and refine the virtual image to improve realism. The refiner 202 can receive the virtual image from the generator 201. The refiner 202 can access auxiliary data, such as image segmentation, depth map, object edges, etc. The refiner 202 can refine (transform) the virtual image into a refined virtual image based on the auxiliary data. For example, the refiner 202 can use the content of the auxiliary data as a hint to improve the realism of the virtual image without changing the annotation. The refiner 202 can output the refined virtual image.

[0033] Discriminator 203 may receive the refined virtual image from refiner 202. Discriminator 203 may classify the refined virtual image as "real" or "synthetic". When the image is classified as "real", discriminator 203 may make the refined virtual image available for training other neural networks. For example, discriminator 203 may make the refined virtual image classified as "real" available for training computer vision neural networks, including those related to autonomous driving.

[0034] When the image is classified as "synthetic", the discriminator 203 may generate feedback parameters for further improving the realism of the refined virtual image. The discriminator 203 may send the feedback parameters to the refiner 202 and / or the generator 201. The refiner 202 and / or the generator 201 may use the feedback parameters to further improve the realism of the previously refined virtual image (which may further refer to the auxiliary data). The further refined virtual image may be sent to the discriminator 203. The virtual image may be further refined (transformed) based on the auxiliary data and / or the feedback parameters until the discriminator 203 classifies the virtual image as "real" (or until no further improvement in realism is possible after performing a specified number of refinements, etc.).

[0035] Figure 3 A flow chart of an exemplary method 300 for refining synthetic data using auxiliary inputs by GAN 200 is shown. Method 300 will be described with respect to components of GAN 200 and data.

[0036] Generator 201 may generate a virtual image 211 representing an image of a road scene (e.g., an image of a road, an image of a highway, an image of an interstate highway, an image of a parking lot, an image of an intersection, etc.). Virtual image 211 includes synthetic image data 212 and annotations 213. Synthetic image data 212 may include pixel values ​​of pixels in virtual image 211. Annotation 213 annotates the synthetic image data with ground truth data of the road scene. Generator 201 may output virtual image 211.

[0037] Method 300 includes accessing synthetic image data representing an image of a road scene, the synthetic image data including annotations, the annotations annotating the synthetic image data with ground truth data of the road scene (301). For example, refiner 202 may access virtual image 211 of a scene that may be encountered during driving (e.g., an intersection, a road, a parking lot, etc.), including synthetic image data 212 and annotations 213. Method 300 includes accessing one or more auxiliary data streams for the image (302). For example, refiner 202 may access one or more of the following from auxiliary data 221: image segmentation 222, depth map 223, and object edges 224.

[0038] Image segmentation 222 can segment virtual image 211 into multiple regions and identify the content of each region, such as leaves from a tree or the side of a green building. Depth map 223 can distinguish how each image region in virtual image 211 looks based on the distance of the object from the camera, such as different levels of detail / texture. Object edges 224 can define transitions between different objects in virtual image 211.

[0039] The method 300 includes refining the synthetic image data using the content of the one or more auxiliary data streams as hints, thereby refining the synthetic image data to improve the realism of the image without changing the annotation (303). For example, the refiner 202 may use the content of one or more of the following as hints: image segmentation 222, depth map 223, and object edges 224 to refine (transform) the virtual image 211 into refined synthetic image data 212. The refiner 202 may refine the synthetic image data 211 into refined synthetic image data 212 without changing the annotation 213. The refined synthetic image data 212 may improve the realism of the scene relative to the synthetic image data 211.

[0040] In one aspect, image segmentation 222 is included in an auxiliary image that encodes pixel-level semantic segmentation of virtual image 211. In another aspect, depth map 223 is included in another auxiliary image that encodes a depth map of the content of virtual image 211. In yet another aspect, object edges 224 are included in yet another auxiliary image that encodes edges between objects in virtual image 211. Thus, refiner 202 can use one or more auxiliary images to refine composite image data 211 into refined composite image data 212.

[0041] In one aspect, the generator 201 generates the auxiliary data 221 as a byproduct of generating the virtual image 211. In another aspect, the auxiliary data 221 is extracted from a dataset of real images used to train the discriminator 203.

[0042] The method 300 includes outputting refined synthetic image data representing a refined image of a road scene (304). For example, the refiner 202 may output a refined virtual image 214 of a scene that may be encountered during driving. The refined virtual image 214 includes the refined synthetic image data 216 and the annotations 213.

[0043] The discriminator 203 may access the refined virtual image 214. The discriminator 203 may perform an image type classification 217 on the refined virtual image 214 using the refined synthetic image data 216 and the annotations 213. The image type classification 217 classifies the refined virtual image 214 as “real” or “synthetic.” If the discriminator 203 classifies the refined virtual image 214 as “real,” the discriminator 203 may make the refined virtual image 214 available for training other neural networks, such as computer vision neural networks, including those related to autonomous driving.

[0044] In another aspect, if the discriminator 203 classifies the refined virtual image 214 as “synthetic”, the discriminator 203 may generate image feedback parameters 218 for further improving the realism of the refined virtual image 214. The discriminator 203 may send the image feedback parameters 218 to the refiner 202 and / or the generator 201. The refiner 202 and / or the generator 201 may use the image feedback parameters 218 to further improve the realism of the refined virtual image 214 (possibly with further reference to the auxiliary data 221). The further refined virtual image may be sent to the discriminator 203. The refiner 202 and / or the generator 201 may further refine (transform) the refined virtual image 214 based on the auxiliary data 221 and / or the image feedback parameters 218 (or additional other feedback parameters). Image refinement may continue until discriminator 203 classifies a further refined virtual image (further refined from refined virtual image 214) as "real" (or until no further improvement in realism is possible after performing a specified number of refinements, etc.).

[0045] Figure 4 An exemplary data flow 400 for refining synthetic data using an auxiliary input through a generative adversarial network is shown. Generator 401 generates virtual image 411, image segmentation image 433, and depth map image 423. Refiner 402 refines (transforms) virtual image 411 into refined virtual image 414 using the content of image segmentation image 433 and depth map image 423 (e.g., as a hint). The realism of refined virtual image 414 can be improved relative to virtual image 411. Discriminator 403 classifies refined virtual image 414 as "real" or "synthetic"

[0046] In one aspect, one or more processors are configured to execute instructions (e.g., computer readable instructions, computer executable instructions, etc.) to perform any of the plurality of described operations. One or more processors can access information from system memory and / or store information in system memory. One or more processors can transform information between different formats, such as, for example, virtual images, synthetic image data, annotations, auxiliary data, auxiliary images, image segmentation, depth maps, object edges, refined virtual images, refined synthetic data, image type classification, image feedback parameters, etc.

[0047] The system memory can be coupled to one or more processors and can store instructions (e.g., computer readable instructions, computer executable instructions, etc.) executed by the one or more processors. The system memory can also be configured to store any of a number of other types of data generated by the described components, such as, for example, virtual images, synthetic image data, annotations, auxiliary data, auxiliary images, image segmentation, depth maps, object edges, refined virtual images, refined synthetic data, image type classification, image feedback parameters, etc.

[0048] In the above disclosure, reference has been made to the accompanying drawings, which form a part of the present invention and in which specific implementations in which the present disclosure may be practiced are shown by way of illustration. It should be understood that other implementations may be utilized and structural changes may be made without departing from the scope of the present disclosure. References in the specification to "one embodiment," "embodiment," "exemplary embodiment," etc. indicate that the described embodiment may include specific features, structures, or characteristics, but each embodiment may not necessarily include the specific features, structures, or characteristics. In addition, these phrases do not necessarily refer to the same embodiment. In addition, when describing specific features, structures, or characteristics in conjunction with an embodiment, it is subject to the knowledge of those skilled in the art that these features, structures, or characteristics are affected in conjunction with other embodiments, whether or not explicitly described.

[0049] The implementation of the system, device and method disclosed herein may include or utilize a special or general-purpose computer, which includes computer hardware, such as one or more processors and system memory as discussed herein. The implementation within the scope of the present disclosure may also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media can be any available media that can be accessed by a general or special-purpose computer system. The computer-readable medium that stores computer-executable instructions is a computer storage medium (device). The computer-readable medium that carries computer-executable instructions is a transmission medium. Therefore, as an example and not limitation, the implementation of the present disclosure may include at least two distinct computer-readable media: a computer storage medium (device) and a transmission medium.

[0050] Computer storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives ("SSD") (e.g., RAM-based), flash memory, phase-change memory ("PCM"), other types of memory, other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general or special purpose computer.

[0051] The implementation of the device, system and method disclosed herein can communicate through a computer network. "Network" is defined as one or more data links capable of transmitting electronic data between a computer system and / or module and / or other electronic device. When information is transmitted or provided to a computer through a network or another communication connection (hardwired, wireless or a combination of hardwired or wireless), the computer correctly regards the connection as a transmission medium. The transmission medium may include a network and / or a data link, which can be used to carry the required program code means in the form of a computer executable instruction or data structure and can be accessed by a general or special computer. The above combination should also be included in the scope of computer-readable media.

[0052] For example, computer executable instructions include instructions and data that cause a general-purpose computer, a special-purpose computer, or a special-purpose processing device to perform a specific function or group of functions when executed on a processor. For example, computer executable instructions can be binary files, intermediate format instructions such as assembly language, or even source code. Although the subject matter is described in a language dedicated to structural features and / or method actions, it should be understood that the subject matter defined in the attached claims is not necessarily limited to the above-mentioned features or actions. Instead, the described features and actions are disclosed as exemplary forms of implementing the claims.

[0053] Those skilled in the art will appreciate that the present disclosure can be practiced in a network computing environment with many types of computer system configurations, including built-in or other vehicle computers, personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablets, pagers, routers, switches, various storage devices, etc. The present disclosure can also be practiced in a distributed system environment, where local and remote computer systems linked by a network (by a hardwired data link, a wireless data link, or a combination of hardwired and wireless data links) all perform tasks. In a distributed system environment, program modules can be located in storage devices of local and remote memories.

[0054] In addition, where appropriate, the functions described herein may be performed in one or more of the following: hardware, software, firmware, digital components, or analog components. For example, one or more application specific integrated circuits (ASICs) may be programmed to perform one or more of the systems and programs described herein. Certain terms are used throughout the specification and claims to refer to specific system components. As will be appreciated by those skilled in the art, components may be referenced by different names. This document is not intended to distinguish between components that are differently named but non-functional.

[0055] It should be noted that the sensor embodiments discussed above may include computer hardware, software, firmware, or any combination thereof to perform at least a portion of their functionality. For example, the sensor may include computer code configured to be executed in one or more processors, and may include hardware logic / circuitry controlled by the computer code. These exemplary devices are provided herein for illustrative purposes, not limiting. As will be known to those skilled in the relevant art, embodiments of the present disclosure may be implemented in other types of devices.

[0056] At least some embodiments of the present disclosure relate to computer program products that include such logic (e.g., in the form of software) stored on any computer-usable medium. Such software, when executed in one or more data processing devices, causes the devices to operate as described herein.

[0057] Although various embodiments of the present disclosure have been described above, it should be understood that they are presented by way of example only and are not limiting. It will be apparent to those skilled in the relevant art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure. Therefore, the breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should only be defined in accordance with the following claims and their equivalents. The foregoing description has been presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. In view of the above teachings, many modifications and variations are possible. In addition, it should be noted that any or all of the above-described alternative implementations may be used in any desired combination to form additional hybrid implementations of the present disclosure.

[0058] According to the present invention, a method for refining training data for a machine learning model is provided, the method comprising: accessing synthetic image data representing a road scene image, the synthetic image data including ground truth data annotations; accessing auxiliary data; using the auxiliary data as a hint to generate refined synthetic image data to improve the realism of the synthetic image data without changing the annotations; and outputting the refined synthetic image data.

[0059] According to an embodiment, the invention is also characterized in that accessing the auxiliary data comprises accessing one or more auxiliary data streams corresponding to the image.

[0060] According to an embodiment, the invention is further characterized in that accessing one or more auxiliary data streams includes accessing one or more of: semantic image segmentation data of the image, depth map data of the image, or edge data of the image.

[0061] According to an embodiment, the invention is further characterized in that accessing the auxiliary data comprises accessing a pixel-level semantic segmentation of the image.

[0062] According to an embodiment, the invention is further characterized in that accessing the auxiliary data comprises accessing a depth map defining different levels of detail of objects in the image based on their distance from the camera.

[0063] According to an embodiment, the invention is further characterized in that accessing the auxiliary data comprises accessing edge data defining edges between objects in the image.

[0064] According to the present invention, a method for refining training data for a machine learning model is provided, the method comprising: accessing synthetic image data representing an image of a road scene, the synthetic image data comprising annotations, the annotations annotating the synthetic image data with ground-truth data of the road scene; accessing one or more auxiliary data streams for the image; using the contents of the one or more auxiliary data streams as hints to refine the synthetic image data, thereby refining the synthetic image data to improve the realism of the image without changing the annotations; and outputting the refined synthetic image data, the refined synthetic image data representing a refined image of the road scene.

[0065] According to an embodiment, the present invention is also characterized by receiving feedback indicating that the refined image lacks sufficient realism, the feedback including one or more parameters for further refining the refined image; using the parameters to further refine the refined synthetic image data, further refining the refined synthetic image data further improving the realism of the refined image without changing the annotation; and outputting the further refined synthetic image data, the further refined synthetic image data representing a further refined image of the road scene.

[0066] According to an embodiment, the invention is further characterized in that accessing one or more auxiliary data streams for the image comprises accessing one or more of: a semantic image segmentation of the image or a depth map of the image.

[0067] According to an embodiment, the invention is further characterized in that accessing one or more auxiliary data streams for the image includes accessing a pixel-level semantic segmentation of the image.

[0068] According to an embodiment, the invention is also characterized in that accessing one or more auxiliary data streams for the image comprises accessing a depth map defining different levels of detail of objects based on their distance from the camera.

[0069] According to an embodiment, the invention is also characterized in that accessing one or more auxiliary data streams for the image includes accessing edge data defining edges between objects in the image.

[0070] According to an embodiment, the invention is also characterized in that the one or more auxiliary data streams are extracted from other image data.

[0071] According to an embodiment, the invention is also characterized in that extracting the one or more auxiliary data streams from other image data includes extracting the auxiliary data streams from a sensor synchronized with the camera data stream.

[0072] According to an embodiment, the present invention is also characterized by using the refined synthetic image data to train the machine learning module, which is used for autonomous driving of the vehicle.

[0073] According to the present invention, a computer system is provided, which has: one or more processors; a system memory coupled to the one or more processors, the system memory storing instructions executable by the one or more processors; and the one or more processors executing the instructions stored in the system memory to refine training data for a machine learning model, including the following: accessing synthetic image data representing a road scene image, the synthetic image data including annotations, the annotations annotating the synthetic image data with ground-truth data of the road scene; accessing one or more auxiliary data streams for the image; using the content of the one or more auxiliary data streams as a hint to refine the synthetic image data, thereby refining the synthetic image data to improve the realism of the image without changing the annotations; and outputting the refined synthetic image data.

[0074] According to an embodiment, the present invention is also characterized in that the one or more processors execute the instructions to access one or more auxiliary data streams, including the one or more processors executing the instructions to access one or more of the following: semantic image segmentation data of the image, depth map data of the image, or edge data of the image.

[0075] According to an embodiment, the present invention is also characterized in that the one or more processors execute the instructions to: receive feedback indicating that the refined image lacks sufficient realism, the feedback including one or more parameters for further refining the refined image; use the parameters to further refine the refined synthetic image data, further refining the refined synthetic image data to improve the realism of the refined image without changing the annotation; and output the further refined synthetic image data, the further refined synthetic image data representing a further refined image of the road scene.

[0076] According to an embodiment, the invention is also characterized in that the one or more processors execute the instructions to extract the auxiliary data stream from the sensor synchronized with the camera data stream.

[0077] According to an embodiment, the present invention is also characterized in that the one or more processors execute the instructions to train the machine learning module using the refined synthetic image data, and the machine learning module is used for autonomous driving of a vehicle.

Claims

1. A method for refining training data for a machine learning model, the method include: accessing synthetic image data representing an image of a road scene, the synthetic image data including ground truth data annotations; Accessing auxiliary data, wherein the accessed auxiliary data includes: semantic image segmentation data of the image, depth map data of the image, and edge data of the image; generating refined synthetic image data using the auxiliary data as a hint using a generative adversarial network (GAN) to improve the realism of the synthetic image data without changing the annotation; wherein generating refined synthetic image data using the auxiliary data as a hint using a generative adversarial network (GAN) comprises: ensuring that correct textures are applied to different regions of the synthetic image through semantic image segmentation data of the image, depth map data of the image, and edge data of the image, wherein the semantic image segmentation data segments the synthetic image into a plurality of regions and identifies the content of each region, the depth map data distinguishes different texture levels of each image region in the synthetic image based on the distance of the object from the camera, and the edge data defines transitions between different objects in the synthetic image; and The refined composite image data is output.

2. A method for refining training data for a machine learning model, the method include: accessing synthetic image data representing an image of a road scene, the synthetic image data comprising annotations annotating the synthetic image data with ground truth data of the road scene; accessing one or more auxiliary data streams for the image, wherein accessing one or more auxiliary data streams for the image comprises: semantic image segmentation of the image, a depth map of the image, and edge data of the image; Refining the synthesized image data using the content of the one or more auxiliary data streams as cues using a generative adversarial network (GAN), thereby refining the synthesized image data to improve the realism of the image without changing the annotation; wherein refining the synthetic image data using the content of the one or more auxiliary data streams as cues using a generative adversarial network (GAN) comprises: ensuring that correct textures are applied to different regions of the synthetic image through semantic image segmentation data of the image, depth map data of the image, and edge data of the image, wherein the semantic image segmentation data segments the synthetic image into a plurality of regions and identifies the content of each region, the depth map data distinguishes different texture levels of each image region in the synthetic image based on the distance of the object from the camera, and the edge data defines transitions between different objects in the synthetic image; and The refined composite image data is output, the refined composite image data representing a refined image of the road scene.

3. The method according to claim 2, further comprising: include: receiving feedback indicating that the refined image lacks realism, the feedback comprising one or more parameters for further refining the refined image; further refining the refined composite image data using the parameters, further refining the refined composite image data further improving the realism of the refined image without changing the annotation; and The further refined composite image data is output, the further refined composite image data representing a further refined image of the road scene.

4. The method of claim 2, further comprising extracting the one or more auxiliary data streams from other image data.

5. The method of claim 4, wherein extracting the one or more auxiliary data streams from other image data comprises extracting auxiliary data streams from a sensor synchronized with a camera data stream.

6. The method of claim 2 further comprising using the refined synthetic image data to train the machine learning module, the machine learning module being used for autonomous driving of a vehicle.

7. A computer system, wherein the computer system include: one or more processors; a system memory coupled to the one or more processors, the system memory storing instructions executed by the one or more processors; as well as The one or more processors execute the instructions stored in the system memory to refine training data for a machine learning model, including the following: accessing synthetic image data representing an image of a road scene, the synthetic image data comprising annotations annotating the synthetic image data with ground truth data of the road scene; accessing one or more auxiliary data streams for the image, wherein the one or more processors executing the instructions to access the one or more auxiliary data streams comprises the one or more processors executing the instructions to access: semantic image segmentation data for the image, depth map data for the image, and edge data for the image; Refining the synthesized image data using the content of the one or more auxiliary data streams as cues using a generative adversarial network (GAN), thereby refining the synthesized image data to improve the realism of the image without changing the annotation; wherein refining the synthetic image data using the content of the one or more auxiliary data streams as cues using a generative adversarial network (GAN) comprises: ensuring that correct textures are applied to different regions of the synthetic image through semantic image segmentation data of the image, depth map data of the image, and edge data of the image, wherein the semantic image segmentation data segments the synthetic image into a plurality of regions and identifies the content of each region, the depth map data distinguishes different texture levels of each image region in the synthetic image based on the distance of the object from the camera, and the edge data defines transitions between different objects in the synthetic image; and The refined composite image data is output.

8. The computer system of claim 7, further comprising the one or more processors executing the instructions to: receiving feedback indicating that the refined image lacks realism, the feedback comprising one or more parameters for further refining the refined image; further refining the refined composite image data using the parameters, further refining the refined composite image data improving the realism of the refined image without changing the annotation; and The further refined composite image data is output, the further refined composite image data representing a further refined image of the road scene.

9. The computer system of claim 7, further comprising the one or more processors executing the instructions to: extracting an auxiliary data stream from the sensor synchronized with the camera data stream; and The refined synthetic image data is used to train a machine learning module for autonomous driving of the vehicle.

Citation Information

Patent Citations

  • Systems and methods for decoding light field image files

    US20130077882A1

  • Auxiliary data for artifacts - aware view synthesis

    US20170188002A1