Generative adversarial network model for small road object detection
By using a generative adversarial network model to understand the semantics of road scenes and predict the likelihood of the presence of small objects, the problem of difficulty in detecting small objects in existing technologies is solved, enabling autonomous vehicles to accurately identify and respond to small objects in a timely manner.
Patent Information
- Application Number
- CN202110338780.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-03-31
- Filing Date
- 2021-03-30
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-03-30
AI Technical Summary
Existing depth detector models face challenges in detecting small objects in road scenes, especially traffic lights, which are difficult to detect due to their distance from the detector and low contrast. This may prevent autonomous vehicles from recognizing and reacting in a timely manner.
A Generative Adversarial Network (GAN) model is used to predict the likelihood of object presence by understanding the semantics of the road scene, rather than relying on object features. The distribution of road scene images is generated by the GAN model and sampled to detect small objects.
This improves the accuracy of detecting small objects, ensuring that autonomous vehicles can promptly identify and respond to important small objects on the road, such as traffic lights, and reduces the risk of false detections and missed detections.
Smart Images

Figure CN113469933B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to small object detection using generative adversarial network models. More specifically, the present disclosure relates to small object detection in road scenes using generative adversarial network models. BACKGROUND
[0002] Road scene understanding is important in autonomous driving applications. Object detection plays an important role in road scene understanding. However, detecting small objects (e.g., traffic lights) in image data is challenging because small objects are often located far away from the detector and have very low contrast compared to their surrounding environment. SUMMARY
[0003] Deep detector models are used to detect objects in image data. However, small road object detection remains a challenge for deep detector models. Accordingly, the present disclosure provides, among other things, systems, methods, and non-transitory computer-readable media for detecting small objects in road scenes.
[0004] The present disclosure provides a system for detecting small objects in road scenes. In some implementations, the system includes a camera and an electronic controller. The camera is coupled to a vehicle and is configured to capture a road scene image. The electronic controller is coupled to the camera. The electronic controller is configured to receive the road scene image from the camera. The electronic controller is further configured to generate a generative adversarial network model (GAN) using the road scene image. The electronic controller is further configured to determine a distribution indicating a likelihood that each location in the road scene image can contain a road object using the GAN model. The electronic controller is further configured to determine a plurality of locations in the road scene image by sampling the distribution. The electronic controller is further configured to detect the road object at one of the plurality of locations in the road scene image.
[0005] The present disclosure also provides a method for detecting small objects in road scenes. The method includes receiving, with an electronic processor, a road scene image from a camera coupled to a vehicle. The method also includes generating, with the electronic processor, a generative adversarial network (GAN) model using the road scene image. The method further includes determining, with the electronic processor, a distribution indicating a likelihood that each location in the road scene image can contain a road object using the GAN model. The method also includes determining, with the electronic processor, a plurality of locations in the road scene image by sampling the distribution. The method further includes detecting, with the electronic processor, the road object at one of the plurality of locations in the road scene image.
[0006] The present disclosure also provides a non-transitory computer-readable medium storing computer-readable instructions that, when executed by an electronic processor of a computer, cause the computer to perform operations. The operations include receiving a road scene image from a camera coupled to a vehicle. The operations also include generating a generative adversarial network (GAN) model using the road scene image. The operations further include determining a distribution using the GAN model, the distribution indicating a likelihood that each location in the road scene image can contain a road object. The operations also include determining a plurality of locations in the road scene image by sampling the distribution. The operations further include detecting the road object at one of the plurality of locations in the road scene image. BRIEF DESCRIPTION OF DRAWINGS
[0007] The accompanying drawings are incorporated in and form a part of the specification. They illustrate various embodiments and, together with the description, serve to explain the principles and advantages of the embodiments and to enable others skilled in the art to make and use these embodiments. In the drawings, like reference numerals refer to like elements throughout the several views.
[0008] Figure 1 is a block diagram of one example of a vehicle equipped with a system for detecting small objects in a road scene in accordance with some embodiments.
[0009] Figure 2 is a block diagram of one example of an electronic controller included in a system Figure 1 illustrated in FIG. 1 in accordance with some embodiments.
[0010] Figure 3 is a block diagram of one example of a generative adversarial network (GAN) model in accordance with some embodiments.
[0011] Figure 4 is a flowchart of one example of a method for detecting small objects in a road scene in accordance with some embodiments.
[0012] In the drawings, the system and method components have been represented where appropriate by conventional symbols in the drawings, showing only those specific details that are pertinent to understanding the embodiments so as not to obscure the disclosure with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein. DETAILED DESCRIPTION
[0013] Figure 1 is a block diagram of one example of a vehicle 100 equipped with a system 102 for detecting small objects in a road image. Figure 1The vehicle 100 is an automobile that includes four wheels 104, 106, 108, and 110. In some implementations, the system 102 is equipped to a vehicle having more or less than four wheels. For example, the system 102 can be equipped to a motorcycle, a truck, a bus, a trailer, etc. Indeed, the vehicle 100 includes additional components such as a propulsion system, a steering system, a braking system, etc. For ease of explanation, these additional components are not illustrated here.
[0014] Figure 1 The system 102 is illustrated here in a configuration that includes fewer or additional components in aspects of the configuration that are different than those illustrated. The camera 112 is coupled to a component of the vehicle 100 that faces in a forward driving direction (e.g., a front bumper, a hood, a rearview mirror, etc.). The camera 112 is configured to capture road scene images of an environment (e.g., a traffic scene) surrounding the vehicle 100. The road scene images can include small road objects. Small road objects include, for example, road markings, traffic signs, traffic lights, roadside objects (e.g., trees and poles), and other common objects found in road scenes. In some implementations, the camera 112 includes one or more cameras, one or more photographic cameras, or a combination thereof. The camera 112 is configured to communicate the captured road scene images to the electronic controller 114. Figure 1
[0015] Figure 2 is a block diagram of an example of the electronic controller 114. Figure 2 The electronic controller 114 is illustrated here in a configuration that includes fewer or additional components in aspects of the configuration that are different than those illustrated. For example, in practice, the electronic controller 114 can include additional components such as one or more power sources, one or more sensors, etc. For ease of explanation, these additional components are not illustrated here. Figure 2
[0016] The input / output interface 206 includes routines for transferring information between components within the electronic controller 114 and components external to the electronic controller 114. For example, the input / output interface 206 allows the electronic processor 202 to communicate with external hardware such as the camera 112. The input / output interface 206 is configured to transmit and receive data via one or more wired couplings (e.g., wires, fiber optics, etc.), wirelessly, or a combination thereof.
[0017] The user interface 208 includes, for example, one or more input mechanisms (e.g., a touch screen, a keypad, buttons, knobs, etc.), one or more output mechanisms (e.g., a display, a printer, a speaker, etc.), or a combination thereof. In some implementations, the user interface 208 includes a touch-sensitive interface (e.g., a touch screen display) that displays visual output generated by software applications executed by the electronic processor 202. The visual output includes, for example, graphical indicators, lights, colors, text, images, graphical user interfaces (GUIs), combinations of the foregoing, etc. The touch-sensitive interface also receives user input using detected physical contact (e.g., detected capacitance or resistance). In some implementations, the user interface 208 is separate from the electronic controller 114.
[0018] The bus 210 connects various components of the electronic controller 114, including, for example, the memory 204, to the electronic processor 202. The memory 204 includes, for example, read-only memory (ROM), random access memory (RAM), electrically erasable programmable read-only memory (EEPROM), other non-transitory computer-readable media, or a combination thereof. In some implementations, the memory 204 is included in the electronic processor 202. The electronic processor 202 is configured to retrieve computer-readable instructions and data from the memory 204 and execute the computer-readable instructions to perform the functionality described herein. Figure 2 The memory 204, which is illustrated in the middle, includes a detection platform 212 and a generative adversarial network (GAN) model 214. The detection platform 212 includes computer-readable instructions that cause the electronic processor 202 to perform, among other things, the methods described herein. For example, the detection platform 212 includes computer-readable instructions that cause the electronic processor 202 to generate the GAN model 214 to detect traffic lights (and other small road objects) in image data.
[0019] Current deep detector models rely on features of the objects being detected. If the objects are too small within the observed scene, the detector model can fail to detect the objects. This can pose a driving hazard for vehicles, especially autonomous vehicles, because important small objects such as traffic lights can not be detected and appropriate actions such as stopping at a red light will not be performed. The semantics of a road scene can provide an efficient method of predicting the likelihood that a road object can exist within the road scene. For example, human perception can predict the likelihood that a road object can exist at a particular location based on the semantics of the road scene without needing to see the road object itself. As described in more detail below, the GAN model 214 is configured to use the semantics of a road scene to predict the existence of a road object rather than relying on features of the road object.
[0020] GANs are a class of machine learning systems in which two neural networks compete against each other in a game. For example, one neural network generates augmented road scene images that appear real, and a second neural network evaluates the realism of the augmented road scene images. The GAN model 214 is trained to understand the semantics of road scenes and determine the likelihood that any location can contain a road object. To understand the semantics of road scenes, in some implementations, the GAN model 214 is trained using inpainted images in which road objects are removed from road scene images, and the GAN model 214 uses the original ground truth of the original image to determine the presence of the removed road objects. In some implementations, the output of the GAN model 214 is a multi-scale prediction in which one scale predicts the center of a road object and a second scale predicts a distribution (or heat map) indicating the likelihood that each location in the road scene image can contain a road object. In some implementations, the GAN model 214 is trained using an intersection over union (“loU”) function, a reconstruction loss function, a GAN loss function, or a combination thereof.
[0021] Figure 3 is a block diagram of one example of the GAN model 214. Figure 3 The GAN model 214 illustrated in FIG. 3 includes an image inpainter 302, a generator 304, a first discriminator 306, and a second discriminator 308. In some implementations, the GAN model 214 includes more than one generator and more or fewer than two discriminators. In addition, in practice, the GAN model 214 can include additional components such as one or more feedback networks, one or more backpropagation networks, one or more encoders, one or more decoders, and the like. For ease of explanation, these additional components are not illustrated here. Figure 3 The GAN model 214 illustrated in FIG. 3 includes an image inpainter 302, a generator 304, a first discriminator 306, and a second discriminator 308. In some implementations, the GAN model 214 includes more than one generator and more or fewer than two discriminators. In addition, in practice, the GAN model 214 can include additional components such as one or more feedback networks, one or more backpropagation networks, one or more encoders, one or more decoders, and the like. For ease of explanation, these additional components are not illustrated here.
[0022] The image inpainter 302 is configured to receive a road scene image and remove one or more traffic lights included therein to generate an inpainted image with the one or more traffic lights removed. In some implementations, the image inpainter 302 sets each pixel in the road scene image that contains a traffic light to a predetermined value (e.g., zero). The inpainted image is input into the generator 304. Figure 3The generator 304 illustrated in the middle includes a compression network 310 and a reconstruction network 312. The compression network 310 is configured to compress the content of the inpainted image into a smaller amount of content. The smaller amount of content output by the compression network 310 is the most meaningful content in the inpainted image and represents an understanding of the key connections between the localization of road objects and their surrounding environment. The reconstruction network 312 is configured to reconstruct the original image starting from the most meaningful content output by the compression network 310. For example, the reconstruction network 312 uses the most meaningful content to select locations to add road objects that are likely to be consistent with the semantics of the road scene. The first discriminator 306 and the second discriminator 308 are configured to compare the reconstructed image generated by the generator 304 to the original image to determine whether the reconstructed image is realistic or not realistic. The first discriminator 306 and the second discriminator 308 compare the reconstructed image at different scales. For example, the first discriminator 306 determines whether the final reconstructed image generated by the generator 304 is realistic or not realistic. In addition, the second discriminator 308 determines whether earlier stages of the reconstructed image generated by the generator 304 are realistic or not realistic. Based on the determinations of the first discriminator 306 and the second discriminator 308, the generator 304 is configured to adjust parameters within the compression network 310 and the reconstruction network 312 that are used to determine locations to add road objects that are likely to be consistent with the semantics of the road scene.
[0023] After multiple iterations of the training described above, the generator 304 is configured to consistently generate reconstructed images that are determined to be realistic by the first discriminator 306, the second discriminator 308, or both. As a result of the training, the first discriminator 306 determines, based on the semantics of the road scene, a distribution of the likelihood that any location in a road scene image can contain a road object. In addition, as a result of the training, the second discriminator 308 determines locations of the centers of road objects (i.e., anchor centers that can be sampled to detect small and occluded road objects) in a road scene image. The determined distribution and anchor centers are used to detect road objects in a road scene image, as will be described in more detail below.
[0024] Figure 4This is a flowchart of an example of a method 400 for detecting small objects in a road scene. In some implementations, the detection platform 212 includes computer-readable instructions that cause the electronic processor 202 to execute all (or any combination thereof) of the boxes of method 400 described below. In box 402, road scene images are received from a camera 112 coupled to the vehicle 100 as described above. For example, the electronic processor 202 receives road scene images from the camera 112 via an input / output interface 206. In some implementations, the camera 112 continuously captures video data of the road scene and transmits the video data as road scene images to the electronic processor 202. In some implementations, the camera 112 captures still images of the road scene and transmits the still images as road scene images to the electronic processor 202. In box 404, the road scene images are used to generate a GAN model 214. For example, as described above regarding... Figure 3 As described above, a GAN model 214 is generated using the repaired image.
[0025] In box 406, GAN model 214 is used to determine the distribution of the probability that each location in the road scene image can contain a road object. For example, electronic processor 202 uses the above-mentioned... Figure 3 The described GAN model 214 determines the distribution. In some implementations, GAN model 214 is configured to determine the distribution based solely on the semantics of a road scene image. For example, a traffic light may be located above a vehicle 100 or on the side of the road on which the vehicle 100 is traveling. Both traffic light instances will be supported by poles or other suspension mechanisms. GAN model 214 may be configured to identify posts or poles designed to hold traffic lights or pedestrian crossings, which may include traffic light intersections. In some implementations, GAN model 214 is configured to determine the distribution based on the semantics of a road scene image and at least one feature of a road object. For example, to detect traffic lights, GAN model 214 is configured to identify three adjacently placed circles (i.e., red, yellow, and green lights).
[0026] In some implementations, the distribution defines a set of anchor points used for sampling to detect road objects. In some implementations, the distribution includes heatmaps illustrating one or more predicted locations of road objects. In some implementations, the distribution includes multiple heatmaps with unique configurations based on random vectors. These random vectors are associated with different factors of the road scene (e.g., lighting, background objects, etc.). Based on these random vectors, each generated heatmap among the multiple heatmaps has a unique configuration.
[0027] At block 408, a plurality of locations in the road scene image is selected by sampling the distribution. For example, the electronic processor 202 samples the distribution to select a plurality of locations having a high likelihood of containing a road object. At block 410, a road object is detected at one of the plurality of locations in the road scene image. For example, the electronic processor 202 is configured to analyze each of the plurality of locations in the road scene image using a deep neural network model and detect a road object in one of the plurality of locations.
[0028] In some implementations, the electronic processor 202 is configured to take at least one action for the vehicle 100 based on detecting the road object at block 410. The action can be tracking the road object in subsequently captured road scene images or performing a driving operation based on the road object. For example, if the road object is a traffic light, the traffic light can be tracked and if the traffic light is yellow or red, the electronic processor 202 can generate a command to the vehicle 100 to stop at the traffic light.
[0029] The GAN model 214 described herein enables detection of small road objects that are not detected using other models. For example, Table 1 illustrates an example of a percentage of traffic lights that are not detected by other models but are detected using the GAN model 214 described herein.
[0030] Table 1 Sampling Efficiency Comparison
[0031]
[0032] Various aspects of the present disclosure can take any one or more of the following example configurations.
[0033] EEE (1) A system for detecting small objects in a road scene, the system comprising: a camera coupled to a vehicle, the camera configured to capture road scene images; and an electronic controller coupled to the camera, the electronic controller configured to: receive a road scene image from the camera, generate a generative adversarial network (GAN) model using the road scene image, determine a distribution indicating a likelihood that each location in the road scene image can contain a road object using the GAN model, determine a plurality of locations in the road scene image by sampling the distribution, and detect a road object at one of the plurality of locations in the road scene image.
[0034] EEE (2) The system of EEE (1), wherein the electronic controller is further configured to: remove the road object from the road scene image to generate a repaired image, and generate the GAN model using the repaired image.
[0035] EEE (3) The system of EEE (1) or EEE (2), wherein the GAN model is configured to determine the distribution based only on semantics of the road scene image.
[0036] EEE (4) The system of EEE (1) or EEE (2), wherein the GAN model is configured to determine the distribution based on semantics of the road scene image and at least one characteristic of the road object.
[0037] EEE (5) The system of any one of EEE (1) to EEE (4), wherein the distribution defines a set of anchor points for sampling the road object.
[0038] EEE (6) The system of any one of EEE (1) to EEE (5), wherein the distribution comprises a heat map that illustrates one or more predicted locations of the road object.
[0039] EEE (7) The system of any one of EEE (1) to EEE (6), wherein the distribution comprises a plurality of heat maps, wherein each of the plurality of heat maps has a unique configuration based on a random vector.
[0040] EEE (8) The system of EEE (7), wherein the random vector is based on a set of factors of the road scene image.
[0041] EEE (9) The system of any one of EEE (1) to EEE (8), wherein the electronic controller is further configured to take at least one action on the vehicle based on detecting the road object.
[0042] EEE (10) A method for detecting small objects in a road scene, the method comprising: receiving, with an electronic processor, a road scene image from a camera coupled to a vehicle; generating, with the electronic processor, a generative adversarial network (GAN) model using the road scene image; determining, with the electronic processor, a distribution using the GAN model, the distribution indicating a likelihood that each location in the road scene image can contain a road object; determining, with the electronic processor, a plurality of locations in the road scene image by sampling the distribution; and detecting, with the electronic processor, the road object at one of the plurality of locations in the road scene image.
[0043] EEE (11) The method of EEE (10), further comprising: removing, with the electronic processor, the road object from the road scene image to generate a repaired image; and generating, with the electronic processor, the GAN model using the repaired image.
[0044] EEE (12) The method of EEE (10) or EEE (11), wherein the GAN model determines the distribution based only on semantics of the road scene image.
[0045] EEE (13) The method of EEE (10) or EEE (11), wherein the GAN model determines the distribution based on semantics of the road scene image and at least one characteristic of the road object.
[0046] EEE (14) The method of any one of EEE (10) to EEE (13), wherein the distribution defines a set of anchor points for sampling the road object.
[0047] EEE (15) The method of any one of EEE (10) to EEE (14), wherein the distribution comprises a heat map that illustrates one or more predicted locations of the road object.
[0048] EEE (16) The method of any one of EEE (10) to EEE (15), wherein the distribution comprises a plurality of heat maps, wherein each of the plurality of heat maps has a unique configuration based on a random vector.
[0049] EEE (17) The method of EEE (16), wherein the random vector is based on a set of factors of the road scene image.
[0050] EEE (18) The method of any one of EEE (11) to EEE (17), further comprising taking, with the electronic processor, at least one action with the vehicle based on detecting the road object.
[0051] EEE (19) A non-transitory computer-readable medium storing computer-readable instructions that, when executed by an electronic processor of a computer, cause the computer to perform operations comprising: receiving a road scene image from a camera coupled to a vehicle; generating a generative adversarial network (GAN) model using the road scene image; determining, using the GAN model, a distribution indicating a likelihood that each location in the road scene image can contain a road object; determining a plurality of locations in the road scene image by sampling the distribution; and detecting the road object at one of the plurality of locations in the road scene image.
[0052] EEE (20) The non-transitory computer-readable medium of EEE (19), wherein the operations further comprise: removing the road object from the road scene image to generate a repaired image; and generating the GAN model using the repaired image.
[0053] Accordingly, the present disclosure provides, among other things, systems, methods, and non-transitory computer-readable media for detecting small objects in a road scene. Various features, advantages, and embodiments are set forth in the following claims.
[0054] Machine learning generally refers to the ability of a computer program to learn without being explicitly programmed. In some implementations, a computer program (e.g., a learning engine) is configured to construct an algorithm based on input. Supervised learning involves presenting a computer program with example inputs and their desired outputs. The computer program is configured to learn the general rules that map inputs to outputs from the training data it receives. Example machine learning engines include decision tree learning, association rule learning, artificial neural networks, classifiers, inductive logic programming, support vector machines, clustering, Bayesian networks, reinforcement learning, representation learning, similarity and metric learning, sparse dictionary learning, and genetic algorithms. Using one or more of the above methods, a computer program can ingest, parse, and understand data and progressively refine the algorithms for data analysis.
[0055] In the foregoing specification, specific embodiments have been described. However, it will be apparent that various modifications and changes can be made within the scope of the application as set forth in the claims below. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense, and all such modifications are intended to be included within the scope of present teachings.
[0056] The benefits, advantages, solutions to problems, and any one or more elements of any of the elements that cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as a critical, required, or essential feature or element of any or all the claims. The application is defined solely by the appended claims including any amendments made during the pendency of this application and all equivalents of the claims as issued.
[0057] Furthermore, relational terms such as first and second, top and bottom, and the like can be used herein solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. The terms "comprises," "comprising," "has," "having," "includes," "including," "contains," "containing" or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises, has, includes, contains a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a," "has... a," "includes... a," or "contains... a" does not, without more constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises, has, includes, contains the element. The terms "a" and "an" are defined as one or more unless explicitly stated otherwise herein. The terms "substantially," "essentially," "approximately," "about" or any other version thereof, are defined as being close to as understood by one of ordinary skill in the art, and in one non-limiting embodiment the term is defined to be within 10%, in another embodiment within 5%, in another embodiment within 1% and in another embodiment within 0.5%. The term "coupled" as used herein is defined as connected, although not necessarily directly, and not necessarily mechanically. A device or structure that is "configured" in a certain way is configured in at least the stated way, but can also be configured in
[0058] The Abstract is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that the Abstract will not be used to interpret or limit the scope or the meaning of the claims. In addition, in the foregoing Detailed Description, it can be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than are explicitly recited in each claim. Rather, inventive subject matter can lie in less than all features of a single disclosed embodiment. Thus, the following claims are hereby expressly incorporated into this Detailed Description, with each claim acting as a separate embodiment of the inventive subject matter.
Claims
1. A system for detecting a small object in a road scene, the system comprising: a camera coupled to a vehicle, the camera configured to capture a road scene image; and an electronic controller coupled to the camera, the electronic controller configured to: receive the road scene image from the camera, generate a generative adversarial network (GAN) model using the road scene image, determine, using the GAN model, a distribution indicating a likelihood that each location in the road scene image can contain a road object based only on semantics of the road scene image, determine a plurality of locations in the road scene image by sampling the distribution, and detect the road object at one of the plurality of locations in the road scene image.
2. The system of claim 1, wherein, the electronic controller is further configured to: remove the road object from the road scene image to generate a inpainted image, and generate the GAN model using the inpainted image.
3. The system of claim 1, wherein the GAN model is configured to determine the distribution based on semantics of the road scene image and at least one characteristic of the road object.
4. The system of claim 1, wherein the distribution defines a set of anchor points for sampling the road object.
5. The system of claim 1, wherein the distribution comprises a heat map that illustrates one or more predicted locations of the road object.
6. The system of claim 1, wherein the distribution comprises a plurality of heat maps, wherein each of the plurality of heat maps has a unique configuration based on a random vector.
7. The system of claim 6, wherein the random vector is based on a set of factors of the road scene image.
8. The system of claim 1, wherein the electronic controller is further configured to take at least one action on the vehicle based on detecting the road object.
9. A method for detecting a small object in a road scene, the method comprising: receiving, with an electronic processor, a road scene image from a camera coupled to a vehicle; generating, with the electronic processor, a generative adversarial network (GAN) model using the road scene image; determining, with the electronic processor, a distribution indicating a likelihood that each location in the road scene image can contain a road object based only on semantics of the road scene image using the GAN model; determining, with the electronic processor, a plurality of locations in the road scene image by sampling the distribution; and detecting, with the electronic processor, the road object at one of the plurality of locations in the road scene image.
10. The method of claim 9, further comprising: removing, with the electronic processor, the road object from the road scene image to generate an inpainted image; and generating, with the electronic processor, the GAN model using the inpainted image.
11. The method of claim 9, wherein the GAN model determines the distribution based on semantics of the road scene image and at least one characteristic of the road object.
12. The method of claim 9, wherein the distribution defines a set of anchor points for sampling the road object.
13. The method of claim 9, wherein the distribution comprises a heat map that illustrates one or more predicted locations of the road object. 14. The method of claim 9, wherein the distribution comprises a plurality of heat maps, wherein each of the plurality of heat maps has a unique configuration based on a random vector.
15. The method of claim 14, wherein the random vector is based on a set of factors of the road scene image.
16. The method of claim 9, further comprising taking at least one action with the vehicle based on detecting the road object using the electronic processor.
17. A non-transitory computer-readable medium storing computer-readable instructions that, when executed by an electronic processor of a computer, cause the computer to perform operations comprising: receiving a road scene image from a camera coupled to a vehicle; generating a generative adversarial network (GAN) model using the road scene image; determining a distribution indicating a likelihood that each location in the road scene image can contain a road object using the GAN model based only on semantics of the road scene image; determining a plurality of locations in the road scene image by sampling the distribution; and detecting a road object at one of the plurality of locations in the road scene image.
18. The non-transitory computer-readable medium of claim 17, wherein the operations further comprise: removing the road object from the road scene image to generate a repaired image; and generating the GAN model using the repaired image.
Citation Information
Patent Citations
Context-based priors for object detection in images
CN107851191A
Traffic sign recognition method, device and equipment and computer readable medium
CN110135301A
Detecting objects using a weakly supervised model
CN110276366A
Refining Synthetic Data With A Generative Adversarial Network Using Auxiliary Inputs
US20190080206A1