Use of Image Augmentation with Simulated Objects for Training Machine Learning Models in Autonomous Driving Applications

By augmenting real-world images with simulated objects, the method addresses the data scarcity and safety challenges in training DNNs for autonomous vehicles, improving detection accuracy and reducing training time, thus enhancing the reliability of autonomous vehicle systems.

JP7702268B2Active Publication Date: 2025-07-03NVIDIA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2021059521
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-01-15
Filing Date
2021-03-31
Publication Date
2025-07-03
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

Conventional systems face challenges in training deep neural networks (DNNs) for autonomous vehicles to accurately detect road debris due to the scarcity of real-world data and the danger of intentionally creating scenarios with road objects, leading to a discrepancy in training data and safety concerns.

Method used

The method involves augmenting real-world images with simulated road objects to enhance training data, using a combination of real-world and simulated data to improve object detection accuracy and reduce training time, while allowing for zero-shot learning to identify unknown objects.

Benefits of technology

This approach enhances the accuracy and speed of training DNNs for object detection, enabling safer and more reliable autonomous vehicle operations by providing a larger, diverse dataset that mimics real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702268000001
    Figure 0007702268000001
  • Figure 0007702268000002
    Figure 0007702268000002
  • Figure 0007702268000003
    Figure 0007702268000003
Patent Text Reader

Abstract

To provide a system for and a method of storing a plenty of detailed information from a real world image by expanding a real world image with an object obtained by simulation of training a machine learning model so that an object can be detected in an input image.SOLUTION: A machine learning model is trained during its arrangement so that it can determine a boundary shape for use in detection of an object or encapsulation of a detected object. The machine learning model is further trained so as to carry out determination of a type of an object encountered on road, calculation of risk degree, and calculation of reliability degree. During the arrangement, detection of an object on a road, determination of its boundary shape, identification of a type of the object on road, and / or calculation of risk degree by the machine learning model can be used, in various autonomous machine applications, for succeeding steps relating to surrounding environment, for example, such as determination of avoidance of an object fallen on a road, travelling on the object fallen on road, or complete stop.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 003,879, filed Apr. 1, 2020, which is incorporated herein by reference in its entirety.

Background Art

[0002] Autonomous and semi-autonomous vehicles utilize machine learning, such as a deep neural network (DNN), to analyze the road surface to guide the position of the vehicle relative to road boundaries, lanes, road debris, road obstacles, road signs, and the like when the vehicle is in operation. For example, a DNN can be used to detect road debris (e.g., animals, cones, construction materials) in the approaching portion of the road when an autonomous vehicle is in operation, which can lead to adjustment of the position of the autonomous vehicle (e.g., steering to avoid driving on a traffic cone in the center of the road). However, training a DNN to accurately detect objects on the road requires a vast amount of training data, computing power, and human time and effort. Additionally, since debris is generally avoided by drivers and / or quickly removed from the roadway, capturing real-world image data of roads with objects, such as debris, is a difficult task. However, thousands of training data instances are required to adequately train a DNN. Thus, there can be a large discrepancy between the amount of useful training data that can be collected that includes debris and the amount of training data required to accurately train a DNN to detect road debris.

[0003] For example, conventional systems often rely on real-world data captured by physical vehicles operating in various environments to generate training data for DNNs. However, this methodology has several problems. For example, since society prioritizes removing objects that pose dangerous risks to drivers from the road, the chances of a physical vehicle encountering a lane with road objects that it needs to drive through or over are limited. However, it is difficult and dangerous to intentionally create lanes with heterogeneous objects, especially on major roads, within the vehicle's path. Similarly, it takes time and effort to test a DNN for the accuracy of detecting objects in order to determine whether to drive over or through an object in such real-world environments. Furthermore, it is unlikely that automobile manufacturers will release fully autonomous vehicles that operate using only DNNs until a high level of safety and accuracy is achieved. As a result, the conflict of interest between safety and accuracy is making it increasingly difficult to generate practical, stable, and reliable autonomous driving systems.

Summary of the Invention

Means for Solving the Problems

[0004] Embodiments of the present disclosure relate to training a machine learning model for detecting objects using a real-world image augmented with simulated objects. Systems and methods are disclosed that preserve rich, detailed, and large amounts of information in a real-world image when training a machine learning model by inserting simulated instances of road debris into the real-world image, which is difficult to systematically or intentionally capture the configuration in a real-world scenario. As such, embodiments of the present disclosure relate to the detection of objects (including, but not limited to, road debris (e.g., cardboard boxes, rocks, wheels, wooden pallets, dead animals, logs, traffic cones, mattresses, etc.) and / or billboards (e.g., road signs, poles, etc.)) for autonomous machines.

[0005] In contrast to conventional systems such as those described above, the system of the present disclosure can train a machine learning model to determine the boundary shape of each object in an input image by detecting objects on the road (such as dead animals or mattresses) and using a real-world image augmented with simulated objects. As a result, while it is cumbersome and dangerous to obtain real-world images of road debris on the driving surface, the machine learning model can be trained with real-world images of a large dataset augmented with various simulated road objects that retain as much detail as possible from the real-world images. As such, by training the machine learning model using real-world images with simulated objects, the model can achieve higher accuracy in detecting objects on the driving surface, reduce the time required to train the machine learning model to achieve a high level of accuracy, and use a large number of images closer to real-world scenarios, which leads to an improvement in the results in decision-making when encountering objects on the road, and can be trained.

[0006] In addition to detecting objects on the road, the boundary shape around the object on the road and the degree of danger of the object on the road can be calculated by the machine learning model to help the system understand the harmful danger of the object on the road and the actions to be performed in response (e.g., steer around it, drive over it, come to a complete stop, etc.), thereby increasing the accuracy of the system's response to objects on the road encountered. To further reduce the execution time of the system's real-time operation, the machine learning model can be trained to detect whether there is an object on the road (regardless of whether the type of the object is accurately identified) using zero-shot learning, thereby removing the constraint of conventional systems that require accurate identification of the object on the road before determining the next step in response to the detected object on the road.

[0007] The present system and method for training a machine learning model to detect an object using a real-world image augmented with a simulated object will be described in detail below with reference to the accompanying drawings.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11

Figure 12

DETAILED DESCRIPTION OF THE INVENTION

[0009] Disclosed are systems and methods for training a machine learning model to detect an object using a real-world image augmented with a simulated object. The present disclosure may be described with respect to an exemplary autonomous vehicle 1000 (an example of which is described herein with respect to FIGS. 10A-10D and which is alternatively referred to herein as "vehicle 1000" or "ego vehicle 1000"), but this is not intended to be limiting. For example, the systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in an advanced driver assistance system (ADAS)), robots, warehouse vehicles, off-road vehicles, flying vessels, boats, passenger vehicles, cars, trucks, buses, first responder vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police vehicles, ambulances, construction vehicles, submarines, drones, and / or other types of vehicles (e.g., unmanned and / or carrying one or more passengers). Additionally, the present disclosure may be described with respect to autonomous driving, but this is not intended to be limiting. For example, the systems and methods described herein may be used in other technical fields such as robotics, aviation systems, marine systems, and / or other processes such as perception, world model management, path planning, obstacle avoidance, and / or others.

[0010] The system of the present disclosure generates training data that includes both real-world data and simulated data for training a machine learning model (e.g., a deep neural network (DNN) such as a convolutional neural network (CNN)) while retaining the real-world data as richly as possible and for testing the machine learning model by using the training data as ground truth data (e.g., annotated labels corresponding to lane boundaries, road boundaries, text, and / or other features). As a result, the accuracy of detecting objects in a real-world environment, such as road debris and / or road signs, is enhanced while providing a large amount of data for training and testing the machine learning model. For example, the image data for training and testing the machine learning model may be real-world images of a road that include synthetic or simulated images of objects on the road. A number of instances of simulated objects may be generated to correspond to several different features (e.g., location, orientation, appearance attributes, lighting, occlusion, and / or environmental conditions) so that a variety of object instances can be used to generate a more robust training set. By randomly sampling these object instances (or the conditions used to modify the instances), various instances of the simulated objects can be inserted into previously captured real-world image data, thereby generating a large amount of image data and ground truth data corresponding to simulated road objects (e.g., road debris) within the real-world image data. Some non-limiting benefits of the system and method of the present disclosure are improved object detection accuracy, improved recognition of road debris that does not easily conform to a specific class of objects, and the ability to classify detected objects (thereby reducing the computational load for in-vehicle inference).

[0011] In an embodiment of the present disclosure, to ultimately assist in determining the path of an autonomous vehicle, a DNN can be used to generate boundary shapes (e.g., squares, rectangles, triangles, polygons) corresponding to objects or debris on the road (e.g., cardboard boxes, rocks, wheels, wooden pallets, dead animals, logs, traffic cones, mattresses, road signs, etc.). Simulated objects on the road can be generated using a simulator (e.g., NVIDIA DriveSim). The simulated objects can be generated from the perspective of the virtual sensors of a passenger vehicle within a virtual environment generated by the simulator. In some examples, multiple instances of each object are generated by simulating the same object under different conditions. For example, a simulated mattress can be modified as if it were at different times, different positions relative to the sun, different orientations or poses, and / or different distances from the virtual sensors within the virtual environment to create instances of the simulated mattress that reflect different environmental and location conditions. A segmentation mask corresponding to each instance of the object can also be generated and used to determine the boundary shape enclosing each instance. For example, the boundary shape can be generated by forming a polygon that tightly encloses all the pixels within the segmentation mask of the instance of the object.

[0012] In some embodiments, real-world images (and their corresponding ground-truth data) can be augmented with instances of simulated objects in order to generate image data for training a machine learning model as well as to generate ground-truth data for testing the machine learning model. For example, a real-world image of a road can have corresponding ground-truth data indicating an object (e.g., a passenger vehicle on the road) in a scene having a boundary shape surrounding each object (e.g., each passenger vehicle on the road has a bounding box around it). In some instances, real-world images are captured using sensors attached to physical passenger vehicles. To generate training images, the real-world images can be augmented by inserting instances of simulated objects, such as cardboard boxes, into the images. In some instances, the instances for insertion into the real-world images are selected by determining a reference random sampling and selecting one or more instances that meet (or most closely match) the sampling criteria. For example, the only instance of a simulated object that meets the position and orientation criteria can be selected for inclusion in the real-world image. Similarly, the criteria that the instances must meet can also be determined by conditions present within the real-world image. For example, the position and time of day of the sun in the real-world image can be identified and used to select instances of simulated objects under similar conditions. In other instances, the conditions can be randomly selected such that the simulated instances of the objects need not match the real-world environmental conditions at the time of image capture. By doing so, in embodiments, the machine learning model can learn to detect previously unknown objects from the various different training data samples it processes.

[0013] When an instance of a simulated object is inserted into a real-world image, one or more conditions can be evaluated to determine whether the augmented real-world image is a final training image. For example, the distance from an instance of a simulated object to the driving surface can be evaluated to determine whether the distance exceeds a threshold. For example, if an instance of a simulated object, such as a cardboard box that adheres to a specified standard, is 152.4 cm (5 feet) above the driving surface and exceeds a threshold of 91.44 cm (3 feet), then the instance can be not selected for inclusion in the real-world image. Similarly, overlapping conditions regarding the placement of an instance of a simulated object relative to an object existing within the real-world image can be considered. For example, the boundary shape of an instance of a simulated object can be compared to the boundary shape of an object within the real-world image (as determined by the corresponding ground truth data of the real-world image) to determine whether the boundary shapes overlap. If the boundary shapes overlap, the instance of the simulated object can be rejected and not inserted into the real-world image. In some examples, the manner in which the boundary shapes overlap (e.g., whether one object is in front of another) can be considered. If all (or sufficient) conditions are met, an instance of a simulated object can be inserted into the real-world image and used to train a machine learning model.

[0014] The boundary shape of the simulated instance of the object and the boundary shape of the object in the real-world image (as shown in the corresponding ground truth data) can then be used as ground truth data to test whether the machine learning model accurately identifies the object or debris. In some examples, the ground truth data can also be used in the real-world image to determine whether the machine learning model accurately classifies the object, e.g., an animal vs. a cardboard box. Even if the model does not accurately classify an object not previously seen, the machine learning model can also be tested via the use of zero-shot learning techniques to determine whether the model can identify a previously unseen object on the road as a potential hazard. For example, a real-world image augmented with a towel (not included in the images used to train the model) can be used to test the model and determine whether the model accurately draws a boundary shape around the towel, regardless of whether the model identifies the object itself as a towel. As a result, due to the vast amount of potential road debris types, the machine learning model can be further used to accurately detect unfamiliar road debris while maintaining the safety of the occupants and enable the autonomous vehicle to pass around the debris.

[0015] Referring to FIG. 1, FIG. 1 is a data flow diagram showing an exemplary process 100 for generating a real-world image augmented with virtual objects for training a machine learning model to detect road debris, according to some embodiments of the present disclosure. It should be understood that this and other configurations described herein are merely examples. Other configurations and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components and in any suitable combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory. In some embodiments, the components, features, and / or functionality described with respect to process 100 of FIG. 1 may be similar to those of vehicle 1000, exemplary computing device 1100 of FIG. 11, and / or exemplary data center 1200 of FIG. 12.

[0016] At a high level, process 100 may include a 3D model generator 102 that generates multiple representations (e.g., virtual instances) of objects (e.g., road debris) that adhere to various criteria in a virtual, simulated environment generated by simulator 104. Simulator 104 may include a virtual vehicle having virtual sensors that record and capture simulation data from the perspective of the virtual sensors as the virtual vehicle drives on a driving surface. In some embodiments, the representations (e.g., virtual instances) as generated by 3D model generator 102 may then be inserted into the virtual environment of simulator 104 and recorded in an image from the perspective of the virtual sensors. A segmentation mask of the representation of the object may also be generated by segmentation mask generator 106 when simulation data is being recorded. Ground truth data, which may include real-world images and annotations of objects on the real-world image's driving surface, may be loaded into ground truth loader 108. Semantic augmenter 110 can augment the image with one or more generated graphic representations of the object onto the loaded ground truth data (e.g., by inserting). In some embodiments, the representations may be selected based on random sampling of criteria and / or criterion matching conditions in the real-world image of the ground truth data. In some embodiments, only those representations that meet constraints regarding a distance threshold from the driving surface and an amount of overlap with other objects within the real-world image are used to augment (e.g., insert into) the real-world image to create a training image. Semantic augmenter 110 can also draw a boundary shape around the inserted representation based on the corresponding segmentation mask to determine whether the inserted representation meets the constraints. After representations that meet the constraints are inserted into the real-world image, the ground truth data can be updated by ground truth data updater 112 to include the augmented real-world image and the corresponding boundary shape around the virtual object.The augmented reality image can be completed by the dataset generator 114 and used by the trainer 116 to train a machine learning model. After being trained, the machine learning model can be deployed via the deployment module 118 and used by the vehicle 1100 to detect debris or objects on the driving surface using the image and to determine the next step in response to the detected road objects. The next step or action may include world model management, route planning, control decisions, obstacle avoidance, and / or other actions of the autonomous or semi-autonomous driving software stack.

[0017] The 3D model generator 102 may be a game-based engine that generates virtual instances of virtual objects. The virtual objects span a variety of potential road debris or object types including, but not limited to, cardboard boxes, rocks, wheels (and wheel parts), wooden pallets, dead animals, ladders, logs, traffic cones, mattresses, road signs, and other objects. The 3D model generator 102 can generate multiple instances of a given virtual object and each virtual instance can be subjected to different sets of conditions such as location, lighting, orientation, pose, occlusion, environmental conditions, and / or appearance attributes. For example, the 3D model generator 102 can generate 20 instances of a simulated construction cone where each representation of the construction cone uses a different material or color tone, providing a large dataset of construction cones for training a machine learning model.

[0018] In some embodiments, the environmental conditions imposed on the virtual instances of the virtual object may include changing time, position relative to the sun, weather conditions (e.g., rain, fog, cloud), visibility distance, and / or different distances from virtual sensors within the virtual environment for domain randomization. In some embodiments, the appearance attributes of the virtual instances may include changing color tones, materials, and textures. For example, the 3D model generator 102 can generate a set of 3D virtual instances of a mattress, where each instance of the mattress is at a different position, exposed to different orientations and positions of the sun. Similarly, each 3D virtual instance of the mattress can be exposed to changing weather conditions, such as a rainy day versus a cloudy day, or changing lighting conditions, such as midday, night, dawn, dusk, etc. In some embodiments, the 3D model generator 102 can generate virtual instances of an object at set intervals or increments of various conditions to provide a large number of virtual instances or other representations that can be used to train a machine learning model. For example, the 3D model generator 102 can generate virtual instances at set distance intervals away from the virtual sensors of a virtual vehicle to ensure that there are a number of virtual instances to choose from. For example, the 3D model generator can generate a virtual instance every 60.96 cm (2 feet) away from the virtual sensor. Similarly, the 3D model generator 102 can generate multiple virtual instances that have the same set of some conditions (e.g., position, location, orientation) but differ in one condition (e.g., weather condition) for each instance to ensure that any possible set of selected criteria is likely to match exactly the generated virtual instances.

[0019] In some embodiments, the 3D model generator 102 can operate with the simulator 104 to determine the application of various conditions to a virtual instance of an object. For example, the simulator 104 can provide information regarding simulated environmental conditions that can be used by the 3D model generator 102 and, in response, generate a virtual instance. For example, information regarding the angle of a virtual sensor of a virtual vehicle in a simulated environment can be used by the 3D model generator 102 to create a virtual instance having a corresponding orientation or pose relative to the angle of the virtual sensor mounted on the virtual vehicle in the simulated environment. In any instance, a sensor model corresponding to a real-world sensor of the vehicle 1100 can be used in a virtual sensor within the simulated environment.

[0020] Referring now to FIG. 2A, FIG. 2A is an example of a simulated image that includes multiple representations (e.g., virtual instances) of a dead deer as a road obstacle, according to some embodiments of the present disclosure. In this example, the simulated image 200 includes a virtual environment 202 and representations (e.g., virtual instances) 204, 206, and 208 of a dead deer. Each virtual instance 204, 206, and 208 of the dead deer has a different position and location and, for the purposes of this example, may also have other varying characteristics such as color tone. A number of virtual instances, each having a different set of conditions, realize a high likelihood that at least one (or exactly one) of the generated virtual instances will meet a selected criterion.

[0021] Returning to FIG. 1, the simulator 104 can provide a virtual environment to capture data of a virtual vehicle driving in the virtual environment from the perspective of virtual sensors mounted on the virtual vehicle. In some embodiments, the virtual environment and the virtual sensors of the virtual vehicle can be adjusted to conform to conditions (e.g., sensor angle and height from the ground) in a real-world image that is used as part of the ground truth data. Virtual instances of objects generated by the 3D model generator 102 can be inserted into the virtual environment of the simulator 104. After being inserted into the virtual environment, the simulator 104 can adjust some properties of the virtual instances. For example, those virtual instances that are not located on the driving surface can be discarded. Similarly, virtual instances can be scaled down or enlarged based on the distance from the virtual sensors of the virtual vehicle such that virtual instances farther from the virtual sensors can be scaled down smaller than virtual instances closer to the sensors. The simulator 104 can also relocate virtual instances based on other conditions, for example, placing those virtual instances that are generated under shaded conditions in a portion of the virtual environment where clouds are blocking the sun.

[0022] After the virtual instance is inserted into the virtual environment, the simulator 104 can record a vehicle that encounters one or more virtual instances of objects driving on the driving surface and falling on the road. As described herein, the data recorded by the simulator 104 can also be used by the segmentation mask generator 106 to generate a segmentation mask of the virtual instance.

[0023] Referring now to FIG. 2A, the simulated image 200 includes, as generated by the simulator 104, a virtual environment 202 and a plurality of representations (e.g., virtual instances) 204, 206, and 208 of a dead deer. As shown in FIG. 2A, the plurality of virtual instances 204, 206, and 208 are integrated into the virtual environment 202 as road debris on the driving surface. The virtual instances 204, 206, and 208 can be further modified or extended to meet the conditions of the virtual environment 202. For example, virtual instances outside the driving surface within the virtual environment 202 are not included. Similarly, the virtual instances 204, 206, and 208 can be scaled in proportion to the distance from the virtual sensors of the virtual vehicle within the virtual environment 202. For example, since the virtual instance 206 is farther from the virtual sensor compared to the virtual instance 204, the virtual instance 206 of the dead deer is scaled down smaller than the virtual instance 204. Similarly, as shown in FIG. 21, the virtual instance 208 is scaled down to be much smaller than the virtual instance 204 and slightly smaller than the virtual instance 206 based on the distance from the virtual sensor. In addition, each virtual instance is oriented so as to lie flat on the driving surface, which has a slight upward slope in the virtual environment 202.

[0024] Returning to FIG. 1, the segmentation mask generator 106 generates a segmentation mask corresponding to a virtual instance of a virtual object. The segmentation mask represents a smaller portion of the image and can indicate pixels regarding that smaller portion of the image. For example, the segmentation mask generator 106 can generate a separate segmentation mask representing the pixels corresponding to each virtual instance of the object. For example, the segmentation mask of a particular virtual instance can indicate the pixels in the image that represent the virtual instance but do not provide data (such as color, appearance, etc.) regarding the virtual instance itself. In some embodiments, the segmentation mask can include pixels encompassed by all virtual instances. In some embodiments, the segmentation mask can be stored as a single file or as a series of files, with each file corresponding to a particular virtual instance. Similarly, the segmentation mask can also be stored as a transparent image file (such as a PNG file type).

[0025] Referring now to FIG. 2B, FIG. 2B is an exemplary segmentation mask corresponding to the virtual instances 204, 206, and 208 of the dead deer of FIG. 2A. The segmentation mask can be represented as a mask representing all of the virtual instances 204, 206, and 208, as shown by segmentation mask 210. Similarly, the segmentation mask can be represented by a series of segmentation masks, where each mask corresponds to a single virtual instance, as shown by segmentation masks 212, 214, and 216. The segmentation masks 212, 214, and 216 each delineate and indicate the contours of the pixels present within each virtual instance 204, 206, and 208, indicating the positions of each virtual instance 204, 206, and 208 within the virtual environment 202. However, the segmentation masks 212, 214, and 216 need not each include information (such as pixel color) regarding the corresponding virtual instances 204, 206, and 208.

[0026] For the pixels of each virtual instance indicated by the segmentation mask, the segmentation mask generator 106 can also generate a boundary shape (e.g., square, rectangle, ellipse, triangle, polygon, etc.) that encloses each virtual instance to determine whether the virtual instance satisfies the constraints within the real-world image. For example, for a given virtual instance of an object, the segmentation mask generator 106 can use the segmentation mask of that virtual instance to determine the pixels corresponding to the virtual instance and draw a boundary polygon (or other shape) around the segmentation mask that encompasses all the pixels. In some embodiments, the boundary shape can be firmly formed around the pixels.

[0027] Referring now to FIG. 3, FIG. 3 is an exemplary visualization 300 of a real-world image having an extended simulated object that cannot satisfy a threshold distance constraint, according to some embodiments of the present disclosure. Visualization 300 also includes ground truth annotations such as boundary shapes around objects within the real-world image. As shown in FIG. 3, the virtual instance 206 of FIG. 2A is inserted into visualization 300. The virtual instance 206 includes a boundary shape 304 that is formulated based on the segmentation mask 214 of FIG. 2B. The boundary shape 304 encapsulates all the pixels that represent the virtual instance 206. FIGS. 4 and 5 similarly show boundary shapes 402 and 502 around the virtual instances 204 and 208 respectively drawn by the segmentation mask generator 106 based on the respective segmentation masks 212 and 216.

[0028] Returning to FIG. 1, the ground truth loader 108 receives and loads ground truth data. The ground truth data may be a real-world image having annotation data, such as boundary shapes and labels (e.g., type of object within a road, degree of danger) within the road. The ground truth data can be generated by manually and / or automatically generating a shape around the object. For example, the boundary shapes or other annotation data used for the ground truth data can be synthetically manufactured (e.g., generated from a computer model or rendering), manufactured in reality (e.g., designed and manufactured from real-world data), machine automated (e.g., using feature analysis and learning to extract features from the data and then generate a boundary shape), manually annotated by a person (e.g., an annotation expert defines the boundary shape of the object), and / or a combination thereof (e.g., a polyline point annotated by a person and a rasterizer generates a complete polygon from the polyline point). For example, the ground truth data may be a real-world image of a road having other passenger vehicles within a road lane having annotation such as boundary shapes around each passenger vehicle on the road. The ground truth data may also include annotations such as boundary shapes around other objects within the view, such as road signs or people crossing the road. Similarly, the ground truth data can include a label to indicate that the object is a towel and can be labeled as non-harmful to the detected object.

[0029] Referring now to FIG. 3, visualization 300 includes a ground truth annotation of a real-world image that includes a boundary shape around an object within the view. For example, visualization 300 includes a real-world image 302 that includes a rich, accurate level of detail that is only possible via real-world image data. Visualization 300 also includes boundary shapes 306, 308, 310, 312, and 314 around various objects on the driving surface prior to the insertion of virtual instance 204. For example, boundary shapes 306, 308, 312, and 314 each encapsulate pixels that represent a particular vehicle. Similarly, boundary shape 310 encapsulates pixels that represent a person crossing the roadway. In this example, boundary shapes 306, 308, 310, 312, and 314 may have been drawn by a person prior to being loaded by the ground truth loader 108 of FIG. 1. For the purposes of this example, an object within real-world image 302 (e.g., a person crossing the street) may be labeled as harmful to indicate that the object should be avoided. Visualizations 400 and 500 of FIGS. 4 and 5 similarly show a real-world image 302 having ground truth annotations such as boundary shapes 306, 308, 310, 312, and 314.

[0030] Returning to FIG. 1, the semantic augmenter 110 can determine which representation of an object to use to augment the real-world image. In some embodiments, the semantic augmenter 110 samples a criterion or condition (e.g., location, lighting, orientation, pose, environmental conditions, and / or appearance attributes) (randomly according to one or more embodiments), and can select one or more representations that match or closely match the sampled criterion. For example, the semantic augmenter 110 can randomly select the orientation and position of a mattress and search for a representation of the mattress that matches the randomly sampled criterion. In some embodiments, the semantic augmenter 110 can select a representation that meets the conditions present in the real-world image that exists within the ground truth data. For example, the semantic augmenter 110 can search for representations that exhibit similar conditions present in the real-world image, such as the position of the sun, time of day, pose, orientation, lighting, occlusion, shadow, and weather conditions. In some embodiments, the semantic augmenter 110 can select criteria for using a hybrid approach having some criteria based on what exists within the real-world image and others selected based on random sampling. In some embodiments, the semantic augmenter 110 can also insert a representation into the real-world image and modify the representation to adhere to the conditions within the real-world image. For example, a representation that closely matches (but does not exactly match) the sampled criterion can be inserted into the real-world image and adjusted to fit the precisely sampled criterion. For example, a representation having a closely matching criterion can be adjusted so that its pose is precisely aligned with the viewpoint of the sensor used to capture the real-world image.

[0031] The semantic augmenter 110 can also use an expression that meets the selected criteria to expand the real-world image and determine whether the expression meets the placement constraints within the real-world image. For example, the expression may have been generated under conditions similar to that of the real-world image (e.g., weather conditions) and that of the randomly sampled criteria, but when inserted into the real-world image, it may not meet the additional distance threshold and duplication constraints. For example, after using the expression to expand the real-world image, the expression may be higher than the driving surface such that the expression appears to be in the air. Similarly, the inserted expression may appear to overlap or be directly above (or below) another object in the road when placed within the real-world image. To minimize the occurrence of these problems, the semantic augmenter 110 can check various constraints before using the augmented real-world image for training the machine learning model.

[0032] The semantic augmenter 110 can determine whether the representations used to augment the real-world image satisfy the distance threshold constraint. For example, the semantic augmenter 110 can verify that the representation satisfies a vertical distance threshold constraint (e.g., the vertical distance from the driving surface of the real-world image). For example, the semantic augmenter 110 can determine whether the inserted representation is 30.48 cm (1 foot), 60.96 cm (2 feet), 91.44 cm (3 feet), 152.4 cm (5 feet), etc. above the driving surface. Similarly, the distance threshold can be a horizontal distance threshold from the boundary line of the driving surface to determine whether the virtual instance is outside the horizontal boundary of the driving surface. For example, the semantic augmenter 110 can check whether the inserted representation is within 30.48 cm (1 foot), 60.96 cm (2 feet), 152.4 cm (5 feet), etc. outside the boundary line of the driving surface. In some embodiments, the distance can be measured in a three-dimensional coordinate system within the real-world image to determine whether the inserted representation satisfies the distance threshold constraint. In some embodiments, the semantic augmenter 110 measures the distance from the driving surface using the corresponding boundary shape of the representation to determine whether the representation satisfies the distance threshold. In some embodiments, representations that cannot satisfy the distance threshold constraint are rejected and not inserted into the real-world image.

[0033] Referring now to FIG. 3, FIG. 3 is an example of a real-world image (with ground-truth annotations) augmented with a representation of a simulated object. The representation (e.g., virtual instance) 206 is used to augment the real-world image 302 and has an associated boundary shape 304. In this example, the distance between the virtual instance 206 (or its boundary shape 304) and the driving surface 303 can be determined within the three-dimensional coordinate system of the real-world image 302 to determine whether the virtual instance 206 (or its boundary shape 304) is above the driving surface 303 by a vertical distance threshold, e.g., 30.48 cm (1 foot). When calculated, the virtual instance 206 is actually 91.44 cm (3 feet) above the driving surface 303 and cannot meet the distance threshold. Similarly, the distance from the virtual instance 206 (or its boundary shape 304) to the horizontal boundary of the driving surface 303 can be checked to see if the virtual instance 206 (or its boundary shape 304) is within, for example, 15.24 cm (6 inches) outside the boundary of the driving surface 303 and meets the horizontal distance threshold. The virtual instance 206 (and its boundary shape 304) is well placed within the boundary of the driving surface 303, and the virtual instance 206 (and its boundary shape 304) easily meets the horizontal distance threshold constraint but cannot meet the vertical distance threshold constraint. Since the visualization 300 does not meet both constraints, the visualization 300 may not be selected for training or testing a machine learning model to detect road debris.

[0034] Referring now to FIG. 4, FIG. 4 is an exemplary visualization 400 of a real-world image having an extended simulated object that satisfies threshold distance constraints and overlap constraints, according to some embodiments of the present disclosure. Visualization 400 also includes ground truth annotations such as boundary shapes around representations of other objects within the real-world image. The virtual instance 204 of FIG. 2A and its associated boundary shape 402 are included in visualization 400. The virtual instance 204 can be checked for distance threshold constraints as described with respect to FIG. 3. In this example, the virtual instance 204 (and its boundary shape 402) is well within a vertical distance threshold of 91.44 cm (3 feet) above the driving surface and a horizontal distance threshold of 15.24 cm (6 inches) outside the boundary line of the driving surface. Thus, the virtual instance 204 can be inserted into the real-world image 302 to generate a training image.

[0035] Referring now to FIG. 5, FIG. 5 is an exemplary visualization 500 of a real-world image having an extended simulated object that satisfies threshold distance constraints and overlap constraints, according to some embodiments of the present disclosure. Visualization 500 includes the virtual instance 208 of FIG. 2A disposed within the real-world image 302 as well as existing ground truth data such as boundary shapes around objects within the real-world image 302. Similar to FIG. 4, the virtual instance 208 and its boundary shape 502 are well within the vertical and horizontal distance thresholds and can be inserted into the real-world image 302 to generate a training image for training a machine learning model to detect road debris.

[0036] The semantic augmenter 110 can also determine whether the virtual instances inserted into the real-world images used for ground truth data satisfy the overlap constraints with other objects in the real-world images. For example, the real-world images used in the ground truth data may include objects on the driving surface that should not overlap with the inserted virtual instances, such as other passenger vehicles and pedestrians. The ground truth data may also include annotations such as the boundary shapes of each object in the real-world image. Using the ground truth annotations of the objects in the real-world image, the semantic augmenter 110 can determine whether to extend the real-world image using the virtual instances.

[0037] If the boundary shape of the virtual instance does not intersect or overlap with the boundary shapes of the objects in the real-world image and the existing ground truth data, the semantic augmenter 110 can determine that the overlap constraints are satisfied. Similarly, if the pixels included within the boundary shape of the virtual instance do not overlap with, or are not included in the same space as, the pixels included within the boundary shape of the objects in the ground truth data, the semantic augmenter 110 can determine that the overlap constraints are satisfied. For example, if a pedestrian is on a crosswalk, a virtual instance of a mattress can be checked to ensure that the virtual instance is not on top of or overlapping with the pedestrian by verifying whether the boundary shape corresponding to the pedestrian intersects the boundary shape of the virtual instance. Similarly, the pixels within each boundary shape can be checked against each other to determine overlap. In some embodiments, the semantic augmenter 110 can determine whether a virtual instance satisfies the overlap constraints based on whether the virtual instance is inserted in front of or behind other objects in the real-world image. In some embodiments, virtual instances that do not satisfy the overlap constraints are rejected and not inserted into the real-world image.

[0038] Referring now to FIG. 4, the virtual instance 204 of FIG. 2A is inserted into the visualization 400. Existing ground truth data includes the boundary shapes 306, 308, 310, 312, and 314 around the objects existing within the real-world image. For example, the boundary shapes 306, 308, 310, 312, and 314 surround the pixels of the passenger vehicle and the pedestrian. A boundary shape 402 encapsulating all the pixels represented by the virtual instance 204 is also shown. To determine whether the overlap constraint is satisfied for the virtual instance 204, the pixels within the boundary line of the boundary shape 402 are checked against the pixels within the boundary lines of the boundary shapes 306, 308, 310, 312, and 314. In this example, the pixels surrounded by the boundary shape 402 overlap or are in the same location as the pixels surrounded by the boundary shapes 310, 312, and 314. However, since the virtual instance 204 and its corresponding boundary shape 402 are in front of the bounding boxes 310, 312, and 314 (and the objects they enclose), the virtual instance 204 satisfies the overlap constraint. Since the virtual instance 204 satisfies the distance threshold constraint (as described herein) and the overlap constraint, the virtual instance 204 is selected and inserted into the real-world image 302 to train the machine learning model.

[0039] Referring now to FIG. 5, visualization 500 includes a virtual instance 208 of FIG. 2A inserted into the real-world image 302. As in FIGS. 3 and 4, FIG. 5 includes ground-truth annotations within the existing ground-truth, such as boundary shapes 306, 308, 310, 312, and 314 around objects existing within the real-world image. The virtual instance 208 and its corresponding boundary shape 502 can be checked to determine whether the virtual instance 208 meets the overlap constraints. The boundary shape 502 and the virtual instance 208 are checked for compliance with overlap constraints that do not require overlap with other boundary shapes within the visualization 500. In this example, the pixels enclosed within the boundary shape 502 do not intersect or overlap with any of the pixels enclosed within the other boundary shapes 306, 308, 310, 312, and 314. As described herein, the virtual instance 208 meets the distance threshold constraint and the overlap constraint. Thus, the real-world image 302 extended with the visualization 500 and the virtual instance 208 can be used as a training image for training a machine learning model to detect road debris.

[0040] Returning to FIG. 1, the ground-truth data updater 112 can update the ground-truth data to reflect that the existing ground-truth data has been extended with virtual instances of objects. For example, after the virtual instances inserted into the real-world image meet the required constraints (e.g., distance threshold and overlap constraints), the ground-truth data updater 112 can update the existing ground-truth data as stored for evaluating the effectiveness of the machine learning model to include the inserted virtual instances and the associated boundary shapes around the virtual instances. The ground-truth data updater 112 can also add labels regarding the risk level and / or object type labels for the inserted virtual instances.

[0041] Referring to FIG. 5, visualization 500, including boundary shapes 502 around virtual instance 208 and boundary shapes 306, 308, 310, 312, and 314 of other objects within real-world image 302, can be used to update existing ground-truth data. For example, the existing ground-truth data can include real-world image 302 and boundary shapes 306, 308, 310, 312, and 314. Ground-truth data updater 112 can update the existing ground-truth data to include inserted virtual instance 208 and boundary shape 502. The ground-truth data can also be updated to include labels regarding the object, for example, that virtual instance 208 is a dead animal with a "high" level of danger. Thus, when a machine learning model is trained using training images of real-world image 302 augmented with virtual instance 204 and corresponding ground-truth data having boundary shapes 306, 308, 310, 312, 314, and 502, the labels exist for evaluating machine learning model effectiveness.

[0042] Returning to FIG. 1, dataset generator 114 can generate a new dataset and / or training images using real-world images augmented with virtual instances of objects. For example, after semantic augmenter 110 determines that the virtual instance meets any required constraints for insertion into the real-world image, dataset generator 114 can generate a dataset for training a machine learning model by modifying the real-world image to include the satisfied virtual instance. In some embodiments, dataset generator 114 can generate a dataset by removing boundary shapes from updated ground-truth data including real-world images augmented with virtual instances.

[0043] Referring to FIG. 6, training image 600 is an example of a training image that includes a real-world image augmented with a representation of a simulated object that satisfies a pre-set constraint according to some embodiments of the present disclosure. To generate training image 600, all ground-truth annotations in the visualization in FIG. 5, such as boundary shapes 306, 308, 310, 312, 314, or 502, have been removed. As shown in FIG. 6, training image 600 includes virtual instance 204 inserted into real-world image 302. Training image 600 can be used to train a machine learning model to detect road debris.

[0044] Returning to FIG. 1, trainer 116 can train a machine learning model using a training image that includes a real-world image augmented with virtual instances. For example, trainer 116 can receive a training image, such as training image 600, provide it to the machine learning model, and evaluate the output of the model (e.g., the boundary shape around the road debris, the degree of danger) by using the corresponding ground-truth data, such as visualization 500 in FIG. 5. Trainer 116 can then determine whether the output of the model for the boundary shape and the degree of danger matches the corresponding ground-truth data and provide the result to the machine learning model. In some embodiments, trainer 116 can implement zero-shot training by inputting an image having an object that was not present in the previous training image into the machine learning model. In some embodiments, even if the type of the object (as identified by the machine learning model) is incorrect (e.g., even if the machine learning model was not trained to specifically classify that particular class or type of object), trainer 116 can evaluate the result of the training using zero-shot training by determining whether the machine learning model identifies the object as road debris.

[0045] The deployment module 118 can deploy a trained machine learning model to provide an autonomous vehicle 1000 with a mechanism for detecting road debris. For example, the deployment module 118 can deploy a trained machine learning model in an autonomous vehicle 1000 driving on a driving surface to detect whether an object is on the driving surface. In some embodiments, the deployment module 118 can also deploy a machine learning model to detect an object type (e.g., cardboard box, dead animal, mattress), whether the object is dangerous, and the percentage confidence in the assessment that the object is dangerous.

[0046] The deployment module 118 can also deploy a model that will be used in determining the next action. For example, based on the output of the model, the vehicle 1000 (e.g., semi-autonomous or autonomous driving software stack) can determine whether to steer around the object, drive over the object, or come to a complete stop, depending on the object type, degree of danger, and / or percentage confidence. For example, the deployment module 118 can deploy a machine learning model and receive an output of a boundary shape and degree of danger around an object on the driving surface, e.g., a small rectangle around a flattened cardboard box in the road and a "low" degree of danger. In response, via the machine learning model, the deployment module 118 can determine whether the next step is to drive over the flattened cardboard box, stop around the cardboard box or move around it to avoid it. However, for an object with a larger boundary shape and / or a higher degree of danger, e.g., a large dead animal or fallen construction material that poses a danger to the driver, the deployment module 118 can instead determine that the response action is to bring the vehicle to a complete stop or move past the dangerous object.

[0047] Referring now to FIG. 7, FIG. 7 is a data flow diagram showing an exemplary process 700 for training a machine learning model to detect road debris, according to some embodiments of the present disclosure. Training image data 702 may include image data that meets a set of constraints, such as real-world images augmented with simulated objects (e.g., road debris, road signs). As described herein, FIG. 1 illustrates the process of augmenting a real-world image with augmented objects, and FIG. 6 is an example of a training image. Training image data 702 may also include images that include objects that the machine learning model has not been trained on. For example, the machine learning model may have been previously trained on images that include cardboard boxes and dead animals. However, training image data 702 may include images that have objects not present in previous training images, such as construction cones and mattresses for implementing zero-shot learning.

[0048] The machine learning model 704 can use one or more images or other data representations (e.g., LIDAR data, RADAR data, etc.) as represented by the training image data 702 as input to generate the output 706. In non-limiting examples, the machine learning model 704 can obtain one or more of the following: an image represented by the training image data 702 (e.g., after preprocessing) as input to generate the output 706, such as the boundary shape 708 and / or the risk level 710. Examples are described herein with respect to the use of a neural network, specifically a CNN as the machine learning model 704, but this is not intended to be limiting. For example, and without limitation, the machine learning model 704 described herein can be any type of machine learning model, such as linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forest, dimensionality reduction algorithm, gradient boosting algorithm, neural network (e.g., autoencoder, convolutional, recurrent, perceptron, long / short term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, adversarial generation, liquid state machine, etc.), and / or other types of machine learning models.

[0049] The output 706 of the machine learning model 704 may include a boundary shape 708, a risk level 710, and / or other output types. To decode the output of the machine learning model 704, in some non-limiting examples, GPU acceleration may be implemented. For example, a parallel processing platform (e.g., NVIDIA's CUDA) may be implemented to parallelize the algorithm through several computational kernels for decoding the output, thereby reducing the execution time. In some embodiments, the output 706 may include additional outputs such as an object type (e.g., a dead animal, a construction material) and a percentage confidence associated with each output.

[0050] The boundary shape 708 can include the shape of each object in the training image data 702, and the shape encloses all the pixels representing each object. The boundary shape 708 can be any shape (e.g., a square, a rectangle, a triangle, a polygon) that captures and encloses the pixels in a given object. For example, as described herein, FIG. 5 is an exemplary visualization of boundary shapes 306, 308, 310, 312, 314, and 502 that enclose the pixels of a given object. In some embodiments, the boundary shape 708 can be loosely or tightly fitted around the object to suppress the number of pixels outside the object captured within the boundary shape 708. As such, the machine learning model 704 can calculate the output of the boundary shape 708 for each object within the training image data 702. A percentage confidence value may also be attached to each boundary shape 708.

[0051] The risk level 710 can include a rating of how harmful each object in the training image data 702 is. The risk level 710 can be a numerical value (e.g., on a scale of 1 to 10), a risk label (e.g., whether it is dangerous or not), and / or a relative value of the harmful threat posed by a particular object on the road (e.g., low, medium, high). For example, a large, dead animal, such as a moose, can receive a risk level 710 indicating a higher level of danger than a flattened cardboard box. The risk level 710 can also be attached with an object type - label that can be used to determine the level of harmful threat posed by the object. In some embodiments, when zero - shot learning is implemented, the risk level 710 can be calculated regardless of the type of the detected object. For example, assigning a high risk level 710 to a large rock can be more valuable for the machine - learning model 704 to determine than a machine - learning model that identifies the object as a rock. Each risk level 710 can be attached with a percentage confidence value.

[0052] The model trainer 720 may be a trainer of the machine learning model 704 that evaluates the output 706 against the ground truth data 714, determines the difference using one or more loss functions, and updates the parameters (e.g., weights and biases) of the machine learning model 704 based on the results. The ground truth data 714 may include annotation data, e.g., boundary shapes 716 and risk levels 718 corresponding to real-world objects and / or simulated objects. For example, the ground truth data 714 may include annotations, e.g., rectangles around each object, labels describing the object type, and labels indicating whether each object is harmful. As described herein, FIG. 5 is an example visualization of ground truth data corresponding to FIG. 6 that includes boundary shapes around each object on a road. Based on the ground truth data 714 and the output 706, the model trainer 720 can adjust the learning algorithm and the training image data 702 to train the machine learning model 704 to formulate boundary shapes around each object on the road to closely match the ground truth data 714. Similarly, using the ground truth data 714, the model trainer 720 can also adjust the learning algorithm of the machine learning model 704 to more accurately label objects with object types and risk levels 710.

[0053] Referring now to FIGS. 8-9, each block of methods 800 and 900 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. Methods 800 and 900 can also be implemented as computer-usable instructions stored on a computer storage medium. Methods 800 and 900 can be provided, for example, as a stand-alone application, service, or hosted service (either stand-alone or in combination with another hosted service), or as a plug-in to another product. Additionally, methods 800 and 900 are described, by way of example, with respect to process 100 of FIG. 1. However, these methods can be executed additionally or alternatively by any one system, or any combination of systems, including but not limited to those described herein.

[0054] FIG. 8 is a flow diagram illustrating a method 800 for training a machine learning model to detect an object in an image using a real-world image augmented with simulated objects, according to some embodiments of the present disclosure. Method 800 includes, at block B802, generating a first image of a virtual instance of an object from the perspective of a virtual sensor of a virtual vehicle within a virtual environment. For example, virtual instances 204, 206, and 208 of a dead animal, such as shown in FIG. 2A, can be generated via a game-based engine from the perspective of a virtual sensor mounted on a virtual vehicle in a simulated environment.

[0055] Method 800 includes, at block B804, determining a boundary shape corresponding to the virtual instance of the object. For example, a boundary shape, such as boundary shape 304, corresponding to a virtual instance, such as virtual instance 206, can be created using a segmentation mask 214 associated with virtual instance 206. In some embodiments, the boundary shape can be tightly or loosely fitted around the virtual instance.

[0056] Method 800 includes, at block B806, inserting a virtual instance of an object into a second image, where the second image is generated using real-world sensors of a real-world vehicle. For example, virtual instance 206 may be inserted into real-world image 302 captured using real-world sensors mounted on a real-world vehicle.

[0057] Method 800 includes, at block B808, determining that a virtual instance of an object is within a threshold distance to a driving surface shown in the second image. For example, inserted virtual instance 206 within visualization 300 is checked to determine whether it meets a vertical distance threshold from the driving surface. Since virtual instance 206 does not meet the vertical distance threshold, visualization 300 will not be used to train a machine learning model. However, as shown in FIG. 4 and described herein, inserted virtual instance 204 meets both a vertical distance threshold constraint and a horizontal distance threshold constraint. Inserted virtual instance 208 of FIG. 5 similarly meets both of these distance threshold constraints as described herein.

[0058] Method 800 includes, at block B810, determining that a boundary shape satisfies an overlap constraint. For example, the boundary shape 402 around the inserted virtual instance 204 is evaluated to determine whether its boundary shape satisfies an overlap constraint with the boundary shapes around other objects on the driving surface. As shown in FIG. 4 and described herein, the boundary shape 402 overlaps with other boundary shapes 310, 312, and 314 and satisfies the overlap constraint condition for inserting the virtual instance 204 into the real-world image 302 to generate a training image. The boundary shape 502 around the virtual instance 208 in FIG. 5 does not overlap with any of the boundary shapes 306, 308, 310, 312, and 314 and thus similarly satisfies the overlap constraint as described herein. Since the boundary shape 502 and the virtual instance 208 within the visualization 500 satisfy both the distance threshold and the overlap constraint, the real-world image 302 extended with the virtual instance 208 can be used to train a machine learning model.

[0059] Method 800 includes, at block B812, training a neural network using a training image and a boundary shape as ground truth data. For example, training image data 702 and ground truth data 714 having boundary shapes 716 and risk levels 718 can be used to train a machine learning model 704, and in some examples, the machine learning model 704 can include a neural network (e.g., a CNN). For example, FIG. 6, which includes the real-world image 302 and the virtual instance 208 that satisfies all the required constraints, can be used as the training image data 702. Similarly, the visualization 500 can be used as the corresponding ground truth data 714 for evaluating the machine learning model.

[0060] Referring to FIG. 9, FIG. 9 is a flowchart showing a method 900 for training a machine learning model to detect an object in a real-world image using a simulated object generated based on a set of criteria according to some embodiments of the present disclosure. Method 900 includes, at block B902, randomly sampling a plurality of criteria to determine a set of criteria. For example, criteria such as object type, location, lighting, orientation, pose, environmental conditions, and / or appearance attributes can be randomly selected. The value of each of these criteria can then be randomly selected as well. For example, a set of criteria such as weather and object type can be randomly selected, and their values, such as fog and cardboard box, can be randomly selected respectively.

[0061] Method 900 includes, at block B904, selecting a virtual instance of an object from a plurality of virtual object instances. For example, a set of virtual instances of an object can be generated by a game-based engine, where each virtual instance has incrementally the values of a given set of criteria. For example, FIG. 2A includes virtual instances 204, 206, and 208 that are of the same object type but have different positions within the virtual environment respectively. From the plurality of virtual instances, those that match (or closely match) the criteria sampled randomly can be selected for insertion into the real-world image. In some embodiments, the virtual instances can be selected, in part, based on the conditions present within the real-world image.

[0062] Method 900 includes, at block B906, generating a boundary shape corresponding to a virtual instance of an object. For example, a virtual instance 206 that meets randomly sampled criteria may be selected for use in extending the real-world image 302. A boundary shape 304 corresponding to the virtual instance 206 may be generated based on the segmentation mask 214. The boundary shape 304 is generated to capture all pixels representing the virtual instance 206 and may be tightly or loosely fitted around the virtual instance 206. Similarly, virtual instances 204 and 208 each have boundary shapes 402 and 502, respectively.

[0063] Method 900 includes, at block B908, inserting a virtual instance of an object into the image. For example, virtual instances 204, 206, and 208 are inserted into the real-world image 302 as shown in FIGS. 3-5 and described herein.

[0064] Method 900 includes, at block B910, determining that a virtual instance of an object is within a threshold distance to a driving surface depicted in a second image. For example, the inserted virtual instances 204, 206, and 208 are evaluated to determine whether each virtual instance meets vertical and horizontal distance threshold constraints. As shown in FIG. 3 and described herein, the virtual instance 206 (and its boundary shape 304) does not meet the vertical distance threshold. As shown in FIGS. 4 and 5 and described herein, the virtual instances 204 and 208 (and their boundary shapes 402 and 502) meet both the vertical distance threshold constraint and the horizontal distance threshold constraint.

[0065] Method 900 includes, at block B912, determining that the boundary shape satisfies the overlap constraint. For example, the boundary shape of each virtual instance is evaluated to determine whether a given boundary shape satisfies the overlap constraint. As shown in FIG. 5 and described herein, since the boundary shape 502 does not overlap with the boundary shapes 306, 308, 310, 312, and 314 of other objects in the real-world image 302, the virtual instance 208 satisfies the overlap constraint. However, the virtual instance 208 and the boundary shape 502 are merely examples that satisfy both the distance threshold constraint and the overlap constraint.

[0066] Method 900 includes, at block B914, training a neural network using a training image and a boundary shape as ground truth data. For example, the visualization 500 including the boundary shape 502 and the training image 600 including the real-world image 302 extended with the virtual instance 208 are used to train a neural network. For example, the training image data 702 and the ground truth data 714 having the boundary shape 716 and the risk level 718 can be used to train the machine learning model 704, where the machine learning model 704 can include a neural network (e.g., a CNN) in some examples.

[0067] Exemplary autonomous vehicle Figure 10A is a diagram of an exemplary autonomous vehicle 1000 according to some embodiments of the present disclosure. The autonomous vehicle 1000 (or referred to herein as "vehicle 1000") can include, but is not limited to, passenger vehicles such as cars, trucks, buses, first responder vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police vehicles, ambulances, boats, construction vehicles, submarines, drones, and / or other types of vehicles (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described in terms of automation levels as defined by the National Highway Traffic Safety Administration (NHTSA), a department of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). The moving vehicle 1000 can have the ability to function according to one or more of automation levels 3 to 5 of the autonomous driving level. For example, the moving vehicle 1000 can have the ability of conditional automation (level 3), highly automated (level 4), and / or fully automated (level 5) according to the embodiments.

[0068] The moving vehicle 1000 can include components such as the chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the moving vehicle. The moving vehicle 1000 can include a propulsion system 1050 such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. The propulsion system 1050 can be connected to the drive train of the moving vehicle 1000 and can include a transmission to enable the propulsion force of the moving vehicle 1000. The propulsion system 1050 can be controlled in response to receiving a signal from the throttle / acceleration device 1052.

[0069] The steering system 1054, which may include a steering wheel, can be used to steer the moving vehicle 1000 (e.g., along a desired path or route) when the propulsion system 1050 is operating (e.g., when the moving vehicle is in motion). The steering system 1054 can receive a signal from the steering actuator 1056. The steering wheel may be an option for a fully automated (Level 5) function.

[0070] The brake sensor system 1046 can be used to operate the vehicle brakes in response to receiving a signal from the brake actuator 1048 and / or a brake sensor.

[0071] The controller 1036, which may include one or more system-on-chips (SoCs) 1004 (FIG. 10C) and / or GPUs, can provide signals (e.g., representations of commands) to one or more components and / or systems of the mobile vehicle 1000. For example, the controller can send signals to operate the mobile vehicle brakes via one or more brake actuators 1048, to operate the steering system 1054 via one or more steering actuators 1056, and to operate the propulsion system 1050 via one or more throttle / acceleration devices 1052. The controller 1036 can include one or more mounted (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or to assist the driver in operating the mobile vehicle 1000. The controller 1036 can include a first controller 1036 for autonomous driving functions, a second controller 1036 for functional safety functions, a third controller 1036 for artificial intelligence functions (e.g., computer vision), a fourth controller 1036 for infotainment functions, a fifth controller 1036 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1036 can process two or more of the aforementioned functions, and two or more controllers 1036 can process a single function and / or any combination thereof. In some embodiments, the controller 1036 can send a signal to operate the vehicle brakes using the brake actuator 1048 in response to a detected object on a high-risk and unavoidable driving surface.

[0072] Controller 1036 can provide signals for controlling one or more components and / or systems of the moving vehicle 1000 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received, for example and without limitation, from a global navigation satellite system sensor 1058 (e.g., a global positioning system sensor), a RADAR sensor 1060, an ultrasonic sensor 1062, a LIDAR sensor 1064, an inertial measurement unit (IMU) sensor 1066 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 1096, a stereo camera 1068, a wide-view camera 1070 (e.g., a fish-eye camera), an infrared camera 1072, a surround camera 1074 (e.g., a 360-degree camera), a long-range and / or mid-range camera 1098, a speed sensor 1044 (e.g., for measuring the speed of the moving vehicle 1000), a vibration sensor 1042, a steering sensor 1040, a brake sensor (e.g., as part of a brake sensor system 1046), and / or other sensor types.

[0073] One or more of the controllers 1036 of the vehicle 1000 receives an input (e.g., represented by input data) from the instrument cluster 1032 of the vehicle 1000 and provides an output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1034, an audible annunciator, a loudspeaker, and / or other components of the vehicle 1000. The output can include information such as vehicle velocity, speed, time, map data (e.g., the HD map 1022 of FIG. 10C), position data (e.g., the position of the vehicle 1000 on a map, etc.), direction, the position of other vehicles (e.g., occupancy grid), information regarding objects and the situation of objects as perceived by the controller 1036. For example, the HMI display 1034 can display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at exit 34B within 3.22 km (2 miles), etc.). The controller 1036 can also receive an output from a machine learning model regarding an action to be performed in response to a detected object on the driving surface (e.g., driving over it or maneuvering around it or coming to a complete stop).

[0074] The mobile vehicle 1000 further includes a network interface 1024 that can communicate via one or more networks using one or more wireless antennas 1026 and / or modems. For example, the network interface 1024 may have the ability to communicate via LTE, WCDMA®, UMTS, GSM, CDMA2000, etc. The wireless antenna 1026 can also use local area networks such as Bluetooth®, Bluetooth LE, Z-Wave, ZigBee, and / or low power wide-area networks (LPWAN) such as LoRaWAN, SigFox to enable communication between objects in the environment (e.g., mobile vehicles, mobile devices, etc.).

[0075] FIG. 10B is an example of the camera positions and fields of view of the exemplary autonomous vehicle 1000 of FIG. 10A according to some embodiments of the present disclosure. The cameras and their respective fields of view are one exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the mobile vehicle 1000.

[0076] The camera type of the camera may include, but is not limited to, a digital camera adapted to be used with components and / or systems of the moving vehicle 1000. The camera can operate at automotive safety integrity level (ASIL) B and / or at another ASIL. Depending on the embodiment, the camera type may have the ability to capture images at any rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc. The camera may have the ability to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, a clear pixel camera, such as a camera having an RCCC, RCCB, and / or RBGC color filter array, may be used in efforts to increase light sensitivity.

[0077] In some examples, one or more of the cameras may be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional mono-camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all of the cameras) may be able to record and provide image data (e.g., video) simultaneously.

[0078] One or more of the cameras may be mounted on mounting components such as custom-designed (3D printed) components to remove stray light and reflections from inside the vehicle that may interfere with the camera's image data capture ability (e.g., reflections from the dashboard reflected in the front windshield mirror). Referring to the side mirror mounting component, the side mirror component may be custom 3D printed such that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera may be integrated within the side mirror. For side view cameras, the camera may also be integrated within four struts at each corner of the cabin.

[0079] A camera having a field of view that includes a portion of the environment in front of the moving vehicle 1000 (e.g., a forward-facing camera) can be used for surround view to assist in identifying the forward path and obstacles and, with the assistance of one or more controllers 1036 and / or a control SoC, in providing information essential for generating an occupancy grid and / or determining a preferred moving vehicle path. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems including other functions such as lane departure warning ("LDW (Lane Departure Warning)"), autonomous cruise control ("ACC (Autonomous Cruise Control)"), and / or traffic sign recognition.

[0080] Various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 1070 that can be used to capture objects entering the view from the surroundings (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in FIG. 10B, any number of wide-view cameras 1070 may be present on the moving vehicle 1000. Additionally, a long-range camera 1098 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. The long-range camera 1098 can also be used for object detection and classification, as well as for basic object tracking.

[0081] One or more stereo cameras 1068 can also be included in the forward-facing configuration. The stereo camera 1068 can include an integrated control unit with an expandable processing unit that can provide a programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet® interface integrated on a single chip. Such a unit can be used to generate a 3D map of the environment of the moving vehicle, including distance estimates for all points in the image. An alternative stereo camera 1068 can include a compact stereo vision sensor that includes two camera lenses (one each on the left and right) and an image processing chip that can measure the distance from the moving vehicle to the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1068 may be used in addition to, or instead of, those described herein.

[0082] A camera (e.g., a side view camera) having a field of view that includes a portion of the environment relative to the side of the moving vehicle 1000 can be used for surround view to provide information for creating and updating an occupancy grid and for generating a side impact collision warning. For example, surround cameras 1074 (e.g., four surround cameras 1074 as shown in FIG. 10B) can be positioned on the moving vehicle 1000. The surround cameras 1074 can include wide view cameras 1070, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be arranged in front of, behind, and on the sides of the moving vehicle. In an alternative arrangement, the moving vehicle may use three surround cameras 1074 (e.g., left, right, and rear), and one or more other cameras (e.g., a forward-facing camera) can be utilized as a fourth surround view camera.

[0083] A camera (e.g., a rear view camera) having a field of view that includes a portion of the environment relative to the rear of the moving vehicle 1000 can be used for parking assistance, surround view, rear collision warning, and for creating and updating an occupancy grid. A wide variety of cameras can be used, including but not limited to cameras suitable as forward-facing cameras (e.g., long range and / or mid-range cameras 1098, stereo cameras 1068), infrared cameras 1072, etc., as described herein.

[0084] Figure 10C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 1000 of FIG. 10A, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by an entity may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.

[0085] Each of the components, features, and systems of the moving vehicle 1000 of FIG. 10C is illustrated as being connected via a bus 1002. The bus 1002 may include a Controller Area Network (CAN) data interface (or referred to as a "CAN bus"). CAN may be a network within the moving vehicle 1000 used to assist in controlling various features and functions of the moving vehicle 1000, such as the operation of brakes, acceleration, brakes, steering, front windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering angle, ground speed, engine revolutions per minute (RPM), button position, and / or other moving vehicle status indicators. The CAN bus may be ASIL B compliant.

[0086] Bus 1002 is described herein as being a CAN bus, but this is not intended to be limiting. For example, in addition to, or as an alternative to, a CAN bus, FlexRay and / or Ethernet® may be used. Additionally, a single line is used to represent bus 1002, but this is not intended to be limiting. For example, any number of buses 1002 may exist that include one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 1002 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1002 may be used for a collision avoidance function and a second bus 1002 may be used for actuation control. In any example, each bus 1002 may communicate with any of the components of the moving vehicle 1000, and two or more buses 1002 may communicate with the same component. In some examples, each SoC 1004, each controller 1036, and / or each computer within the moving vehicle may have access to the same input data (e.g., input from the sensors of the moving vehicle 1000) and may be connected to a common bus such as a CAN bus.

[0087] The moving vehicle 1000 may include one or more controllers 1036, such as those described herein with respect to FIG. 10A. The controller 1036 may be used for various functions. The controller 1036 may be coupled to any of the various other components and systems of the moving vehicle 1000 and may be used for the control of the moving vehicle 1000, the artificial intelligence of the moving vehicle 1000, the infotainment for the moving vehicle 1000, and / or the like. As described herein, the controller 1036 may be able to determine an adjustment to the position of the vehicle 1000 to avoid or overcome an object detected on the driving surface.

[0088] The mobile vehicle 1000 may include a system-on-chip (SoC) 1004. The SoC 1004 may include a CPU 1006, a GPU 1008, a processor 1010, a cache 1012, an acceleration device 1014, a data store 1016, and / or other components and features not shown. The SoC 1004 may be used to control the mobile vehicle 1000 within various platforms and systems. For example, the SoC 1004 may be coupled in a system (such as the system of the mobile vehicle 1000) having an HD map 1022 that can obtain map refreshes and / or updates via a network interface 1024 from one or more servers (such as the server 1078 in FIG. 10D).

[0089] The CPU 1006 may include a CPU cluster or CPU complex (or also referred to as a "CCPLEX"). The CPU 1006 may include multiple cores and / or an L2 cache. For example, in some embodiments, the CPU 1006 may include 8 cores within a coherent multiprocessor configuration. In some embodiments, the CPU 1006 may include 4 dual-core clusters, each cluster having a dedicated L2 cache (such as a 2MB L2 cache). The CPU 1006 (such as a CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of clusters of the CPU 1006 to become active at any given time.

[0090] The CPU 1006 can implement a power management capability that includes one or more of the following features: individual hardware blocks can be automatically clock-gated when in an idle state to save dynamic power, each core clock can be gated when the core is not actively executing instructions by the execution of WFI / WFE instructions, each core can be independently power-gated, each core cluster can be independently clock-gated when all cores are clock-gated or power-gated, and / or each core cluster can be independently power-gated when all cores are power-gated. The CPU 1006 can further implement an enhanced algorithm for managing power states, where the allowed power states and the expected wake-up times are specified and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing cores can support a simplified power state input sequence in software where the work is offloaded to microcode.

[0091] The GPU 1008 may include an integrated GPU (or referred to herein as "iGPU"). The GPU 1008 can be programmable and can be efficient for parallel workloads. In some examples, the GPU 1008 can use an enhanced tensor instruction set. The GPU 1008 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having at least 96 KB of storage capacity), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, the GPU 1008 may include at least eight streaming microprocessors. The GPU 1008 can use a compute application programming interface (API). Additionally, the GPU 1008 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0092] The GPU 1008 can be power-optimized for the best performance in automotive and embedded use cases. For example, the GPU 1008 can be manufactured on FinFET (Fin field-effect transistor). However, this is not intended to be limiting, and the GPU 1008 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores partitioned into multiple blocks. By way of non-limiting example, for instance, 64 PF32 cores and 32 PF64 cores may be partitioned into 4 processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths for efficient execution of workloads having a mix of compute and addressing operations. The streaming microprocessor can include independent thread scheduling capabilities to enable higher fine-grained synchronization and cooperation among parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to simplify programming while improving performance.

[0093] In some examples, GPU 1008 may include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide for a peak memory bandwidth of 900 GB / second. In some examples, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.

[0094] GPU 1008 can include unified memory technology that includes access counters to enable more accurate movement of those memory pages to the processor that most frequently accesses the memory pages, thereby improving the efficiency of the storage ranges shared between processors. In some examples, address translation service (ATS) support may be used to enable the GPU 1008 to directly access the CPU 1006 page table. In such examples, when the GPU 1008 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 1006. In response, the CPU 1006 can examine its page table for the virtual-to-real mapping of the address and send the translation back to the GPU 1008. As such, unified memory technology can enable a single unified virtual address space for the memory of both the CPU 1006 and the GPU 1008, thereby simplifying the GPU 1008 programming and porting of applications to the GPU 1008.

[0095] In addition, GPU 1008 may include an access counter that can record the frequency of access of GPU 1008 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that accesses that page most frequently.

[0096] SoC 1004 may include any number of caches 1012, including those described herein. For example, cache 1012 may include an L3 cache that is available to both CPU 1006 and GPU 1008 (e.g., connected to both CPU 1006 and GPU 1008). Cache 1012 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include more than 4 MB, although smaller cache sizes may be used, depending on the embodiment.

[0097] SoC 1004 may include an arithmetic logic unit (ALU) that can be utilized when executing processing for any of the various tasks or operations of vehicle 1000 (e.g., processing DNN). In addition, SoC 1004 may include a floating point unit (FPU) (or other mass coprocessor or numeric coprocessor type) for performing mathematical operations within the system. For example, SoC 1004 may include one or more FPUs integrated as execution units within CPU 1006 and / or GPU 1008.

[0098] SoC 1004 may include one or more acceleration devices 1014 (e.g., hardware acceleration devices, software acceleration devices, or combinations thereof). For example, SoC 1004 may include a hardware acceleration cluster that may include optimized hardware acceleration devices and / or large on-chip memories. A large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 1008 and to offload some of the tasks of the GPU 1008 (e.g., to free up more cycles of the GPU 1008 for other tasks). As an example, the acceleration device 1014 may be used for target workloads that are stable enough to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs)). In this specification, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).

[0099] Acceleration device 1014 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) configured to provide an additional 10 trillion operations per second for deep learning applications and inferences. The TPU may be an acceleration device configured and optimized to execute image processing functions (e.g., CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating-point operations, as well as for inferences. The design of the DLA can provide more performance per millimeter than a general-purpose GPU and greatly exceed the performance of a CPU. The TPU can execute several functions, including, for example, a single-instance convolution function and a post-processor function that support INT8, INT16, and FP16 data types for both features and weights.

[0100] The DLA can rapidly and efficiently execute neural networks, particularly CNNs, with processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection and identification and detection using data from a microphone, CNNs for face recognition and moving vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety-related events. The DLA can also execute a neural network to determine the risk of detected pieces of road debris and the confidence percentage of all outputs of the neural network.

[0101] The DLA can execute any function of the GPU 1008, and by using the inference acceleration device, for example, a designer can target either the DLA or the GPU 1008 for any function. For example, the designer can focus on processing CNN and floating-point operations on the DLA and leave other functions to the GPU 1008 and / or other acceleration devices 1014.

[0102] The acceleration device 1014 (for example, a hardware acceleration cluster) may include, or be referred to herein as, a programmable vision accelerator (PVA). The PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0103] The RISC core can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC core can use any of several protocols, depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0104] DMA may enable components of the PVA to access system memory independently of the CPU1006. DMA can support any number of features used to provide optimization for the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing up to six dimensions or more, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0105] The vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem can operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.

[0106] Each vector processor may include an instruction cache and may be connected to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in some embodiments, multiple vector processors included in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA can execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms sequentially on an image or a portion of an image. In particular, any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system security.

[0107] The acceleration device 1014 (e.g., a hardware acceleration cluster) may include a computer vision network on chip and SRAM for providing high-bandwidth, low-latency SRAM for the acceleration device 1014. In some examples, the on-chip memory may include at least 4 MB of SRAM consisting of, for example and without limitation, 8 field-configurable memory blocks that may be accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA can access the memory via a backbone that provides high-speed access to the PVA and the DLA to the memory. The backbone may include a computer vision network on chip that interconnects the PVA and the DLA to the memory (e.g., using the APB).

[0108] The computer vision network on chip may include an interface that determines that both the PVA and the DLA are operable and providing valid signals before any control signal / address / data transmission. Such an interface can provide separate phases and separate channels for transmitting control signal / address / data, as well as burst-type communication for continuous data transfer. This type of interface may comply with the ISO26262 or IEC61508 standard, although other standards and protocols may be used.

[0109] In some examples, SoC 1004 may include a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and scale of objects (e.g., within a world model) for generating real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for SONAR system simulation, for general wave propagation simulation, for comparison against LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.

[0110] Accelerator 1014 (e.g., a hardware accelerator cluster) has various uses for autonomous driving. The PVA may be a programmable vision accelerator that can be used in extremely important processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are suitable for algorithms that require predictable processing at low power and low latency. In other words, the PVA functions well with small data sets, as well as with semi-dense or dense normal calculations, with low latency, low power, and a predictable execution time. Therefore, since the PVA is efficient in object detection and integer calculations, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms.

[0111] For example, according to one embodiment of the present technology, PVA is used to perform computer stereo vision. A semi-global matching-based algorithm may be used in some examples, but this is not intended to be limiting. A number of applications for level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). PVA can perform computer stereo vision functions with inputs from two monocular cameras.

[0112] In some examples, PVA may be used to perform high-density optical flow. By processing raw RADAR data (e.g., using 4D fast Fourier transform) to provide processed RADAR. In other examples, PVA is used for flight depth processing time, for example, by processing the raw time of flight data to provide the processed time of flight data.

[0113] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a measure of the reliability of each object detection. Such reliability values can be interpreted as probabilities or as providing the relative "weight" of each detection compared to other detections. This reliability value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a reliability threshold and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the moving vehicle to automatically execute emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network that regresses the reliability value. The neural network can receive as its input at least some subset of parameters such as bounding box dimensions, ground plane estimation obtained (for example, from another subsystem), the azimuth, distance, and 3D position estimation of the object obtained from the neural network and / or other sensors (for example, LIDAR sensor 1064 or RADAR sensor 1060), and the output of an inertial measurement unit (IMU) sensor 1066 that correlates with the above, among others.

[0114] SoC 1004 may include a data store 1016 (e.g., a memory). The data store 1016 may be on-chip memory of the SoC 1004 and can store neural networks to be executed on the GPU and / or DLA. In some examples, the data store 1016 may have a capacity large enough to store multiple instances of neural networks for redundancy and security. The data store 1012 may include an L2 or L3 cache 1012. References to the data store 1016 may include references to memory related to the PVA, DLA, and / or other accelerators 1014 as described herein.

[0115] SoC 1004 may include one or more processors 1010 (e.g., embedded processors). The processor 1010 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC 1004 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low-power state transitions, management of the SoC 1004 heat and temperature sensors, and / or management of the SoC 1004 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1004 can use the ring oscillator to detect the temperature of the CPU 1006, GPU 1008, and / or accelerator 1014. If the temperature is determined to exceed a threshold, the boot and power management processor can enter a temperature fault routine and place the SoC 1004 in a lower power state and / or put the vehicle 1000 in a chauffeur safe stop mode (e.g., safely stop the vehicle 1000).

[0116] Processor 1010 may further include a set of embedded processors capable of performing the functions of an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.

[0117] Processor 1010 may further include an always-on processor engine that provides the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (such as timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0118] Processor 1010 may further include a safety cluster engine that includes a dedicated processor subsystem for processing the safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (such as timers, interrupt controllers, etc.), and / or routing logic. In the safety mode, the two or more cores can operate in a lockstep mode and function as a single core having comparison logic for detecting any differences between their operations.

[0119] Processor 1010 may further include a real-time camera engine that may include a dedicated processor subsystem for processing real-time camera management.

[0120] Processor 1010 may further include a high dynamic range signal processor that includes an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0121] The processor 1010 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the post-video processing functions required by the video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction with the wide-view camera 1070, with the surround camera 1074, and / or with the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of the advanced SoC configured to identify in-cabin events and respond appropriately. The in-cabin system can perform lip reading to activate cellular service and make calls, compose emails, change the destination of the moving vehicle, activate or change the infotainment system and settings of the moving vehicle, or provide voice-activated web surfing. Certain functions are available to the driver only when operating in autonomous mode and are otherwise disabled.

[0122] The video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs within the video, the noise reduction reduces the weight of the information provided by adjacent frames and appropriately weights the spatial information. When an image or a portion of an image does not contain motion, the temporal noise reduction performed by the video image synthesizer can reduce the noise in the current image using information from the previous image.

[0123] The video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. The video image synthesizer can further be used for user interface synthesis when the operating system desktop is in use, and the GPU 1008 is not required to continuously render new surfaces. Even when the power of the GPU 1008 is turned on and 3D rendering is actively performed, the video image synthesizer can be used to offload the GPU 1008 to improve performance and responsiveness.

[0124] SoC 1004 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for the camera and related pixel input functions to receive video and inputs from the camera. SoC 1004 may further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.

[0125] SoC1004 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC1004 can be used to process data from cameras (connected via, for example, Gigabit Multimedia Serial Link and Ethernet®), sensors (such as LIDAR sensor 1064, RADAR sensor 1060, etc. that can be connected via Ethernet®), data from bus 1002 (such as the speed, steering wheel position, etc. of moving vehicle 1000), and data from GNSS sensor 1058 (connected via, for example, Ethernet® or CAN bus). SoC1004 may include its own DMA engine and may further include a dedicated high-performance large-capacity storage controller that can be used to free the CPU1006 from routine data management tasks.

[0126] SoC1004 may be an end-to-end platform with a flexible architecture that extends to automation levels 3 - 5, thereby leveraging and efficiently using computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack together with deep learning tools, providing an integrated functional safety architecture. SoC1004 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when accelerator 1014 is coupled with CPU1006, GPU1008, and data store 1016 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0127] Therefore, the present technology provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be configured to be executed on a CPU using a high-level programming language such as the C programming language to execute a variety of processing algorithms over a variety of visual data. However, a CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time composite object detection algorithms, which are the requirements of in-vehicle ADAS applications and the actual level 3-5 autonomous vehicle requirements.

[0128] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein enables multiple neural networks to be executed simultaneously and / or sequentially and enables the results to be combined to enable level 3-5 autonomous driving functionality. For example, a CNN executed on a DLA or a dGPU (e.g., GPU 1020) can include text and word recognition that enables a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained on. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module executed on the CPU complex.

[0129] As another example, multiple neural networks can be executed simultaneously as required for level 3, 4, or 5 operation. For example, along with the electro-optical warning sign consisting of "Caution: Flashing light indicates a frozen state" can be interpreted independently or collectively by several neural networks. The sign itself can be identified as a traffic sign by the first deployed neural network (e.g., a trained neural network), and the text "Flashing light indicates a frozen state" can be interpreted by a second deployed neural network that informs the route planning software of the moving vehicle (preferably running on a CPU complex) that a frozen state exists when a flashing light is detected. The flashing light can be identified by informing the route planning software of the moving vehicle of the presence (or absence) of the flashing light and operating a third deployed neural network through multiple frames. All three neural networks can be executed simultaneously, for example, within the DLA and / or on the GPU 1008.

[0130] In some examples, a CNN for face recognition and moving vehicle owner identification can use data from a camera sensor to identify the presence of a regular driver and / or owner of the moving vehicle 1000. A always-on sensor processing engine can be used to unlock and illuminate the moving vehicle when the owner approaches the driver's side door, and to stop the operation of the moving vehicle when the owner leaves the moving vehicle in security mode. In this way, the SoC 1004 provides security against theft and / or carjacking.

[0131] In another example, the CNN for emergency vehicle detection and identification can detect and identify an emergency vehicle siren using data from microphone 1096. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 1004 uses a CNN for environmental and urban sound classification, as well as for visual data classification. In a preferred embodiment, the CNN executed on the DLA is trained to identify the relative end speed of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the moving vehicle is operating, as identified by GNSS sensor 1058. Thus, for example, when operating in Europe, the CNN will attempt to detect European sirens, and when in the United States, the CNN will attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine that, with the assistance of ultrasonic sensor 1062, decelerates the moving vehicle, stops it at the side of the road, parks the moving vehicle, and / or idles the moving vehicle until the emergency vehicle has passed.

[0132] The moving vehicle can include a CPU 1018 (e.g., a discrete CPU or dCPU) coupled to the SoC 1004 via a high-speed interconnect (e.g., PCIe). The CPU 1018 can include, for example, an X86 processor. The CPU 1018 can be used to perform any of a variety of functions, including, for example, mediating potential inconsistencies between the ADAS sensors and the SoC 1004, and / or monitoring the status and health of the controller 1036 and / or the infotainment SoC 1030.

[0133] The mobile vehicle 1000 may include a GPU 1020 (e.g., an individual GPU or a dGPU) that can be connected to the SoC 1004 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1020 can provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the mobile vehicle 1000.

[0134] The mobile vehicle 1000 may further include a network interface 1024 that may include one or more wireless antennas 1026 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 1024 can be used to enable wireless connections with the cloud via the Internet (e.g., with server 1078 and / or other network devices), with other mobile vehicles, and / or with computing devices (e.g., a passenger's client device). To communicate with other mobile vehicles, a direct link can be established between two mobile vehicles and / or an indirect link can be established (e.g., through a network and via the Internet). A direct link can use and provide a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1000 information regarding mobile vehicles in proximity to the mobile vehicle 1000 (e.g., mobile vehicles in front of, beside, and / or behind the mobile vehicle 1000). This functionality may be part of the cooperative adaptive cruise control function of the mobile vehicle 1000.

[0135] Network interface 1024 may include a SoC that provides modulation and demodulation functions and enables the controller 1036 to communicate via a wireless network. The network interface 1024 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. Frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless functions for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0136] The moving vehicle 1000 may further include a data store 1028 that may include storage outside the chip (e.g., outside the SoC 1004). The data store 1028 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least 1 bit of data.

[0137] The vehicle 1000 may further include a GNSS sensor 1058. The GNSS sensor 1058 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) supports mapping, perception, occupancy grid generation, and / or route planning functions. For example, any number of GNSS sensors 1058 may be used, including but not limited to GPS with a USB connector having Ethernet (registered trademark) to a serial (RS-232) bridge.

[0138] The moving vehicle 1000 may further include a RADAR sensor 1060. The RADAR sensor 1060 can be used by the moving vehicle 1000 for long-range moving vehicle detection even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 1060 can use CAN and / or bus 1002 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 1060) using access to Ethernet (registered trademark) to access raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 1060 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0139] The RADAR sensor 1060 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for an adaptive cruise control function. The long-range RADAR system can provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. The RADAR sensor 1060 can help distinguish between static and moving objects and can be used by an ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor may include a monostatic multi-modal RADAR having a plurality (e.g., six or more) of fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example having six antennas, the four central antennas can create a focused beam pattern designed to record the surroundings of the moving vehicle 1000 at high speed while minimizing interference from traffic in adjacent lanes. The other two antennas can widen the field of view and enable rapid detection of moving vehicles entering or leaving the lane of the moving vehicle 1000.

[0140] As an example, a mid - range RADAR system can include a range up to 1060m (front) or 80m (rear), and a field of view up to 42 degrees (front) or 1050 degrees (rear). A short - range RADAR system can include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the moving vehicle.

[0141] The short - range RADAR system can be used in an ADAS system for blind - spot detection and / or lane - change assist.

[0142] The moving vehicle 1000 can further include ultrasonic sensors 1062. The ultrasonic sensors 1062, which can be positioned at the front, rear, and / or sides of the moving vehicle 1000, can be used for parking assist and / or for creating and updating occupancy grids. A variety of ultrasonic sensors 1062 can be used, and different ultrasonic sensors 1062 can be used for different ranges of detection (e.g., 2.5m, 4m). The ultrasonic sensors 1062 can operate at a functional safety level of ASIL B.

[0143] The moving vehicle 1000 can include a LIDAR sensor 1064. The LIDAR sensor 1064 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1064 can also be at a functional safety level of ASIL B. In some examples, the moving vehicle 1000 can include multiple (e.g., 2, 4, 6, etc.) LIDAR sensors 1064 that can use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).

[0144] In some examples, the LIDAR sensor 1064 may have the ability to provide a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 1064 may have, for example, an accuracy of 2 cm to 3 cm, support for 1000 Mbps Ethernet® connections, and a advertised range of about 1000 m. In some examples, one or more non-protruding LIDAR sensors 1064 may be used. In such examples, the LIDAR sensor 1064 may be implemented as a small device that can be incorporated into the front, rear, sides, and / or corners of the moving vehicle 1000. In such examples, the LIDAR sensor 1064 may have a range of 200 m even for low-reflectivity objects and can provide up to a 120-degree horizontal and 35-degree vertical field of view. The LIDAR sensor 1064 attached to the front may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0145] In some examples, LIDAR technologies such as 3D flash LIDAR can also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the area around a moving vehicle up to about 200 m. The flash LIDAR unit includes a receptor that records the laser pulse travel time and the reflected light on each pixel, corresponding in sequence to the range from the moving vehicle to an object. Flash LIDAR can enable high-precision and distortion-free images of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of the moving vehicle 1000. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a blower. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser light in the form of 3D range point clouds and co-recorded intensity data. By using flash LIDAR, and since the flash LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1064 can be less susceptible to the effects of motion blur, vibration, and / or shock.

[0146] The moving vehicle can further include an IMU sensor 1066. In some examples, the IMU sensor 1066 can be positioned at the center of the rear axle of the moving vehicle 1000. The IMU sensor 1066 can include, for example, but is not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, in a 6-axis application, the IMU sensor 1066 can include an accelerometer and a gyroscope, while in a 9-axis application, the IMU sensor 1066 can include an accelerometer, a gyroscope, and a magnetometer.

[0147] In some embodiments, the IMU sensor 1066 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 1066 may enable the moving vehicle 1000 to estimate its heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 1066. In some examples, the IMU sensor 1066 and the GNSS sensor 1058 may be combined in a single integrated unit.

[0148] The moving vehicle may include a microphone 1096 disposed within and / or around the moving vehicle 1000. The microphone 1096 may be used, among other things, for emergency vehicle detection and identification.

[0149] The mobile vehicle may further include any number of camera types, including stereo camera 1068, wide view camera 1070, infrared camera 1072, surround camera 1074, long distance and / or medium distance camera 1098, and / or other camera types. The cameras can be used to capture image data around the entire outer surface of the mobile vehicle 1000. The type of cameras used is determined according to the embodiments and requirements of the mobile vehicle 1000, and any combination of camera types can be used to achieve the necessary coverage around the mobile vehicle 1000. Additionally, the number of cameras can vary according to the embodiments. For example, the mobile vehicle may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, and / or another number of cameras. As an example, the cameras can support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet (registered trademark). Each camera is described in more detail herein in relation to FIGS. 10A and 10B.

[0150] The mobile vehicle 1000 may further include a vibration sensor 1042. The vibration sensor 1042 can measure the vibration of components of the mobile vehicle, such as the axles. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 1042 are used, the difference in vibration can be used to determine the friction or slipperiness of the road surface (e.g., when the difference in vibration is between a power-driven axle and a freely rotating axle).

[0151] The moving vehicle 1000 may include an ADAS system 1038. In some examples, the ADAS system 1038 may include a SoC. The ADAS system 1038 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0152] The ACC system may use a RADAR sensor 1060, a LIDAR sensor 1064, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. The longitudinal ACC monitors and controls the distance to the moving vehicle directly in front of the moving vehicle 1000 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. The lateral ACC performs distance holding and advises the moving vehicle 1000 to change lanes when necessary. The lateral ACC is related to other ADAS applications such as LCA and CWS.

[0153] CACC can use information from other moving vehicles that can be received via a wireless link from other moving vehicles via network interface 1024 and / or wireless antenna 1026, or indirectly via a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the immediately preceding vehicle (e.g., the vehicle immediately in front of vehicle 1000 in the same lane as vehicle 1000), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of an I2V information source and a V2V information source. Given information about the vehicle ahead of vehicle 1000, CACC can be made more reliable and has the potential to make the traffic flow smoother and reduce road congestion.

[0154] The FCW system is designed to warn the driver of a hazard so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 1060 connected to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically connected to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide an alert in the form of an acoustic, visual alert, vibration, and / or quick brake pulse.

[0155] The AEB system can detect an imminent forward collision with another moving vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a forward-facing camera and / or RADAR sensor 1060 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a danger, it usually first warns the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or imminent collision braking.

[0156] The LDW system provides visual, audible, and / or tactile warnings, such as vibrations of the steering wheel or seat, to warn the driver when the moving vehicle 1000 crosses a lane dividing line. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. The LDW system can use a front-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically connected to driver feedback, such as a display, speaker, and / or vibrating component.

[0157] The LKA system is a modified form of the LDW system. The LKA system provides steering input or brakes to correct the moving vehicle 1000 if the moving vehicle 1000 begins to drift out of the lane.

[0158] The BSW system detects and warns the driver of a moving vehicle in the blind spots of a motor vehicle. The BSW system can provide visual, audible, and / or tactile warnings to indicate that a merge or lane change is not safe. The system can provide additional warnings when the driver uses the turn indicator. The BSW system can use a rear-facing camera and / or RADAR sensor 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0159] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1000 is backing up. Some RCTW systems include AEB to ensure that vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 1060 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0160] Conventional ADAS systems allow the driver to be warned and to determine whether a safety condition actually exists and act accordingly. Thus, conventional ADAS systems, while not usually catastrophic, tend to produce false positive results that can annoy and distract the driver. However, in the case of the autonomous vehicle 1000, when the results conflict, the moving vehicle 1000 itself must decide whether to accept the results from the primary computer or the secondary computer (e.g., the first controller 1036 or the second controller 1036). For example, in some embodiments, the ADAS system 1038 may be a backup and / or secondary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can execute diverse software that is redundant in hardware components to detect malfunctions in perception and dynamic driving tasks. The output from the ADAS system 1038 can be provided to the supervisory MCU. When the outputs from the primary computer and the secondary computer conflict, the supervisory MCU needs to determine how to resolve the conflict to ensure safe operation.

[0161] In some examples, the primary computer can be configured to provide a reliability score to the supervisory MCU that indicates the reliability of the primary computer in the selected result. If the reliability score exceeds a threshold, the supervisory MCU can follow the instructions of the primary computer regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold and the primary and secondary computers indicate different results (e.g., conflict), the supervisory MCU can mediate between the computers to determine an appropriate result.

[0162] The supervisory MCU may be configured to execute a neural network that is trained and configured to determine, based on outputs from the primary computer and the secondary computer, a state in which the secondary computer provides a false alarm. Thus, the neural network within the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network within the supervisory MCU can learn when the FCW is identifying metallic objects such as manhole covers or gratings in a drain that are not actually dangerous but trigger an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network within the supervisory MCU can learn to ignore the LDW when a person on a bicycle or a pedestrian is present and lane departure is actually the safest maneuver. In an embodiment that includes a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise components of the SoC1004 and / or be included as components of the SoC1004.

[0163] In other examples, the ADAS system 1038 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. As such, the secondary computer can use classical computer vision rules (if-then), and the presence of a neural network within the supervisory MCU can improve reliability, safety, and performance. For example, diverse implementations and intentional non-identities make the overall system more fault-tolerant, particularly against faults caused by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.

[0164] In some examples, the output of the ADAS system 1038 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 1038 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network that is trained as described herein and thus reduces the risk of misjudgment.

[0165] The mobile vehicle 1000 may further include an infotainment SoC 1030 (e.g., an in-vehicle infotainment system (IVI) in the mobile vehicle). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more individual components. The infotainment SoC 1030 may include a combination of hardware and software used to provide the mobile vehicle 1000 with audio (e.g., music, mobile devices, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calls), network connections (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, wireless data systems, fuel level, total mileage, brake fuel level, oil level, opening / closing doors, air filter information, and other mobile vehicle-related information). For example, the infotainment SoC 1030 may include radio, disk player, navigation system, video player, USB and Bluetooth connections, car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control device, hands-free voice control, heads-up display (HUD), HMI display 1034, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 1030 may be further used to provide information (e.g., visual and / or audible) to the user of the mobile vehicle, such as information from the ADAS system 1038, autonomous driving information such as planned mobile vehicle operations, trajectories, surrounding environment information (e.g., intersection information, mobile vehicle information, road information, etc.), and / or other information.

[0166] The infotainment SoC 1030 may include GPU functionality. The infotainment SoC 1030 can communicate with other devices, systems, and / or components of the moving vehicle 1000 via a bus 1002 (e.g., CAN bus, Ethernet®, etc.). In some examples, the GPU of the infotainment system can be connected to the supervisory MCU such that the infotainment SoC 1030 can execute some self-driving functions in the event that the primary controller 1036 (e.g., the primary and / or backup computer of the moving vehicle 1000) fails. In such examples, the infotainment SoC 1030 can put the moving vehicle 1000 into the chauffeur safe stop mode as described herein.

[0167] The moving vehicle 1000 may further include an instrument cluster 1032 (e.g., digital dash, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 1032 may include a controller and / or a supercomputer (e.g., an individual controller or supercomputer). The instrument cluster 1032 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, gear shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting control device, safety system control device, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1030 and the instrument cluster 1032. In other words, the instrument cluster 1032 may be included as part of the infotainment SoC 1030, and vice versa.

[0168] FIG. 10D is a system diagram of communication between a cloud-based server of FIG. 10A and an exemplary autonomous vehicle 1000 according to some embodiments of the present disclosure. System 1076 may include a server 1078, a network 1090, and a moving vehicle including moving vehicle 1000. Server 1078 may include a plurality of GPUs 1084(A)-1084(H) (collectively referred to herein as GPU 1084), PCIe switches 1082(A)-1082(H) (collectively referred to herein as PCIe switch 1082), and / or CPUs 1080(A)-1080(B) (collectively referred to herein as CPU 1080). The GPUs 1084, CPUs 1080, and PCIe switches may be interconnected by high-speed interconnects, such as, but not limited to, an NVLink interface 1088 and / or a PCIe connection 1086 developed by NVIDIA, for example. In some examples, the GPUs 1084 are connected via an NVLink and / or an NVSwitch SoC, and the GPUs 1084 and PCIe switches 1082 are connected via a PCIe interconnect. Eight GPUs 1084, two CPUs 1080, and two PCIe switches are shown, but this is not intended to be limiting. Depending on the embodiment, each server 1078 may include any number of GPUs 1084, CPUs 1080, and / or PCIe switches. For example, server 1078 may include eight, sixteen, thirty-two, and / or more GPUs 1084, respectively.

[0169] Server 1078 can receive, via network 1090, from a moving vehicle, image data representing an image indicating an unexpected or changed road condition, such as a recently started road construction. Server 1078 can transmit, via network 1090, to the moving vehicle, neural network 1092, an updated neural network 1092, and / or map information 1094 including information regarding traffic and road conditions. The update of map information 1094 may include an update of HD map 1022, such as information regarding a construction site, a depression, a detour, a flood, and / or other obstacles. In some examples, neural network 1092, the updated neural network 1092, and / or map information 1094 may have resulted from new training and / or experiences represented in data received from any number of moving vehicles in the environment, and / or based on training performed in a data center (e.g., using server 1078 and / or other servers).

[0170] Server 1078 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by a moving vehicle and / or (e.g., using a game engine) generated in a simulation. In some examples, the training data is tagged (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., when the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques, including but not limited to, for example, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and variations or combinations thereof. After the machine learning model is trained, the machine learning model can be used by the moving vehicle (e.g., transmitted to the moving vehicle via network 1090), and / or the machine learning model can be used by server 1078 to remotely monitor the moving vehicle.

[0171] In some examples, server 1078 can receive data from a moving vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 1078 can include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1084, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1078 can include a deep learning infrastructure that uses only CPU-powered data centers.

[0172] The deep learning infrastructure of server 1078 can have the ability of high-speed real-time inference, and use that ability to evaluate and verify the condition of the processors, software, and / or related hardware within moving vehicle 1000. For example, the deep learning infrastructure can receive periodic updates from moving vehicle 1000, such as sequences of images and / or objects in which moving vehicle 1000 is located within those sequences of images (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can execute its own neural network to identify objects and compare them with those identified by moving vehicle 1000. If the results do not match and the infrastructure concludes that the AI within moving vehicle 1000 is not functioning properly, server 1078 can send a signal to moving vehicle 1000 that infers control, notifies the passengers, and commands the fail-safe computer of moving vehicle 1000 to complete a safe parking operation.

[0173] For inference, server 1078 can include GPUs 1084 and one or more programmable inference acceleration devices (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time responsiveness. In other examples, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors can be used for inference.

[0174] Exemplary computing device FIG. 11 is a block diagram of an example of a computing device 1100 suitable for use in the implementation of some embodiments of the present disclosure. The computing device 1100 may include an interconnect system 1102 that indirectly or directly connects the following devices: memory 1104, one or more central processing units (CPUs) 1106, one or more graphics processing units (GPUs) 1108, a communication interface 1110, an input / output (I / O) port 1112, an input / output component 1114, a power supply 1116, one or more presentation components 1118 (e.g., a display), and one or more logic units 1120. In at least one embodiment, the computing device 1100 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). As a non-limiting example, one or more of the GPUs 1108 may include one or more vGPUs, one or more of the CPUs 1106 may include one or more vCPUs, and / or one or more of the logic units 1120 may include one or more virtual logic units. As such, the computing device 1100 may include individual components (e.g., all GPUs dedicated to the computing device 1100), virtual components (e.g., a portion of a GPU dedicated to the computing device 1100), or a combination thereof.

[0175] Although the various blocks of FIG. 11 are shown as being connected via an interconnect system 1102 with lines, this is not intended to be limiting and is merely for clarity. For example, in some embodiments, a presentation component 1118 such as a display device may be considered an I / O component 1114 (e.g., if the display is a touch screen). As another example, the CPU 1106 and / or GPU 1108 may include memory (e.g., the memory 1104 may represent a storage device in addition to the memory of the GPU 1108, CPU 1106, and / or other components). In other words, the computing device of FIG. 11 is merely illustrative. Categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types are all intended to be within the scope of the computing device of FIG. 11 and thus are not distinguished.

[0176] The interconnect system 1102 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 1102 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 1106 may be directly connected to the memory 1104. Further, the CPU 1106 may be directly connected to the GPU 1108. If there are direct or point-to-point connections between components, the interconnect system 1102 may include a PCIe link for implementing the connection. In these examples, the PCI bus need not be included in the computing device 1100.

[0177] The memory 1104 may include any of a variety of computer-readable media. The computer-readable media may be any available media that can be accessed by the computing device 1100. The computer-readable media may include both volatile and non-volatile media, and removable and non-removable media. By way of example, but not limitation, the computer-readable media may comprise computer storage media and communication media.

[0178] A computer storage medium can include both volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules, and / or other data types. For example, memory 1104 can store computer readable instructions, such as an operating system, that represent (e.g., a program and / or program elements). A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and that can be accessed by computing device 1100. In this specification, a computer storage medium does not include a signal per se.

[0179] A computer storage medium can implement computer readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transport mechanism and includes any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the foregoing should also be included within the scope of computer readable media.

[0180] The CPU 1106 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to execute one or more of the methods and / or processes described herein. The CPU 1106 may each include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores having the ability to process multiple software threads simultaneously. The CPU 1106 may include any type of processor and may include different types of processors depending on the type of computing device 1100 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 1100, the processor may be an Advanced RISC Machines (ARM) processor implemented using reduced instruction set computing (RISC), or an x86 processor implemented using complex instruction set computing (CISC). The computing device 1100 may include one or more CPUs 1106 within one or more microprocessors or auxiliary coprocessors, such as a computing coprocessor.

[0181] In addition to or instead of the CPU 1106, the GPU 1108 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 1100 to execute one or more of the methods and / or processes described herein. One or more of the GPU 1108 may be an integrated GPU (e.g., it may be with one or more of the CPU 1106), and / or one or more of the GPU 1108 may be a discrete GPU. In an embodiment, one or more of the GPU 1108 may be one or more coprocessors of the CPU 1106. The GPU 1108 may be used by the computing device 1100 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, the GPU 1108 may be used for general-purpose computing on the GPU (GPGPU). The GPU 1108 may include hundreds or thousands of cores having the ability to process hundreds or thousands of software threads simultaneously. The GPU 1108 can generate pixel data for an output image in response to rendering commands (e.g., rendering commands from the CPU 1106 received via a host interface). The GPU 1108 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, e.g., GPGPU data. The display memory may be included as part of the memory 1104. The GPU 1108 may include two or more GPUs operating in parallel (e.g., via a link). The link can be directly connected to the GPUs (e.g., using NVLINK), or the GPUs can be connected via a switch (e.g., using NVSwitch). When coupled together, each GPU 1108 can generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., the first GPU for the first image and the second GPU for the second image). Each GPU can include its own memory or share memory with other GPUs.

[0182] In addition to and / or instead of CPU 1106 and / or GPU 1108, logic unit 1120 may be configured to execute at least some of the computer-readable instructions to control one or more of computing devices 1100 to execute one or more of the methods and / or processes described herein. In an example, CPU 1106, GPU 1108, and / or logic unit 1120 may execute any combination of methods, processes, and / or portions thereof discretely or in parallel. One or more of logic units 1120 may be part of and / or integrated with one or more of CPU 1106 and / or GPU 1108, and / or one or more of logic units 1120 may be discrete components relative to and / or otherwise external to CPU 1106 and / or GPU 1108. In an example, one or more of logic units 1120 may be one or more coprocessors of one or more of CPU 1106 and / or one or more of GPU 1108.

[0183] Examples of the logic unit 1120 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel visual cores (PVC), vision processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application-specific integrated circuits (ASIC), floating-point units (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.

[0184] The communication interface 1110 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 1100 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 1110 can include components and functions for enabling communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0185] The I / O port 1112 can enable the computing device 1100 to be logically connected to other devices, including some of which may be built-in (e.g., integrated) into the computing device 1100, such as I / O components 1114, presentation components 1118, and / or other components. Exemplary I / O components 1114 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, and the like. The I / O components 1114 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 1100 (as will be described in more detail later). The computing device 1100 can include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 1100 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 1100 to render immersive augmented reality or virtual reality.

[0186] The power supply device 1116 can include a hard-wired power supply device, a battery power supply device, or a combination thereof. The power supply device 1116 can provide power to the computing device 1100 to enable the components of the computing device 1100 to operate.

[0187] The presentation component 1118 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 1118 can receive data from other components (e.g., the GPU 1108, the CPU 1106, etc.) and output the data (e.g., as an image, a video, an audio, etc.).

[0188] Exemplary data center FIG. 12 shows an exemplary data center 1200 that may be used in at least one embodiment of the present disclosure. The data center 1200 may include a data center infrastructure layer 1210, a framework layer 1220, a software layer 1230, and / or an application layer 1240.

[0189] As shown in FIG. 12, the data center infrastructure layer 1210 may include a resource orchestrator 1212, grouped computing resources 1214, and node computing resources (“node C.R.”) 1216(1) to 1216(N), where “N” represents any integer of natural numbers. In at least one embodiment, the node C.R. 1216(1) to 1216(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors or graphics processing units (GPU), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and / or cooling modules, etc., but are not limited thereto. In some embodiments, one or more of the node C.R. 1216(1) to 1216(N) may correspond to a server having one or more of the aforementioned computing resources. Additionally, in some embodiments, the node C.R. 1216(1) to 12161(N) may include one or more virtual components, such as vGPU, vCPU, and / or the like, and / or one or more of the node C.R. 1216(1) to 1216(N) may correspond to a virtual machine (VM).

[0190] In at least one embodiment, the grouped computing resources 1214 can include a separate group of node C.R.s 1216 housed within one or more racks (not shown), or multiple racks housed in data centers at various geographical locations (also not shown). Separate groups of node C.R.s 1216 within the grouped computing resources 1214 can include grouped computing, network, memory, or storage resources that can be configured or assigned to support one or more workloads. In at least one embodiment, some node C.R.s 1216 that include CPUs, GPUs, and / or other processors can be grouped within one or more racks to provide computing resources for supporting one or more workloads. One or more racks can also include any number of power modules, cooling modules, and / or network switches in any combination.

[0191] The resource orchestrator 1222 can configure or otherwise control one or more node C.R.s 1216(1)-1216(N) and / or the grouped computing resources 1214. In at least one embodiment, the resource orchestrator 1222 can include a software design infrastructure ("SDI") management entity of the data center 1200. The resource orchestrator 1222 can include hardware, software, or some combination thereof.

[0192] In at least one embodiment, as shown in FIG. 12, the framework layer 1220 may include a job scheduler 1232, a configuration manager 1234, a resource manager 1236, and / or a distributed file system 1238. The framework layer 1220 may include a framework to support software 1232 of the software layer 1230 and / or one or more applications 1242 of the application layer 1240. The software 1232 or the application 1242 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1220 may be of a type of free and open-source software web application framework, such as Apache Spark (trademark) (hereinafter “Spark”), which may use the distributed file system 1238 for large-scale data processing (e.g., “big data”), but is not limited thereto. In at least one embodiment, the job scheduler 1232 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1200. The configuration manager 1234 may have the ability to configure different layers, such as the software layer 1230 and the framework layer 1220 including Spark and the distributed file system 1238 to support large-scale data processing. The resource manager 1236 may have the ability to manage the mapped or allocated clustered or grouped computing resources for the support of the distributed file system 1238 and the job scheduler 1232. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 714 in the data center infrastructure layer 1210. The resource manager 1036 may coordinate with the resource orchestrator 1212 to manage these mapped or allocated computing resources.

[0193] In at least one embodiment, the software 1232 included in the software layer 1230 may include at least a portion of nodes C.R. 1216(1) through 1216(N), the grouped computing resources 1214, and / or software used by the distributed file system 1238 of the framework layer 1220. One or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.

[0194] In at least one embodiment, the application 1242 included in the application layer 1240 may include at least a portion of nodes C.R. 1216(1) through 1216(N), the grouped computing resources 1214, and / or one or more types of applications used by the distributed file system 1238 of the framework layer 1220. One or more types of applications may include, but are not limited to, machine learning applications including any number of genomics applications, cognitive computing, and training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0195] In at least one embodiment, any one of the configuration manager 1234, the resource manager 1236, and the resource orchestrator 1212 can implement any number and type of self-rewriting actions based on any amount and type of data obtained in any technically possible manner. Self-rewriting actions can free the data center operator of the data center 1200 by making potentially bad configuration decisions and perhaps avoiding underutilized and / or poorly performing portions of the data center.

[0196] Data center 1200 may include tools, services, software, or other resources for training one or more machine learning models or predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters by a neural network architecture that uses the software and / or computing resources described above with respect to data center 1200. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1200 by using weight parameters calculated via one or more training techniques, which are not limited to, for example, those described herein.

[0197] In at least one embodiment, data center 1200 may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) for training and / or performing inference using the aforementioned resources. Further, the aforementioned one or more software and / or hardware resources may be configured as services that enable a user to train or perform inference of information, such as image recognition, speech recognition, or other artificial intelligence services.

[0198] Exemplary network environment A network environment suitable for use in the implementation of embodiments of the present disclosure may include one or more client devices, a server, network attached storage (NAS), other backend devices, and / or other device types. The client devices, server, and / or other device types (e.g., each device) may be implemented as one or more instances of the computing device 1100 of FIG. 11. For example, each device may include similar components, features, and / or functionality of the computing device 1100. Additionally, when backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may be included as part of the data center 1200, an example of which is described in further detail herein with respect to FIG. 12.

[0199] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or a network of networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the public switched telephone network (PSTN), and / or one or more private networks. When the network includes a wireless telecommunications network, components, such as base stations, communication towers, or access points (and other components), may provide a wireless connection.

[0200] Compatible network environments may include one or more peer-to-peer network environments (in which case, a server may or may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to the server may be implemented on any number of client devices.

[0201] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of a core network server and / or an edge server among the servers. The framework layer may include a framework to support software in the software layer and / or one or more applications in the application layer. The software or application may each include web-based service software or an application. In an embodiment, one or more of the client devices may use web-based service software or an application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be of the type of a free and open-source software web application framework, but is not limited thereto, and may use a distributed file system for large-scale data processing (e.g., "big data").

[0202] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed to multiple locations from a central or core server (such as one or more data centers that may be distributed across a state, region, country, the world, etc.). When the connection to a user (such as a client device) is relatively close to an edge server, the core server may delegate at least a portion of the functionality to the edge server. The cloud-based network environment may be private (e.g., restricted to a single organization), public (e.g., available to multiple organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0203] The client device may include at least some of the components, features, and functionality of the exemplary computing device 1100 described herein with respect to FIG. 11. By way of example, and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0204] This disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions such as program modules, which may be executed by a computer or other machine, such as a mobile information terminal or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs specific tasks or implements specific abstract data types. This disclosure may be implemented in a variety of configurations, including handheld devices, household appliances, general-purpose computers, more specialized computing devices, etc. This disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked via a communication network.

[0205] As used herein, the description of "and / or" with respect to two or more elements should be construed to mean one element only, or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. In addition, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0206] The subject matter of this disclosure is described with specificity in order to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors intend that the claimed subject matter may be otherwise implemented, including in combination with other current or future technologies, in different steps or combinations of steps similar to those described in this document. Further, the terms "step" and / or "block" may be used herein to imply different elements of a method being used, but these terms should not be construed as implying any particular order among the various steps disclosed herein except where the order of individual steps is explicitly recited and when so recited.

Claims

1. A processor comprising a processing circuit for executing a neural network, wherein the neural network is at least partially trained using a training image and a boundary shape corresponding to the training image as ground truth data, and the training image and the boundary shape are generating a first image including simulated objects from the perspective of a virtual sensor of a virtual vehicle within a virtual environment; determining the boundary shape corresponding to the simulated object; inserting a representation of the simulated object into a second image, the second image being generated using a real-world sensor of a real-world vehicle; determining that a position corresponding to the representation of the simulated object is within a threshold distance to a driving surface depicted in the second image; and determining that the boundary shape satisfies an overlap constraint when compared with the ground truth data corresponding to the second image and at least partially generated thereby; and at least partially generated thereby, wherein executing the neural network comprises calculating data indicating one or more boundary shapes corresponding to one or more objects depicted in the third image, at least partially based on the third image generated using the neural network and using sensors of an autonomous machine. The processor according to claim 1, further comprising:

2. wherein the neural network is further at least partially trained by incorporating a plurality of representations of the simulated object into the virtual environment, wherein the plurality of representations are generated using a game engine, and each of the plurality of representations corresponds to at least one of different locations, different orientations, different appearance attributes, different sets of lighting conditions, different sets of occlusion conditions, or different environmental conditions. The processor according to claim 1.

3. wherein the boundary shape is further at least partially generated by generating a segmentation mask corresponding to the first image of the simulated object, and determining the boundary shape is at least partially based on the segmentation mask. The processor according to claim 1.

4. ​ ​ ​ ​ Determining that a position corresponding to the representation of the simulated object is within the threshold distance to the driving surface, Determining that the position corresponding to the representation of the simulated object is within a threshold vertical distance from the driving surface, or Determining that the position corresponding to the representation of the simulated object is within a threshold horizontal distance from one or more boundary lines of the driving surface The processor according to claim 1, comprising at least one of the above.

5. The processor according to claim 1, wherein the neural network is trained using zero-shot learning.

6. Determining that the boundary shape satisfies the overlapping constraint includes determining that a first pixel within the boundary shape is not included within a second pixel determined from the ground truth data. The processor according to claim 1.

7. Further executing the neural network, Calculating data indicating the classification of the object using the neural network and based at least in part on the third image The processor according to claim 1, comprising:

8. Inserting the representation of the simulated object into the second image, Determining the pose of the representation of the simulated object with respect to the virtual sensor, and Inserting the representation of the simulated object into the second image in the pose with respect to the real-world sensor The processor according to claim 1, comprising:

9. The processor according to claim 1, wherein the processing circuit further performs one or more operations based at least in part on one or more outputs of the neural network generated while executing the neural network.

10. Sampling a plurality of criteria to determine a set of criteria; Selecting a representation of a simulated object from a plurality of representations based at least in part on the set of criteria; Generating a boundary shape corresponding to the representation of the simulated object; Training images, Using the representation of the simulated object to expand the image, Determining that a position corresponding to the representation of the simulated object is within a threshold distance to a driving surface depicted in the image, and determining that the boundary shape satisfies an overlap constraint when compared to ground truth data corresponding to the image at least partially by generating; training a neural network using the training image and the boundary shape; and including using the neural network and at least partially based on an image generated using sensors of an autonomous machine, a method in which data indicating one or more boundary shapes corresponding to one or more objects depicted in the generated image is calculated. **Claim 11** The method according to claim 10, wherein the plurality of criteria includes at least one of an object type, an object location, an object orientation, an occlusion condition, an environmental condition, an appearance attribute, or an illumination condition. **Claim 12** The step of generating the boundary shape includes generating a segmentation mask corresponding to another representation of the simulated object captured using a virtual sensor in a virtual environment, and generating the boundary shape to include each of the pixels of the segmentation mask corresponding to the another representation of the simulated object. The method according to claim 10. **Claim 13** The method according to claim 10, wherein the plurality of representations are generated using a game engine and at least partially based on the plurality of criteria. **Claim 14** The method according to claim 10, wherein the neural network is trained using zero-shot learning. **Claim 15** One or more processing units, and one or more memory devices storing instructions that, when executed using the one or more processors, cause the one or more processors to perform operations, wherein the operations include generating a first image including a representation of a simulated object from the perspective of a virtual sensor of a virtual vehicle in a virtual environment; determining a boundary shape corresponding to the simulated object; inserting the training image by inserting the representation of the simulated object into a second image, the second image being generated using real-world sensors of a real-world vehicle. Determining that a position within the second image corresponding to the representation of the simulated object is within a threshold distance to a driving surface depicted in the second image, and determining that the boundary shape satisfies an overlap constraint when compared to ground truth data corresponding to the second image at least partially generating by training a neural network using the training image and the boundary shape including a system, wherein data indicating one or more boundary shapes corresponding to one or more objects depicted in the third image is calculated based at least in part on the third image generated using the neural network and sensors of an autonomous machine. **Claim 16** wherein the operations further include incorporating a plurality of representations of the simulated object within the virtual environment including wherein the plurality of representations are generated using a game engine wherein each of the plurality of representations corresponds to at least one of different locations, different orientations, different appearance attributes, different lighting conditions, different occlusion conditions, or different environmental conditions The system according to claim 15. **Claim 17** wherein the operations further include generating a segmentation mask corresponding to the representation of the simulated object depicted in the first image including wherein determining the boundary shape is at least partially based on the segmentation mask The system according to claim 15. **Claim 18** wherein determining that the representation of the simulated object is within the threshold distance to the driving surface includes determining that a position corresponding to the representation of the simulated object is within a threshold vertical distance from the driving surface, or determining that a position corresponding to the representation of the simulated object is within a threshold horizontal distance from one or more boundary lines of the driving surface The system according to claim 15, including at least one of. **Claim 19** wherein determining that the boundary shape satisfies the overlap constraint includes determining that a first pixel within the boundary shape is not included within a second pixel determined from the ground truth data **Claim 20** wherein the system A system for performing a simulation operation, A system for performing a deep learning operation, A system implemented using an edge device, A system incorporating one or more virtual machines (VMs), A system at least partially implemented in a data center, or A system at least partially implemented using cloud computing resources The system according to claim 15, included in at least one of

Citation Information

Patent Citations

  • Image synthesizing device, image synthesizing method, and program

    JP2019200774A

  • Information processing method and information processing system

    JP2020038605A