Sensor Fusion for Autonomous Machine Applications Using Machine Learning

The machine learning-based sensor fusion system addresses the challenges of inaccurate sensor data in autonomous driving by using a multi-sensor fusion network to combine data from multiple models and additional channels, resulting in improved accuracy and precision for object detection and tracking.

JP7689806B2Active Publication Date: 2025-06-09NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021560671
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-06-21
Filing Date
2021-06-21
Publication Date
2025-06-09
Estimated Expiration
2041-06-21

AI Technical Summary

Technical Problem

Existing autonomous driving systems face challenges in accurately detecting static and dynamic objects, obstacles, and environmental features due to noisy or inaccurate sensor data, especially at the boundaries or overlapping regions of different sensors' fields of view.

Method used

A machine learning-based sensor fusion system that uses a multi-sensor fusion network, such as a deep neural network (DNN), to combine data from multiple individual machine learning models. This system processes outputs from various source models and internal values like feature extractor layers, and incorporates additional channels like position priority and velocity images to improve accuracy and precision.

Benefits of technology

The system reduces noise and improves the accuracy and precision of sensor fusion outputs, leading to more reliable object tracking, planning, and obstacle avoidance in autonomous machines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689806000001
    Figure 0007689806000001
  • Figure 0007689806000002
    Figure 0007689806000002
  • Figure 0007689806000003
    Figure 0007689806000003
Patent Text Reader

Abstract

In various instances, a multi-sensor fusion machine learning model—e.g., a deep neural network (DNN)—can be deployed to fuse data from multiple individual machine learning models. As such, the multi-sensor fusion network can use outputs from multiple machine learning models as inputs to generate a fused output representing data from the respective sensor fields of view or perception fields that feed into the machine learning model, while taking into account learned associations between boundaries or overlapping regions of the various fields of view of the source sensors. In this way, the fusion network can be trained to account for multiple instances of the same object appearing in different input representations, reducing the likelihood that the fused output will contain duplicate, inaccurate, or noisy data about objects or features in the environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to sensor fusion for autonomous machine applications using machine learning.

Background Art

[0002] The ability to safely detect static and dynamic objects, obstacles, hazards, waiting conditions, road markings, signs, and / or other features of the environment is an important issue for any autonomous or semi-autonomous driving system. For example, detecting the positions of various static and / or dynamic features or objects within the environment of a three-dimensional (3D) space is particularly difficult and can result in noisy or more inaccurate data - especially with respect to the boundaries or overlapping regions between the fields of view of different sensors.

[0003] In some systems, a deep neural network (DNN) can be deployed to generate 3D information corresponding to the field of view of the relevant sensors that provide sensor data to the DNN. For example, for an object or landmark that appears near the image boundary, where the object or landmark disappears and reappears within the sensor's field of view - for example, resulting in a rapid change between detection and non-detection of the same object or landmark - these perception systems can result in noisy results. As a result, when tracking an object or landmark across frames, the results can be inaccurate or imprecise and can lead to a degradation in system performance.

[0004] To explain these potential drawbacks of individual DNNs, some conventional systems combine the outputs from multiple DNNs using handcrafted or rule-based algorithms. However, when the fields of view overlap between adjacent sensors that provide sensor data to their respective DNNs, the fused predictions along the boundaries or overlapping regions can be inconsistent or inaccurate. For example, a DNN running on a sensor with a 120-degree field of view might predict that an object 10 meters from the sensor is 9 meters away, while another DNN running on a sensor with a 60-degree field of view might predict that the object is 11 meters away. In such an example, the individual DNNs are only 1 meter off from the actual distance of the object, but the outputs are 2 meters apart from each other, which can be a sufficient distance separation to include two separate double objects in the fused output for a single object actually represented in the data. This incorrect determination that two separate objects exist can then be trusted by the underlying system, and the false detection can propagate through the machine's object tracking, planning, control, obstacle avoidance, and / or other operations.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Patent Document 2

Summary of the Invention

Means for Solving the Problems

[0006] Embodiments of the present disclosure relate to machine learning-based sensor fusion for autonomous machine applications. A system and method for fusing data from multiple individual machine learning models using a multi-sensor fusion machine learning model - for example, a deep neural network (DNN) - are disclosed. For example, multiple machine learning models can be trained to generate outputs that can be used directly or after post-processing as inputs to a multi-sensor fusion network. Each output received as an input by the multi-sensor fusion network can correspond to the same type of representation - for example, a rasterized image from the same perspective (e.g., looking down from above, egocentric perspective). As such, the multi-sensor fusion network can use these inputs to generate a fusion output that represents data from the field of view or perception field of each sensor type while explaining the learned associations between the various field-of-view boundaries or overlapping regions of the source sensors. In this way, the fusion network can be trained to explain multiple instances of the same object that appear in different input representations, so that the likelihood that the fusion output contains duplicate, inaccurate, or noisy data regarding an object or feature in the environment can be reduced.

[0007] In some embodiments, in addition to or alternatively to the use of the outputs from various source machine learning models, the internal values of each machine learning model - for example, the output values from one or more feature extractor layers inside various DNNs - can be provided as inputs to the multi-sensor fusion network. In such an example, the selected layers of the multi-sensor fusion network and the individual machine learning models can be trained together in an end-to-end training process. As such, the updates of the weights and biases as a result of one or more loss functions can be backpropagated not only through the layers of the multi-sensor fusion network but also through the layers of each source machine learning model (e.g., the feature extractor layer).

[0008] To further improve the accuracy and precision of the multi-sensor fusion network, one or more additional channels can be provided as input to the multi-sensor fusion network - for example, in each iteration. For example, a position priority channel can be used to represent distance estimation uncertainty - for example, as a probability distribution function (PDF) - at one or more pixels or points corresponding to the field of view or sensor field of one or more source sensors. In some embodiments, a velocity image channel can be used as input to the multi-sensor fusion network to represent the associated velocity of detected objects in the environment. For example, different sensor or machine learning model outputs can include velocity information - for example, in the x and / or y directions - and the associated velocity of an object can be used to fuse objects (for example, if the velocity predictions exactly match) or to represent two or more objects (for example, if the velocity predictions of neighboring objects are greater than a threshold difference). Additionally, spatial and / or temporal association channels based on instance and / or appearance can be used as additional input to the multi-sensor fusion network to help identify object appearances across different inputs and / or similar object appearances over time - for example, by using a recurrent neural network to track objects over time.

[0009] As a result, the underlying machine learning model that provides data to the multi-sensor fusion network can be optimized to generate outputs that help improve the accuracy of the fusion output of the multi-sensor fusion network. Additional inputs - for example, position priority channels, velocity image channels, etc. - can be used to further improve the accuracy and precision of the multi-sensor fusion network, especially with respect to the boundary lines or overlapping regions between adjacent sensor fields of view or cognitive fields. By reducing noise and improving the accuracy and precision of the multi-sensor fusion network, downstream processes that rely on these outputs of the multi-sensor fusion network - representing the position, velocity, pose, appearance, etc. of static and dynamic objects or features within the environment - can also benefit from improved performance and effectiveness.

[0010] The present system and method for machine learning-based sensor fusion for autonomous machine applications will be described in detail below with reference to the accompanying drawings.

Brief Description of the Drawings

[0011]

Figure 1A

Figure 1B

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 4C

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8A

Figure 8B

Figure 8C

Figure 8D

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0012] A system and method for machine learning-based sensor fusion for autonomous machine applications are disclosed. This disclosure may be described with respect to an exemplary autonomous vehicle 800 (alternatively referred to herein as "vehicle 800", "autonomous machine 800", or "ego vehicle 800", examples of which are described with respect to FIGS. 8A - 8D), although this is not intended to be limiting. For example, the systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), manned and unmanned robots or robotics platforms, warehouse vehicles, off-road vehicles, vehicles coupled to one or more trailers, flying vessels, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, aircraft, construction vehicles, submarines, drones, and / or any other vehicle or machine type, although not limited thereto. Additionally, this disclosure may be described with respect to sensor fusion for autonomous or semi-autonomous machine operation, although this is not intended to be limiting, and the systems and methods described herein may be used in extended reality, virtual reality, mixed reality, robotics, security and surveillance, autonomous or semi-autonomous machine applications, and / or any other technical space where sensor fusion may be used.

[0013] Referring to FIGS. 1A - 1B, FIGS. 1A - 1B are exemplary data flow diagrams corresponding to processes 100A and 100B (collectively referred to as "process 100") for multi - sensor fusion according to some embodiments of the present disclosure. It should be understood that these and other configurations described herein are merely described as examples. Other configurations and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those illustrated, and some elements may be omitted entirely. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor that executes instructions stored in a memory. In some embodiments, processes 100A and 100B may be executed using the components, features, and / or functionality of machine 800 of FIGS. 8A - 8D, exemplary computing device 900 of FIG. 9, and / or data center 1000 of FIG. 10.

[0014] Process 100 may include generating, accessing, and / or receiving sensor data from one or more sensors. Sensor data may be received, as non - limiting examples, from one or more sensors of a vehicle (e.g., vehicle 800 of FIGS. 8A - 8D described herein). The sensor data may be 2 - dimensional (2D) and / or 3 - dimensional (3D) signals (S 1 ~S n ) 104A - 104N corresponding to camera sensors, used by vehicle 800 within process 100, 3D RADAR signals (S RADAR ) 108 of RADAR sensors (which may be calculated using only RADAR data and / or using a DNN), 3D ultrasonic signals (S USS)112. The 3D LiDAR signal (S of the LiDAR sensor that can be calculated using only LiDAR data and / or can be calculated using a DNN) LiDAR ) (not shown) can be used to calculate and / or to calculate other output signals using additional or alternative sensor data (from any sensor modality) and / or the corresponding DNN. In some embodiments, the sensor data is an output internal to the DNN itself, e.g., the feature output (F 1 ~F n ) 102A~102N of a DNN using a camera signal, the RADAR feature output (F RADAR ) 106 of a RADAR sensor, the ultrasonic feature output (F USS ) 110 of an ultrasonic sensor, the LiDAR feature output (F LiDAR ) of a LiDAR sensor, and / or other feature outputs of other sensor types can be used to calculate. As such, the sensor data can be calculated using any number of sensors and using any number of different sensor modalities, the sensor data can be used directly (e.g., converting raw LiDAR data to a point cloud with or without preprocessing), and / or can be used after processing by a DNN or another type of machine learning model.

[0015] During training (described in more detail herein with respect to FIGS. 6A-7), sensor data can be generated and / or pre-generated and included in a training data set using one or more data collection vehicles that generate sensor data for training a DNN, such as a DNN related to feature output and / or 3D signals. The sensor data used during training can be additionally or alternatively generated using simulated sensor data (e.g., sensor data generated using one or more virtual sensors of a virtual vehicle in a virtual environment) and / or augmented sensor data (e.g., sensor data generated using one or more data collection vehicles and modified with virtual data). After being trained and deployed in vehicle 800, the sensor data can be generated by one or more sensors of vehicle 800 and processed by a DNN to calculate various output signals and / or to calculate feature outputs from one or more feature extractor layers of the respective DNN.

[0016] As a non-limiting example, sensor data can include data from any of the sensors of vehicle 800, such as, referring to FIGS. 8A - 8C, RADAR sensor 860, ultrasonic sensor 862, LIDAR sensor 864, stereo camera 868, wide view camera 870 (e.g., a fish-eye camera), infrared camera 872, surround camera 874 (e.g., a 360-degree camera), long-range and / or mid-range camera 878, and / or other sensor types. As another example, sensor data can include virtual (e.g., simulated or augmented) sensor data generated from any number of sensors of a virtual vehicle or other virtual objects in a virtual (e.g., test) environment. In such an example, the virtual sensors can correspond to virtual vehicles or other virtual objects in a simulated environment (e.g., used to test, train, and / or verify DNN performance), and the virtual sensor data can represent sensor data captured by the virtual sensors in the simulated or virtual environment. As such, by using virtual sensor data, the DNNs described herein can be tested, trained, and / or verified using simulated or augmented data in a simulated environment, enabling testing of more extreme scenarios outside of the real-world environment where the safety of such tests can be lower.

[0017] In some embodiments, the sensor data can include image data representing an image, image data representing a video (e.g., a snapshot of a video), and / or sensor data representing a representation of the sensor's field of perception (e.g., a depth map, a point cloud, a distance image of a LiDAR sensor, etc., a value graph of an ultrasonic sensor, etc.). When the sensor data includes image data, for example, without limitation, a compressed image as a frame derived from a compressed video format such as JPEG (Joint Photographic Experts Group) or luminance / chrominance (YUV) format, H.264 / AVC (Advanced Video Coding) or H.265 / HEVC (High Efficiency Video Coding), a raw image such as one derived from a red clear blue (RCCB), red clear (RCCC), or other type of image sensor, and / or any type of image data format such as other formats can be used. Additionally, in some examples, the sensor data can be used within process 100 without preprocessing (e.g., in a raw or captured format), while in other examples, the sensor data can undergo preprocessing (e.g., noise balancing, demosaicing, scaling, trimming, expanding, white balancing, tone curve adjustment, point cloud generation, projection image generation (e.g., for generating a LiDAR distance image), etc., using, for example, a sensor data preprocessor (not shown)). As used herein, the sensor data can refer to unprocessed sensor data, preprocessed sensor data, or a combination thereof.

[0018] In an example where a LiDAR sensor is used to generate, for example, a point cloud, the point cloud from one or more LiDAR sensors can be integrated - for example, after performing motion correction of the ego machine 800 and / or performing data synchronization to address motion-related LiDAR and time synchronization issues. Similarly, combinations of sensor data representations (e.g., point clouds, depth maps, etc.) of different sensor types - e.g., RADAR, ultrasonic, etc. - can also undergo similar preprocessing before being directly used and / or processed by one or more machine learning models to compute a 3D signal.

[0019] For training, as described in more detail herein, sensor data can include the original image (e.g., as captured by one or more image sensors), downsampled images, upsampled images, trimmed or region of interest (ROI) images, images augmented in other ways, and / or combinations thereof. Similarly, if the sensor data corresponds to data other than image data, the sensor data used for training can include different orientations, projections, reference points, augmentations, trimmings, filterings, and / or the like. During training of the DNN-fusion DNN 120, ground truth data can be generated. In some embodiments, the ground truth data can be automatically generated, for example, in an unsupervised manner (e.g., using photometric consistency loss between the outputs of sensors having at least partially overlapping fields of view or cognitive fields), generated in a supervised manner using annotation data, and / or generated in a semi-supervised manner. When annotations are used to generate the ground truth data, the annotations can be generated within a drawing program (e.g., an annotation program), a computer aided design (CAD) program, a labeling program, another type of program suitable for generating annotations, and / or, in some instances, can be hand-drawn. In any instance, the annotation data can be synthetically produced (e.g., generated from a computer model or rendering), physically produced (e.g., designed and manufactured from real-world data), machine automated (e.g., using feature analysis and learning to extract features from data and then generate labels), manually annotated by a person (e.g., a labeler or annotation expert defines the position of the label), and / or combinations thereof (e.g., identifying the center point or origin and dimensions of an area and the machine generates polygons and / or labels for the intersection areas).

[0020] In an embodiment, each feature output or 3D signal may correspond to a respective sensor pipeline or stream. For example, a first sensor pipeline can include a first camera that can generate image data that can be processed by a first DNN to generate a feature output F 1 and / or 3D signal 104A, and a second sensor pipeline can include a second camera that can generate image data that can be processed by a second DNN to generate a feature output F 2 and / or 3D signal 104B. A third sensor pipeline can include a first RADAR sensor that can generate RADAR data that can be processed directly - for example, using a sensor data preprocessor - and / or processed using a DNN to generate a feature output F RADAR and / or 3D signal 108, and so on. Depending on the embodiment, any number of sensor pipelines may be used.

[0021] When process 100 uses 3D signals, the 3D signals can be calculated by each respective DNN in the same format or can be converted to the same format - for example, using a post-processor. For example, in some non-limiting embodiments, the 3D signals used as input to the fusion DNN 120 can include a rasterized image (e.g., an aerial view looking down from above, a projected image, e.g., a distance image, a survey image, etc.) generated from a particular perspective that encodes any number of input channels. In embodiments where the egomachine 800 is at the center of the representation, the rasterized image - including the fusion output rasterized image calculated using the fusion DNN - can be egocentric. In other embodiments, the input 3D signals can be generated from the perspective of the egomachine 800 and the fusion output 122 can be generated from an egocentric perspective. For example, the input channels can indicate the shape, orientation, and / or classification of objects or features in the environment. In such instances, the rasterized image can include, for example, a boundary shape or cuboid corresponding to a dynamic actor, lane markers (e.g., lane dividers, road dividers, solid lines, dashed lines, double lines, yellow lines, white lines, etc.), waiting conditions (e.g., crosswalks, stop lines, etc.), and / or lines corresponding to other driving surface features, boundary lines or encoded values of pixels corresponding to drivable free space (e.g., an area of the environment that the egomachine 800 can traverse), and / or other objects or features.

[0022] The input channels included in the rasterized input image (or other input representation) can include the starting altitude of an object as a channel and / or the ending altitude of an object as a channel, an object or feature position or occupancy channel (e.g., in the case of a rasterized image of a top-down view, the pixel position can indicate the x or y pixel position corresponding to a feature or object that indicates the lateral or vertical position relative to the ego machine 800, and the pixel can be encoded with one or more altitude channels indicating the rise of the object or feature), one or more velocity channels (e.g., the pixels of the image can be encoded with the velocity in the x and / or y directions), one or more orientation channels corresponding to an object (e.g., encoded as an angle), and / or one or more classification channels corresponding to an object or feature, and / or additional or alternative channels. Additionally, these channels can be rasterized for each different input image corresponding to each sensor pipeline. For example, for the camera sensor pipeline including the 3D signal 104A, with respect to the input representation of each sensor pipeline, there can be an occupancy channel, a velocity channel, a classification channel, etc., and for the RADAR sensor pipeline including the 3D signal 108, there can be an occupancy channel, a velocity channel, a classification channel, etc. In some embodiments, the different input signals to the fusion DNN 120 can include different channel types, the same channel type, or a combination thereof. As a non-limiting example, the first sensor pipeline can generate a 3D signal 104A including an occupancy channel and a velocity channel, while the 3D signal 112 can include only the occupancy channel.

[0023] Non-limiting examples of input channels of different sensor types that may be included in an input 2D or 3D signal are shown in FIG. 1B, where various input signals 104, 108, 112, 114, 116, and / or 118 may include a rasterized image representing one or more of the illustrated input channel types. For example, other examples of 3D signals that include a rasterized image are shown in FIGS. 2A and 2B. For example, FIG. 2A may represent a first rasterized image 202A from a first sensor pipeline that includes a sensor having a field of view or perception field in front of the ego machine 800 and a second rasterized image 202B from a second sensor pipeline that includes a sensor having a field of view or perception field behind the ego machine 800. The rasterized images 202A and 202B may include representations of lane boundary lines 204A, 204B, etc., object detections 206A, 206B, 206C, etc. (e.g., corresponding to vehicles, pedestrians, animals, cyclists, debris, robots, etc.), and / or indications of free space (e.g., pixels may be encoded as free space or not free space). As illustrated with respect to detection 206, the detection may include a boundary shape indicating the position, shape, and / or orientation of the detected object (e.g., if points within the boundary shape indicate a direction of travel or a directed location). Object detections 206, lane boundary lines 204, and / or other features or objects encoded in the rasterized image 202 may also include classification information, velocity information, and / or other information. As described herein, additional or alternative features or objects, such as standby conditions, road profile information (e.g., depressions, perturbations, speed bumps, etc.), and / or other information describing the surrounding environment, may be included in the rasterized image 202 in addition to those shown in FIGS. 2A-2B.

[0024] In some embodiments, the classification information corresponding to an object or feature may be represented by an intensity value. For example, if the classification corresponds to an x of the intensity value, the overlapping area of two objects will be x + 1, three objects will be x + 2, and so on. However, if there is an overlap in the object, the boundary line of the object may be encoded with an intensity value indicating the boundary line, for example, without limitation, with a maximum intensity value of 255, and the intensity value ranges from 0 to 255. As a result, the boundary line between the detected objects can be more easily identified.

[0025] Different input 3D signals may correspond to different sensors with different fields of view or perception fields. For example, as shown in FIGS. 4A - 4C and also in FIG. 8B, various different sensor configurations may be used to calculate 3D signals corresponding to a part or all around the ego - machine 800. For example, the fields of view or perception fields (e.g., the fields of view 404A, 404B, 404C, 404D, 404E shown in FIGS. 4A - 4C) may together constitute a 360 - degree field of view around the machine 800 or a field of view less than 360 degrees around the machine 800, depending on the embodiment. Additionally, in some embodiments, one or more of the 3D signals and thus the associated sensors may include a 360 - degree field of view or perception field around the machine 800. For example, a rotating LiDAR sensor or a 360 - surround camera can include a 360 - degree field of view or perception field, while a camera can include a 30 - degree, 60 - degree, 120 - degree, and / or other fields of view (e.g., up to 360 degrees, e.g., in the case of a surround camera).

[0026] When the sensor pipeline is adapted to cameras and / or other sensor types that do not directly compute depth, the 3D signal may further include predicted depth values represented in a rasterized image or other input representation. For example, an individual DNN used in a sensor data pipeline may be trained to predict depth using a single camera image, or may be trained to predict depth from the inputs of two or more sensors having overlapping fields of view (e.g., fields of view 404A and 404B that may include an overlapping region 406 as shown in FIG. 4A). In some embodiments, the location-prioritized image 114 may be generated as an additional input to the fusion DNN 120 that instructs the fusion DNN 120 to fuse a probability distribution function 252 (represented as a 2D Gaussian representation in the embodiment) indicating the predicted depth - or distribution of potential depths - of an object based on the predicted depth from the 3D signal. For example, as shown in FIG. 2C, one or more sensors of the ego machine 800 can detect objects 240A, 240B, 240C, 240D, etc., as indicated by circles within the location-prioritized image 114A. In such an instance, when the object location prediction is based on sensor data generated using a monocular camera or other sensor type where the accuracy of the location prediction may not be ideal, the distribution of potential locations may be fed to the fusion DNN 120 to assist in generating a more accurate prediction in the fusion output 122. As such, the ellipse or probability distribution function (PDF) 252 representation corresponding to each object 240 may indicate the potential locations where the object 240 may be located - for example, having corresponding confidence values. For different sensor modalities, the corresponding ellipse or PDF 252 may be of different shapes. For example, since camera predictions may be along a ray (e.g., ray 402 in FIG. 4A), the ellipse or PDF may be longer and narrower in shape (e.g., representing a bimodal distribution along the ray direction), while a RADAR sensor may have a shorter but wider representation to account for its respective inaccuracies or the RADAR sensor, etc. For example, FIG. 2D may show different ellipses or PDFs 252 corresponding to RADAR sensor predictions within the field of view 250A of the RADAR sensor.In addition, the shape of the ellipse or PDF can vary depending on where the detection is located within the field of view of each sensor. For example, referring to FIG. 2E, an ellipse or PDF that is near the edge of field of view 250B or 250C or further from the sensor may be different from the center of field of view 250 or closer to the sensor - for example, at the edge, the prediction may be more inaccurate than at the center, and thus the shape of ellipse or PDF 252 may be larger to exhibit greater variability. Similar representations are shown in FIG. 2D with respect to field of view 250A.

[0027] Pixels or points (e.g., corresponding to a digitized 3D location) corresponding to a detected object as predicted using a sensor pipeline to generate a location-priority image 114 - e.g., location-priority image 114A of FIG. 2C or location-priority image 114B of FIG. 2E - can be used to obtain the corresponding learned or determined ellipse or PDF corresponding to that point or pixel within the field of view of each sensor. For example, FIG. 2D may represent a subset of the ellipses or PDFs of field of view 250A of a particular sensor. The figure includes several ellipses or PDFs 252A, 252B, 252C, etc., but this is not intended to be limiting. For example, in some embodiments, each pixel or point representing the field of view may have a corresponding ellipse or PDF 252. In other embodiments, any number of pixels or points may have a corresponding ellipse or PDF 252. These ellipses or PDFs 252 can be learned or determined based on past predictions of a particular sensor and the corresponding errors of those predictions. As such, when an object is detected with respect to a particular sensor at a particular point or pixel, the corresponding known ellipse or PDF 252 can be obtained and inserted into location-priority image 114.

[0028] In some embodiments, instead of or in addition to including one or more of the channels within the input signal of each pipeline, as described herein, one or more of the channels may be used to generate distinct input representations - for example, velocity image 116, instance / appearance image 118, etc. For example, the velocity information generated using one or more of the sensor pipelines may be used to generate one or more velocity images 116 and / or one or more instance / appearance images 118 - for example, temporal instance / appearance images and / or spatial instance / appearance images. The velocity image 116 may include one or more channels corresponding to velocity in the x-direction, y-direction, etc. In some embodiments, as described herein, the velocity image 116 may be generated based on the output of one or more of the sensor pipelines. For example, one or more of the 3D signals from the sensor pipeline may include velocity information, and instead of or in addition to encoding this information in the rasterized image used as an input to the fusion DNN 120, the velocity information may be used to generate one or more distinct inputs corresponding to the velocity image 116. As such, the velocity information can assist the fusion DNN 120 in determining whether neighboring objects or features are the same object (e.g., similar or the same velocity) or different objects (e.g., different velocities), and thus how to represent the objects in the fusion output 122 - for example, as a single object or as two or more objects.

[0029] As another example, an instance / appearance image 118 corresponding to a spatial association of instances and / or appearances may be used. For example, one or more of the sensor pipelines can generate an output indicating a feature or an object descriptor. In such an example, one or more of the machine learning models within one or more of the sensor pipelines may be trained to produce an N-dimensional vector of each boundary shape / 2D cuboid / 3D cuboid of an object - for example, in an embodiment N may be equal to 3. As a result, an N-dimensional vector can be generated for one or more detected objects or features. As such, one or more vectors generated using one or more sensor pipelines can be compared to each other to determine how many object instances are present within a given frame. If a first vector and a second vector are sufficiently similar (e.g., within a threshold similarity), they can be determined to correspond to the same object, and this information can be represented in the instance / appearance image 118. Similarly, an instance and / or appearance image can be generated for a temporal association between objects or features. For example, a recurrent neural network can be used to obtain input vectors or instance / appearance information across frames and to generate an instance / appearance image 118 representing object instances across time. In any example, the instance / appearance image 118 can be used as an additional input for the fusion DNN 120 when calculating the fusion output 122. For example, the instance / appearance information can assist the fusion DNN 120 in determining whether adjacent objects or features are the same object or different objects, and thus how to represent the objects in the fusion output 122 - for example, as a single object or as two or more objects.

[0030] In each iteration of the fusion DNN 120, 3D signals 104, 108, 112, etc., the position-priority image 114, the velocity image 116, and / or the instance / appearance image 118 can be provided as inputs to the fusion DNN 120. The fusion DNN 120 can process the inputs to generate a fusion output 122 - examples of which are shown in FIGS. 3A-3B. One or more of the fusion DNN 120 and / or the DNN or machine learning model used to generate the 3D signal can be, for example and without limitation, any type of machine learning model, such as linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forest, dimensionality reduction algorithm, gradient boosting algorithm, neural network (e.g., autoencoder, convolutional, recurrent, perceptron, long / short term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, adversarial generation, liquid state machine, etc.), and / or other types of machine learning models.

[0031] In an embodiment, when one or more of the fusion DNN 120 and / or the DNN or machine learning model used to generate the 3D signal includes a convolutional neural network (CNN), one or more of the layers can include an input layer. The input layer can hold values related to various inputs. For example, when the input is an image (e.g., a rasterized image corresponding to various channels from a 3D signal), the input layer can hold values representing the raw pixel values of the image as a volume (e.g., width, W, height, H, and color channels, C (e.g., RGB), e.g., 32×32×3), and / or a batch size, B (e.g., when batch processing is used).

[0032] One or more layers may include convolutional layers. A convolutional layer can compute the output of neurons connected to local regions (e.g., the input layer) within the input layer, and each neuron computes the dot product between their weights and the small region to which they are connected in the input volume. The result of the convolutional layer may be another volume having one of the dimensions based on the number of filters applied (e.g., the width, height, and number of filters, e.g., if the number of filters is 12, 32×32×12).

[0033] One or more of the layers may include a ReLU (rectified linear unit) layer. The ReLU layer can apply, for example, an element-wise activation function that thresholds at zero, e.g., max(0, x). The resulting volume of the ReLU layer may be the same as the volume of the input to the ReLU layer.

[0034] One or more of the layers may include a pooling layer. The pooling layer can perform a downsampling operation along spatial dimensions (e.g., height and width) that can result in a volume smaller than the input to the pooling layer (e.g., from a 32×32×12 input volume to 16×16×12). In some examples, the fused DNN 120 and / or other DNNs may not include a pooling layer. In such examples, a strided convolutional layer may be used instead of a pooling layer. In some examples, the feature extractor layer (e.g., F in FIG. 6B) of the fused DNN 120 and / or other DNNs 1 , F 2 etc.) may include alternating convolutional and pooling layers or may not include a pooling layer at all.

[0035] One or more of the layers may include a fully-connected layer. Each neuron within the fully-connected layer may be connected to each neuron in the previous volume. The fully-connected layer can compute class scores, and the resulting volume can be 1×1×number of classes. In some examples, the fully-connected layer may not be used overall by the fused DNN 120 and / or other DNNs in an effort to increase the number of processing operations and reduce computational resource requirements. In such examples, if the fully-connected layer is not used, the fused DNN 120 and / or other DNNs may be referred to as a fully convolutional network.

[0036] One or more of the layers may include a transposed convolutional layer in some examples. However, the use of the term "transposed convolutional" can be misleading and is not intended to be limiting. For example, a transposed convolutional layer may also be referred to as a transposed convolution layer or a fractionally-strided convolution layer. A transposed convolutional layer may be used to perform upsampling on the output of the previous layer. For example, a transposed convolutional layer can be used to upsample to a spatial resolution equal to the spatial resolution of the input to the fused DNN 120 and / or other DNNs (e.g., the spatial resolution of a 3D signal), or may be used to upsample to the input spatial resolution of the next layer.

[0037] The input layer, convolutional layer, pooling layer, ReLU layer, transposed convolutional layer, and fully-connected layer are discussed in this specification in relation to the fused DNN 120, but this is not intended to be limiting. For example, additional or alternative layers such as normalization layers, SoftMax layers, and / or other layer types may be used.

[0038] Different orders and numbers of layers of the fusion DNN 120 may be used according to the embodiments. Additionally, for example, some of the layers may include parameters (e.g., weights and / or biases), while at the same time, other layers, such as ReLU layers and pooling layers, may not be included. In some examples, the parameters may be learned by the fusion DNN 120 during training. Further, some of the layers may include additional hyperparameters (e.g., learning rate, stride, epoch, kernel size, number of filters, type of pooling in the pooling layer, etc.) - for example, convolutional layers, transposed convolutional layers, and pooling layers, while other layers, such as ReLU layers, may not be included. Various activation functions, including but not limited to ReLU, leaky ReLU, sigmoid, hyperbolic tangent (tanh), ELU (exponential linear unit), etc., may be used. The parameters, hyperparameters, and / or activation functions should not be limited and may vary according to the embodiments.

[0039] In some embodiments, as described with respect to FIG. 6B, one or more of the DNNs used to generate 3D signals may include a trunk, T, a feature extractor layer, F, and / or a head, H, that may be used to generate 3D signals 104, 108, 112, etc. In an example, in addition to or alternatively to using 3D signals 104, 108, 112, etc. generated from individual input sources as described herein, the fusion DNN 120 may directly receive the output layer of the feature extractor (e.g., F 1 ~F n ) of FIG. 1A as an input for generating the fusion output 122. In such an example, during training, the feature extractor layer of the individual DNN may be directly connected to one or more layers of the fusion DNN 120 such that backpropagation from the fusion DNN 120 may survive to the individual DNN to train both the individual DNN and the fusion DNN 120.

[0040] The fusion DNN 120 may generate a fusion output 122 using the output of one or more layers of an individual DNN, 3D signals 104, 108, 112, etc., a position - priority image 114, a velocity image 116, and / or an instance appearance image 118. In an embodiment, the fusion output 122 may be similar to a 3D signal - for example, the fusion output 122 may include data that represents or can be used to generate a rasterized image (e.g., similar to those shown in FIGS. 3A - 3B). For example, the fusion output may be an aerial view image, a projection image, such as a distance image, a perspective image, etc., that encodes any number of output channels (e.g., similar to one or more of the input channels described with respect to FIGS. 1A - 1B). In an embodiment where the ego - machine 800 is at the center of the representation, the fusion output rasterized image calculated using the fusion DNN may be egocentric. The fusion output 122 may include a fused representation of one or more fields of view or sensor fields of one or more input sources. For example, if the rasterized image of FIG. 2A includes the field of view of a forward - facing camera and the rasterized image of FIG. 2B includes the field of view of a rear - facing camera, the fusion output 300A of FIG. 3A may represent both the rasterized image of FIG. 2A and the rasterized image of FIG. 2B. FIG. 3B may include another visualization of the fusion output 300B that includes objects 206 (including object 206D) and lanes 204 (including lanes 204C and 204D).

[0041] The illustrations of FIGS. 2A, 2B, and 3A are for illustrative purposes only and are not intended to be limiting. For example, the 3D signal may represent a smaller and / or larger portion of the surrounding environment of the ego - vehicle 800 depending on the embodiment and may represent the fusion of any number of input sources. Further, one or more of the 3D signals may include a field of view or perception field that overlaps with the field of view or perception field of one or more other 3D signals so that the fusion DNN 120 can learn to more accurately predict distances, for example, using stereo - camera functionality.

[0042] The fused output 122 can be used by an autonomous driving software stack ("drive stack") 124 to perform one or more operations by a vehicle 800 (and / or other ego-machine type). For example, the drive stack 124 may include a world model manager that can be used to generate, update, and / or define a world model. The world model manager can use information generated by and received from the cognitive components of the drive stack 124. The cognitive components may include an obstacle perception device, a path perception device, a waiting perception device, a map perception device, and / or other cognitive components. For example, the world model can be defined based at least in part on affordances of obstacles, paths, and waiting conditions that can be recognized in real time or near real time by an obstacle perception device, a path perception device, a waiting perception device, and / or a map perception device. The world model manager can continuously update the world model based on newly generated and / or received inputs (e.g., data) from the obstacle perception device, the path perception device, the waiting perception device, the map perception device, and / or other components of the ego-machine 800. For example, the world model manager and / or the cognitive components can use the fused output 122 to perform one or more operations.

[0043] The world model can be used to help inform the planning components, control components, obstacle avoidance components, and / or actuation components of the drive stack 124. The obstacle perception device can perform obstacle perception based on where the vehicle 800 can travel or has the ability to travel and how fast the vehicle 800 can travel without colliding with an obstacle (e.g., an object, e.g., a building, an entity, a vehicle, etc.) sensed by the vehicle 800 (and, e.g., represented in the fused output 122).

[0044] The route recognition device can execute route recognition, for example, by recognizing the nominal routes available in a specific situation. In some examples, the route recognition device can further consider lane changes for route recognition. The lane graph can represent one or more routes available to vehicle 800 and can be as simple as a single route on the arterial road entrance lane. In some examples, the lane graph can include a route to a desired lane and / or indicate available changes along the arterial road (or other road type), or can include nearby lanes, lane changes, intersections, corners, cloverleaf intersections, merges, and / or other information.

[0045] The standby recognition device can be responsible for determining the restrictions on vehicle 800 as a result of rules, customs, and / or implementation considerations. For example, the rules, customs, and / or implementation considerations can be related to traffic signals, multi-way stops, yields, merges, toll booths, gates, police or other emergency responders, road workers, stopped buses or other vehicles, one-way bridge traffic control, ferry entrances, etc. In some examples, the standby recognition device can be responsible for determining the longitudinal restrictions on vehicle 800 that require the vehicle to wait or decelerate until some conditions are met. In some examples, the standby conditions can be due to potential obstacles that may not be directly perceivable by the obstacle recognition device (for example, by using sensor data from the sensor, as obstacles may sometimes block the sensor's field of view), such as cross traffic at an intersection. As a result, the standby recognition device can provide situation awareness by eliminating the risk of obstacles that may not always be immediately perceivable through recognized and / or learned rules and customs. Therefore, the standby recognition device can be utilized to identify potential obstacles and implement one or more controls (such as deceleration, stop, etc.) that may not have been possible relying only on the obstacle recognition device.

[0046] The map recognition device may include a mechanism for determining, thereby, the behavior and, in some examples, the specific examples of which conventions apply in a particular locale.

[0047] The planning components may include, among several components, features, and / or functionalities, a route planner, a lane planner, a behavior planner, and a behavior selector. The route planner can generate a planned route that can be composed of GNSS waypoints (e.g., GPS waypoints) using information from, among several information, the map recognition device, the map manager, and / or the localization manager. The waypoints can represent a specific distance into the future of the vehicle 800 that can be used as a target for the lane planner, e.g., several blocks, several kilometers / miles, several meters / feet, etc.

[0048] The lane planner can use a lane graph (e.g., a lane graph from a route recognition device that can be generated, at least in part, using the fusion output 122), the object pose within the lane graph (e.g., by the localization manager), and / or the target points and directions at a future distance from the route planner as input. The target points and directions can be mapped to the most matching drivable points and directions in the lane graph (e.g., based on GNSS and / or compass bearings). A graph search algorithm can then be executed on the lane graph from the current end in the lane graph to find the shortest path to the target point.

[0049] The behavior planner can determine the feasibility of the vehicle 800's basic behaviors, such as staying in a lane or changing lanes left or right, so that the executable behavior can match up with the most desirable behavior output from the lane planner. For example, if the desired behavior is determined to be unsafe and / or unavailable, a default behavior can be selected instead (for example, the default behavior may be to stay in the lane when it is not safe to change to the desired behavior or lane).

[0050] The control component can closely follow the trajectory or route (horizontal and vertical) received from the behavior selector of the planning component as much as possible and within the scope of the capabilities of the vehicle 800.

[0051] The obstacle avoidance component can assist the autonomous vehicle 800 in avoiding collisions with objects (e.g., moving and stationary objects). In some examples, the obstacle avoidance component can be used independently of the components, features, and / or functionality of the vehicle 800 that are required to follow traffic rules and drive carefully. In such examples, the obstacle avoidance component can ignore traffic laws, road rules, and norms of careful driving in order to ensure that no collision occurs between the vehicle 800 and any object. As such, the obstacle avoidance layer may be a separate layer from the rules of the road layer, and the obstacle avoidance layer can ensure that the vehicle 800 is only performing safe actions from the perspective of obstacle avoidance. On the other hand, the rules of the road layer can ensure that the vehicle follows traffic laws and customs and observes legal and customary priorities.

[0052] In any instance, one or more of the layers, components, features, and / or functionality of the drive stack 124 may use the fusion output 122 to generate outputs related to world model management, planning, control, actuation, collision or obstacle avoidance, and / or the like to assist the ego machine 800 in environmental navigation.

[0053] Referring now to FIG. 5, each block of method 500 described herein includes a computational process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions may be implemented by a processor executing instructions stored in memory. Method 500 may also be implemented as computer-usable instructions stored on a computer storage medium. Method 500 may be provided, by way of example, as a stand-alone application, service or hosted service (either stand-alone or in combination with another hosted service), or a plug-in to another product. Further, method 500 is described, by way of example, with respect to processes 100A and 100B of FIGS. 1A-1B. However, method 500 may be implemented additionally or alternatively within any one process, or any combination of processes and systems, including but not limited to those described herein.

[0054] FIG. 5 is a flow diagram showing a method 500 of using a multi-sensor fusion network to calculate a fusion output using a plurality of input channels, according to some embodiments of the present disclosure. Method 500 includes receiving, at block B502, first data that at least represents a first rasterized image generated at least in part based on first sensor data generated using a first sensor of a first type. For example, first 3D signals 104, 108, 112, etc. may be generated using a first sensor pipeline.

[0055] Method 500 includes, at block B504, receiving second data representing at least a second rasterized image generated at least in part based on second sensor data generated using a second sensor of a second type. For example, second 3D signals 104, 108, 112, etc. may be generated using a second sensor pipeline.

[0056] Method 500 includes, at block B506, calculating third data representing a fused rasterized image using a fusion DNN based at least in part on the first data and the second data. For example, fusion DNN 120 can calculate a fusion output 122 using a first 3D signal, a second 3D signal, and / or one or more other 3D signals. In some embodiments, as described herein, in addition to or instead of using 3D signals as inputs, fusion DNN 120 can use the outputs of one or more feature extractor layers of one or more individual DNNs corresponding to one or more of the sensor pipelines.

[0057] Method 500 includes, at block B508, executing one or more operations using an autonomous machine based at least in part on the third data. For example, the drive stack 124 of ego machine 800 can use the fusion output 122 to execute one or more operations such as world model management, planning, control, actuation, obstacle avoidance, and / or the like.

[0058] Referring to FIGS. 6A-6B, FIGS. 6A-6B both represent a data flow diagram of a process 600 for training a multi-sensor fusion network and one or more associated source networks according to some embodiments of the present disclosure. For example, referring to FIG. 6B, an individual sensor pipeline may include an individual or source DNN, such as one or more trunk layers 604 (e.g., trunk layers 604A, 604B, and 604N), one or more feature extractor layers 606 (e.g., feature extractor layers 606A, 606B, and 606N), and / or one or more heads or output layers 608 (e.g., output heads 608A, 608B, and 608N). The sensor pipeline may include a sensor 602 (e.g., sensors 602A, 602B, and 602N) that generates sensor data - such as any of the sensor data described herein with respect to FIGS. 1A-1B - that is processed using the source DNN to generate 3D signals (e.g., 3D signals 104A, 104B, and 104N). For example, the sensor may include a camera, a LiDAR sensor, a RADAR sensor, an ultrasonic sensor, and / or other sensor types that may be used to generate sensor data for an individual DNN. The 3D signals output by the source DNN may be compared to ground truth data 614 using another loss function to calculate a loss 612 corresponding to the source DNN. The ground truth data 614 may be generated using map data (e.g., from an HD map, or other map type, such as those used for localization) that can indicate the position of static features or objects such as lane boundaries, waiting conditions, signs, stationary objects, and / or the like. In some embodiments, the ground truth 614 may be generated using a 2D or 3D ground truth generation pipeline described in U.S. Non-Provisional Patent Application No. 17 / 187,350, filed Feb. 26, 2021, which is hereby incorporated by reference in its entirety. The ground truth data 614 may include a rasterized image including any of the channels described with respect to the 3D signals 104, 108, 112, etc. of FIGS. 1A-1B.For example, the output 3D signal 104 of the source DNN may correspond to a rasterized image or may be used to generate a rasterized image, which may be compared to a ground truth rasterized image to calculate losses 622 (e.g., losses 622A, 622B, and 622N). These losses may be used, for example, in backpropagation to update the parameters (e.g., weights and biases) of the source DNN to assist in training the DNN until they converge to an acceptable level of accuracy or precision.

[0059] In addition to or alternatively to calculating the loss 622 using the ground truth data 614, the photometric consistency loss can be calculated at 610 to compare the output 3D signals 104 from two or more source DNNs having sensors 602 with at least partially overlapping fields of view. For example, similar to the generation of a stereo camera disparity map, the photometric consistency loss 624 (e.g., the loss 624A between the 3D signal 104A and the 3D signal 104B and / or between the 3D signal 104B and the 3D signal 104A, the loss 624B between the 3D signal 104N and the 3D signal 104K and / or between the 3D signal 104K and the 3D signal 104N, etc.) can be calculated to compare the outputs of two or more source DNNs to ensure consistency in at least a portion of the overlapping fields of view or sensor fields (e.g., the overlap 406 in FIG. 4A). As such, with respect to the 3D signals 104A and 104B, the coordinate transformer 618 transforms the 3D signal 104A into the coordinate space of the 3D signal 104B to generate the 3D signal 616A, and then compares the transformed 3D signal 616A with the 3D signal 104B to determine the consistency - or lack thereof - between the signals, thereby calculating the consistency loss 624A. As such, if the overlapping regions are inconsistent - e.g., an object in the transformed 3D signal 616A is different from the same object within the 3D signal 104B (e.g., in terms of position, depth, orientation, class, etc.) - the loss 624A may be higher and the source DNN corresponding to the 3D signal 104A and / or the 3D signal 104B may be penalized (e.g., the parameters may be updated). This process can be similarly performed by using the coordinate transformer 618 to transform the 3D signal 104B into the coordinate space of the 3D signal 104A to generate the 3D signal 616B, by using the coordinate transformer 618 to transform the 3D signal 104N into the coordinate space of the 3D signal 104K to generate the 3D signal 616C, by using the coordinate transformer 618 to transform the 3D signal 104K into the coordinate space of the 3D signal 104N to generate the 3D signal 616D, etc. As such, by calculating the loss 622 and / or the loss 624, the source DNNs can be trained until they reach an acceptable level of accuracy or precision.

[0060] The benefit of training the source DNN separately from the fusion DNN 120 is that the sensor data or simulated sensor data used in training does not need to rely on a realistic rendering of the input image. For example, 3D signals 104, 108, 112, etc. can represent rasterized images generated using various channels, so the rasterized images used for training can be generated without requiring realistic sensor data - for example, since the input to the fusion DNN 120 is a rasterized image.

[0061] Similarly, and with respect to FIGS. 6A and 6B, the fusion output 122 of the fusion DNN 120 can be compared to the ground truth data 614 corresponding to the fusion output in order to calculate the loss 630. The ground truth data 614 for the fusion output 122 can be generated using similar techniques or data as the ground truth data 614 of the source DNN. As such, the loss 630 can be used to update the parameters of the fusion DNN 120 until the fusion DNN 120 converges to an acceptable level of accuracy or precision. Other input types - for example, the position priority channel 114, the velocity image 116, and / or the instance / appearance image 118 - are not shown in FIGS. 6A - 6B, but this is not intended to be limiting, and in some embodiments, one or more of these input channels can also be supplied as input to the fusion DNN 120 in each training iteration.

[0062] For example, with respect to FIG. 6C, the fused output 646 can include an object 648, and the ground truth data 614 can indicate the actual position of the object as the ground truth object 650. In such an example, the calculated loss 630 can represent this difference, and the parameters of the fused DNN 120 can be updated. In an example, if the 3D signal from the source DNN corresponds to a rasterized image (e.g., including less or the same amount of the surrounding environment of the ego machine 800), the loss 622 can be calculated similarly using the corresponding rasterized image of the 3D signal.

[0063] In some embodiments, as described herein, the input to the fused DNN 120 can correspond to the feature output 102 - e.g., the output of the feature extractor layer 606 of the source DNN - such that the layers of the source DNN can correspond to separate input trunks or streams of the layers of the fused DNN 120. In such an example, the connections between the nodes of the feature extractor layer 606 of the source DNN and the nodes of the layer (e.g., the input layer) of the fused DNN 120 can be connected. As such, the fused DNN 120 and the source DNN can be trained simultaneously such that the loss calculated for the fused DNN 120 can be backpropagated to the source DNN.

[0064] Referring now to FIG. 6, each block of method 700 described herein can include a computing process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. Method 700 can also be implemented as computer-usable instructions stored on a computer storage medium. Method 700 can be provided, for example, by a stand-alone application, a service, or a hosted service (either stand-alone or in combination with another hosted service), or a plug-in to another product. Further, method 700 is described, by way of example, with respect to process 600 of FIGS. 6A-6B. However, method 700 can be executed, additionally or alternatively, within any one process by any one system, or any combination of processes and systems, including but not limited to those described herein.

[0065] FIG. 7 is a flow diagram showing a method 700 for training a multi-sensor fusion network to compute a fused output using a plurality of input channels, according to some embodiments of the present disclosure. Method 700 includes generating a plurality of trained DNNs by calculating one or more losses with respect to the outputs of individual DNNs and calculating one or more consistency losses with respect to the outputs of two or more individual DNNs at block B702. For example, a source DNN can be trained using ground truth data 614 to generate a loss 622 and can be trained using a photometric consistency loss 624.

[0066] Method 700 includes, at block B704, generating a trained fusion DNN by calculating one or more losses with respect to a fusion output, the fusion output being calculated using outputs of a plurality of trained DNNs. For example, after being trained, the source DNNs can calculate outputs, and these outputs can be used as inputs to the fusion DNN 120 for training the fusion DNN 120. For example, the fusion DNN 120 can calculate a fusion output 122, and the fusion output 122 can be compared to ground truth data 614 for calculating a loss 630. In some examples, the fusion DNN 120 can be trained using simulated or manufactured data - such as rasterized images - additionally or alternatively without using the actual outputs of the source DNNs.

[0067] Method 700 includes, at block B706, deploying the plurality of trained DNNs and the fusion DNN in an ego machine. For example, after being trained, the fusion DNN 120 and the source DNNs can be deployed in the ego machine 800. Exemplary Autonomous Vehicle

[0068] FIG. 8A is a diagram of an exemplary autonomous vehicle 800 according to some embodiments of the present disclosure. The autonomous vehicle 800 (or referred to herein as "vehicle 800") can include, but is not limited to, a passenger vehicle, such as a car, truck, bus, first responder vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police vehicle, ambulance, boat, construction vehicle, submarine, drone, vehicle connected to a trailer, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). The autonomous vehicle is generally described in terms of the level of automation as defined by the National Highway Traffic Safety Administration (NHTSA), a department of the United States Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). The moving vehicle 800 can have the ability to function according to one or more of automation levels 3 to 5 of the autonomous driving level. For example, the moving vehicle 800 can have the ability of conditional automation (level 3), high automation (level 4), and / or full automation (level 5) depending on the embodiment.

[0069] The moving vehicle 800 can include components such as a chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the moving vehicle. The moving vehicle 800 can include a propulsion system 850, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. The propulsion system 850 can be connected to the drive train of the moving vehicle 800 and can include a transmission to enable the propulsion force of the moving vehicle 800. The propulsion system 850 can be controlled in response to receiving a signal from the throttle / acceleration device 852.

[0070] The steering system 854, which may include a steering wheel, can be used to steer the moving vehicle 800 (e.g., along a desired path or route) when the propulsion system 850 is operating (e.g., when the moving vehicle is in motion). The steering system 854 can receive signals from a steering actuator 856. The steering wheel may be an option for a fully automated (Level 5) function.

[0071] The brake sensor system 846 can be used to operate the vehicle brakes in response to receiving signals from a brake actuator 848 and / or a brake sensor.

[0072] The controller 836, which may include one or more system-on-chips (SoCs) 804 (FIG. 8C) and / or GPUs, can provide signals (e.g., representations of commands) to one or more components and / or systems of the moving vehicle 800. For example, the controller can send signals to operate the moving vehicle brakes via one or more brake actuators 848, to operate the steering system 854 via one or more steering actuators 856, and to operate the propulsion system 850 via one or more throttle / acceleration devices 852. The controller 836 can include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or to assist the driver in operating the moving vehicle 800. The controller 836 can include a first controller 836 for autonomous driving functions, a second controller 836 for functional safety functions, a third controller 836 for artificial intelligence functions (e.g., computer vision), a fourth controller 836 for infotainment functions, a fifth controller 836 for redundancy in an emergency, and / or other controllers. In some examples, a single controller 836 can process two or more of the foregoing functions, and two or more controllers 836 can process a single function and / or any combination thereof.

[0073] Controller 836 can provide signals for controlling one or more components and / or systems of mobile vehicle 800 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received from, for example and without limitation, a global navigation satellite system sensor 858 (e.g., a global positioning system sensor), a RADAR sensor 860, an ultrasonic sensor 862, a LIDAR sensor 864, an inertial measurement unit (IMU) sensor 866 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 896, a stereo camera 868, a wide-view camera 870 (e.g., a fish-eye camera), an infrared camera 872, a surround camera 874 (e.g., a 360-degree camera), a long-range and / or mid-range camera 898, a speed sensor 844 (e.g., for measuring the speed of mobile vehicle 800), a vibration sensor 842, a steering sensor 840, a brake sensor (e.g., as part of a brake sensor system 846), and / or other sensor types.

[0074] One or more of the controllers 836 of the vehicle 800 may receive an input (e.g., represented by input data) from the instrument cluster 832 of the vehicle 800 and provide an output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 834, an audible annunciator, a loudspeaker, and / or other components of the vehicle 800. The output may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 822 of FIG. 8C), position data (e.g., the position of the vehicle 800 on a map, etc.), direction, the position of other vehicles (e.g., occupancy grid), information regarding objects and the situation of objects as perceived by the controller 836, etc. For example, the HMI display 834 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B within 3.22 km (2 miles), etc.).

[0075] The vehicle 800 further includes a network interface 824 that can communicate via one or more networks using one or more wireless antennas 826 and / or a modem. For example, the network interface 824 may have the ability to communicate via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. The wireless antenna 826 may also use local area networks such as Bluetooth (registered trademark), Bluetooth LE, Z-Wave, ZigBee, etc., and / or low power wide-area networks (LPWAN) such as LoRaWAN, SigFox, etc., to enable communication between objects (e.g., vehicles, mobile devices, etc.) in the environment.

[0076] FIG. 8B is an example of the camera positions and fields of view of the exemplary autonomous vehicle 800 of FIG. 8A according to some embodiments of the present disclosure. The cameras and their respective fields of view are one exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the moving vehicle 800.

[0077] The camera type of the camera may include, but is not limited to, a digital camera adapted to be used with the components and / or systems of the moving vehicle 800. The camera can operate at automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have the ability to capture images at any rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may have the ability to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to increase light sensitivity.

[0078] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional mono-camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all of the cameras) can record and provide image data (e.g., video) simultaneously.

[0079] One or more of the cameras can be mounted on mounting components such as custom-designed (3D printed) parts to remove stray light and reflections from within the vehicle that may interfere with the camera's image data capture capabilities (e.g., reflections from the dashboard reflected in the front windshield mirror). Referring to the side mirror mounting component, the side mirror component can be custom 3D printed so that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera can be integrated within the side mirror. For a side view camera, the camera can also be integrated within four struts located at each corner of the cabin.

[0080] A camera having a field of view that includes a portion of the environment in front of the moving vehicle 800 (e.g., a forward-facing camera) can be used for surround view to assist in identifying the forward path and obstacles and, with the help of one or more controllers 836 and / or a control SoC, assist in providing information essential for generating an occupancy grid and / or determining a preferred moving vehicle path. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems including other functions such as lane departure warning (LDW), autonomous cruise control (ACC), and / or traffic sign recognition.

[0081] A variety of cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 870 that can be used to capture objects entering the view from the surroundings (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in FIG. 8B, any number of wide-view cameras 870 may be present on the moving vehicle 800. Additionally, a long-range camera 898 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. The long-range camera 898 can also be used for object detection and classification, as well as basic object tracking.

[0082] One or more stereo cameras 868 can also be included in the forward-facing configuration. The stereo camera 868 can include an integrated control unit with an expandable processing unit that can provide a programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet® interface integrated on a single chip. Such a unit can be used to generate a 3D map of the environment of the moving vehicle, including distance estimates for all points in the image. An alternative stereo camera 868 can include a compact stereo vision sensor that includes two camera lenses (one each for left and right) and an image processing chip that can measure the distance from the moving vehicle to the target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 868 may be used in addition to, or instead of, those described herein.

[0083] A camera (e.g., a side-view camera) having a field of view that includes a portion of the environment relative to the side of the moving vehicle 800 can be used for surround view to provide information for creating and updating an occupancy grid and for generating a side-impact collision warning. For example, surround cameras 874 (e.g., four surround cameras 874 as shown in FIG. 8B) can be positioned on the moving vehicle 800. The surround cameras 874 can include wide-view cameras 870, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be disposed in front of, behind, and on the sides of the moving vehicle. In an alternative arrangement, the moving vehicle may use three surround cameras 874 (e.g., left, right, and rear), and one or more other cameras (e.g., a forward-facing camera) can be utilized as a fourth surround-view camera.

[0084] A camera (e.g., a rear-view camera) having a field of view that includes a portion of the environment relative to the rear of the moving vehicle 800 can be used for parking assistance, surround view, rear-collision warning, and for creating and updating an occupancy grid. A wide variety of cameras can be used, including but not limited to cameras suitable as forward-facing cameras (e.g., long-range and / or mid-range cameras 898, stereo cameras 868), infrared cameras 872, etc., as described herein.

[0085] FIG. 8C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 800 of FIG. 8A, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in a memory.

[0086] Each of the components, features, and systems of the moving vehicle 800 of FIG. 8C is illustrated as being connected via a bus 802. The bus 802 may include a Controller Area Network (CAN) data interface (or referred to as a "CAN bus"). CAN may be a network within the moving vehicle 800 used to assist in controlling various features and functions of the moving vehicle 800, such as the operation of brakes, acceleration, brakes, steering, front windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other moving vehicle status indicators. The CAN bus may be ASIL B compliant.

[0087] Bus 802 is described herein as being a CAN bus, but this is not intended to be limiting. For example, in addition to, or as an alternative to, the CAN bus, FlexRay and / or Ethernet® may be used. Additionally, a single line is used to represent bus 802, but this is not intended to be limiting. For example, any number of buses 802 may exist that include one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 802 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 802 may be used for a collision avoidance function and a second bus 802 may be used for actuation control. In any example, each bus 802 may communicate with any of the components of the vehicle 800, and two or more buses 802 may communicate with the same component. In some examples, each SoC 804, each controller 836, and / or each computer within the vehicle may have access to the same input data (e.g., input from sensors of the vehicle 800) and may be connected to a common bus such as a CAN bus.

[0088] Vehicle 800 may include one or more controllers 836, such as those described herein with respect to FIG. 8A. Controller 836 may be used for various functions. Controller 836 may be coupled to any of the various other components and systems of vehicle 800 and may be used for the control of vehicle 800, the artificial intelligence of vehicle 800, the infotainment for vehicle 800, and / or the like.

[0089] The mobile vehicle 800 may include a system-on-chip (SoC) 804. The SoC 804 may include a CPU 806, a GPU 808, a processor 810, a cache 812, an accelerator 814, a data store 816, and / or other components and features not shown. The SoC 804 may be used to control the mobile vehicle 800 within various platforms and systems. For example, the SoC 804 may be coupled in a system (such as the system of the mobile vehicle 800) having an HD map 822 that can obtain map refreshes and / or updates via a network interface 824 from one or more servers (such as server 878 of FIG. 8D).

[0090] The CPU 806 may include a CPU cluster or CPU complex (or also referred to as a "CCPLEX"). The CPU 806 may include multiple cores and / or an L2 cache. For example, in some embodiments, the CPU 806 may include 8 cores within a coherent multi-processor configuration. In some embodiments, the CPU 806 may include 4 dual-core clusters, each cluster having a dedicated L2 cache (such as a 2MB L2 cache). The CPU 806 (such as a CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of clusters of the CPU 806 to become active at any given time.

[0091] The CPU 806 can implement a power management capability that includes one or more of the following features: individual hardware blocks can be automatically clock-gated when in an idle state to conserve dynamic power, each core clock can be gated when the core is not actively executing instructions by the execution of WFI / WFE instructions, each core can be independently power-gated, each core cluster can be independently clock-gated when all cores are clock-gated or power-gated, and / or each core cluster can be independently power-gated when all cores are power-gated. The CPU 806 can further implement an enhanced algorithm for managing power states, where the allowed power states and the expected wake-up times are specified and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing cores can support a simplified power state input sequence in software where the work is offloaded to microcode.

[0092] The GPU 808 may include an integrated GPU (or referred to herein as "iGPU"). The GPU 808 can be programmable and can be efficient for parallel workloads. In some examples, the GPU 808 can use an enhanced tensor instruction set. The GPU 808 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having a storage capacity of 512 KB). In some embodiments, the GPU 808 may include at least eight streaming microprocessors. The GPU 808 can use a compute application programming interface (API). Additionally, the GPU 808 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0093] The GPU808 can be power-optimized for the best performance in automotive and embedded use cases. For example, the GPU808 can be fabricated on FinFET (Fin field-effect transistor). However, this is not intended to be limiting, and the GPU808 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores partitioned into multiple blocks. By way of example and not limitation, for instance, 64 PF32 cores and 32 PF64 cores may be partitioned into 4 processing blocks. In such an example, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, 2 mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor may include independent parallel integer and floating-point data paths for efficient execution of workloads having a mix of compute and addressing operations. The streaming microprocessor may include independent thread scheduling capabilities to enable higher fine-grained synchronization and cooperation among parallel threads. The streaming microprocessor may include a combined L1 data cache and shared memory unit to simplify programming while improving performance.

[0094] In some examples, GPU 808 may include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide for a peak memory bandwidth of 900 GB / second. In some examples, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.

[0095] GPU 808 can include unified memory technology that includes access counters to enable more accurate movement of those memory pages to the processor that most frequently accesses the memory pages, thereby improving the efficiency of the memory ranges shared among the processors. In some examples, address translation service (ATS) support may be used to enable the GPU 808 to directly access the CPU 806 page table. In such examples, when the GPU 808 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 806. In response, the CPU 806 can examine its page table for the virtual-to-physical mapping of the address and send the translation back to the GPU 808. As such, unified memory technology can enable a single unified virtual address space for the memory of both the CPU 806 and the GPU 808, thereby simplifying GPU 808 programming and porting of applications to the GPU 808.

[0096] In addition, the GPU 808 may include an access counter that can record the frequency of access of the GPU 808 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses that page.

[0097] The SoC 804 may include any number of caches 812, including those described herein. For example, the cache 812 may include an L3 cache that is available to both the CPU 806 and the GPU 808 (e.g., connected to both the CPU 806 and the GPU 808). The cache 812 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may include more than 4 MB, although smaller cache sizes may be used, depending on the embodiment.

[0098] The SoC 804 may include an arithmetic logic unit (ALU) that can be utilized when executing processing for any of the various tasks or operations of the vehicle 800 (e.g., processing DNN). In addition, the SoC 804 may include a floating point unit (FPU) (or other mass co-processor or numeric co-processor type) for performing mathematical operations within the system. For example, the SoC 104 may include one or more FPUs integrated as execution units within the CPU 806 and / or the GPU 808.

[0099] The SoC 804 may include one or more acceleration devices 814 (e.g., a hardware acceleration device, a software acceleration device, or a combination thereof). For example, the SoC 804 may include a hardware acceleration cluster that may include an optimized hardware acceleration device and / or a large on-chip memory. The large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 808 and to offload a portion of the tasks of the GPU 808 (e.g., to free up more cycles of the GPU 808 for executing other tasks). As an example, the acceleration device 814 may be used for target workloads that are stable enough to be suitable for acceleration (e.g., perception, convolutional neural network (CNN)). In this specification, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).

[0100] Accelerator 814 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that can be configured to provide an additional 10 trillion operations per second for deep learning applications and inferences. The TPU may be an accelerator configured and optimized to execute image processing functions (e.g., those of CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating point operations, as well as for inferences. The design of the DLA can provide more performance per millimeter than a general-purpose GPU and greatly exceed the performance of a CPU. The TPU can execute several functions, including, for example, a single instance convolution function and a post-processor function that support INT8, INT16, and FP16 data types for both features and weights.

[0101] The DLA can rapidly and efficiently execute neural networks, particularly CNNs, with processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection and identification and detection using data from a microphone, CNNs for face recognition and moving vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety related events.

[0102] The DLA can execute any function of the GPU 808, and by using the inference accelerator, for example, a designer can target either the DLA or the GPU 808 for any function. For example, a designer can focus on processing CNNs and floating point operations on the DLA and leave other functions to the GPU 808 and / or other accelerators 814.

[0103] The acceleration device 814 (e.g., a hardware acceleration cluster) may include, or may be referred to herein as, a computer vision acceleration device and a programmable vision accelerator (PVA). The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0104] The RISC core can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC core can use any of several protocols depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0105] DMA can enable the components of the PVA to access the system memory independent of the CPU806. DMA can support any number of features used to provide optimization for the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, DMA can support addressing up to six dimensions or more, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0106] The vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem can operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.

[0107] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently from other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may be able to execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms sequentially on an image or portions of an image. In particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system safety.

[0108] The accelerator 814 (e.g., a hardware acceleration cluster) may include a computer vision network on chip and SRAM to provide high bandwidth, low latency SRAM for the accelerator 814. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and without limitation, 8 field-configurable memory blocks that may be accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high speed access to the PVA and the DLA to the memory. The backbone may include a computer vision network on chip that interconnects the PVA and the DLA to the memory (e.g., using an APB).

[0109] The computer vision network-on-chip may include an interface that determines that both the PVA and the DLA are operable and provide valid signals before any control signal / address / data transmission. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. This type of interface can comply with the ISO26262 or IEC61508 standards, although other standards and protocols may be used.

[0110] In some examples, the SoC804 may include a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and scale of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison against LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing-related operations.

[0111] The acceleration device 814 (e.g., a hardware acceleration device cluster) has various applications for autonomous driving. The PVA may be a programmable vision acceleration device that can be used in extremely important processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are suitable for areas of algorithms that require predictable processing at low power and low latency. In other words, the PVA functions well in normal calculations of semi-high density or high density, even on small data sets, with low latency and low power and a predictable execution time. Therefore, since the PVA is efficient in object detection and integer calculations, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms.

[0112] For example, according to one embodiment of the present technology, the PVA is used to execute computer stereo vision. A semi-global matching-based algorithm may be used in some examples, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require motion estimation / stereo matching on-the-fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). The PVA can execute computer stereo vision functions with inputs from two monocular cameras.

[0113] In some examples, the PVA can be used to execute high-density optical flow. By processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used in the time for flight depth processing, for example, by processing the raw time of flight data to provide the processed time of flight data.

[0114] The DLA can be used to execute any type of network to enhance control and driving safety, including, for example, a neural network that outputs a measurement of the reliability of each object detection. Such reliability values can be interpreted as a probability or as providing the relative "weight" of each detection compared to other detections. This reliability value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a reliability threshold and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the moving vehicle to automatically execute emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. The DLA can execute a neural network that regresses the reliability value. The neural network can receive as its input at least some subset of parameters such as bounding box dimensions, ground plane estimation obtained (e.g., from another subsystem), the azimuth, distance, 3D position estimation of an object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 864 or RADAR sensor 860), and the output of an inertial measurement unit (IMU) sensor 866 that correlates with the moving vehicle 800, among others.

[0115] The SoC 804 can include a data store 816 (e.g., memory). The data store 816 can be the on-chip memory of the SoC 804 and can store neural networks to be executed by the GPU and / or DLA. In some examples, the data store 816 can have a capacity large enough to store multiple instances of the neural network for redundancy and safety. The data store 812 can include an L2 or L3 cache 812. References to the data store 816 can include references to memory related to the PVA, DLA, and / or other accelerators 814 as described herein.

[0116] SoC804 may include one or more processors 810 (e.g., embedded processors). The processor 810 may include a boot and power management processor which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC804 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low power state transitions, management of the SoC804 heat and temperature sensors, and / or management of the SoC804 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC804 can use the ring oscillator to detect the temperature of the CPU806, GPU808, and / or accelerator 814. If the temperature is determined to have exceeded a threshold, the boot and power management processor can enter a temperature fault routine and place the SoC804 in a lower power state and / or put the vehicle 800 into a chauffeur safe stop mode (e.g., safely stop the vehicle 800).

[0117] The processor 810 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor having dedicated RAM.

[0118] The processor 810 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0119] The processor 810 may further include a safety cluster engine that includes a dedicated processor subsystem for handling the safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, support peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In the safety mode, the two or more cores can operate in a lockstep mode and function as a single core with comparison logic for detecting any differences between their operations.

[0120] The processor 810 may further include a real-time camera engine that may include a dedicated processor subsystem for handling real-time camera management.

[0121] The processor 810 may further include a high dynamic range signal processor that may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0122] The processor 810 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the post-video processing functions required by the video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction with the wide-view camera 870, the surround camera 874, and / or the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of a high-level SoC configured to identify in-cabin events and respond appropriately. The in-cabin system can perform lip reading to enable cellular service, make calls, compose emails, change the destination of the moving vehicle, activate or change the infotainment system and settings of the moving vehicle, or provide voice-activated web surfing. Certain functions are only available to the driver when operating in autonomous mode and are otherwise disabled.

[0123] The video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, when motion occurs in the video, the noise reduction reduces the weight of the information provided by adjacent frames and appropriately weights the spatial information. When an image or a portion of an image does not contain motion, the temporal noise reduction performed by the video image synthesizer can reduce the noise in the current image using information from the previous image.

[0124] The video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. The video image synthesizer may further be used for user interface synthesis when the operating system desktop is in use, and the GPU 808 is not required to continuously render new surfaces. Even when the power of the GPU 808 is turned on and 3D rendering is actively performed, the video image synthesizer may be used to offload the GPU 808 to improve performance and responsiveness.

[0125] SoC 804 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that may be used for receiving video and inputs from a camera, and an input / output controller that may be controlled by software and may be used for receiving I / O signals that are not committed to a specific role.

[0126] SoC 804 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 804 may be used to process data from a camera (e.g., connected via gigabit multimedia serial link and Ethernet (R)), sensors (e.g., LIDAR sensor 864, RADAR sensor 860, etc. that may be connected via Ethernet (R)), data from bus 802 (e.g., speed, steering wheel position, etc. of moving vehicle 800), data from GNSS sensor 858 (e.g., connected via Ethernet (R) or CAN bus). SoC 804 may include its own DMA engine and may further include a dedicated high-performance large-capacity storage controller that may be used to free the CPU 806 from routine data management tasks.

[0127] SoC804 may be an in-vehicle platform with a flexible architecture that extends to automation levels 3 to 5, thereby leveraging and efficiently using computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible and reliable driving software stack together with deep learning tools, providing a comprehensive functional safety architecture. SoC804 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when accelerator 814 is coupled to CPU806, GPU808, and data store 816 can provide a fast and efficient platform for level 3 to 5 autonomous vehicles.

[0128] Accordingly, the present technology provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be implemented using a high-level programming language such as the C programming language and executed on a CPU to perform a variety of processing algorithms across a variety of visual data. However, a CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time composite object detection algorithms, which are requirements for in-vehicle ADAS applications and actual level 3 to 5 autonomous vehicles.

[0129] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the techniques described herein enable multiple neural networks to be executed simultaneously and / or sequentially, and enable results to be combined to enable level 3-5 autonomous driving functions. For example, a CNN executed on a DLA or a dGPU (e.g., GPU820) can include text and word recognition that enables a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained on. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module executed on the CPU complex.

[0130] As another example, multiple neural networks can be executed simultaneously as required for level 3, 4, or 5 driving. For example, along with the electric light, a warning sign consisting of "Caution: Flashing light indicates a frozen state" can be interpreted independently or collectively by several neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing light indicates a frozen state" can be interpreted by a second deployed neural network that informs the route planning software of the moving vehicle (preferably executed on the CPU complex) that a frozen state exists when the flashing light is detected. The flashing light can be identified by informing the route planning software of the moving vehicle of the presence (or absence) of the flashing light and operating a third deployed neural network through multiple frames. All three neural networks can be executed simultaneously, such as within the DLA and / or on the GPU808.

[0131] In some examples, the CNN for face recognition and moving vehicle owner identification can use data from the camera sensor to identify the presence of the regular driver and / or owner of the moving vehicle 800. The always-on sensor processing engine can be used to unlock and turn on the moving vehicle when the owner approaches the driver's side door, and, in security mode, to stop the operation of the moving vehicle when the owner leaves the moving vehicle. In this way, the SoC 804 provides security against theft and / or carjacking.

[0132] In another example, the CNN for emergency vehicle detection and identification can use data from the microphone 896 to detect and identify emergency vehicle sirens. In contrast to conventional systems that use a general classifier to detect sirens and manually extract features, the SoC 804 uses a CNN for environmental and urban sound classification, as well as for visual data classification. In a preferred embodiment, the CNN executed on the DLA is trained to identify the relative end speed of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the moving vehicle is operating, as identified by the GNSS sensor 858. Therefore, for example, when operating in Europe, the CNN will attempt to detect European sirens, and when in the United States, the CNN will attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine to slow down the moving vehicle, stop it at the side of the road, park the moving vehicle, and / or idle the moving vehicle, with the assistance of the ultrasonic sensor 862 until the emergency vehicle has passed.

[0133] The mobile vehicle may include a CPU 818 (e.g., an individual CPU, or dCPU) that can be connected to the SoC 804 via a high-speed interconnect (e.g., PCIe). The CPU 818 may include, for example, an X86 processor. The CPU 818 can be used to perform any of a variety of functions, such as mediating potential inconsistencies between ADAS sensors and the SoC 804, and / or monitoring the status and health of the controller 836 and / or the infotainment SoC 830.

[0134] The mobile vehicle 800 may include a GPU 820 (e.g., an individual GPU, or dGPU) that can be connected to the SoC 804 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 820 can provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the mobile vehicle 800.

[0135] The mobile vehicle 800 may further include a network interface 824 that may include one or more wireless antennas 826 (e.g., one or more wireless antennas for different communication protocols such as cellular antennas, Bluetooth antennas, etc.). The network interface 824 may be used to enable wireless connections with the cloud via the Internet (e.g., with server 878 and / or other network devices), with other mobile vehicles, and / or with computing devices (e.g., passenger client devices). To communicate with other mobile vehicles, a direct link may be established between two mobile vehicles and / or an indirect link may be established (e.g., through a network and via the Internet). The direct link may use and provide a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 800 information regarding mobile vehicles in the vicinity of mobile vehicle 800 (e.g., mobile vehicles in front of, beside, and / or behind mobile vehicle 800). This function may be part of the cooperative adaptive cruise control function of mobile vehicle 800.

[0136] The network interface 824 may include a system-on-chip (SoC) that provides modulation and demodulation functions and enables the controller 836 to communicate via a wireless network. The network interface 824 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In some examples, the radio frequency front end functions may be provided by separate chips. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0137] The moving vehicle 800 may further include a data store 828 that may include storage outside the chip (e.g., outside the SoC 804). The data store 828 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least 1 bit of data.

[0138] The vehicle 800 may further include a GNSS sensor 858. The GNSS sensor 858 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) supports mapping, perception, occupancy grid generation, and / or route planning functions. For example, any number of GNSS sensors 858 may be used, including but not limited to GPS with an Ethernet® to serial (RS-232) bridge USB connector.

[0139] The moving vehicle 800 may further include a RADAR sensor 860. The RADAR sensor 860 can be used by the moving vehicle 800 for long-range moving vehicle detection even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 860 can use CAN and / or bus 802 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 860) using access to Ethernet® for accessing raw data. A wide variety of RADAR sensor types may be used. For example, and without limitation, the RADAR sensor 860 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0140] The RADAR sensor 860 can include different configurations, such as long range with a narrow field of view, short range with a wide field of view, and short-range side coverage. In some examples, the long-range RADAR can be used for an adaptive cruise control function. The long-range RADAR system can provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. The RADAR sensor 860 can help distinguish between static and moving objects and can be used by an ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor can include a monostatic multi-modal RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example with six antennas, the four central antennas can create a focused beam pattern designed to record the surroundings of the moving vehicle 800 at high speed while minimizing interference from traffic in adjacent lanes. The other two antennas can widen the field of view and enable rapid detection of moving vehicles entering or leaving the lane of the moving vehicle 800.

[0141] As an example, the mid-range RADAR system can include a range up to 860 m (front) or 80 m (rear), and a field of view up to 42 degrees (front) or 850 degrees (rear). The short-range RADAR system can include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the moving vehicle.

[0142] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.

[0143] The mobile vehicle 800 may further include an ultrasonic sensor 862. The ultrasonic sensor 862, which may be positioned at the front, rear, and / or sides of the mobile vehicle 800, may be used for parking assist and / or for creating and updating an occupancy grid. A wide variety of ultrasonic sensors 862 may be used, and different ultrasonic sensors 862 may be used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensor 862 can operate at a functional safety level of ASIL B.

[0144] The mobile vehicle 800 may include a LIDAR sensor 864. The LIDAR sensor 864 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 864 may also be at a functional safety level of ASIL B. In some examples, the mobile vehicle 800 may include a plurality (e.g., 2, 4, 6, etc.) of LIDAR sensors 864 that can use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).

[0145] In some examples, the LIDAR sensor 864 may have the ability to provide a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 864 may have an accuracy of, for example, 2 cm to 3 cm, support an 800 Mbps Ethernet (registered trademark) connection, and may have an advertised range of about 800 m. In some examples, one or more non-protruding LIDAR sensors 864 may be used. In such examples, the LIDAR sensor 864 may be implemented as a small device that can be incorporated into the front, rear, sides, and / or corners of the mobile vehicle 800. In such examples, the LIDAR sensor 864 may have a range of 200 m even for low-reflectivity objects and can provide up to a 120-degree horizontal and 35-degree vertical field of view. The LIDAR sensor 864 attached to the front may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0146] In some examples, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the area around a moving vehicle up to about 200 m. The flash LIDAR unit includes a receptor that records the laser pulse travel time and the reflected light on each pixel, corresponding in sequence to the range from the moving vehicle to an object. Flash LIDAR can enable high-precision and distortion-free images of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors may be deployed, one on each side of the moving vehicle 800. Available 3D flash LIDAR systems include solid-state 3D steering array LIDAR cameras (e.g., non-scanning LIDAR devices) that have no moving parts other than a blower. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser light in the form of 3D range point clouds and co-recorded intensity data. By using flash LIDAR, and since the LIDAR sensor 864 is a solid-state device with no moving parts, the LIDAR sensor 864 may be less susceptible to the effects of motion blur, vibration, and / or shock.

[0147] The moving vehicle may further include an IMU sensor 866. In some examples, the IMU sensor 866 may be positioned at the center of the rear axle of the moving vehicle 800. The IMU sensor 866 may include, for example, but is not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, in a 6-axis application, the IMU sensor 866 may include an accelerometer and a gyroscope, while in a 9-axis application, the IMU sensor 866 may include an accelerometer, a gyroscope, and a magnetometer.

[0148] In some embodiments, the IMU sensor 866 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 866 may enable the moving vehicle 800 to estimate its heading without the need for input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 866. In some examples, the IMU sensor 866 and the GNSS sensor 858 may be combined in a single integrated unit.

[0149] The moving vehicle may include a microphone 896 disposed within and / or around the moving vehicle 800. The microphone 896 may be used, among other things, for emergency vehicle detection and identification.

[0150] The mobile vehicle may further include any number of camera types, including stereo camera 868, wide view camera 870, infrared camera 872, surround camera 874, long range and / or medium range camera 898, and / or other camera types. The cameras may be used to capture image data around the entire outer surface of the mobile vehicle 800. The type of camera used is determined according to the embodiments and requirements of the mobile vehicle 800, and any combination of camera types may be used to achieve the required coverage around the mobile vehicle 800. Additionally, the number of cameras may vary according to the embodiment. For example, the mobile vehicle may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, and / or another number of cameras. As an example, the cameras may support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet (registered trademark). Each camera is described in more detail herein with reference to FIGS. 8A and 8B.

[0151] The mobile vehicle 800 may further include a vibration sensor 842. The vibration sensor 842 can measure the vibration of components of the mobile vehicle, such as an axle. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 842 are used, the difference in vibration may be used to determine the friction or slipperiness of the road surface (e.g., when the difference in vibration is between a power-driven axle and a freely rotating axle).

[0152] The moving vehicle 800 may include an ADAS system 838. In some examples, the ADAS system 838 may include a SoC. The ADAS system 838 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0153] The ACC system may use a RADAR sensor 860, a LIDAR sensor 864, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the moving vehicle directly in front of the moving vehicle 800 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance holding and advises the moving vehicle 800 to change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0154] CACC can use information from other moving vehicles that can be received via a wireless link from other moving vehicles via network interface 824 and / or wireless antenna 826, or indirectly via a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the immediately preceding moving vehicle (e.g., the moving vehicle immediately in front of moving vehicle 800 in the same lane as moving vehicle 800), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of an I2V information source and a V2V information source. Given information about the moving vehicle ahead of moving vehicle 800, CACC can become more reliable and has the potential to make traffic flow smoother and reduce road congestion.

[0155] The FCW system is designed to warn the driver of a danger so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide an alert in the form of an acoustic, visual alert, vibration, and / or quick brake pulse.

[0156] The AEB system can detect an imminent frontal collision with another moving vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a forward-facing camera and / or RADAR sensor 860 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a danger, it usually first warns the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or imminent collision braking.

[0157] The LDW system provides visual, audible, and / or tactile warnings, such as vibrations of the steering wheel or seat, to warn the driver when the moving vehicle 800 crosses a lane demarcation line. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. The LDW system can use a front-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically connected to driver feedback, such as a display, speaker, and / or vibrating component.

[0158] The LKA system is a modified form of the LDW system. The LKA system provides steering input or brakes to correct the moving vehicle 800 if the moving vehicle 800 begins to veer out of the lane.

[0159] The BSW system detects and warns the driver of a moving vehicle in the blind spot of a motor vehicle. The BSW system can provide visual, audible, and / or tactile warnings to indicate that a merge or lane change is not safe. The system can provide additional warnings when the driver uses the turn indicator. The BSW system can use a rear-facing camera and / or RADAR sensor 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0160] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 800 is backing up. Some RCTW systems include AEB to ensure that vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 860 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0161] Conventional ADAS systems allow the driver to be warned and to determine whether a safety condition actually exists and act accordingly. Thus, conventional ADAS systems, while not usually catastrophic, tend to produce false positive results that can annoy and distract the driver. However, in the case of the autonomous vehicle 800, when the results conflict, the moving vehicle 800 itself must determine whether to accept the results from the primary computer or the secondary computer (e.g., the first controller 836 or the second controller 836). For example, in some embodiments, the ADAS system 838 may be a backup and / or secondary computer for providing perception information to the backup computer rationality module. The backup computer rationality monitor can execute diverse software that is redundant in hardware components to detect malfunctions in perception and dynamic driving tasks. The output from the ADAS system 838 can be provided to the supervisory MCU. When the outputs from the primary computer and the secondary computer conflict, the supervisory MCU needs to determine how to resolve the conflict to ensure safe operation.

[0162] In some examples, the primary computer can be configured to provide a reliability score to the supervisory MCU that indicates the reliability of the primary computer in the selected result. If the reliability score exceeds a threshold, the supervisory MCU can follow the instructions of the primary computer regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold and the primary and secondary computers indicate different (e.g., conflicting) results, the supervisory MCU can mediate between the computers to determine an appropriate result.

[0163] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on outputs from the primary computer and the secondary computer, a state in which the secondary computer provides a false alarm. Accordingly, the neural network within the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot be trusted. For example, when the secondary computer is a RADAR-based FCW system, the neural network within the supervisory MCU can learn when the FCW identifies a metal object, such as a manhole cover or a drain grate, that is not actually dangerous and triggers an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network within the supervisory MCU can learn to ignore the LDW when a person on a bicycle or a pedestrian is present and lane departure is actually the safest operation. In an embodiment that includes a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise components of the SoC804 and / or may be included as components of the SoC804.

[0164] In other examples, the ADAS system 838 may include a secondary computer that performs ADAS functions using conventional rules of computer vision. As such, the secondary computer can use classical computer vision rules (if-then), and the presence of the neural network within the supervisory MCU can improve reliability, safety, and performance. For example, the various implementation forms and intentional non-identities make the overall system more fault-tolerant against faults caused especially by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have a greater confidence that the overall result is correct and the bug in the software or hardware on the primary computer has not caused a critical error.

[0165] In some examples, the output of the ADAS system 838 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 838 indicates a forward collision warning due to an object directly ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network that is trained as described herein and thus reduces the risk of misjudgment.

[0166] The mobile vehicle 800 may further include an infotainment SoC 830 (e.g., an in-vehicle infotainment system (IVI) in a mobile vehicle). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more individual components. The infotainment SoC 830 may include a combination of hardware and software used to provide the mobile vehicle 800 with audio (e.g., music, mobile devices, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calls), network connections (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation systems, rear parking assistance, wireless data systems, fuel level, total travel distance, brake fuel level, oil level, opening / closing of doors, air filter information, and other mobile vehicle-related information). For example, the infotainment SoC 830 may include radio, disk player, navigation system, video player, USB and Bluetooth connections, car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, heads-up display (HUD), HMI display 834, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 830 may be further used to provide information (e.g., visual and / or audible) to the user of the mobile vehicle, such as information from the ADAS system 838, autonomous driving information such as planned mobile vehicle operations, trajectories, surrounding environment information (e.g., intersection information, mobile vehicle information, road information, etc.), and / or other information.

[0167] The Infotainment SoC 830 may include GPU functionality. The Infotainment SoC 830 can communicate with other devices, systems, and / or components of the mobile vehicle 800 via a bus 802 (e.g., CAN bus, Ethernet®, etc.). In some examples, the Infotainment SoC 830 can be coupled to a supervisory MCU such that the GPU of the Infotainment system can perform some self-driving functions in the event that the primary controller 836 (e.g., the primary and / or backup computer of the mobile vehicle 800) fails. In such examples, the Infotainment SoC 830 can put the mobile vehicle 800 into a chauffeur safe stop mode as described herein.

[0168] The mobile vehicle 800 may further include an instrument cluster 832 (e.g., digital dash, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 832 may include a controller and / or a supercomputer (e.g., an individual controller or supercomputer). The instrument cluster 832 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, gear shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting control device, safety system control device, navigation information, etc. In some examples, information may be displayed and / or shared between the Infotainment SoC 830 and the instrument cluster 832. In other words, the instrument cluster 832 may be included as part of the Infotainment SoC 830, and vice versa.

[0169] FIG. 8D is a system diagram of communication between a cloud-based server of FIG. 8A and an exemplary autonomous vehicle 800 according to some embodiments of the present disclosure. The system 876 may include a server 878, a network 890, and a moving vehicle including the moving vehicle 800. The server 878 may include a plurality of GPUs 884(A)-884(H) (collectively referred to herein as GPUs 884), PCIe switches 882(A)-882(H) (collectively referred to herein as PCIe switches 882), and / or CPUs 880(A)-880(B) (collectively referred to herein as CPUs 880). The GPUs 884, CPUs 880, and PCIe switches may be interconnected by high-speed interconnects such as, but not limited to, an NVLink interface 888 and / or a PCIe connection 886 developed by NVIDIA, for example. In some examples, the GPUs 884 are connected via an NVLink and / or an NVSwitch SoC, and the GPUs 884 and the PCIe switches 882 are connected via a PCIe interconnect. Eight GPUs 884, two CPUs 880, and two PCIe switches are shown, but this is not intended to be limiting. Depending on the embodiment, each server 878 may include any number of GPUs 884, CPUs 880, and / or PCIe switches. For example, the server 878 may include eight, sixteen, thirty-two, and / or more GPUs 884, respectively.

[0170] Server 878 can receive, via network 890, from a moving vehicle, image data representing an image indicating an unexpected or changed road condition, such as a recently started road construction. Server 878 can transmit, via network 890, to the moving vehicle, neural network 892, an updated neural network 892, and / or map information 894 including information regarding traffic and road conditions. The update of the map information 894 can include an update of the HD map 822, such as information regarding a construction site, a depression, a detour, a flood, and / or other obstacles. In some examples, the neural network 892, the updated neural network 892, and / or the map information 894 may have resulted from new training and / or experience represented in data received from any number of moving vehicles in the environment and / or based on training performed in a data center (e.g., using server 878 and / or other servers).

[0171] Server 878 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by a moving vehicle and / or (e.g., using a game engine) generated in a simulation. In some examples, the training data is tagged (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., when the neural network does not require supervised learning). The training can be performed according to any one or more classes of machine learning techniques including, but not limited to, for example, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and their variants or combinations. After the machine learning model is trained, the machine learning model can be used by the moving vehicle (e.g., transmitted to the moving vehicle via network 890), and / or the machine learning model can be used by server 878 to remotely monitor the moving vehicle.

[0172] In some examples, server 878 can receive data from a moving vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 878 can include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 884, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 878 can include a deep learning infrastructure that uses only CPU-powered data centers.

[0173] The deep learning infrastructure of server 878 can have the ability of high-speed real-time inference, and use that ability to evaluate and verify the condition of the processors, software, and / or related hardware within moving vehicle 800. For example, the deep learning infrastructure can receive periodic updates from moving vehicle 800, such as a sequence of images and / or objects in which the moving vehicle 800 is located within the sequence of images (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can execute its own neural network to identify objects and compare them with those identified by moving vehicle 800. If the results do not match and the infrastructure concludes that the AI within moving vehicle 800 is not functioning properly, server 878 can send a signal to moving vehicle 800 that infers control, notifies the passengers, and commands the fail-safe computer of moving vehicle 800 to complete a safe parking operation.

[0174] For inference, server 878 can include GPU 884 and one or more programmable inference acceleration devices (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time responsiveness. In other examples, such as when less performance is required, a server powered by a CPU, FPGA, and other processors can be used for inference.

[0175] Exemplary computing device FIG. 9 is a block diagram of an example of a computing device 900 suitable for use in implementing some embodiments of the present disclosure. The computing device 900 may include an interconnect system 902 that indirectly or directly connects the following devices: a memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, an input / output (I / O) port 912, input / output components 914, a power supply 916, one or more presentation components 918 (e.g., a display), and one or more logic units 920. In at least one embodiment, the computing device 900 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). By way of non-limiting example, one or more of the GPUs 908 may include one or more vGPUs, one or more of the CPUs 906 may include one or more vCPUs, and / or one or more of the logic units 920 may include one or more virtual logic units. As such, the computing device 900 may include individual components (e.g., all GPUs dedicated to the computing device 900), virtual components (e.g., a portion of a GPU dedicated to the computing device 900), or a combination thereof.

[0176] Although the various blocks of FIG. 9 are shown as being connected via an interconnect system 902 with lines, this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presentation component 918 such as a display device may be considered an I / O component 914 (e.g., if the display is a touch screen). As another example, the CPU 906 and / or GPU 908 may include memory (e.g., memory 904 may represent a storage device in addition to the memory of the GPU 908, CPU 906, and / or other components). In other words, the computing device of FIG. 9 is merely illustrative. Categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types are all intended to be within the scope of the computing device of FIG. 9 and thus are not distinguished.

[0177] The interconnect system 902 may represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 902 may include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 906 may be directly connected to the memory 904. Further, the CPU 906 may be directly connected to the GPU 908. If there are direct or point-to-point connections between components, the interconnect system 902 may include a PCIe link for making the connections. In these examples, the PCI bus need not be included in the computing device 900.

[0178] The memory 904 may include any of a variety of computer-readable media. A computer-readable media may be any available media that can be accessed by the computing device 900. The computer-readable media may include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, the computer-readable media may comprise computer storage media and communication media.

[0179] A computer storage medium can include both volatile and nonvolatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 904 can store computer-readable instructions such as an operating system (e.g., representing a program and / or program elements). A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 900. In this specification, a computer storage medium does not include the signal itself.

[0180] A computer storage medium can implement computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristic sets or that changes in such a way as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0181] The CPU 906 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to execute one or more of the methods and / or processes described herein. The CPU 906 may each include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores having the ability to process multiple software threads simultaneously. The CPU 906 may include any type of processor and, depending on the type of computing device 900 implemented, may include different types of processors (e.g., a processor having fewer cores for a mobile device and a processor having more cores for a server). For example, depending on the type of computing device 900, the processor may be an Advanced RISC Machines (ARM) processor implemented using reduced instruction set computing (RISC), or an x86 processor implemented using complex instruction set computing (CISC). The computing device 900 may include one or more CPUs 906 within one or more microprocessors or auxiliary coprocessors, such as a compute coprocessor.

[0182] In addition to or instead of CPU 906, GPU 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 900 to execute one or more of the methods and / or processes described herein. One or more of GPU 908 may be an integrated GPU (e.g., may be with one or more of CPU 906), and / or one or more of GPU 908 may be a discrete GPU. In an embodiment, one or more of GPU 908 may be one or more coprocessors of CPU 906. GPU 908 may be used by computing device 900 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 908 may be used for general-purpose computing on GPU (GPGPU). GPU 908 may include hundreds or thousands of cores having the ability to process hundreds or thousands of software threads simultaneously. GPU 908 can generate pixel data for an output image in response to rendering commands (e.g., rendering commands from CPU 906 received via a host interface). GPU 908 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of memory 904. GPU 908 may include two or more GPUs operating in parallel (e.g., via a link). The link can connect directly to the GPUs (e.g., using NVLINK), or can connect the GPUs via a switch (e.g., using NVSwitch). When coupled together, each GPU 908 can generate pixel data or GPGPU data for different portions of the output or different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.

[0183] In addition to and / or instead of CPU 906 and / or GPU 908, logic unit 920 may be configured to execute at least some of the computer-readable instructions to control one or more of computing devices 900 to execute one or more of the methods and / or processes described herein. In an example, CPU 906, GPU 908, and / or logic unit 920 may execute any combination of methods, processes, and / or portions thereof discretely or in parallel. One or more of logic units 920 may be part of and / or integrated with one or more of CPU 906 and / or GPU 908, and / or one or more of logic units 920 may be discrete components with respect to and / or otherwise external to CPU 906 and / or GPU 908. In an example, one or more of logic units 920 may be one or more coprocessors of one or more of CPU 906 and / or one or more of GPU 908.

[0184] Examples of the logic unit 920 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel visual cores (PVC), vision processing units (VPU), data processing units (DPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic logic units (ALU), application-specific integrated circuits (ASIC), floating-point units (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.

[0185] The communication interface 910 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 900 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 910 can include components and functions for enabling communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0186] The I / O port 912 can enable the computing device 900 to be logically connected to other devices, including some of which may be built-in (e.g., integrated) into the computing device 900, such as I / O components 914, presentation components 918, and / or other components. Exemplary I / O components 914 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, and the like. The I / O components 914 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 900 (as will be described in more detail later). The computing device 900 can include a depth camera, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof, for gesture detection and recognition. Additionally, the computing device 900 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 900 to render immersive augmented reality or virtual reality.

[0187] The power supply device 916 can include a hard-wired power supply device, a battery power supply device, or a combination thereof. The power supply device 916 can provide power to the computing device 900 to enable the components of the computing device 900 to operate.

[0188] The presentation component 918 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 918 can receive data from other components (e.g., GPU 908, CPU 906, etc.) and output the data (e.g., as an image, video, sound, etc.).

[0189] Exemplary data center FIG. 10 shows an exemplary data center 1000 that may be used in at least one embodiment of the present disclosure. The data center 1000 may include a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and / or an application layer 1040.

[0190] As shown in FIG. 10, the data center infrastructure layer 1010 may include a resource orchestrator 1012, grouped computing resources 1014, and node computing resources (“node C.R.”) 1016(1) to 1016(N), where “N” represents any integer of natural numbers. In at least one embodiment, the node C.R. 1016(1) to 1016(N) may include any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors or graphics processing units (GPUs), etc.), memory devices (e.g., dynamic random access memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and / or cooling modules, etc., but are not limited thereto. In some embodiments, one or more of the node C.R. 1016(1) to 1016(N) may correspond to a server having one or more of the aforementioned computing resources. Additionally, in some embodiments, the node C.R. 1016(1) to 1016(N) may include one or more virtual components, such as vGPUs, vCPUs, and / or the like, and / or one or more of the node C.R. 1016(1) to 1016(N) may correspond to a virtual machine (VM).

[0191] In at least one embodiment, the grouped computing resources 1014 can include a separate group of node C.R.s 1016 housed within one or more racks (not shown), or multiple racks housed in data centers at various geographical locations (also not shown). Separate groups of node C.R.s 1016 within the grouped computing resources 1014 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s 1016 that include CPUs, GPUs, and / or other processors can be grouped within one or more racks to provide computing resources for supporting one or more workloads. One or more racks can also include any number of power modules, cooling modules, and / or network switches in any combination.

[0192] The resource orchestrator 1022 can configure or otherwise control one or more node C.R.s 1016(1)-1016(N) and / or the grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1022 can include a software design infrastructure ("SDI") management entity of the data center 1000. The resource orchestrator 1022 can include hardware, software, or some combination thereof.

[0193] In at least one embodiment, as shown in FIG. 10, the framework layer 1020 may include a job scheduler 1032, a configuration manager 1034, a resource manager 1036, and / or a distributed file system 1038. The framework layer 1020 may include a framework to support software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. The software 1032 or the application 1042 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1020 may be of a type of free and open source software web application framework, such as Apache Spark (trademark) (hereinafter “Spark”), which may use the distributed file system 1038 for large-scale data processing (e.g., “big data”), but is not limited thereto. In at least one embodiment, the job scheduler 1032 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1000. The configuration manager 1034 may have the ability to configure different layers, such as the software layer 1030 and the framework layer 1020 including Spark and the distributed file system 1038 to support large-scale data processing. The resource manager 1036 may have the ability to manage the mapped or assigned clustered or grouped computing resources for the support of the distributed file system 1038 and the job scheduler 1032. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 1014 in the data center infrastructure layer 1010. The resource manager 1036 can coordinate with the resource orchestrator 1012 to manage these mapped or assigned computing resources.

[0194] In at least one embodiment, the software 1032 included in the software layer 1030 may include at least a portion of nodes C.R. 1016(1) to 1016(N), the grouped computing resources 1014, and / or software used by the distributed file system 1038 of the framework layer 1020. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.

[0195] In at least one embodiment, the application 1042 included in the application layer 1040 may include at least a portion of nodes C.R. 1016(1) to 1016(N), the grouped computing resources 1014, and / or one or more types of applications used by the distributed file system 1038 of the framework layer 1020. The one or more types of applications may include, but are not limited to, machine learning applications including any number of genomics applications, cognitive computing, and training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0196] In at least one embodiment, any one of the configuration manager 1034, the resource manager 1036, and the resource orchestrator 1012 can implement any number and type of self-rewriting actions based on any amount and type of data obtained in any technically possible manner. The self-rewriting actions can free the data center operator of the data center 1000 by making potentially bad configuration decisions and possibly avoiding underutilized and / or poorly performing portions of the data center.

[0197] The data center 1000 may include tools, services, software, or other resources for training one or more machine learning models or predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, the machine learning model may be trained by calculating weight parameters by a neural network architecture that uses the software and / or computing resources described above with respect to the data center 1000. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1000 by using weight parameters calculated via one or more training techniques, which are not limited to, for example, those described herein.

[0198] In at least one embodiment, the data center 1000 can use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) for training and / or performing inferences using the aforementioned resources. Further, the aforementioned one or more software and / or hardware resources may be configured as services that enable a user to train or perform inferences on information, such as image recognition, speech recognition, or other artificial intelligence services.

[0199] Exemplary network environment A network environment suitable for use in implementing embodiments of the present disclosure may include one or more client devices, servers, network-attached storage (NAS), other backend devices, and / or other device types. The client devices, servers, and / or other device types (e.g., each device) may be implemented as one or more instances of the computing device 900 of FIG. 9, and for example, each device may include similar components, features, and / or functionality of the computing device 900. Additionally, when a backend device (e.g., a server, NAS, etc.) is implemented, the backend device may be included as part of the data center 1000, an example of which is described in more detail herein with respect to FIG. 10.

[0200] The components of the network environment may communicate with each other via a network that may be wired, wireless, or both. The network may include multiple networks, or a network of networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the public switched telephone network (PSTN), and / or one or more private networks. When the network includes a wireless telecommunications network, components, such as base stations, communication towers, or access points (and other components), may provide a wireless connection.

[0201] A compatible network environment may include one or more peer-to-peer network environments (in which case, a server may or may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to a server may be implemented on any number of client devices.

[0202] In at least one embodiment, the network environment may include one or more cloud-based network environments, distributed computing environments, combinations thereof, and the like. A cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of a server, which may include one or more core network servers and / or edge servers. The framework layer may include a framework to support software of the software layer and / or one or more applications of the application layer. The software or application may each include web-based service software or an application. In an embodiment, one or more of the client devices may use web-based service software or an application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be of a type of free and open source software web application framework, which may use a distributed file system for large-scale data processing (e.g., "big data"), but is not limited thereto.

[0203] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed to multiple locations from a central or core server (such as one or more data centers that may be distributed across a state, region, country, the world, etc.). When the connection to a user (such as a client device) is relatively close to an edge server, the core server may delegate at least a portion of the functionality to the edge server. The cloud-based network environment may be private (such as restricted to a single organization), public (such as available to multiple organizations), and / or a combination thereof (such as a hybrid cloud environment).

[0204] A client device may include at least some of the components, features, and functionality of the exemplary computing device 900 described herein with respect to FIG. 9. By way of example, and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0205] The present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a mobile information terminal or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, household appliances, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0206] As used herein, a description of "and / or" with respect to two or more elements should be construed to mean either only one element, or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0207] The subject matter of this disclosure is described with specificity to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors intend that the claimed subject matter may be practiced otherwise, including in combination with other current or future technologies, in different steps or in a combination of steps similar to those described in this document. Further, the terms "step" and / or "block" may be used herein to imply different elements of a method being used, but these terms should not be construed as implying any particular order among the various steps disclosed herein except where the order of individual steps is explicitly recited and so described.

Claims

1. One or more circuits for receiving first data representing a plurality of outputs of a plurality of deep neural networks (DNNs), wherein at least one of the plurality of outputs has a respective field of view different from the field of view corresponding to one or more other sensors of a plurality of sensors of an autonomous machine, and for calculating second data representing a fusion of the plurality of outputs using a fusion DNN and based at least in part on the first data, and for performing one or more operations using the autonomous machine based at least in part on the second data. A processor comprising:

2. The processor of claim 1, wherein calculating the second data is further based at least in part on third data representing at least one probability distribution function corresponding to at least one point of at least one of the plurality of outputs, the at least one point corresponding to a detected object, and the at least one probability distribution function corresponding to one or more potential positions of the detected object.

3. The processor of claim 1, wherein calculating the second data is further based at least in part on third data representing one or more velocity representations including an encoded value corresponding to at least one of a velocity in the x direction or a velocity in the y direction.

4. The processor of claim 1, wherein calculating the second data is further based at least in part on third data representing one or more representations corresponding to at least one of an object instance or an object appearance determined using the plurality of outputs.

5. The processor of claim 1, wherein each of the plurality of outputs includes a rasterized image representing one or more objects.

6. The processor of claim 5, wherein the one or more objects include at least one of a vehicle, a pedestrian, a person on a bicycle, a motor vehicle driver, a lane marker, a road boundary marker, a free space boundary, or a waiting line.

7. The first output of the plurality of outputs corresponds to a first field of view, the second output of the plurality of outputs corresponds to a second field of view different from the first field of view, and the fusion of the plurality of outputs corresponds to both the first field of view and the second field of view. The processor according to claim 1.

8. The processor according to claim 7, wherein the first field of view and the second field of view at least partially overlap.

9. The first data further represents one or more additional outputs generated using a LiDAR sensor, a RADAR sensor, or an ultrasonic sensor, and the one or more additional outputs are generated using another DNN or without using another DNN. The processor according to claim 1.

10. The first output of the plurality of outputs includes a first representation of an object, the second output of the plurality of outputs includes a second representation of the object, and the fusion of the plurality of outputs includes a fused representation of the object. The processor according to claim 1.

11. The processor is included in at least one of a control system of an autonomous machine, a cognitive system of an autonomous machine, a system for performing a simulation operation, a system for performing a deep learning operation, a system implemented using an edge device, a system implemented using a robot, a system incorporating one or more virtual machines (VMs), a system at least partially implemented in a data center, or a system at least partially implemented using cloud computing resources. The processor according to claim 1.

12. One or more processing devices and one or more memory units storing instructions that cause the one or more processing devices to perform operations when executed by the one or more processing devices, the operations including receiving first data representing at least a first rasterized image generated at least in part based on first sensor data generated using a first deep neural network (DNN) and a first sensor, the first rasterized image including at least a first object; receiving second data representing at least a second rasterized image generated at least in part based on second sensor data generated using a second deep neural network (DNN) and a second sensor, the second rasterized image including at least a second object; calculating third data representing a fused rasterized image including both the first object and the second object using a fusion DNN and based at least in part on the first data and the second data; and performing one or more operations using an autonomous machine based at least in part on the third data. Claim 13 The system of claim 12, wherein the first sensor and the second sensor include one of an image sensor, a LiDAR sensor, a RADAR sensor, or an ultrasonic sensor. Claim 14 The system of claim 12, wherein the first sensor and the second sensor include at least partially overlapping fields of view, the first rasterized image includes a first representation of a third object, the second rasterized image includes a second representation of the third object, and the fused rasterized image includes a fused representation of the third object. Claim 15 The operation further includes receiving fourth data representing at least one probability distribution function corresponding to at least one pixel of at least one of the first rasterized image or the second rasterized image, wherein the at least one pixel corresponds to at least one detected object of the first object or the second object, the at least one probability distribution function corresponds to one or more potential positions of the detected object, and calculating the third data is further at least partially based on the fourth data, the system according to claim 12.

16. The operation further includes receiving fourth data representing one or more velocity representations including encoded values corresponding to at least one of a velocity in the x direction or a velocity in the y direction, and calculating the third data is further at least partially based on the fourth data, the system according to claim 12.

17. The system is included in at least one of a control system of an autonomous machine, a cognitive system of an autonomous machine, a system for performing a simulation operation, a system for performing a deep learning operation, a system implemented using an edge device, a system implemented using a robot, a system incorporating one or more virtual machines (VMs), a system at least partially implemented in a data center, or a system at least partially implemented using cloud computing resources, the system according to claim 12.

18. Receiving first data representing at least a first rasterized image generated at least in part based on first sensor data generated using a first sensor of a first type, wherein the first rasterized image includes at least a first object; receiving second data representing at least a second rasterized image generated at least in part based on second sensor data generated using a second sensor of a second type different from the first type, wherein the second rasterized image includes at least a second object; calculating third data representing a fused rasterized image including both the first object and the second object using a fused deep neural network (DNN) and based at least in part on the first data and the second data; and performing one or more operations using an autonomous machine and based at least in part on the third data. **Claim 19** The method of claim 18, wherein the first type and the second type include one of an image sensor, a LiDAR sensor, a RADAR sensor, or an ultrasonic sensor. **Claim 20** The method of claim 18, wherein the first type includes an image sensor, the first rasterized image is generated using a deep neural network (DNN), the second type includes one of a LiDAR sensor, a RADAR sensor, or an ultrasonic sensor, and the second rasterized image is generated without using a DNN.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Ground truth data generation for deep neural network perception in autonomous driving applications

    US20220277193A1

  • Information processing device, information processing method, program, mobile body control device, and mobile body

    WO2020116195A1