Prediction of Time Information in Autonomous Machine Applications

By using sequential deep neural networks combined with sensors and image data, predicting time information in autonomous machine applications, solving the problem of prior art sensitivity to non-rigid motion and observation direction, and achieving high-accurate time information prediction.

CN111695717BActive Publication Date: 2025-06-20NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010009841.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-07-17
Filing Date
2020-01-06
Publication Date
2025-06-20
Estimated Expiration
2040-05-01

AI Technical Summary

Technical Problem

The prior art is sensitive to non-rigid motion and observation direction when predicting time information, it is difficult to capture or create a moving understanding of objects relative to a static environment, and relies on a fixed time baseline, resulting in inaccurate predictions.

Method used

Sequential deep neural networks (DNNs), such as recurrent neural networks (RNNs), are used to combine sensor data and image data, and generate automatic ground real-life data through cross-sensor fusion, and train DNNs to predict collision time (TTC), 2D and 3D object movement without sensor data input.

Benefits of technology

It realizes the generation of accurate time information predictions in deployment, improves the accuracy and computing efficiency of the system, and can handle object movements in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111695717B_ABST
    Figure CN111695717B_ABST
Patent Text Reader

Abstract

In various examples, ground truth data generated by correlating sensor data with image data representing an image sequence (e.g., via cross-sensor fusion) can be used to train a sequential deep neural network (DNN). In deployment, the sequential DNN can leverage sensor correlations to compute various predictions using only the image data. The predictions can include the speed of an object in the field of view of the ego vehicle in world space, the current and future positions of the object in image space, and / or the time to collision (TTC) between the object and the ego vehicle. These predictions can be used as part of a perception system to understand and respond to the current physical environment of the ego vehicle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims priority to U.S. Provisional Application No. 62 / 819,412, filed on Mar. 15, 2019, which is hereby incorporated by reference in its entirety. BACKGROUND OF THE DISCLOSURE

[0003] The ability to accurately predict time information is crucial for generating a world model for obstacle analysis and for control determination for an assisted vehicle. For example, time information prediction can be generated as part of the perception layer of an autonomous driving software stack and can be used by a world model manager - in addition to mapping, localization, and other perception functions - to help the vehicle understand the environment. Time information prediction can also be a fundamental input to advanced driver assistance systems (ADAS), such as automatic emergency braking (AEB) and adaptive cruise control (ACC).

[0004] Some conventional systems have used conventional computer vision algorithms (such as tracking algorithms) to estimate changes in local appearance to extract local motion information that can be used to generate a time information estimate. However, as a non-limiting example, these computer vision algorithms are sensitive to non-rigid motion and viewing direction (e.g., rotation of an obstacle), which can result in poor predictions - especially for object types such as pedestrians. The background of an object or obstacle in the environment also cannot be considered using conventional computer vision algorithms, and thus mainstream algorithms generally cannot capture or create an understanding of how an object or obstacle moves relative to its static environment. In addition, these algorithms are implemented using simple internal mechanisms that rely on a fixed time baseline to calculate time information - thus not separately considering different objects, scenarios, and / or conditions in the environment and resulting in less useful and accurate predictions. SUMMARY OF THE DISCLOSURE

[0005] Embodiments of the present disclosure relate to time information prediction for perception in autonomous machine applications. Systems and methods are disclosed that relate to using a sequential deep neural network (DNN) (e.g., a recurrent neural network (RNN)) to predict time to collision (TTC), two-dimensional (2D) object motion, and three-dimensional (3D) object motion used in autonomous machine applications.

[0006] Compared with traditional systems (such as those described above), the system of the present disclosure uses the correlation between sensor data and image data to generate ground truth data for training a sequential DNN to predict - in deployment - TTC, 2D object motion, and 3D object motion, without sensor data as input. As a result, the sequential DNN can generate predictions only from images, without the need for dense optical flow or motion of the vehicle during deployment to generate predictions - as required by traditional systems. To generate accurate results during deployment without sensor data input, the system of the present disclosure utilizes an automatic ground truth data generation pipeline using cross-sensor fusion, dedicated preprocessing steps for filtering out inconsistent ground truth or training data, an increase in the time of training data to increase the robustness of the training dataset, and postprocessing steps for smoothing predictions.

[0007] Additionally, depending on the implementation of the sequential DNN, the system implements stateless or stateful training and inference methods, which improve the accuracy and computational efficiency of the system. For example, a stateless inference method is described, which utilizes previous computations or performs parallel computations to improve the runtime efficiency and accuracy of the predictions of the sequential DNN. As another example, a stateful training method is described, which reduces the likelihood of overfitting to the training dataset by randomizing the sequence length, stride, and / or orientation of the frames in a mini-batch of the training dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present system and method for predicting temporal information for perception in autonomous vehicle applications are described in detail below with reference to the accompanying drawings, in which:

[0009] Figure 1 is a data flow diagram showing an example process for making temporal predictions of objects in an environment according to some embodiments of the present disclosure;

[0010] Figures 2 - 3 is an example visualization of temporal information prediction of a sequential deep neural network (DNN) according to some embodiments of the present disclosure;

[0011] Figure 4 is a flowchart showing a method 400 for making temporal information predictions using a sequential DNN according to some embodiments of the present disclosure;

[0012] Figure 5A is an example data flow diagram for stateless inference according to some embodiments of the present disclosure;

[0013] Figure 5B is a flowchart showing a method for stateless inference using a buffer according to some embodiments of the present disclosure;

[0014] Figure 6AAn example data flow diagram for stateless inference according to some embodiments of the present disclosure;

[0015] Figure 6B A flowchart showing a method of stateless inference using parallel processing according to some embodiments of the present disclosure;

[0016] Figure 7 A data flow diagram showing an example process for training a sequential DNN to predict temporal information about an object in an environment according to some embodiments of the present disclosure;

[0017] Figure 8 An example visualization of a cross-sensor fusion technique for generating ground truth according to some embodiments of the present disclosure;

[0018] Figure 9 A flowchart showing a method of generating automatic ground truth using cross-sensor fusion according to some embodiments of the present disclosure;

[0019] Figure 10 An example illustration of temporal augmentation according to some embodiments of the present disclosure;

[0020] Figure 11 A flowchart showing a method for temporal augmentation according to some embodiments of the present disclosure;

[0021] Figure 12 An example illustration of a training dataset for state training according to some embodiments of the present disclosure;

[0022] Figure 13 A flowchart showing a method for state training according to some embodiments of the present disclosure;

[0023] Figure 14A An illustration of an exemplary autonomous vehicle according to some embodiments of the present disclosure;

[0024] Figure 14B According to some embodiments of the present disclosure Figure 14A An example of the camera positions and fields of view of an exemplary autonomous vehicle;

[0025] Figure 14C According to some embodiments of the present disclosure Figure 14A A block diagram of an exemplary system architecture of an exemplary autonomous vehicle;

[0026] Figure 14D According to some embodiments of the present disclosure for a cloud-based server and Figure 14A A system diagram of the communication between an exemplary autonomous vehicle; and

[0027] Figure 15is a block diagram of an exemplary computing device suitable for implementing some embodiments of the present disclosure. Detailed Description

[0028] The systems and methods disclosed herein relate to predicting temporal information for perception in autonomous machine applications. Although the present disclosure may be described with reference to an exemplary autonomous vehicle 1400 (optionally referred to herein as "vehicle 1400" or "autonomous vehicle 1400"), which is described by way of example Figures 14A - 14D this is not limiting. For example, the systems and methods described herein may be used by non-autonomous vehicles, semi-autonomous vehicles (e.g., in one or more adaptive driver assistance systems (ADAS)), robots, vans, off-road vehicles, aircraft, boats, shuttles, emergency response vehicles, motorcycles, electric or motorized bicycles, airplanes, engineering vehicles, underwater vehicles, drones, and / or other vehicle types, but are not limited thereto. Additionally, although the present disclosure may be described with reference to autonomous driving, this is not restrictive. For example, the systems and methods described herein may be used in robotics, aviation systems, marine systems, and / or other technical fields, such as for perception, world model management, path planning, obstacle avoidance, and / or other processes.

[0029] Temporal Prediction Using Sequential Deep Neural Networks

[0030] Image data associated with sensor data (e.g., LIDAR data, RADAR data, etc.) can be used to train the deep neural network (DNN) of the present disclosure, e.g., by using cross-sensor fusion techniques. The DNN may include a recurrent neural network (RNN), which may have one or more long short-term memory (LSTM) layers and / or gated recurrent unit (GRU) layers. As a result, a combination of image data and sensor data can be used to train the DNN to predict temporal information, such as time to collision (TTC), two-dimensional (2D) motion, and / or three-dimensional (3D) motion corresponding to objects in the environment. Compared to traditional systems and as a result of the training methods described herein, the DNN can be configured to generate accurate predictions of temporal information in deployment using only image data (e.g., the RNN in deployment can learn to generate temporal information with the accuracy of supplementary sensor data without the need for supplementary sensor data).

[0031] Image data for training a DNN can be represented as a sequence of images, and the sequence of images can be applied to the DNN. The DNN can then predict data as output that represents the current location of an object and one or more future locations of the object (e.g., as bounding box locations), and / or data representing the speed of the object. A ratio change between the bounding shapes at the current and future locations of the object can be used to generate an urgency prediction corresponding to the reciprocal of the TTC (e.g., 1 / TTC). By training the DNN to predict the reciprocal of the TTC, the output pushed to infinity (e.g., where the relative motion between the ego vehicle and another vehicle is zero) can be estimated to be close to zero (e.g., effectively removed), resulting in more efficient processing and availability of the predicted values. Additionally, during training, using the reciprocal of the TTC is more beneficial for gradient descent-based optimization methods. The image data can include timestamps (e.g., corresponding to the frame rate), and in addition to the ratio change, the system can also use the timestamps (or frame rate). For example, the timestamps (or frame rate) and the ratio change can be used to generate an urgency or TTC (e.g., TTC can be expressed or represented as being equal to or approximately equal to Δt / Δs, where t is the time between capturing two frames and s is the ratio change between the bounding shapes of the object in the two frames). In some examples, the DNN can implicitly consider Δt such that the output of the DNN can be used directly with minimal post-processing (e.g., it is not necessary to use the frame rate in post-processing calculations to generate a final prediction of the TTC). In other examples, the DNN can be trained to output a Δs value, and post-processing can be used to use the value of Δt to determine the TTC (e.g., where the DNN does not implicitly consider Δt).

[0032] In some examples, in addition to the size information (e.g., the lengths of the sides of the bounding shape), the current and future locations can be output by the DNN as the location of the origin or center of the bounding shape. In other examples, in addition to the size information, the current location can be output as the location of the origin or center of the bounding shape, and the future location can be output as a translation of the origin or center of the current bounding shape, in addition to the ratio change value of the bounding shape size. Speed information for an object can be calculated by looking at the timestamps on the images in the sequence of images and the change in the location of the object.

[0033] Once the output is predicted by the DNN, the output can be decoded. For example, in the case of calculating 1 / TTC, the 1 / TTC value can be converted to TTC. In the case of predicting the current and future positions, the output can be used to generate a boundary shape (e.g., associating the boundary shape with corresponding pixel positions in the image). Additionally, in some examples, temporal smoothing can be applied to one or more outputs. Temporal smoothing can include a state estimator (such as a Kalman filter). Temporal smoothing can be applied in image space, or can be applied to the 3D world space relative to the ego vehicle, to the 3D world space relative to a fixed origin in the world space, or to the bird's-eye view in the 2D world space.

[0034] Now refer to Figure 1 , Figure 1 FIG. is a data flow diagram showing an example process 100 for temporal prediction of objects in an environment according to some embodiments of the present disclosure. At a high level, process 100 can include one or more sequential deep neural networks (DNNs) 104 that receive one or more inputs, such as image data 102, and generate one or more outputs 110, such as time to collision (TTC) 112, object two-dimensional (2D) motion 114, and / or object three-dimensional (3D) motion 116. The image data 102 can include image data generated by one or more cameras of an autonomous vehicle (e.g., vehicle 1400, as described herein at least with respect to Figures 14A - 14D ). In some embodiments, in addition to or as an alternative to the image data 102, the sequential DNN 104 can use LIDAR data from one or more LIDAR sensors 1464, RADAR data from one or more RADAR sensors 1460, SONAR data from one or more SONAR sensors, data from one or more ultrasonic sensors 1462, audio data from one or more microphones 1496, etc.

[0035] The sequential DNN 104 can be trained to generate the TTC 112, object 2D motion 114, and / or object 3D motion 116. These outputs 110 can be used by a world model manager (e.g., the world model manager layer of an autonomous driving software stack), a perception component (e.g., the perception layer of an autonomous driving software stack), a planning component (e.g., the planning layer of an autonomous driving software stack), an obstacle avoidance component (e.g., the obstacle or impact avoidance layer of an autonomous driving software stack), and / or other components or layers of the autonomous driving software stack to assist the autonomous vehicle 1400 in performing one or more operations in the environment (e.g., world model management, obstacle avoidance, path planning, etc.).

[0036] In some embodiments, the image data 102 may include data representing the fields of view of one or more cameras of the vehicle (such as the stereo camera 1468, the wide field of view camera 1470 (e.g., a fish-eye camera), the infrared camera 1472, the surround camera 1474 (e.g., a 360-degree camera), the long-range and / or mid-range camera 1498, and / or Figures 14A - 14D ) of other camera types of the autonomous vehicle 1400). In some non-limiting examples, one or more forward-facing cameras and / or one or more side-facing cameras may be used. The image data 102 captured from the forward-facing and / or side-facing perspectives can be used for perception during navigation - for example, within a lane, by changing lanes, by turning, by passing through an intersection, etc. - because the forward-facing camera and / or side-facing camera may include a field of view that includes the current driving lane of the vehicle 1400, the adjacent lanes of the vehicle 1400, the oncoming lanes of other objects, the pedestrian area, and / or the boundaries of the driving surface. In some examples, more than one camera or other sensors (such as LIDAR sensors, RADAR sensors, etc.) may be used to combine multiple fields of view (such as the fields of view of the long-range camera 1498, the surround camera 1474, the forward-facing stereo camera 1468, and / or Figure 14B the forward-facing wide field of view camera 1470).

[0037] In some examples, the image data 102 may be captured in one format (such as RCCB, RCCC, RBGC, etc.) and then converted (e.g., during the preprocessing of the image data) to another format. In some other examples, the image data may be provided as input to an image data preprocessor (not shown) to generate preprocessed image data. Many types of images or formats can be used as input, such as compressed images (such as Joint Photographic Experts Group (JPEG), Red Green Blue (RGB), or Luminance / Chroma (YUV) formats), compressed images as frames from compressed video formats (e.g., H.264 / Advanced Video Coding (AVC) or H.265 / High Efficiency Video Coding (HEVC)), and raw images (e.g., from Red Clear Blue (RCCB), Red Clear (RCCC), or other types of imaging sensors). In some examples, different formats and / or resolutions may be used for training the sequential DNN 104 rather than for inference (e.g., during the deployment of the sequential DNN in the autonomous vehicle 1400).

[0038] The image data pre-processor may use the image data 102 representing one or more images (or other data representations) and load the image data into memory in the form of a multi-dimensional array / matrix (or referred to as a tensor, or more specifically, an input tensor in some examples). The array size may be calculated and / or represented as W x H x C, where W represents the image width of the pixels, H represents the height of the pixels, and C represents the number of color channels. Without loss of generality, other types and orders of the input image components are possible. Additionally, the batch size B may be used as a dimension (e.g., an additional fourth dimension) when using batch processing. Batch processing may be used for training and / or inference. Thus, the input tensor may represent an array of dimensions W x H x C x B. Any ordering of the dimensions is possible, which may depend on the specific hardware and software used to implement the sensor data pre-processor. The ordering may be chosen to maximize the training and / or inference performance of the sequential DNN 104.

[0039] In some embodiments, the image data pre-processor may employ a pre-processing image pipeline to process the raw images acquired by the camera and included in the image data 102 to produce pre-processed image data, which may represent the input image to the input layer of the sequential DNN 104. Examples of suitable pre-processing image pipelines may use the raw RCCB Bayer (e.g., 1-channel) type of the image from the camera and convert the image into an RCB (e.g., 3-channel) planar image stored in a fixed precision (e.g., 16 bits per channel) format. The pre-processing image pipeline may include decompression, noise reduction, demosaicing, white balancing, histogram calculation, and / or adaptive global tone mapping (e.g., in that order or in an alternative order).

[0040] In the case where the image data pre-processor employs noise reduction, it may include bilateral denoising in the Bayer domain. In the case where the image data pre-processor employs demosaicing, it may include bilinear interpolation. In the case where the image data pre-processor employs histogram calculation, it may involve calculating the histogram of the C channels and, in some examples, may be combined with decompression or noise reduction. In the case where the image data pre-processor employs adaptive global tone mapping, it may include performing an adaptive gamma-log transform. This may include calculating the histogram, obtaining the middle tone level, and / or estimating the maximum luminance using the middle tone level.

[0041] The sequential DNN 104 can include a recurrent neural network (RNN), a gated recurrent unit (GRU) DNN, a long short-term memory (LSTM) DNN, and / or other types of sequential DNNs 104 (e.g., a DNN that uses sequential data (such as an image sequence) as input). Although examples are described herein with respect to using a sequential DNN, particularly an RNN, LSTM, and / or GRU as the sequential DNN 104, this is not restrictive. By way of example and not limitation, the sequential DNN 104 can more broadly include any type of machine learning model, such as those using linear regression, logistic regression, decision trees, support vector machines (SVMs), Bayes, k-nearest neighbor (Knn), K-means clustering, random forests, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptrons, LSTMs, GRUs, Hopfield, Boltzmann, deep belief, deconvolutional, generative adversarial, liquid state machines, etc.) machine learning models, and / or other types of machine learning models.

[0042] The sequential DNN 104 can generate or predict output data that can undergo decoding 106 and, in some embodiments, can undergo post-processing 108 (e.g., temporal smoothing) to ultimately generate an output 110. TTC 112 information can be output for each detected object in the image data 102 as the value of 1 / TTC expressed in units of 1 / second. For example, with respect to Figure 2In visualization 200, sequential RNN 104 can predict a value of 3.07 1 / s for bus 202, a value of 0.03 1 / s for vehicle 206, and a value of 1.98 1 / s for pedestrian 210. Sequential DNN 104 can also predict the boundary shapes of each object in the image, such as boundary shape 204 for bus 202, boundary shape 208 for vehicle 206, and boundary shape 212 for pedestrian 210. In visualization 200, the boundary shapes can be represented by boundaries with different display variations to indicate urgency (e.g., for the crossing of vehicle 1400 and the upcoming path of the object). For example, boundary shape 204 can be indicated by a first dashed line type to indicate medium urgency, boundary shape 212 can be indicated by a second dashed line type to indicate high urgency, and boundary shape 208 can be represented by a solid line to indicate low urgency. Although shown as dashed and solid lines in visualization 200, this is not restrictive. In some embodiments, urgency can be represented by a display using different colors, patterns, display frequencies, shapes, line types, and / or other visual indicators that can be used to indicate relative urgency. As a result of the output of sequential RNN 104 including the value of 1 / TTC, decoding can include retrieving the value and converting the value to TTC 112 (e.g., by dividing 1 by the value). In some embodiments, the value of 1 / TTC or TTC 112 can undergo post-processing, as described in more detail herein, e.g., to smooth the value over time. According to an embodiment, vehicle 1400 can use the value of 1 / TTC or can use the value of TTC 112 to perform one or more operations, such as but not limited to world model management, path planning, and / or obstacle avoidance.

[0043] Although TTC includes "collision", this is not meant literally. For example, TTC can be an indicator of when vehicle 1400 (e.g., ego vehicle) and an object in the environment will meet along a horizontal plane (e.g., where the object in the environment will meet a plane that is perpendicular to the ground and parallel to, e.g., the rear axle of vehicle 1400), but should not be interpreted as an indication of an actual collision. For example, referring to Figure 2 , bus 202 can be traveling in the opposite direction of vehicle 1400 (e.g., the camera of the ego vehicle captures the image used in visualization 200) on the opposite side of the road. Thus, as long as vehicle 1400 and bus 202 both stay in their respective lanes, no collision will occur. Similarly, with respect to vehicle 1400 and vehicle 206, as long as both vehicles stay in their lanes, no collision will occur - just an overlap or crossing of the two vehicles of a plane that extends horizontally between the two vehicles (e.g., with respect to Figure 2 , from the right side of the page to the left side of the page).

[0044] For each detected object in each image, the object 2D motion 114 can output the position of the origin or center of the bounding shape in the 2D image space and dimensions of the bounding shape (e.g., the length and width of the bounding box measured in 2D pixel distances) that is the current position of the object. In some examples, to determine one or more future positions of the object, the sequential DNN 104 can output translation values (e.g., x and y translations in pixel distances) for the origin and changes in the scaled dimensions of the dimensions in the 2D image space. In this way, the values of the current bounding shape and the translation values and the changes in the scaled dimensions can be used to determine (e.g., during decoding 106) the position in the 2D image space and the dimensions in the 2D image space of the future bounding shape. In other examples, to determine one or more future positions of the object, the future bounding shape origin and the future dimensions of the future bounding shape can be output by the sequential DNN 104. In either example, decoding 106 can be performed on the output of the sequential DNN 104 to determine the positions in the 2D image space of the current and future bounding shapes of each object in order to track or determine the object 2D motion 114. As described herein, in some examples, the values output by the sequential DNN 104 and / or the decoded values can be post-processed, such as smoothing the values over time.

[0045] As a non-limiting example, and as Figure 3 shown, the visualization 300 includes a bus 302, a vehicle 308, and a pedestrian 314 having respective current bounding shapes (represented by solid lines) and future bounding shapes (represented by dashed lines). For example, the bus 302 has a current bounding shape 304 and a future bounding shape 306 generated based on the output of the object 2D motion 116 of the sequential DNN 104. The car 308 has a current bounding shape 310 and a future bounding shape 312 generated based on the output of the object 2D motion 116 of the sequential DNN 104. Similarly, the pedestrian 314 has a current bounding shape 316 and a future bounding shape 318 generated based on the output of the object 2D motion 116 of the sequential DNN 104. The object 2D motion 114 can be used to determine the future positions of the objects in the field of view of the camera and for each image in the image sequence. The object 2D motion - including the object future positions - can be used to help the autonomous driving software stack determine where the objects might be in the future (e.g., merge, change lanes, stay in lane, turn, etc.), thus assisting in obstacle avoidance, path planning, world model generation, and / or other processes or operations of the vehicle 1400.

[0046] For each detected object in each image, an object 3D motion 116 can be output as a three-dimensional array including the velocity in the x-direction, the velocity in the y-direction, and the velocity in the z-direction in world space. In this way, cross-sensor fusion of LIDAR data and RADAR data with image data can be used to train the sequential DNN 104 to generate a 3D velocity vector as an output, which corresponds to the object 3D motion 116 of one or more objects in the image. In this way, the sequential DNN 104 can be trained to learn the 3D velocity of an object using the change in the position of the object within the image sequence and the timestamps (e.g., to determine the time between each consecutive image). Thus, the sequential DNN 104 is capable of predicting the 3D velocity of an object in a deployment using the sensor data (e.g., LIDAR data, SONAR data, etc.) used during training without the need to receive the sensor data as an input. The object 3D motion 116 can be used to assist the obstacle avoidance, path planning, world model generation, and / or other processes or operations of the vehicle 1400 in an assisted autonomous driving software stack. As described herein, in some examples, the values output by the sequential DNN 104 and / or the decoded values can undergo post-processing, such as smoothing the values over time.

[0047] As described herein, the post-processing 108 can include temporal smoothing. For example, the temporal smoothing can include a state estimator (such as a Kalman filter). The temporal smoothing can be applied in the image space, or can be applied to the 3D world space relative to the vehicle 1400, the 3D world space relative to some fixed origin in the world space, or the bird's-eye view in the 2D world space.

[0048] Now refer to Figure 4 , each block of the method 400 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The method 400 can also be embodied as computer-usable instructions stored on a computer storage medium. The method 400 can be provided by a stand-alone application, a service, or a hosted service (stand-alone or in combination with another hosted service) or a plug-in of another product, to name just a few examples. Additionally, as an example, the process 100 of reference Figure 1 describes the method 400. However, the method 400 can additionally or alternatively be performed by any one system or any combination of systems in any one process or any combination of processes, including but not limited to those described herein.

[0049] Figure 4is a flowchart showing a method 400 for predicting time information using a sequential DNN according to some embodiments of the present disclosure. At block B402, method 400 includes receiving image data representing a sequence of images. For example, image data 102 can be received, where image data 102 represents a sequence of images (e.g., captured by one or more cameras on vehicle 1400).

[0050] At block B404, method 400 includes applying the image data to a sequential deep neural network (DNN). For example, image data 102 can be applied to sequential DNN 104. As described herein, since sequential DNN 104 can be trained by applying cross-sensor fusion or other sensor data to image data related techniques, sequential DNN 104 can predict time information using only image data 102 as input.

[0051] At block B406, method 400 includes calculating, via the sequential DNN, a first output corresponding to the speed of an object depicted in one or more images of the image sequence in world space. For example, sequential DNN 104 can predict object 3D motion 116, as described herein.

[0052] At block 408, method 400 includes calculating, via the sequential DNN, a second output corresponding to the current position and the future position of an object in image space. For example, sequential DNN 104 can predict object 2D motion 114, as described herein.

[0053] At block B410, method 400 includes calculating, via the sequential DNN, a third output corresponding to the time to collision (TTC). For example, sequential DNN 104 can predict the reciprocal of TTC (e.g., 1 / TTC), and the 1 / TTC value can be converted to TTC 112. In other examples, sequential DNN 104 can be trained to directly predict TTC 112, or the system can use the reciprocal of TTC.

[0054] In some embodiments, one or more of blocks B406, B408, and B410 can be executed in parallel (e.g., using parallel processing). As a result, the running time can be reduced while maintaining accuracy.

[0055] At block B412, method 400 includes performing one or more operations using at least one of the first output, the second output, or the third output. For example, vehicle 1400 (or its autonomous driving software stack) can use TTC 112, object 2D motion 114, and / or object 3D motion 116 to perform one or more operations. One or more operations can include, but are not limited to, world model management, path planning, vehicle control, and / or obstacle or collision avoidance.

[0056] Stateless Inference for Sequential Deep Neural Networks

[0057] The stateless inference applications described herein can be used for any inference method of any sequential model and are not limited to the specific embodiments described herein. However, during the deployment of the sequential DNN 104, stateless inference can be used in some examples of the present disclosure. For example, when training the sequential DNN 104 in a stateless mode (e.g., resetting the hidden state at each image sequence), the method of providing input data (e.g., image data 102) may require data from multiple images (e.g., feature maps) to be processed by the sequential DNN 104 to determine time-based predictions. Traditional systems use brute-force methods, where each feature map of each image under consideration is calculated at each iteration across each layer of the DNN. However, these traditional DNNs require increased computational power and ultimately reduce the ability of the system to operate in real time, such as in autonomous vehicle applications.

[0058] Compared to these traditional systems, the present system can utilize previously computed feature maps to reduce the computational load of the system at each iteration during stateless inference. For example, the feature map of the first frame can be computed, and then the feature maps of each successive frame can be computed. Many of these feature maps can be stored in a buffer and provided as input to an RNN with the feature map corresponding to the latest frame. Once the feature map is provided as input, the latest feature map can be added to the buffer, and the oldest or second-oldest feature map in the buffer can be removed from the buffer. This process can be repeated for each frame in the sequence to reduce the processing requirements for generating new feature maps for each frame at each iteration.

[0059] As an example, and with reference to Figure 5A , an image 502 captured at time T (e.g., an image represented by the Figure 1 image data 102) can be input into a feature extractor 504 to generate a feature map F T 508A for the image 502. The feature map F T corresponding to the image 502 can be fed into the sequential layer 510 of the DNN (e.g., GRU, LSTM, and / or other sequential layer types), such as the Figure 1 sequential DNN 104. In addition to the feature map F T 508A from the image 502, the feature maps F T-1 to F T-(N-1) (508B - 508D) can be fed into the sequential layer 510 of the DNN. The feature maps F T-1 to F T-(N-1)Each of (508B - 508D) may correspond to a feature map 508 from an image captured before (or after, in some embodiments) image 502. Thus, the sequential layer 510 may use N feature maps 508 from N images to generate a prediction, where the number of images N may be a hyperparameter of the DNN (e.g., Figure 1 sequential DNN 104).

[0060] As a non - limiting example, the number N of feature maps 508 may be four, such that four feature maps 508 are input to the sequential layer 510 at any one time to generate a prediction 512 at time T. In such an example, the circular buffer 506 (and / or another memory or storage device type) may be used to store each of the three previous feature maps (e.g., F T-1 to F T-3 ), such that the feature maps do not need to be regenerated again in each iteration of the DNN. As a result - and contrary to traditional brute - force methods - processing time is saved because previous feature map predictions from the feature extractor 504 are reused to generate the prediction 512 at time T.

[0061] Now referring to Figure 5B , each block of the method 514 described herein includes a computational process that may be performed using any combination of hardware, firmware, and / or software. For example, various functions may be performed by a processor executing instructions stored in a memory. The method 514 may also be embodied as computer - usable instructions stored on a computer storage medium. The method 514 may be provided by an independent application, service, or hosted service (independent or in combination with another hosted service) or a plug - in of another product, to name just a few examples. Additionally, as an example, reference Figure 5A describes the method 514. However, the method 514 may additionally or alternatively be performed by any one system or any combination of systems in any one process or any combination of processes, including but not limited to those described herein.

[0062] Figure 5B is a flow chart showing a method 514 for stateless inference using a buffer according to some embodiments of the present disclosure. At block B516, the method 514 includes generating a first feature map, each first feature map corresponding to one of N frames. For example, feature maps 508B - 508D may be generated.

[0063] At block B518, the method 514 includes storing the first feature map of each of the N frames in a buffer. For example, feature maps 508B - 508D may be stored in the buffer 506.

[0064] At block B520, method 514 includes generating a second feature map of N + 1 frames. For example, feature map 508A can be generated.

[0065] At block B522, method 514 includes applying the first feature map and the second feature map to the deployed neural network. For example, feature maps 508A - 508D can be applied to sequential DNN 104.

[0066] At block B524, method 514 includes removing the earliest feature map of the first feature map corresponding to the earliest frame among the N frames from the buffer. For example, when feature map 504A is added, feature map 508D can be removed from buffer 506. As a result, the buffer can always store the feature maps 508 of N frames.

[0067] At block B526, method 514 includes adding the second feature map to the buffer. For example, feature map 508A can be added to buffer 506.

[0068] Compared with traditional systems, another process of this system includes generating a feature map for each frame of a frame sequence during stateless inference. The number of frames in the frame sequence can be equal to the number of layers (or groups of layers) in sequential DNN 104 used to calculate the hidden state based on the feature map. For each layer (or group of layers), in addition to any feature maps from the previous frames in the frame sequence, the feature map of the corresponding frame in the frame sequence can also be applied. For example, in the case of using four feature maps, the first layer (or group of layers) of sequential DNN 104 can calculate the hidden state corresponding to the first feature map of the latest frame, the second layer (or group of layers) of sequential DNN 104 can calculate the hidden state corresponding to the first feature map of the latest frame and the second feature map of the second latest frame, and so on, such that each layer (or group of layers) always calculates the hidden state using (at least) one additional feature map from one additional frame rather than the previous layer (or group of layers) in the sequence. Additionally, the hidden state from each layer (or group of layers) can be passed to the next layer (or group of layers) in the layer sequence. Each layer (or group of layers) can calculate the hidden state in parallel, thereby reducing the running time of the system when predicting the generation time information.

[0069] As an example, and referring to Figure 6A , the hidden state of the first set of sequential layers 608A (for example, one or more layers can be included in sequential layer group 608) can be set to zero, as shown in 614. The image 602 captured at time T (for example, the image represented by the image data 102 of Figure 1 ) can be input into the feature extractor 604 to generate the feature map F T 606A for the image 602. The feature map F corresponding to the image 602 TEach set of sequential layers 608 (e.g., GRU, LSTM, and / or other sequential layer types) of the 606A is fed into the DNN 104, e.g., Figure 1 of the sequential DNN 104. In addition to the feature map F T from the image 502, the 606A, the feature maps F T-1 to F {T-(N-1)} can be fed into each set of sequential layers 608 of the DNN (e.g., in a non-limiting example of Figure 6A the set of sequential layers 608 - 608D). Each of the feature maps F T-1 to F {T-(N-1)} 606A - 606D can correspond to a feature map 606 from an image captured before (or after, in some embodiments) the image 602 (e.g., the image 602 and each other image can correspond to an image sequence). Thus, the set of sequential layers 608 can generate updated values of the hidden state - in some embodiments in parallel - using multiple feature maps 606 from multiple images and using the hidden state of the previous set of sequential layers 608 in the sequential set of sequential layers 608. In some examples, the number of images can be a hyperparameter of the DNN (e.g., Figure 1 the sequential DNN 104), such that the number sequence of the set of sequential layers 608 corresponds to the number of images.

[0070] As a non-limiting example, the number of feature maps 606 can be four, such that four feature maps 606 are used as inputs to the set of sequential layers 608 (e.g., the first feature map F T 608A inputs to the set of sequential layers 608A, the first feature map F T 606A and the second feature map F {T-(N-3)} 606B are input to the set of sequential layers 608B, the first feature map F {T} 606A, the second feature map feature map F {T} 606A, and the third feature map F {T-(N-2)} 606C are input to the set of sequential layers 608C, and the first feature map F T 606A, the second feature map F {T-(N-3)} 606B, the third feature map F{T-F {T-(N-2)} 606C, and the fourth feature map F T-(N-1) 606D are input to the set of sequential layers 608D). Additionally, in such an example, the hidden state can be set to zero at the first set of sequential layers 608A, and the computed hidden state of the first set of sequential layers 608A (e.g., using the first feature map F T 606A) is passed to the second set of sequential layers 608B, and the computed hidden state of the second set of sequential layers 608B (e.g., using the first feature map F T606A, the second feature map F {T-(N-3)} 606B and the hidden states from the first set of sequential layers 608A) can be passed to the third set of sequential layers 608C, and so on. In some embodiments, the feature maps 606 can be input to the sequential layer group 608 such that the sequential layer group 608 can compute the hidden states in parallel (e.g., the first feature map F T 606A can be applied to each sequential layer group for the first time, and for the second time after the first time, the second feature map F {T-(N-3)} 606B can be applied to each sequential layer group 608 after the first set of sequential layers 608A, and so on. The output of the last set of sequential layers 608 (e.g., Figure 6A the sequential layer group 608D in the illustration) can be provided as an input to an additional layer 610 (e.g., a convolutional layer, a deconvolutional layer, etc.) of the DNN to generate a prediction 612 of the image 602. As a result, compared to traditional brute-force methods, processing time is saved for parallel computing of the hidden states of the sequential layer group 608.

[0071] Regarding Figure 6A The process described can be depicted according to the following equation. For example, the hidden state variable for a sequential layer (e.g., the sequential layer group 608) can be represented as h T N , which is the state computed at time T using N images. Here, h T —1 is the initial state zero, and in some embodiments, it can be zero for all T. When performing stateless inference with N images, the hidden state that can be used when computing the output prediction (e.g., prediction 612) at time T can be h T N . h T N can be a function of N + 1 feature maps F T and can be represented according to equation (1) below on an N + 1 input image sequence:

[0072] h T N = f(F T , F T-1 , …, F T-(N-1) ) (1)

[0073] Thus, at each time step, h T N can be computed. Additionally, since the layer for computing the hidden state h can be recursive (e.g., LSTM, GRU, etc.), the function g can be used to represent the current hidden state as a function of the current input and the previous hidden state, as shown in equation (2) below:

[0074] h T N = g(F T , h T-1 N-1 ) (2)

[0075] Therefore, write out all applications of g (e.g., expand and reproduce in a timely manner) to give the following equations (3)-(6):

[0076] h T- (N - 1) -1 = 0 (3)

[0077] h T-N 0 = g(F T-N , h T-(N-1) -1 ) (4)

[0078] …

[0079] h T-1 N-1 = g(F T-1 , h T-2 N-2 ) (5)

[0080] h T N = g(F T , h T-1 N-1 ) (6)

[0081] Therefore, in a non - restrictive example, applications of g can be shown, such as the following equation (7):

[0082] h T N = g(F T , g(F T-1 , g(F T-2 , …, g(F T-N , 0)))) (7)

[0083] In addition, to calculate the output at all time instances, when calculating the input feature map F T , the N hidden states of N image sequences of different lengths can be calculated in parallel at each moment. Thus, at each moment, according to a non - restrictive example, N states can be calculated, the following equations (8)-(10):

[0084] h T 0 = g(F T , 0) (8)

[0085] …

[0086] h T N-1 = g(F T , h T-1 N-2 )(9)

[0087] h T N = g(F T , h T-1 N-1 )(10)

[0088] It can be shown as Figure 6A N + 1 sets of sequential layers 608 in

[0089] In some embodiments, depending on the recursive function g, some or all of the calculations can be independent of the previous hidden state, thus allowing further computational savings and increased runtime efficiency. For a non-limiting example, in the case of using a GRU layer including six convolutions, three forward-facing convolutional layers can be applied separately to the input (e.g., one for the z gate, one for the r gate, and one for the next hidden state). These values can be calculated once and the results can be reused N times.

[0090] Now referring to Figure 6B , each block of the method 616 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be executed by a processor that executes instructions stored in a memory. The method 616 can also be embodied as computer-usable instructions stored on a computer storage medium. The method 616 can be provided by an independent application, service, or hosted service (independent or in combination with another hosted service) or a plug-in of another product, to name just a few examples. Additionally, by way of example, the method 616 is described with respect to Figure 6A . However, the method 616 can additionally or alternatively be executed by any one system or any combination of systems in any one process or any combination of processes, including but not limited to those described herein.

[0091] Figure 6B is a flowchart showing a method 616 for stateless inference using parallel computing according to some embodiments of the present disclosure. At block B618, the method 616 includes initializing the values of the hidden states of each set of layers in a sequence of sets of layers of a deployed neural network. For example, the hidden states of the sequential layers 608A - 608D can be initialized to zero, random numbers, or states obtained through an optimization process.

[0092] At block B620, the method 616 includes generating a feature map of the current frame of a frame sequence. For example, the feature map 606 can be generated by the feature extractor 604 for the current frame.

[0093] In block B622, method 616 includes applying the feature map of the corresponding frame in the frame sequence and the previous feature maps of the frames in the frame sequence before the corresponding frame to each group of layers in the sequence of groups of layers. For example, sequential layer 608A may have feature map 606A applied to it, sequential layer 608B may have feature map 606B (which may include feature map 606A) applied to it, and so on.

[0094] In block B624, method 616 includes generating updated values of the hidden state in parallel and by each group of layers in the sequence of groups of layers. For example, when feature map 606 is applied to sequential layer 608, sequential layer 608 may compute a hidden value.

[0095] In block B626, method 616 includes passing the updated values of the hidden state through each group of layers in the sequence of groups of layers to the next group of layers in the sequence of layers. For example, the value of the hidden state computed by sequential layer 608A may be passed to sequential layer 608B, and the value of the hidden state computed by sequential layer 608B may be passed to sequential layer 608C, and so on.

[0096] In block B628, method 616 includes computing and using the hidden state of each group of layers in the sequence of groups of layers to compute one or more outputs by the deployed neural network. For example, the deployed neural network (e.g., sequential DNN 104) may use the hidden state of each sequential layer 608 to generate prediction 612.

[0097] Training a sequential deep neural network for temporal prediction

[0098] During training, the sequential DNN may use ground truth data generated from a combination of training image data and training sensor data (e.g., using cross-sensor fusion). For example, the training image data may be used to generate bounding boxes corresponding to the positions of objects in each image of the image sequence. The training sensor data may be used to determine velocity information corresponding to each object. The training sensor data may represent LIDAR data from one or more LIDAR sensors, RADAR data from one or more RADAR sensors, SONAR data from one or more SONAR sensors, and / or other training sensor data from one or more other sensor types.

[0099] To generate ground truth data for 2D motion and 3D motion, training sensor data can be associated or fused with training image data. For example, for LIDAR data, calibration information between a camera and a LIDAR sensor can be used to associate LIDAR data points with pixels in the image space. For RADAR data, the RADAR data can detect objects in its field of view, and the objects can be correlated with objects detected in an image represented by the image data. Similar techniques can be used for other sensor types. In any example, information from LIDAR data, RADAR data, and / or other data can be utilized to enhance pixels corresponding to detected objects in the image (e.g., as detected and identified using object detection techniques). For example, one or more pixels corresponding to an object (or corresponding to the boundary shape of the object) can be associated with velocity data determined using training sensor data. For LIDAR data, for example, depth information of an object can be determined on two or more images in a sequence such that a change in the depth value can indicate velocity. For RADAR data, as another example, velocity data of an object in the field of view of the RADAR sensor can be determined. Thus, since object detection can be automatically generated with respect to the image (e.g., using a convolutional neural network, using an object detection computer vision algorithm, using an object tracking algorithm, etc.), and the training sensor data can be automatically associated with the training image data, ground truth data for training a sequential DNN can be automatically generated.

[0100] Reference Figure 7 , Figure 7 A data flow diagram of a process 700 for training a sequential deep neural network 104 according to some embodiments of the present disclosure is shown. Training image data 702 (which can be similar to image data 102 but can correspond to training data for training the sequential DNN) and / or training sensor data 704 can be used to generate ground truth for training the sequential DNN 104. The training sensor data 704 can include LIDAR data from one or more LIDAR sensors 1464, RADAR data from one or more RADAR sensors 1460, SONAR data from one or more SONAR sensors (e.g., ultrasonic sensor 1462), and / or other types of training sensor data from one or more other sensor types. The training image data 702 can represent an image sequence, and the training sensor data 704 can include sensor data generated at least partially simultaneously with the training image data 702. Thus, the training sensor data 704 (e.g., using timestamps) can be associated with the respective images of the image sequence to generate ground truth.

[0101] As described herein, ground truth generation 706 may include automatic ground truth generation using a combination of training image data 702 and training sensor data - for example but not limited to via cross-sensor fusion. As a result of automatically generating labels or annotations representing the ground truth, in some embodiments, manual labeling for generating ground truth data may not be required. However, in some examples, manual labeling may be performed additionally or alternatively to the automatic labeling. Thus, in any example, the labels may be synthetically generated (e.g., generated from a computer model or rendering), generated from real and / or raw data (e.g., designed and generated from real-world data), machine automated (e.g., using feature analysis and learning to extract features from data and then generate labels), human annotated (e.g., a labeler or annotation expert, defining the location of the label), and / or a combination thereof.

[0102] In cases where automatic ground truth generation is performed, training sensor data 704 may be associated or fused with training image data 702. Vehicle 1400 may be equipped with different types of sensors as well as cameras (including but not limited to Figures 14A - 14C the sensors and cameras shown in). For example, in addition to multiple cameras, many RADAR sensors 1460, LIDAR sensors 1464, and / or other sensor types may be positioned on vehicle 1400 such that there is an overlap between the field of view of the cameras and the field of view or sensing area of the sensors. In some embodiments, the spatial layout of the sensors may be calibrated via a self-calibration algorithm, and the synchronization of the sensors may be controlled to exhibit temporal alignment of the sensor captures. This may help to accurately propagate the training sensor data 704 into the image space or camera domain for automatic ground truth generation 706.

[0103] For example, and with respect to Figure 8 the visualization 800, for LIDAR data, LIDAR data points 820 may be associated with pixels in the image space using calibration information between the camera and LIDAR sensor 1464 that generated the training image data 102. For RADAR data, the RADAR data may represent the speed of objects (e.g., bus 802, vehicle 808, and pedestrian 814) in the field of view (or sensing area) of the RADAR sensor 1460. Then, the objects represented by the RADAR data may be associated with the objects detected in the image represented by the training image data 702 (e.g., using an algorithm designed for visual and RADAR object fusion). Different techniques may be used for other sensor types to associate the training sensor data 704 from the sensors with the training image data 702 in the image space. Pixels corresponding to the objects detected in the image (e.g., as detected and identified using object detection techniques, and by Figure 8The boundary shapes 804, 810, and 816 (shown in ) can be enhanced using information from LIDAR data, RADAR data, and / or other training sensor data 704.

[0104] For example, one or more pixels corresponding to an object (or corresponding to the boundary shape of an object) can be associated with velocity data determined using training sensor data 704. For LIDAR data, for example, depth information of an object can be determined on two or more images in a sequence such that a change in the depth value can indicate velocity. For RADAR data, as another example, velocity data of an object in the field of view of the RADAR sensor 1460 (or sensor field) can be determined. Thus, since object detection on an image can be automatically generated (e.g., using a convolutional neural network, using an object detection computer vision algorithm, using an object tracking algorithm, etc.), and the training sensor data 704 can be automatically associated with the training image data 702, the ground truth for training the sequential DNN 104 can be automatically generated. Then, the training sensor data 704 can be used to determine the position of an object in world space and the two-dimensional or three-dimensional velocity of each object in world space. As Figure 8 shown, the position and velocity data 806 corresponding to the bus 802, the position and velocity data 812 corresponding to the vehicle 808, and the position and velocity data 818 corresponding to the pedestrian 814 can be determined using the training sensor data 704, and can be associated with the objects in image space and used to train the sequential DNN 104 to predict position and velocity information from only image data in deployment. The position and velocity data 806 can represent the x, y, and z positions of the bus 802 in world space (e.g., relative to the origin of the vehicle 1400). Additionally, the position and velocity data 806 can represent the 3D motion 712 of the object, which can correspond to a velocity Vx in the x direction, a velocity Vy in the y direction, and / or a velocity Vz in the z direction in world space.

[0105] When generating ground truth data, an object detection deep neural network (e.g., a convolutional DNN), an object detection computer vision algorithm, an object tracking algorithm, and / or another object detection method can be used to identify bus 802, vehicle 808, and pedestrian 814. For example, for each image in an image sequence, bus 802, vehicle 808, and pedestrian 814 can be detected individually. In other examples, once bus 802, vehicle 808, and / or pedestrian 814 are detected in an image (e.g., using an object detection algorithm, a deep neural network, etc.), an object tracking algorithm can be used to track and identify bus 802, vehicle 808, and / or pedestrian 814 in the remaining images of the image sequence in which bus 802, vehicle 808, and / or pedestrian 814 are present. Thus, boundary shapes 804, 810, and / or 816 can be generated for each object in each image in which an object is present. The transformation from pixels in image space to positions in world space can be known to the system, and thus - in some embodiments - the positions in image space can be used to determine the positions of each object in world space. For each object, object 2D motion 710 can be generated by leveraging the position and size of the boundary shape associated with the object across a sequence of frames. For example, the scale change Δs can be determined by analyzing the differences in the boundary shapes associated with the object in different frames. The movement of an object within image space (e.g., the change in the position of the boundary shape of the object across the frames of a frame sequence) can be used to generate the ground truth data for object 2D motion 710.

[0106] To determine the time to collision (TTC) 708 of the ground truth, the scale change Δs can be used in some non-limiting embodiments. For example, for an ideal camera with a constant object angle, Equation (11) can be used:

[0107] TTC = Δt / Δs (11)

[0108] where Δs is the scale change and Δt is the time between two consecutive frames for which the scale change Δs is calculated. As a result, the prediction of the sequential DNN 104 can depend on the sampling interval of the input images. In some examples, the sequential DNN 104 can be trained to implicitly assume a time (e.g., Δt) baseline such that the value output by the sequential DNN 104 can be directly used to determine the TTC (e.g., as described herein, the sequential DNN 104 can be trained to output 1 / TTC, and thus the value of 1 / TTC can be directly used to calculate the TTC). In practice, in cases where an ideal camera and / or a constant object angle are not possible, the difference between the actual camera and / or constant object angle used may not have a significant impact, and thus Equation (11) can still be used. However, in some embodiments, the following Equation (12) can be a more accurate representation of the TTC:

[0109] TTC ≈ Δt / Δs (12)

[0110] However, directly using TTC as the ground truth for training the sequential DNN 104 can result in unbounded values (e.g., [-∞, -a] ∪ [b, ∞]). Thus, when the vehicle 1400 (e.g., the ego vehicle and another object are stationary relative to each other), the value of TTC can become infinite. In this way, the TTC 708 used as the ground truth when training the sequential DNN 104 can be the reciprocal TTC (e.g., 1 / TTC). For example, 1 / TTC produces well-bounded values (e.g., [-1 / a, 1 / b]), which can allow the use of gradient descent-based methods for optimization. Thus, the following equation (13) can be used to generate the ground truth for training the network, and then 1 / TTC can be converted to TTC using post-processing:

[0111]

[0112] In some examples, preprocessing 714 can be applied to the data to ensure that the data used for the ground truth is accurate. In some embodiments, the preprocessing 714 on the training image data 702 and / or the training sensor data 704 can include the preprocessing described herein with respect to the image data 102. Additionally, the preprocessing 714 can include clearing or removing noise from the training sensor data 704 and / or the training image data 702. For example, one or more filters can be used to filter out ground truth data that may not be inaccurate or may be inconsistent. For example, unless there are consistent tracking results on the images of an image sequence of the same object (e.g., for the bus 802 in Figure 8 ), those tracking results may not be used. In such an example, with respect to tracking the object in those images, images of at least two or more consecutive detections of the object can be ignored. As a result, these incorrect or inaccurate detections can be removed from the ground truth in an effort to reduce noise in the predictions of the sequential DNN 104.

[0113] As another example, during ground truth generation, inconsistent vision (e.g., images) and RADAR associations and / or vision and LIDAR associations (and / or vision and other sensor type associations) in sequential images can be removed. In such an example, when generating the ground truth, training sensor data 704 corresponding to objects with inconsistencies can be ignored, such that only consistent associations between the image and LIDAR data (e.g., LIDAR points projected into the image space) and / or the image and RADAR data are used for ground truth generation regarding any one object. Thus, if the similarity between the value associated with an object in the image space determined from the LIDAR data and the value associated with the object in the image space determined from the RADAR data is not within a threshold, at least one of the LIDAR data or the RADAR data can be ignored.

[0114] In some examples, cross-sensor consistency may be required in the ground truth dataset. For example, for ground truth data from RADAR sensor 1460 and LIDAR sensor 1464 (and / or other sensor types), it may be required that the ground truth be consistent among multiple sensors before using the ground truth data. This can include consistency of data from two or more of the same sensor type and / or consistency of data from two or more different sensor types, where there are similar objects in their field of view - or field of sense.

[0115] Data augmentation 716 can be used on the training image data 702 and / or the training sensor data 704 to increase the size of the training set and / or reduce the likelihood of overfitting in successive DNN 104 predictions. For example, the images represented in the training image data 702 can be cropped, rotated, downsampled, upsampled, skewed, shifted, and / or otherwise augmented to train the sequential DNN 104. Additionally, the ground truth data associated with the images can be similarly augmented such that the ground truth labels and annotations are accurate with respect to the augmented images. In some examples, temporal augmentation can be used to train the sequential DNN 104 at intervals of the images within a sequence different from the images (e.g., every other image, every fourth image, etc.), and / or the image sequence can be trained in reverse order, at least as described in more detail herein Figure 10 and 11 is described.

[0116] Now refer to Figure 9, each block of method 900 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be executed by a processor that executes instructions stored in a memory. Method 900 can also be embodied as computer-usable instructions stored on a computer storage medium. Method 900 can be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service), or a plug-in of another product, to name just a few examples. Additionally, by way of example, reference Figure 7 describes method 900. However, method 900 can additionally or alternatively be performed by any one system or any combination of systems in any one process or any combination of processes, including but not limited to those described herein.

[0117] Figure 9 is a flowchart showing method 900 for generating automatic ground truth using cross-sensor fusion according to some embodiments of the present disclosure. In block B902, method 900 includes receiving image data representing a sequence of images generated over a period of time. For example, training image data 702 can be received, where the training image data 702 is generated over a period of time.

[0118] In block B904, method 900 includes determining the positions of objects in the images of the image sequence. For example, object detection algorithms, object tracking algorithms, and / or object detection neural networks can be used to determine the positions of objects in the images of the image sequence.

[0119] In block B906, method 900 includes receiving sensor data generated by at least one of a LIDAR sensor or a RADAR sensor over a period of time. For example, training sensor data 704 can be received over a period of time.

[0120] In block B908, method 900 includes correlating the sensor data with the image data. For example, using one or more techniques (including but not limited to cross-sensor fusion), the training sensor data 704 can be correlated with the training image data 702.

[0121] In block B910, method 900 includes analyzing the sensor data corresponding to the positions of the objects in the images to determine the speeds of the objects during the period in which the images are generated. For example, the training sensor data 704 corresponding to the objects can be used to determine the positions of the objects.

[0122] In block B912, method 900 includes using the speeds and positions of the objects to automatically generate ground truth data corresponding to the images. For example, the speeds and positions of the objects can be used to automatically generate ground truth data corresponding to the images by, for example, associating the speeds with one or more pixels corresponding to the objects in the images.

[0123] At block B914, method 900 includes training a sequential deep neural network (DNN) using ground truth data. For example, the sequential DNN 104 can be trained using ground truth data.

[0124] Temporal augmentation for training a sequential deep neural network

[0125] In addition to motion prediction for an autonomous vehicle perception system, the temporal augmentation described herein can be a process applicable to any ground truth generation and / or training for any sequential DNN regarding any domain. However, regarding the present disclosure, the temporal augmentation can be used at least for scale change Δs, speed prediction, boundary shape position, and / or TTC prediction.

[0126] In some embodiments, temporal augmentation of training data can be used to train the sequential DNN 104. For example, an image sequence (e.g., each frame in consecutive order) can be used in the order input to the sequential DNN 104 during training. However, since the computation of ground truth can be based on temporal information (e.g., frame rate), the image sequence can be used in different frame intervals (e.g., every frame, every other frame, every third frame, etc.) as input to the sequential DNN 104. Additionally, the images can be input in the forward order or can be input in the reverse order. In this way, changing the frame interval, changing the order, or a combination thereof can be used to augment the training images, increasing the training data size and the robustness of the system. For example, since the frame interval can be changed, the sequential DNN 104 can be trained to interpret different transformations of an object across frames (e.g., when frames are used in sequence, the object may move slowly, and when every other frame or every third frame is used, the object may move quickly). In such an example, training examples can be generated from ground truth data corresponding to a hazardous traffic situation that may be unsafe in the real world.

[0127] Thus, and with reference to Figure 10 , the temporal augmentation can be used for each K images in an image sequence, where K can be any positive (e.g., time forward along line 1012) or negative (e.g., time backward along line 1010) number. For example, ground truth data with scale changes can be generated for every frame 1002, every two frames 1002, every four frames 1002, etc. in a forward or backward facing sequence. For example, each frame 1002 (e.g., 1002A - 1002I) can be from a sequence of frames captured at times T - 4 to T + 4. Brackets 1004 can indicate the position where K equals 1, brackets 1006 can indicate the position where K equals 2, and brackets 1008 can indicate the position where K equals 4. Although Figure 10Only an enhancement of K with 1, 2, and 4 is shown herein, but this is not restrictive. Any number can be used for K to increase the training data over time. By adjusting the value of K, the size of the training data set can be increased, and the robustness of the system can be increased.

[0128] In traditional sequential model training, a time batch can be used by forming a sequence of N consecutive frames that satisfy the following relationship (14):

[0129] {Image T-i}, i ∈ {0, 1, …, N−1} (14)

[0130] However, in the temporal augmentation technique of the present disclosure, instead of using a fixed time step increment of 1 as in traditional systems, the batching can alternatively be represented by the following relationship (15):

[0131]

[0132] In some embodiments, K can be randomly sampled from a discrete probability distribution over integers. For example, K can be uniformly distributed over {−4, −2, −1, 1, 2, 4}, as Figure 10 shown.

[0133] Depending on what the sequential DNN 104 (or any other sequential model in a different application) is going to predict, some ground truth values may need to be augmented according to the change in K, while other ground truth values may not have a similar augmentation. As an example, and for the current state of the art, the boundary shape at time T may not need to be adjusted, the object distance at time T may not need to be adjusted, and the object orientation at time T may not need to be adjusted. However, some dynamic or temporal attributes may need to be adjusted, such as TTC, Δs, velocity, and / or the boundary shape after X steps, as shown in (16) - (19) below:

[0134] TTC ← TTC / K (16)

[0135] Δs ← Δs * K (17)

[0136] Velocity ← Velocity * K (18)

[0137] Bounding shape after X steps ← Bounding shape after X * K steps (19)

[0138] Now referring to Figure 11, each block of the method 1100 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The method 1100 can also be embodied as computer-usable instructions stored on a computer storage medium. The method 1100 can be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service) or a plug-in of another product, to name just a few examples. Additionally, by way of example, reference Figure 10 describes the method 1100. However, the method 1100 can additionally or alternatively be performed by any one system or any combination of systems in any one process or any combination of processes, including but not limited to those described herein.

[0139] Figure 11 is a flowchart showing a method 1100 for temporal enhancement according to some embodiments of the present disclosure. At block B1102, the method 1100 includes calculating a first time difference between a first time of a first image of a captured image sequence and a second time of a second image of the captured image sequence. For example, a first time difference between a first time of a first frame 1002E of a captured frame sequence and a second frame 1002F can be determined.

[0140] At block B1104, the method 1100 includes calculating a second time difference between the first capture time and a third time of a third image of the captured image sequence. For example, a second time difference between a first time of capture of a first frame 1002E and a third frame 1002G can be determined.

[0141] At block B1106, the method 1100 includes generating a first boundary shape for an object in the first image, generating a second boundary shape for an object in the second image, and generating a third boundary shape for an object in the third image. For example, a boundary shape (e.g., Figure 8 the boundary shape 804) can be generated for an object in each of the three frames 1002E - 1002G.

[0142] At block B1108, the method 1100 includes calculating a first scale change between the first boundary shape and the second boundary shape and a second scale change between the first boundary shape and the third boundary shape. For example, a scale change s between the first boundary shape and the second boundary shape and the first boundary shape and the second boundary shape can be calculated.

[0143] In block B1110, method 1100 includes generating first ground truth data corresponding to a second image using a first ratio between a first time difference and a first ratio change, and generating second ground truth data corresponding to a third image using a second ratio between a second time difference and a second ratio change. For example, a ratio of Δt / Δs can be used to generate ground truth data for the second frame 1002F and the third frame 1002G. As a result, the first frame 1002E can be used to generate ground truth at different strides (e.g., per frame, every other frame, every fourth frame, etc.) in order to increase the robustness of the training dataset and train the sequential DNN 104 (or any DNN) for situations where it may be otherwise difficult to generate training data (e.g., dangerous situations).

[0144] In block B1112, method 1100 includes using the first ground truth data and the second ground truth data to train a neural network. For example, the first ground truth data and the second ground truth data can be used to train a neural network (e.g., sequential DNN 104) to predict various temporal information using a more robust training dataset.

[0145] State Training of Sequential Deep Neural Networks

[0146] The state training applications described herein can be used for any training method of any sequential model and are not limited to the specific embodiments described herein. In contrast to stateless training (described herein) which can be performed more efficiently with sparsely labeled data, state training and inference can be performed more efficiently with densely labeled data. For example, densely labeled data can allow multiple gradient calculations to be made using continuously computed predictions.

[0147] During inference (e.g., during deployment), the sequential model can operate in a fully stateful mode where the hidden state can only be reset at startup. As a result, each time instance may require a single forward pass computation, as shown in equation (20) below:

[0148] {y t ,h t} = DNN{x t ,h t-1}

[0149] where y is the prediction, x is the input (e.g., an image), h represents the hidden state, t is the time index, and DNN represents a single feedforward neural network with two inputs and two outputs. Although the DNN can contain multiple hidden states, the hidden states can be jointly represented by a single state variable h in equation (20).

[0150] In some embodiments of the present disclosure, state training can be used for the sequential DNN 104 to address and remedy this contrast between training and inference. For example, a balance between completely uncorrelated mini - batches and completely sequential training data inputs can be achieved by randomizing the effective sequence length during training. For example, the sequential DNN can be trained to maintain or store the hidden state across frames of a sequence without resetting the hidden state. For example, each mini - batch can include a grid of frames or images, where each row of the mini - batch includes a feature vector. To prevent the sequential DNN 104 from overfitting, the image sequences across mini - batches can vary. For example, from a first mini - batch to a second mini - batch, the image sequence can continue (e.g., frames 1, 2, 3 in the first mini - batch and then frames 4, 5, 6 in the second mini - batch). As another example, from a first mini - batch (e.g., in different rows) to a second mini - batch (in different rows), the image sequence may not continue (e.g., frames 3, 5, 7 in the first mini - batch and then frames 3, 2, 1 in the second mini - batch).

[0151] By interleaving the order, frame intervals, and sequences across mini - batches, the sequential DNN 104 can be trained to predict outputs without overfitting to the training dataset. For example, the recurrent layers of the sequential DNN 104 may fail to adapt and overfit to a particular fixed number of context frames, and the relevant quantities between optimizer iterations can be adjusted such that fast convergence is maintained. Additionally, the training data can be made to include examples with a large amount of context, which can be well - adapted for accurate prediction during inference in a fully repeating or stateful mode. Another benefit of the state training method described herein is the efficient use of computational and memory resources. For example, each mini - batch can be used not only to update the weights and biases backward in time within the time frame of a single mini - batch, but each mini - batch can also be used to set the network hidden state to a fully realistic hidden state for the next mini - batch and optimizer iteration.

[0152] Figure 12 Examples include randomizing the effective length, stride, and / or direction of mini - batches in the training dataset. For example, the training dataset 1200 can include multiple mini - batches MB0(1202A) - MB7(1202H), where each mini - batch 1202 can include a mini - batch size (e.g., four in the Figure 12 example) corresponding to multiple examples 1204 (e.g., Figure 12in Examples 1204A - 1204D). Additionally, each Example 1204 within each Mini - Batch 1202 can include a sequence length of frames (e.g., where each frame is indicated by a frame index 1206, such as 1206A - 1206C). The frame sequence and associated ground - truth labels and other metadata within an Example 1204 of a single Mini - Batch 1202 can be referred to herein as a feature vector. The frames in the frame sequence can be at different strides (e.g., stride randomization), such that Example 1204A at Mini - Batch 1202A can be at a stride of - 1 (e.g., 3, 2, 1), and 1204D at Example Mini - Batch 1202A can be at a stride of + 2 (e.g., 1, 3, 5). Additionally, although each Example 1204 within each Mini - Batch 1202 can include only a set sequence length (e.g., Figure 12 is 3 in), the sequences of Mini - Batches 1202 within an Example can have a longer sequence length (e.g., a main sequence length). Thus, the main sequence length can also be random (e.g., main sequence randomization), such that - as a non - limiting example and with respect to Figure 12 - the main sequence length can be three, six, nine, and / or some other factor of three. For example, in exemplary Cross 1204D within Mini - Batches 1202A - 1202C, the main sequence length 1208A can be nine, across Examples 1204A within Mini - Batches 1202F - 1202G, the main sequence length 1208B can be six, and for Mini - Batch 1202F within Example 1204D, the main sequence length 1208C can be three. To determine the frame indices to input into a Mini - Batch to generate the training dataset 1200, a randomization and / or probability algorithm can be used.

[0153] During training, the training dataset 1200 can be input into the sequential DNN 104 (when the sequential DNN 104 is trained for the stateful mode). During this training, and as a non-limiting example, the hidden state of the mini-batch 1202F at example 1204A can be stored after the sequential DNN 104 processes frames 12, 11, and 10. Then, when the next mini-batch 1202G is received at example 1204A, the stored hidden state can be used when calculating the updated hidden states for frames 9, 8, and 7. In this way, when example 1204 includes consecutive frames across mini-batches, the hidden state can be stored (e.g., not reset to random values or zero values). As another non-limiting example, the hidden state of the mini-batch 1202F at example 1204D can be calculated after the sequential DNN 104 processes frames 5, 3, and 1. Then, when the next mini-batch 1202G is received at example 1204A, the hidden state of the previous mini-batch 1202F from example 1204D can be reset to random values or zero values when calculating the hidden states for frames 14, 13, and 12. Thus, when example 1204 does not include consecutive frames across mini-batches, the hidden state can be reset. In some examples, although the hidden state can be reset after example 1204D of the mini-batch 1202F, the hidden state can be stored anyway. This can be beneficial if example 1204D at mini-batch 1202E is serially adjacent to example 1204D at mini-batch 1202F. By interleaving and randomizing the primary sequence length, stride, and / or direction of the frames across examples 1204 and mini-batches 1202, a balance between completely uncorrelated mini-batches and a completely consecutive training dataset can be achieved - resulting in the sequential DNN 104 being more accurate in deployment while also allowing the use of optimization methods based on stochastic gradient descent.

[0154] Now refer to Figure 13 , each block of the method 1300 described herein includes a computational process that can be performed using any combination of hardware, firmware, and / or software. For example, various functions can be performed by a processor executing instructions stored in a memory. The method 1300 can also be embodied as computer-usable instructions stored on a computer storage medium. The method 1300 can be provided by a stand-alone application, service, or hosted service (stand-alone or in combination with another hosted service) or a plug-in for another product, to name just a few examples. Additionally, as an example, refer to Figure 12 for a description of the method 1300. However, the method 1300 can additionally or alternatively be performed by any one system or any combination of systems in any one process or any combination of processes, including but not limited to those described herein.

[0155] Figure 13is a flowchart showing a method 1300 for state training according to some embodiments of the present disclosure. At block B1302, method 1300 includes selecting a first image sequence of first feature vectors of a first mini-batch and a second image sequence of second feature vectors of a second mini-batch immediately following the first mini-batch, with the first image in the first image sequence and the second image sequence being continuously sorted with a first stride. For example, using a randomization or probabilistic algorithm, the first image sequence (e.g., the frames at example 1204A of mini-batch 1202F) and the second image sequence (e.g., the frames at example 1204A of mini-batch 1202G) can be selected, where the frames are continuously sorted in reverse order with a stride of one (e.g., 12, 11, 10, 9, 8, 7).

[0156] At block B1304, method 1300 includes selecting a third image sequence of third feature vectors of the first mini-batch and a fourth image sequence of fourth feature vectors of the second mini-batch, with the second image in the third image sequence and the fourth image sequence being non-continuously sorted. For example, using a randomization or probabilistic algorithm, the third image sequence (e.g., the frames at example 1204C of mini-batch 1202A) and the fourth image sequence (e.g., the frames at example 1204C of mini-batch 1202B) can be selected, where the frames are non-continuously sorted (e.g., 5, 3, 1, 1, 3, 5).

[0157] At block B1306, method 1300 includes applying the first mini-batch and the second mini-batch to a neural network during training. For example, the first mini-batch (e.g., mini-batch 1202A) and the second mini-batch (e.g., mini-batch 1202B) can be applied to a neural network (e.g., sequential DNN 104) during training. As a result, the main sequence length, stride, and / or direction of the image sequences can be randomized to help prevent overfitting to the dataset during training.

[0158] The methods for state training described herein can be equally applicable to the training of any sequential model and can thus find applications in other domains or technical fields other than motion prediction for autonomous vehicle perception systems. For example, when the training data contains labels for all frame indices, these methods can be easily applied because the frame indices can be used for loss and gradient calculations. These methods are also applicable to cases where every Nth frame is labeled in a regular pattern. For example, the data can be iterated such that the labeled frames always appear as the last frame of each mini-batch.

[0159] Exemplary autonomous vehicle

[0160] Figure 14AIllustrated is an exemplary autonomous vehicle 1400 in accordance with some embodiments of the present disclosure. The autonomous vehicle 1400 (optionally referred to herein as "vehicle 1400") may include, but is not limited to, passenger vehicles such as cars, trucks, buses, emergency vehicles, shuttles, electric or motorized bicycles, motorcycles, fire trucks, police cars, ambulances, boats, construction vehicles, underwater vehicles, drones, and / or other types of vehicles (e.g., driverless and / or accommodating one or more passengers). Autonomous vehicles are generally described according to levels of automation, as defined by the National Highway Traffic Safety Administration (NHTSA), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) in "Taxonomy and Definitions for Terms Related to Driving Automation Systems for Highway Motor Vehicles" (Standard No. J3016 - 201806, issued on June 15, 2018, Standard No. J3016 - 201609, issued on September 30, 2016, and prior and future versions of this standard). The vehicle 1400 may have functionality according to one or more of levels 3 to 5 of the autonomous driving level. For example, the vehicle 1400 may have conditional automation (level 3), high automation (level 4), and / or full automation (level 5), depending on the embodiment.

[0161] The vehicle 1400 may include components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. The vehicle 1400 may include a propulsion system 1450, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or other types of propulsion systems. The propulsion system 1450 may be connected to the driveline of the vehicle 1400, which may include a transmission to effect the propulsion of the vehicle 1400. The propulsion system 1450 may be controlled in response to receiving a signal from the throttle / accelerator 1452.

[0162] When the propulsion system 1450 is operating (e.g., when the vehicle is in motion), a steering system 1454, which may include a steering wheel, may be used to maneuver the vehicle 1400 (e.g., along a desired path or route). The steering system 1454 may receive signals from a steering actuator 1456. For full automation (level 5) functionality, the steering wheel may be optional.

[0163] A brake sensor system 1446 may be used to operate vehicle actuators in response to receiving signals from a brake actuator 1448 and / or a brake sensor.

[0164] May include one or more system - on - chips (SoC) 1404 ( Figure 14C) and / or the controller 1436 of the GPU can provide signals (e.g., representing commands) to one or more components or systems of the vehicle 1400. For example, the controller can send signals to operate the vehicle brakes via one or more brake actuators 1448, to operate the steering system 1454 via one or more steering actuators 1456, and to operate the propulsion system 1450 via one or more throttles / accelerators 1452. The controller 1436 can include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing signals) to drive the vehicle 1400 autonomously and / or assist the driver. The controller 1436 can include a first controller 1436 for autonomous driving functions, a second controller 1436 for functional safety functions, a third controller 1436 for artificial intelligence functions (e.g., computer vision), a fourth controller 1436 for infotainment functions, a fifth controller 1436 for redundancy in emergencies, and / or other controllers. In some examples, a single controller 1436 can handle two or more of the above functions, two or more controllers 1436 can handle a single function, and / or any combination thereof.

[0165] The controller 1436 can provide signals for controlling one or more components and / or systems of the vehicle 1400 in response to sensor data received from one or more sensors (e.g., sensor inputs). The sensor data can be received from, for example but not limited to, global navigation satellite system sensors 1458 (e.g., global positioning system sensors), RADAR sensors 1460, ultrasonic sensors 1462, LIDAR sensors 1464, inertial measurement unit (IMU) sensors 1466 (e.g., accelerometers, gyroscopes, magnetic compasses, magnetometers, etc.), microphones 1496, stereo cameras 1468, wide field of view cameras 1470 (e.g., fish-eye cameras), infrared cameras 1472, surround cameras 1474 (e.g., 360-degree cameras), remote and / or mid-range cameras 1498, speed sensors 1444 (e.g., for measuring the speed of the vehicle 1400), vibration sensors 1442, steering sensors 1440, brake sensors (e.g., as part of the brake sensor system 1446), and / or other sensor types.

[0166] One or more controllers 1436 can receive inputs (e.g., represented by input data) from the instrument cluster 1432 of the vehicle 1400 and provide outputs (e.g., represented by output data, display data, etc.) via the human-machine interface (HMI) display 1434, audible signalers, speakers, and / or via other components of the vehicle 1400. The outputs can include information such as vehicle speed, speed, time, map data (e.g., Figure 14CHD map 1422), location data (e.g., the location of vehicle 1400, e.g., on a map), direction, the locations of other vehicles (e.g., occupied grids), information about objects, and the status of objects sensed by controller 1436, etc. For example, HMI display 1434 can display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.), and / or information about driving a motor vehicle being manufactured, in the process of being manufactured, or to be manufactured (e.g., changing lanes now, exit 34B is two miles away, etc.).

[0167] Vehicle 1400 also includes a network interface 1424, which can communicate through one or more networks using one or more wireless antennas 1426 and / or a modem. For example, network interface 1424 can be capable of communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, etc. Wireless antenna 1426 can also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.), using local area networks such as Bluetooth, Bluetooth LE, Z-Wave, ZigBee, etc., and / or low-power wide area networks (LPWAN), such as LoRaWAN, SigFox, etc.

[0168] Figure 14B is according to some embodiments of the present disclosure Figure 14A An example of the camera positions and fields of view of the exemplary autonomous vehicle 1400 according to the present disclosure. The cameras and the corresponding fields of view are an exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on vehicle 1400.

[0169] The camera types for the cameras can include, but are not limited to, digital cameras that can be suitable for components and / or systems of vehicle 1400. The cameras can operate at automotive safety integrity level (ASIL) B and / or another ASIL. Depending on the embodiment, the camera type can have any image capture rate, such as 60 frames per second (fps), 1420 fps, 240 fps, etc. The cameras may be able to use rolling shutters, global shutters, other types of shutters, or combinations thereof. In some examples, the color filter array can include a red clear clear (RCCC) color filter array, a red clear clear blue (RCCB) color filter array, a red blue green clear (RBGC) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, can be used to increase sensitivity.

[0170] In some examples, one or more cameras can be used to perform Advanced Driver Assistance System (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional monocular camera can be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. One or more cameras (e.g., all cameras) can record and provide image data (e.g., video) simultaneously.

[0171] One or more cameras can be installed in a mounting assembly, such as a custom-designed (3D printed) assembly, to remove stray light and reflections inside the vehicle (e.g., reflections from the dashboard, reflections in the windshield mirror), which may interfere with the camera's image data capture ability. Referring to the wing mirror mounting assembly, the wing mirror assembly can be custom 3D printed such that the camera mounting plate matches the shape of the wing mirror. In some examples, the camera can be integrated into the wing mirror. For side-view cameras, the camera can also be integrated within the four struts at each corner of the cabin.

[0172] A camera having a field of view including a portion of the environment in front of vehicle 1400 (e.g., a forward-facing camera) can be used for surround viewing to help identify forward-facing paths and obstacles, and, with the help of one or more controllers 1436 and / or a control SoC, assist in providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems, including lane departure warning ("LDW"), adaptive cruise control ("ACC"), and / or other functions such as traffic sign recognition.

[0173] Various cameras can be used in a forward-facing configuration, including for example a monocular camera platform including a CMOS (Complementary Metal Oxide Semiconductor) color imager. Another example can be a wide field of view camera 1470, which can be used to sense objects entering the field of view from the periphery (e.g., pedestrians, cross traffic, or bicycles). Although Figure 14B only one wide field of view camera is shown, any number of wide field of view cameras 1470 can be on vehicle 1400. Additionally, a tele camera 1498 (e.g., a long field of view stereo camera pair) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. The tele camera 1498 can also be used for object detection and classification, and basic object tracking.

[0174] One or more stereo cameras 1468 may also be included in a forward-facing configuration. The stereo camera 1468 may include an integrated control unit that includes a scalable processing unit that can provide programmable logic (FPGA) and a multi-core microprocessor with an integrated CAN or Ethernet interface on a single chip. Such a unit can be used to generate a 3-D map of the vehicle environment, including distance estimates for all points in the image. An alternative stereo camera 1468 may include a compact stereo vision sensor that may include two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from the vehicle to a target object, and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. Other types of stereo cameras 1468 may be used in addition to or as an alternative to those described herein.

[0175] Cameras with a field of view that includes a portion of the environment of the side of the vehicle 1400 (e.g., side view cameras) can be used for surround viewing, providing information for creating and updating an occupancy grid, and for generating side collision warnings. For example, surround cameras 1474 (e.g., four surround cameras 1474 as Figure 14B shown) may be positioned on the vehicle 1400. The surround cameras 1474 may include wide field of view cameras 1470, fisheye cameras, 360-degree cameras, and the like. For example, four fisheye cameras may be positioned at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 1474 (e.g., left, right, and rear), and may utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround viewing camera.

[0176] Cameras with a field of view that includes a portion of the environment of the rear of the vehicle 1400 (e.g., rear view cameras) can be used for parking assistance, surround view, rear collision warning, and for creating and updating an occupancy grid. A variety of cameras may be used, including but not limited to cameras that are also suitable as forward-facing cameras (e.g., long-range and / or mid-range cameras 1498, stereo cameras 1468, infrared cameras 1472, etc.), as described herein.

[0177] Figure 14C is in accordance with some embodiments of the present disclosure Figure 14ABlock diagram of an exemplary system architecture of an exemplary autonomous vehicle 1400. It should be understood that this and other arrangements described herein are presented only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, function groupings, etc.) may be used in addition to or instead of those shown, and some elements may be omitted entirely. Moreover, many of the elements described herein are functional entities that may be implemented as discrete or distributed components or in conjunction with other components, and may be implemented in any suitable combination and location. The various functions performed by the entities described herein may be executed by hardware, firmware, and / or software. For example, the various functions may be executed by a processor executing instructions stored in a memory.

[0178] Figure 14C Each of the components, features, and systems of vehicle 1400 in is shown as being connected via bus 1402. Bus 1402 may include a Controller Area Network (CAN) data interface (optionally referred to herein as the "CAN bus"). CAN may be a network within vehicle 1400 that aids in controlling various features and functions of vehicle 1400, such as actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle state indicators. The CAN bus may conform to ASIL B.

[0179] Although bus 1402 is described herein as a CAN bus, this is not restrictive. For example, FlexRay and / or Ethernet may be used in addition to or instead of the CAN bus. Additionally, although a single line is used to represent bus 1402, this is not restrictive. For example, any number of buses 1402 may exist, which may include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 1402 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1402 may be used for collision avoidance functions, and a second bus 1402 may be used for actuation control. In any example, each bus 1402 may communicate with any component of vehicle 1400, and two or more buses 1402 may communicate with the same component. In some examples, each SoC 1404, each controller 1436, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of vehicle 1400) and may be connected to a common bus, such as the CAN bus.

[0180] Vehicle 1400 may include one or more controllers 1436, such as those described herein with respect to Figure 14A the controllers described. The controller 1436 may be used for various functions. The controller 1436 may be coupled to any of a variety of other components and systems of the vehicle 1400 and may be used to control the vehicle 1400, the artificial intelligence of the vehicle 1400, the infotainment of the vehicle 1400, and so on.

[0181] Vehicle 1400 may include a system-on-chip (SoC) 1404. The SoC 1404 may include a CPU 1406, a GPU 1408, a processor 1410, a cache 1412, an accelerator 1414, a data store 1416, and / or other components and features not shown. The SoC 1404 may be used to control the vehicle 1400 in a variety of platforms and systems. For example, the SoC 1404 may be combined with an HD map 1422 in a system (e.g., a system of the vehicle 1400), and the HD map 1422 may obtain map refreshes and / or updates from one or more servers (e.g., Figure 14D the server 1478) via a network interface 1424.

[0182] The CPU(s) 1406 may include a CPU cluster or CPU complex (optionally referred to herein as "CCPLEX"). The CPU 1406 may include multiple cores and / or L2 caches. For example, in some embodiments, the CPU 1406 may include eight cores in a coherent multi-processor configuration. In some embodiments, the CPU 1406 may include four dual-core clusters, each of which has a dedicated L2 cache (e.g., a 2MB L2 cache). The CPU 1406 (e.g., CCPLEX) may be configured to support simultaneous cluster operations such that any combination of the clusters of the CPU 1406 is active at any given time.

[0183] The CPU 1406 can implement power management capabilities including one or more of the following features: individual hardware blocks can be automatically clock-gated during idle to save dynamic power; each core clock can be gated when the core is not actively executing instructions due to the execution of WFI / WFE instructions; each core can be independently power-gated; each core cluster can be independently clock-gated when all inner cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. The CPU 1406 can further implement enhanced algorithms for managing power states, where allowed power states and expected wake-up times are specified, and the hardware / microcode determines the optimal power states for the core, cluster, and CCPLEX inputs. The processing core can support a simplified power state input sequence in software, where the work is offloaded to the microcode.

[0184] The GPU 1408 can include an integrated GPU (optionally referred to herein as "iGPU"). The GPU 1408 can be programmable and can be efficient for parallel workloads. In some examples, the GPU 1408 can use an enhanced tensor instruction set. The GPU 1408 can include one or more streaming microprocessors, where each streaming microprocessor can include an L1 cache (e.g., an L1 cache with a storage capacity of at least 96KB), and two or more streaming microprocessors can share an L2 cache (e.g., an L2 cache with a storage capacity of 512KB). In some embodiments, the GPU 1408 can include at least eight streaming microprocessors. The GPU 1408 can use a compute application programming interface (API). Additionally, the GPU 1408 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0185] The GPU 1408 can be power optimized for optimal performance in automotive and embedded use cases. For example, the GPU 1408 can be fabricated on fin field-effect transistors (FinFETs). However, this is not restrictive, and other semiconductor manufacturing processes can be used to fabricate the GPU 1408. Each streaming microprocessor can include multiple mixed-precision processing cores divided into multiple blocks. For example but not limited to, 64 PF32 cores and 32 PF64 cores can be divided into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR COREs for deep learning matrix arithmetic, an L0 instruction cache, a warp scheduler, a scheduling unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths to provide efficient execution of workloads through hybrid computing and addressing computations. The streaming microprocessor can include independent thread scheduling capabilities to enable finer-grained synchronization and collaboration between parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0186] The GPU 1408 can include high-bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem to provide a peak memory bandwidth of approximately 900GB / second in some examples. In some examples, synchronous graphics random access memory (SGRAM), such as graphics double data rate type five synchronous random access memory (GDDR5), can be used in addition to or as an alternative to HBM memory.

[0187] The GPU 1408 can include unified memory technology that includes access counters to allow more accurate migration of memory pages to the processors that access them most frequently, thereby improving the efficiency of the memory range shared between processors. In some examples, address translation service (ATS) support can be used to allow the GPU 1408 to directly access the CPU 1406 page table. In such an example, when the GPU 1408 memory management unit (MMU) experiences a miss, an address translation request can be sent to the CPU 1406. In response, the CPU 1406 can look at its table of pages for the virtual-to-physical mapping of the address and send the translation back to the GPU 1408. Thus, the unified memory technology can allow a single unified virtual address space for the memories of the CPU 1406 and the GPU 1408, thereby simplifying GPU 1408 programming and application porting to the GPU 1408.

[0188] Additionally, GPU 1408 may include an access counter that can track the frequency of GPU 1408's access to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses the pages.

[0189] SoC 1404 may include any number of caches 1412, including those described herein. For example, cache 1412 may include an L3 cache that can be used by both CPU 1406 and GPU 1408 (e.g., which connects both CPU 1406 and GPU 1408). Cache 1412 may include a write-back cache that can track the state of lines, e.g., by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). Depending on the embodiment, the L3 cache may include 4MB or more, although smaller cache sizes may be used.

[0190] SoC 1404 may include one or more accelerators 1414 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 1404 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memories. The large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster may be used to supplement GPU 1408 and offload some of the tasks of GPU 1408 (e.g., to free up more cycles of GPU 1408 to perform other tasks). As an example, accelerator 1414 may be used for target workloads (e.g., perception, convolutional neural network (CNN), etc.) that are stable enough to be suitable for acceleration. The term "CNN" as used herein may include all types of CNNs, including region-based or region-based convolutional neural networks (RCNNs) and fast RCNNs (e.g., for object detection).

[0191] Accelerator 1414 (e.g., the hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) that may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. The TPU may be an accelerator configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA may also be optimized for a specific set of neural network types and floating-point operations, as well as inference. The design of the DLA may provide higher performance per millimeter than a general-purpose GPU and far exceed the performance of a CPU. The TPU may perform several functions, including single-instance convolution functions, e.g., supporting INT8, INT16, and FP16 data types for features and weights, as well as post-processor functions.

[0192] For any of a variety of functions, the DLA can quickly and efficiently execute neural networks, especially CNNs, on processed or unprocessed data, including, for example, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, recognition, and detection using data from microphones; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or security-related events.

[0193] The DLA can perform any function of the GPU 1408, and by using an inference accelerator, for example, a designer can target the DLA or the GPU 1408 for any function. For example, a designer can concentrate the processing of CNNs and floating-point operations on the DLA and leave other functions to the GPU 1408 and / or other accelerators 1414.

[0194] The accelerator 1414 (e.g., a hardware acceleration cluster) can include a programmable vision accelerator (PVA), which may alternatively be referred to herein as a computer vision accelerator. The PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA can include, for example but not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0195] The RISC cores can interact with image sensors (e.g., the image sensors of any of the cameras described herein), image signal processors, etc. Each RISC core can include any number of memories. According to an embodiment, the RISC cores can use any of a number of protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include instruction caches and / or tightly coupled RAM.

[0196] The DMA can enable the components of the PVA to access system memory independently of the CPU 1406. The DMA can support any number of features for optimizing the supply to the PVA, including but not limited to, supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which can include block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0197] A vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, the PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. The vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core can include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can improve throughput and speed.

[0198] Each vector processor can include an instruction cache and can be coupled to dedicated memory. As a result, in some examples, each vector processor can be configured to execute independently of other vector processors. In other examples, the vector processors included in a particular PVA can be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA can execute the same computer vision algorithm but on different regions of an image. In other examples, the vector processors included in a particular PVA can execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on sequential images or portions of an image. Any number of PVAs can be included in a hardware acceleration cluster, and any number of vector processors can be included in each PVA. Additionally, the PVA can include additional error correction code (ECC) memory to enhance overall system security.

[0199] Accelerator 1414 (e.g., a hardware acceleration cluster) can include an on-chip computer vision network and SRAM to provide high-bandwidth, low-latency SRAM for accelerator 1414. In some examples, the on-chip memory can include at least 4MB of SRAM, which includes, for example but not limited to, eight field-configurable memory blocks that can be accessed by both the PVA and the DLA. Each pair of memory blocks can include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVA and the DLA can access the memory through a backbone network that provides high-speed access to the memory for the PVA and the DLA. The backbone can include an on-chip computer vision network that interconnects the PVA and the DLA to the memory (e.g., using APB).

[0200] The on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. Although other standards and protocols may be used, this type of interface may comply with the ISO 26262 or IEC 61508 standards.

[0201] In some examples, the SoC 1404 may include a real-time ray tracing hardware accelerator, such as that described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. The real-time ray-tracing hardware accelerator can be used to quickly and efficiently determine the position and extent of an object (e.g., within a world model), generate real-time visualization simulations, for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulating SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.

[0202] The accelerator 1414 (e.g., a hardware accelerator cluster) has a wide range of uses for autonomous driving. The PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. The functions of the PVA are well-suited for algorithm domains that require predictable processing, low power, and low latency. In other words, the PVA performs well on semi-dense or dense regular computations, even on small data sets that require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are effective in object detection and integer math operations.

[0203] For example, according to one embodiment of the present technology, the PVA is used to perform computer stereo vision. In some examples, an algorithm based on semi-global matching may be used, but this is not restrictive. Many applications for 3-5 level autonomous driving require dynamic estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0204] In some instances, the PVA can be used to perform dense optical flow. The processed RADAR is provided by processing raw RADAR data (e.g., using a 4D fast Fourier transform). In other examples, e.g., by processing raw time-of-flight data to provide processed time-of-flight data, the PVA is used for time-of-flight depth processing.

[0205] DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence metric for each object detection. Such confidence values can be interpreted as probabilities or provide a relative "weight" for each detection compared to other detections. This confidence value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a threshold for the confidence and only consider detections above the threshold as true positive detections. In an automatic emergency braking (AEB) system, false positive detections would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can run a neural network to regress the confidence value. The neural network can take at least some subset of parameters as its input, such as bounding box dimensions, an obtained ground plane estimate (e.g., from another subsystem), the output of an inertial measurement unit (IMU) sensor 1466, which is related to the orientation, distance, 3D position estimate, etc. of the vehicle 1400 of an object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 1464 or RADAR sensor 1460).

[0206] SoC 1404 can include data storage 1416 (e.g., memory). The data storage 1416 can be on-chip memory of the SoC 1404, which can store neural networks to be executed on the GPU and / or DLA. In some examples, the data storage 1416 can be large enough to store multiple instances of the neural network for redundancy and safety. The data storage 1412 can include an L2 or L3 cache 1412. References to the data storage 1416 can include references to memory associated with the PVA, DLA, and / or other accelerators 1414, as described herein.

[0207] SoC 1404 may include one or more processors 1410 (e.g., embedded processors). The processor 1410 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions as well as associated security implementations. The boot and power management processor may be part of the SoC 1404 boot sequence and may provide runtime power management services. The boot power and management processor may provide clock and voltage programming, assist in system low-power state transitions, manage the SoC 1404 heat and temperature sensors, and / or manage the power state of the SoC 1404. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1404 may use the ring oscillator to detect the temperatures of the CPU 1406, GPU 1408, and accelerator 1414. If it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine and place the SoC 1404 in a lower power state and / or place the vehicle 1400 in a driver-to-safe stop mode (e.g., safely stop the vehicle 1400).

[0208] The processor 1410 may also include a set of embedded processors that can be used as an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio through multiple interfaces, as well as a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0209] The processor 1410 may also include an always-on processor engine, which may provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0210] The processor 1410 may also include a security cluster engine, which includes a dedicated processor subsystem to handle the security management of automotive applications. The security cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a security mode, two or more cores may operate in a lockstep mode and act as a single core with comparison logic to detect any differences between their operations.

[0211] The processor 1410 may also include a real-time camera engine, which may include a dedicated processor subsystem for handling real-time camera management.

[0212] The processor 1410 may also include a high dynamic range signal processor, which may include an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0213] The processor 1410 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required for a video playback application to produce the final image of the player window. The video image synthesizer may perform lens distortion correction on the wide field of view camera 1470, the surround camera 1474, and / or the in-cabin surveillance camera sensors. The in-cabin surveillance camera sensors are preferably monitored by a neural network running on another instance of the advanced SoC and are configured to identify and respond accordingly in in-cabin events. The in-cabin system may perform lip reading to activate cellular services and make calls, indicate emails, change the destination of the vehicle, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions may only be available to the driver when the vehicle is operating in autonomous mode and will otherwise be disabled.

[0214] The video image synthesizer may include enhanced temporal noise reduction for spatial and temporal noise reduction. For example, in the case of motion occurring in the video, the noise reduction appropriately weights the spatial information, thereby reducing the weight of the information provided by adjacent frames. In the case where an image or a portion of the image does not include motion, the temporal noise reduction performed by the video image synthesizer may use information from a previous image to reduce the noise in the current image.

[0215] The video image synthesizer may also be configured to perform stereo correction on the input stereo lens frames. When the operating system desktop is in use, the video image synthesizer may also be used for user interface synthesis, and the GPU 1408 does not need to continuously render new surfaces. Even when the GPU 1408 is powered on and active for 3D rendering, the video image synthesizer may be used to offload the GPU 1408 to improve performance and responsiveness.

[0216] The SoC 1404 may also include a Mobile Industry Processor Interface (MIPI) camera serial interface for receiving video and inputs from cameras, high-speed interfaces, and / or video input blocks that may be used for cameras and related pixel input functions. The SoC 1404 may also include an input / output controller, which may be software-controlled and may be used to receive I / O signals that are not committed to a specific role.

[0217] SoC 1404 may also include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 1404 can be used to process data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet connections), sensors (e.g., LIDAR sensor 1464, RADAR sensor 1460, etc., which can be connected via Ethernet), data from bus 1402 (e.g., the speed of vehicle 1400, steering wheel position, etc.), and data from GNSS sensor 1458 (e.g., via Ethernet or CAN bus connection). SoC 1404 may also include dedicated high-performance large-capacity storage controllers, which can include their own DMA engines and can be used to free the CPU 1406 from routine data management tasks.

[0218] SoC 1404 can be an end-to-end platform with a flexible architecture spanning automation levels 3 - 5, thus providing an integrated functional safety architecture that leverages and effectively utilizes computer vision and ADAS technologies to achieve diversity and redundancy, provides a platform for a flexible and reliable driving software stack, and deep learning tools. SoC 1404 can be faster, more reliable, and even more energy-efficient and space-saving than traditional systems. For example, accelerator 1414, when combined with CPU 1406, GPU 1408, and data storage 1416, can provide a fast and effective platform for level 3 - 5 autonomous vehicles.

[0219] Therefore, this technology provides capabilities and functions that traditional systems cannot achieve. For example, computer vision algorithms can be executed on the CPU, which can be configured using a high-level programming language (such as the C programming language) to perform various processing algorithms across multiple visual data. However, the CPU generally cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In particular, many CPUs cannot execute complex object detection algorithms in real time, which is a requirement for in-vehicle ADAS applications and also for actual level 3 - 5 autonomous vehicles.

[0220] Compared with traditional systems, the technology described herein allows multiple neural networks to be executed simultaneously and / or sequentially by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, and combines the results to enable level 3 - 5 autonomous driving functions. For example, a CNN executed on the DLA or dGPU (e.g., GPU 1420) can include text and word recognition, allowing a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained on. The DLA can also include neural networks that are capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing that semantic understanding to the path planning module running on the CPU complex.

[0221] As another example, multiple neural networks can be run simultaneously, such as required for level 3, level 4, or level 5 driving. For example, a warning sign that includes "Caution: Flashing lights indicate icy conditions" along with electric lights can be interpreted independently or jointly by several neural networks. The sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), and the text "Flashing lights indicate icy conditions" can be interpreted by a second deployed neural network, which notifies the vehicle's path planning software (preferably executed on the CPU Complex) that when a flash is detected, there is an icy condition. The flash can be recognized by operating a third deployed neural network over multiple frames, thereby notifying the vehicle's path planning software of the presence (or absence) of the flash. All three neural networks can run simultaneously, for example, within the DLA and / or on the GPU 1408.

[0222] In some examples, the CNNs used for face recognition and owner recognition can use data from the camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 1400. When the vehicle owner approaches the driver's door and turns on the vehicle lights, the always-on sensor processing engine can be used to unlock the vehicle and, in a secure mode, disable the vehicle when the owner leaves the vehicle. In this way, the SoC 1404 provides security against theft and / or carjacking.

[0223] In another example, the CNN used for emergency vehicle detection and recognition can use data from the microphone 1496 to detect and recognize emergency vehicle sirens. Compared to traditional systems that use general classifiers to detect sirens and manually extract features, the SoC 1404 uses the CNN to classify ambient and city sounds, as well as classify visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative closing speed of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by the GNSS sensor 1458. Thus, for example, when operating in Europe, the CNN will seek to detect European sirens, while in the United States, the CNN will seek to identify only North American sirens. Once an emergency vehicle is detected, a control program can be used to execute emergency vehicle safety routines, causing the vehicle to slow down, pull over to the side of the road, park the vehicle, and / or idle the vehicle with the assistance of the ultrasonic sensors 1462 until the emergency vehicle has passed.

[0224] The vehicle may include a CPU 1418 (e.g., a discrete CPU or dCPU), which may be coupled to the SoC 1404 via a high-speed interconnect (e.g., PCIe). For example, the CPU 1418 may include an X86 processor. For example, the CPU 1418 may be used to perform any of a variety of functions, including arbitrating inconsistent results between the ADAS sensors and the SoC 1404, and / or monitoring the status and health of the controller 1436 and / or the infotainment SoC 1430.

[0225] The vehicle 1400 may include a GPU 1420 (e.g., a discrete GPU or dGPU), which may be coupled to the SoC 1404 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1420 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of the vehicle 1400.

[0226] The vehicle 1400 may also include a network interface 1424, which may include one or more wireless antennas 1426 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). The network interface 1424 may be used to wirelessly connect to other vehicles and / or to computing devices (e.g., client devices) via the Internet with a cloud (e.g., with the server 1478 and / or other network devices). To communicate with other vehicles, a direct link may be established between two vehicles and / or an indirect link may be established (e.g., across a network and through the Internet). A vehicle-to-vehicle communication link may be used to provide a direct link. The vehicle-to-vehicle communication link may provide the vehicle 1400 with information about vehicles in the vicinity of the vehicle 1400 (e.g., vehicles in front of, to the side of, and / or behind the vehicle 1400). This function may be part of the cooperative adaptive cruise control function of the vehicle 1400.

[0227] The network interface 1424 may include an SoC that provides modulation and demodulation functions and enables the controller 1436 to communicate via a wireless network. The network interface 1424 may include a radio frequency front end for upconverting from baseband to radio frequency and downconverting from radio frequency to baseband. The frequency conversion may be performed by well-known processes and / or may be performed using a superheterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0228] Vehicle 1400 may also include a data store 1428, which may include off-chip (e.g., outside of SoC 1404) storage. The data store 1428 may include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard drives, and / or other components and / or devices that can store at least one bit of data.

[0229] Vehicle 1400 may also include a GNSS sensor 1458. The GNSS sensor 1458 (e.g., GPS and / or assisted GPS sensors) to assist with mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 1458 may be used, including, for example but not limited to, a GPS using a USB connector with an Ethernet to serial (RS-232) bridge.

[0230] Vehicle 1400 may also include a RADAR sensor 1460. Vehicle 1400 may use the RADAR sensor 1460 for remote vehicle detection even in dark and / or adverse weather conditions. The RADAR functional safety level may be ASIL B. The RADAR sensor 1460 may use CAN and / or bus 1402 (e.g., to transmit data generated by the RADAR sensor 1460) to control and access object tracking data and, in some examples, may access Ethernet to access raw data. Various RADAR sensor types may be used. For example but not limited to, the RADAR sensor 1460 may be suitable for front, rear, and side RADAR use. In some examples, a pulsed Doppler RADAR sensor is used.

[0231] The RADAR sensor 1460 may include different configurations, such as long range with a narrow field of view, short range with a wide field of view, short range side coverage, etc. In some examples, the long-range RADAR may be for an adaptive cruise control function. The long-range RADAR system may provide a wide field of view achieved through two or more independent scans, e.g., within a range of 250 meters. The RADAR sensor 1460 may help distinguish between stationary and moving objects and may be used by the ADAS system for emergency braking assistance and forward collision warning. The long-range RADAR sensor may include a multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In an example with six antennas, the central four antennas may create a focused beam pattern that is designed to record the surrounding environment of vehicle 1400 at a higher speed with minimal interference from traffic in adjacent lanes. The other two antennas may extend the field of view so that vehicles entering or leaving the lane of vehicle 1400 can be quickly detected.

[0232] As an example, a mid-range RADAR system can include, for example, a range of up to 1460 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 1450 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can create two crossbeams that continuously monitor the rear of the vehicle and the blind spots alongside the vehicle.

[0233] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assistance.

[0234] Vehicle 1400 can also include ultrasonic sensors 1462. The ultrasonic sensors 1462 can be positioned at the front, rear, and / or sides of the vehicle 1400 and can be used for parking assistance and / or creating and updating an occupancy grid. A wide variety of ultrasonic sensors 1462 can be used, and different ultrasonic sensors 1462 can be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensors 1462 can operate at a functional safety level of ASIL B.

[0235] Vehicle 1400 can include a LIDAR sensor 1464. The LIDAR sensor 1464 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1464 can be of functional safety level ASIL B. In some examples, vehicle 1400 can include multiple LIDAR sensors 1464 (e.g., two, four, six, etc.) that can use Ethernet (e.g., for providing data to a gigabit Ethernet switch).

[0236] In some examples, the LIDAR sensor 1464 is capable of providing a list of objects and the distance to the objects from a 360-degree field of view. Commercial LIDAR sensors 1464 can have a notification range of approximately 1400 m, an accuracy of 2 cm - 3 cm, and, for example, support a 1400 Mbps Ethernet connection. In some examples, one or more non-protruding LIDAR sensors 1464 can be used. In such examples, the LIDAR sensor 1464 can be implemented as a small device that can be embedded in the front, rear, sides, and / or corners of the vehicle 1400. In such examples, the LIDAR sensor 1464 can provide a horizontal field of view of up to 1420 degrees and a vertical field of view of 35 degrees and can provide a range of 200 meters even for low-reflectivity objects. The front LIDAR sensor 1464 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0237] In some examples, LIDAR technology can also be used, such as 3D flash laser LIDAR. 3D flash laser LIDAR uses a laser flash lamp as the emission source, illuminating approximately 200 meters around the vehicle. The flash laser LIDAR unit includes a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle to the object. Flash laser LIDAR can allow for the generation of a highly accurate and distortion-free image of the surrounding environment for each laser flash. In some examples, four flash laser LIDAR sensors can be deployed, one on each side of the vehicle 1400. Available 3D flash laser LIDAR systems include solid-state 3D staring array LIDAR cameras, which have no moving parts other than a fan (e.g., non-scanning LIDAR devices). The flash laser LIDAR device can use 5-nanosecond-level I (eye-safe) laser pulses per frame and can capture the reflected laser in the form of 3D range point clouds and co-registered intensity data. By using flash laser LIDAR, and since flash laser LIDAR is a solid-state device with no moving parts, the LIDAR sensor 1464 may be less susceptible to motion blur, vibration, and / or shock.

[0238] The vehicle can also include an IMU sensor 1466. In some examples, the IMU sensor 1466 can be located at the center of the rear axle of the vehicle 1400. The IMU sensor 1466 can include, for example but not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, such as in a six-axis application, the IMU sensor 1466 can include an accelerometer and a gyroscope, while in a nine-axis application, the IMU sensor 1466 can include an accelerometer, a gyroscope, and a magnetometer.

[0239] In some embodiments, the IMU sensor 1466 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS), which combines microelectromechanical systems (MEMS) inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. Thus, in some examples, the IMU sensor 1466 can enable the vehicle 1400 to estimate its heading without the need for input from a magnetic sensor by directly observing and correlating the speed changes from GPS to the IMU sensor 1466. In some examples, the IMU sensor 1466 and the GNSS sensor 1458 can be combined in a single integrated unit.

[0240] The vehicle can include microphones 1496 placed in and / or around the vehicle 1400. The microphones 1496 can be used for emergency vehicle detection and identification, etc.

[0241] The vehicle may also include any number of camera types, including a stereo camera 1468, a wide field of view camera 1470, an infrared camera 1472, a surround camera 1474, a long-range and / or mid-range camera 1498, and / or other camera types. The cameras can be used to capture image data around the entire perimeter of the vehicle 1400. The camera types used depend on the embodiment and requirements of the vehicle 1400, and any combination of camera types can be used to provide the necessary coverage around the vehicle 1400. Additionally, according to an embodiment, the number of cameras can vary. For example, the vehicle may include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. By way of example and not limitation, the cameras can support Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet. Each camera is described in more detail herein with respect to Figure 14A and Figure 14B Each camera is described in more detail herein with respect to

[0242] The vehicle 1400 may also include a vibration sensor 1442. The vibration sensor 1442 can measure the vibration of a component of the vehicle (such as an axle). For example, a change in vibration can indicate a change in the road surface. In another example, when two or more vibration sensors 1442 are used, the difference between the vibrations can be used to determine the friction or slip of the road surface (e.g., when the difference in vibration is between a powered drive axle and a freely rotating axle).

[0243] The vehicle 1400 may include an ADAS system 1438. In some examples, the ADAS system 1438 may include a SoC. The ADAS system 1438 may include autonomous / adaptive / automated cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keeping assist (LKA), blind spot warning (BSW), rear cross traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0244] The ACC system can use RADAR sensors 1460, LIDAR sensors 1464, and / or cameras. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle immediately in front of the vehicle 1400 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance keeping and advises the vehicle 1400 to change lanes when necessary. Lateral ACC is related to other ADAS applications, such as LCA and CWS.

[0245] The CACC uses information from other vehicles, which can be received from other vehicles via the network interface 1424 and / or the wireless antenna 1426 via a wireless link, or indirectly via a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. Generally, the V2V communication concept provides information about the immediately preceding vehicle (e.g., the vehicle immediately before vehicle 1400 and in the same lane as vehicle 1400), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of the I2V and V2V information sources. Given information about the vehicle ahead of vehicle 1400, the CACC can be more reliable and it has the potential to improve traffic flow and reduce road congestion.

[0246] The FCW system is designed to warn the driver of a hazard so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 1460, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component. The FCW system can provide warnings, such as in the form of a sound, visual warning, vibration, and / or rapid braking pulse.

[0247] The AEB system detects an impending forward collision with another vehicle or other object, and if the driver does not take corrective action within a specified time or distance parameter, the AEB system can automatically apply the brakes. The AEB system can use a forward-facing camera and / or RADAR sensor 1460, coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a hazard, it typically first warns the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system can include technologies such as dynamic brake support and / or braking for an impending collision.

[0248] The LDW system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibration, to warn the driver when vehicle 1400 crosses a lane marking. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn signal. The LDW system can use a forward-side camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.

[0249] The LKA system is a variant of the LDW system. If vehicle 1400 starts to leave the lane, the LKA system provides steering input or braking to correct vehicle 1400.

[0250] The BSW system detects and warns the driver of vehicles in the vehicle's blind spot. The BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. When the driver uses a turn signal, the system can provide an additional warning. The BSW system can use a rear-facing camera and / or RADAR sensor 1460, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration component.

[0251] When the vehicle 1400 is backing up, the RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the rear-facing camera's range. Some RCTW systems include AEB to ensure that the vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 1460, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibration element.

[0252] Traditional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but are generally not catastrophic because the ADAS system warns the driver and allows the driver to determine whether a safe condition truly exists and act accordingly. However, in an autonomous vehicle 1400, in the case of conflicting results, the vehicle 1400 itself must decide whether to heed the results from the primary computer or the secondary computer (e.g., the first controller 1436 or the second controller 1436). For example, in some embodiments, the ADAS system 1438 can be a backup and / or secondary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can run various software redundantly on hardware components to detect faults in perception and dynamic driving tasks. The output from the ADAS system 1438 can be provided to the supervisory MCU. If the outputs of the primary computer and the secondary computer conflict, the supervisory MCU must determine how to reconcile the conflict to ensure safe operation.

[0253] In some examples, the primary computer can be configured to provide a confidence score to the supervisory MCU, indicating the primary computer's confidence in the selected result. If the confidence score exceeds a threshold, the supervisory MCU can follow the primary computer's instructions, regardless of whether the secondary computer provides conflicting or inconsistent results. In the case where the confidence score does not meet the threshold and the primary computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.

[0254] The monitoring MCU can be configured to run a neural network that is trained and configured to determine the conditions under which the auxiliary computer provides a false alarm based on the outputs from the host computer and the auxiliary computer. Thus, the neural network in the monitoring MCU can learn when the output of the auxiliary computer is trustworthy and when it is not. For example, when the auxiliary computer is a RADAR-based FCW system, the neural network in the monitoring MCU can learn when the FCW system identifies a metallic object that is actually not dangerous, such as a drain grate or manhole cover that triggers an alarm. Similarly, when the auxiliary computer is a camera-based LDW system, the neural network in the monitoring MCU can learn to override the LDW when there is a bicyclist or pedestrian present and the lane departure is actually the safest maneuver. In embodiments that include a neural network running on the monitoring MCU, the monitoring MCU can include at least one of a DLA or a GPU suitable for running a neural network with associated memory. In a preferred embodiment, the monitoring MCU can be included as and / or included in components of the SoC 1404.

[0255] In other examples, the ADAS system 1438 can include an auxiliary computer that uses traditional computer vision rules to perform ADAS functions. In this way, the auxiliary computer can use classical computer vision rules (if - then), and the presence of the neural network in the monitoring MCU can improve reliability, safety, and performance. For example, the diverse implementations and deliberate non-identities make the overall system more fault-tolerant, especially faults caused by software (or software-hardware interface) functions. For example, if there is a software vulnerability or error in the software running on the host computer, and the different software code running on the auxiliary computer provides the same overall result, the monitoring MCU may be more confident that the overall result is correct and that the vulnerability in the software or hardware on the host computer will not cause a major error.

[0256] In some examples, the output of the ADAS system 1438 can be fed into the perception block of the host computer and / or the dynamic driving task block of the host computer. For example, if the ADAS system 1438 indicates a forward collision warning due to an object immediately in front, the perception block can use this information when identifying the object. In other examples, as described herein, the auxiliary computer can have its own trained neural network, thereby reducing the risk of false alarms.

[0257] Vehicle 1400 may also include an infotainment system SoC 1430 (e.g., in-vehicle infotainment (IVI)). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more discrete components. The infotainment SoC 1430 may include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assist, radio data system, vehicle-related information such as fuel level, total distance covered, brake fuel level, oil level, door open / closed, air filter information, etc.) to vehicle 1400. For example, the infotainment system SoC 1430 may be a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-vehicle, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, head-up display (HUD), HMI display 1434, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment system SoC 1430 may also be used to provide information (e.g., visual and / or auditory) to a user of the vehicle, such as information from the ADAS system 1438, autonomous driving information such as planned vehicle maneuvers, trajectories, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0258] The infotainment SoC 1430 may include GPU functionality. The infotainment system SoC 1430 may communicate with other devices, systems, and / or components of vehicle 1400 via a bus 1402 (e.g., CAN bus, Ethernet, etc.). In some examples, the infotainment system SoC 1430 may be coupled to a supervisory MCU such that the GPU of the infotainment system can perform some self-driving functions in the event of a failure of the main controller 1436 (e.g., the primary and / or backup computer of vehicle 1400). In such examples, the infotainment system SoC 1430 may place vehicle 1400 in a driver-to-safe stop mode as described herein.

[0259] The vehicle 1400 may also include an instrument cluster 1432 (e.g., a digital instrument panel, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1432 may include a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). The instrument cluster 1432 may include a set of instruments, such as a speedometer, a fuel level, an oil pressure, a tachometer, an odometer, a turn indicator, a shift position indicator, a seat belt warning light, a parking brake warning light, an engine-functional light, an airbag (SRS) system information, a lighting control, a safety system control, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment system SoC 1430 and the instrument cluster 1432. In other words, the instrument cluster 1432 may be included as part of the infotainment system SoC 1430, and vice versa.

[0260] Figure 14D According to some embodiments of the present disclosure, Figure 14A 14. System diagram of communication between cloud-based servers of an exemplary autonomous vehicle 1400. System 1476 may include servers 1478, network 1490, and vehicles, including vehicle 1400. Server 1478 may include multiple GPUs 1484(A)-1484(H) (collectively referred to herein as GPUs 1484), PCIe switches 1482(A)-1482(H) (collectively referred to herein as PCIe switches 1482), and / or CPUs 1480(A)-1480(B) (collectively referred to herein as CPUs 1480). GPUs 1484, CPUs 1480, and PCIe switches may be interconnected with a high-speed interconnect, such as, but not limited to, NVLink interface 1488 developed by NVIDIA and / or PCIe connection 1486. ​​In some examples, GPU 1484 is connected via NVLink and / or NVSwitch SoC and GPU 1484 and PCIe switch 1482 are connected via PCIe interconnect. Although eight GPUs 1484, two CPUs 1480, and two PCIe switches are shown, this is not limiting. Depending on the embodiment, each server 1478 may include any number of GPUs 1484, CPUs 1480, and / or PCIe switches. For example, the servers 1478 may each include eight, sixteen, thirty-two, and / or more GPUs 1484.

[0261] Server 1478 can receive image data representing unexpected or changed road conditions, such as recently started road works, from a vehicle via network 1490. Server 1478 can send, via network 1490, to the vehicle, neural network 1492, updated neural network 1492, and / or map information 1494, information including traffic and road conditions. Updates to map information 1494 can include updates to HD map 1422, such as information about construction sites, potholes, curves, floods, and / or other obstacles. In some examples, neural network 1492, updated neural network 1492, and / or map information 1494 can be generated by new training and / or experiences represented in data received from any number of vehicles in the environment, and / or based on training performed in a data center (e.g., using server 1478 and / or other servers).

[0262] Server 1478 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by vehicles and / or can be generated in simulations (e.g., using a game engine). In some examples, the training data is labeled (e.g., neural networks benefit from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is unlabeled and / or not preprocessed (e.g., neural networks do not require supervised learning). Once the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1490), and / or the machine learning model can be used by server 1478 to remotely monitor the vehicle.

[0263] In some examples, server 1478 can receive data from a vehicle and apply the data to a state-of-the-art real-time neural network for real-time intelligent inference. Server 1478 can include a deep learning supercomputer powered by GPU 1484 and / or a dedicated AI computer, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1478 can include a deep learning infrastructure of a data center powered only by a CPU.

[0264] The deep learning infrastructure of server 1478 can be capable of performing fast, real-time inference and can use this ability to evaluate and verify the health of processors, software, and / or associated hardware. For example, the deep learning infrastructure can receive periodic updates from vehicle 1400, such as a series of images and / or objects that vehicle 1400 has located within the image sequence (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them to the objects identified by vehicle 1400, and if the results do not match and the infrastructure determines that the AI in vehicle 1400 has malfunctioned, server 1478 can send a signal to vehicle 1400 instructing the fail-safe computer in vehicle 1400 to take control, notify the passengers, and complete a safe parking maneuver.

[0265] For inference, server 1478 can include GPU 1484 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT). The combination of a GPU-powered server and inference acceleration can enable real-time response. In other examples, such as when performance is less important, a server powered by a CPU, FPGA, and other processors can be used for inference.

[0266] Exemplary computing device

[0267] Figure 15 is a block diagram of an exemplary computing device 1500 suitable for implementing some embodiments of the present disclosure. Computing device 1500 can include a bus 1502 that directly or indirectly couples the following devices: a memory 1504, one or more central processing units (CPUs) 1506, one or more graphics processing units (GPUs) 1508, a communication interface 1510, input / output (I / O) ports 1512, input / output components 1514, a power supply 1516, and one or more presentation components 1518 (e.g., a display).

[0268] Although Figure 15 the various blocks are shown connected by lines via bus 1502, this is not restrictive and is for clarity only. For example, in some embodiments, a presentation component 1518 (such as a display device) can be considered an I / O component 1514 (e.g., if the display is a touchscreen). As another example, CPU 1506 and / or GPU 1508 can include memory (e.g., in addition to the memory of GPU 1508, CPU 1506, and / or other components, memory 1504 can represent a storage device). In other words, Figure 15The computing device is merely illustrative. There is no distinction between categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types, as expected within the scope of the computing device in Figure 15 the computing device.

[0269] Bus 1502 may represent one or more buses, such as an address bus, a data bus, a control bus, or a combination thereof. Bus 1502 may include one or more bus types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or other types of buses.

[0270] Memory 1504 may include any of a variety of computer-readable media. Computer-readable media can be any available media that can be accessed by computing device 1500. Computer-readable media can include volatile and non-volatile media, as well as removable and non-removable media. By way of example and not limitation, computer-readable media may include computer storage media and communication media.

[0271] Computer storage media can include volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 1504 may store computer-readable instructions (e.g., representing programs and / or program elements such as an operating system). Computer storage media can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other storage technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic tape cartridges, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by computing device 1500. As used herein, computer storage media does not itself include signals.

[0272] A communication medium can embody computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave or other transmission mechanism, and includes any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a way as to encode information in the signal. By way of example, and not limitation, communication media can include wired media, such as a wired network or direct wired connection, and wireless media, such as acoustic, RF, infrared, and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0273] The CPU 1506 can be configured to execute computer-readable instructions to control one or more components of the computing device 1500 to perform one or more of the methods and / or processes described herein. Each CPU 1506 can include one or more cores (e.g., one, two, four, eight, twenty-eight, seventy-two, etc.) capable of processing multiple software threads simultaneously. The CPU 1506 can include any type of processor and can include different types of processors depending on the type of computing device 1500 being implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 1500, the processor can be an ARM processor implemented using reduced instruction set computing (RISC) or an x86 processor implemented using complex instruction set computing (CISC). In addition to one or more microprocessors or auxiliary co-processors (e.g., a math co-processor), the computing device 1500 can include one or more CPUs 1506.

[0274] The computing device 1500 can use the GPU 1508 to render graphics (e.g., 3D graphics). The GPU 1508 can include hundreds or thousands of cores capable of processing hundreds or thousands of software threads simultaneously. The GPU 1508 can generate pixel data for an output image in response to a rendering command (e.g., received from the CPU 1506 via a host interface). The GPU 1508 can include a graphics memory for storing pixel data, such as a display memory. The display memory can be included as part of the memory 1504. The GPU 708 can include two or more GPUs operating in parallel (e.g., via a link). When combined, each GPU 1508 can generate pixel data for different parts of an output image or for different output images (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or can share memory with other GPUs.

[0275] In an example where computing device 1500 does not include GPU 1508, CPU 1506 can be used to render graphics.

[0276] Communication interface 1510 can include one or more receivers, transmitters, and / or transceivers that enable computing device 700 to communicate with other computing devices via an electronic communication network (including wired and / or wireless communication). Communication interface 1510 can include components and functionality for enabling communication over any of a plurality of different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth, Bluetooth LE, ZigBee, etc.), wired networks (e.g., communication over Ethernet), low-power wide area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0277] I / O port 1512 can enable computing device 1500 to be logically coupled to other devices (including I / O components 1514, presentation components 1518, and / or other components), some of which may be built into (e.g., integrated within) computing device 1500. Exemplary I / O components 1514 include microphones, mice, keyboards, joysticks, gamepads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. I / O components 1514 can provide a natural user interface (NUI) for processing air gestures, voice, or other physiological inputs generated by a user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of speech recognition, stylus recognition, face recognition, biometric recognition, on-screen and near-screen gesture recognition, air gestures, head and eye tracking, and touch recognition associated with the display of computing device 1500 (as described in more detail below). Computing device 1500 can include a depth camera, such as a stereo camera system, an infrared camera system, an RGB camera system, touchscreen technology, and combinations thereof, for gesture detection and recognition. Additionally, computing device 1500 can include an accelerometer or gyroscope capable of detecting motion (e.g., as part of an inertial measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by computing device 1500 to present immersive augmented reality or virtual reality.

[0278] Power supply 1516 can include hardwired power, battery power, or a combination thereof. Power supply 1516 can provide power to computing device 1500 to enable the components of computing device 1500 to operate.

[0279] The presentation component 1518 may include a display (e.g., a monitor, a touch screen, a television screen, a head-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 1518 may receive data from other components (e.g., the GPU 1508, the CPU 1506, etc.) and output data (e.g., as an image, a video, a sound, etc.).

[0280] The present disclosure may be described in the general context of computer code or machine-usable instructions, including computer-executable instructions, such as program modules, executed by a computer or other machine, such as a personal data assistant or other handheld device. Generally, program modules, including routines, programs, objects, components, data structures, etc., refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be practiced in a variety of system configurations, including handheld devices, consumer electronics, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote-processing devices linked through a communications network.

[0281] As used herein, the recitation of “and / or” with respect to two or more elements should be construed as merely indicating one element or a combination of elements. For example, “element A, element B, and / or element C” may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or element A, B, and C. In addition, “at least one of element A or element B” may include at least one of element A, at least one of element B, or at least one of element A and element B. In addition, “at least one of element A and element B” may include at least one of element A, at least one of element B, or at least one of element A and element B.

[0282] The subject matter of the present disclosure is specifically described herein to meet statutory requirements. However, the description itself is not intended to limit the scope of the present disclosure. Instead, the inventors have contemplated that the claimed subject matter may also be embodied in other ways, to include different steps or combinations of steps similar to the steps described herein, in conjunction with other present or future technologies. Additionally, although the terms “step” and / or “block” may be used herein to represent different elements of the methods employed, these terms should not be construed as implying any particular order among or between the various steps disclosed herein, unless and when the order of individual steps is explicitly described.

Claims

1. A method, comprising: Receive image data representing an image sequence generated by an image sensor of a vehicle; Apply the image data to a sequential deep neural network DNN; Use the sequential DNN and calculate at least in part based on the image data: A first output corresponding to the speed of an object depicted in one or more images of the image sequence in world space; A second output corresponding to the current position and the future position of the object in image space; And The reciprocal of the time to collision TTC before the vehicle and the object are predicted to intersect; And Perform one or more operations by the vehicle at least in part based on the first output, the second output, and the reciprocal of the TTC.

2. The method according to claim 1, further comprising: Calculate the TTC using the reciprocal of the TTC, wherein performing the one or more operations is also at least in part based on the TTC.

3. The method according to claim 1, further comprising generating a visualization of the reciprocal of the TTC, the visualization including a boundary shape corresponding to the object and a visual indicator corresponding to the boundary shape, the visual indicator representing the TTC.

4. The method according to claim 3, wherein, When the TTC is below a threshold time, the visual indicator includes a first type of visual indicator, and when the TTC exceeds the threshold time, the visual indicator includes a second type of visual indicator.

5. The method according to claim 1, wherein, Train the sequential DNN using ground truth data generated by associating training sensor data with training image data representing a training sequence of images, the training sensor data including at least one of LIDAR data from one or more LIDAR sensors, RADAR data from one or more RADAR sensors, SONAR data from one or more SONAR sensors, or ultrasonic data from one or more ultrasonic sensors.

6. The method according to claim 5, wherein, Use cross-sensor fusion to associate the training sensor data with the training image data.

7. The method according to claim 1, wherein, The current position includes a first pixel position within an image at a first origin of a first boundary shape corresponding to the object, and the future position includes a second pixel position within an image at a second origin of a second boundary shape corresponding to the object.

8. The method according to claim 7, wherein, The current position further includes first size information corresponding to the first boundary shape, and the future position further includes second size information corresponding to the second boundary shape.

9. The method according to claim 1, wherein, The current position includes a first pixel position within an image at a first origin of a first boundary shape corresponding to the object, and the future position includes a translation value of a second pixel position within an image at a second origin of a second boundary shape corresponding to the object relative to the first pixel position.

10. The method according to claim 9, wherein, The current position further includes first size information corresponding to the first boundary shape, and the future position further includes scale information regarding the first size information corresponding to the first boundary shape, the scale information being used to determine second size information corresponding to the second boundary shape.

11. An automated vehicle system, comprising: One or more processing units for: Apply an image sequence represented by image data to a sequential deep neural network DNN; Use the sequential DNN and calculate at least in part based on the image sequence: A first output corresponding to the speed of an object depicted in one or more images of the image sequence in world space; A second output corresponding to a current position of the object and a future position of the object in the image space; and a reciprocal of a time to collision TTC before the autonomous vehicle and the object are predicted to intersect; and perform one or more operations for controlling the autonomous vehicle based at least in part on the first output, the second output, and the reciprocal of the TTC.

12. The system according to claim 11, wherein, The one or more processing units are further configured to: calculate the TTC using the reciprocal of the TTC, wherein the performance of the one or more operations is further based at least in part on the TTC.

13. The system according to claim 11, wherein, calculate the current position with respect to an image of the image sequence, and calculate the future position of the object with respect to the image and at least in part based on data representing the object within one or more previous images of the image sequence.

14. The system according to claim 13, wherein, The data includes a position and a velocity of the object within the one or more previous images.

15. A processor, comprising: One or more circuits configured to determine one or more control decisions of the ego machine based at least in part on a time to collision TTC measurement between the ego machine and an object, the TTC measurement being generated at least in part based on an output of a neural network including a reciprocal of the TTC measurement, the output being calculated at least in part based on a sequence of frames represented by sensor data generated using one or more sensors of the ego machine.

16. The processor according to claim 15, wherein the determination of the one or more control decisions is further at least partially based on another output of the neural network indicating the rate of the object.

17. The processor according to claim 15, wherein the determination of the one or more control decisions is further at least partially based on another output of the neural network indicating the current position of the object and the future position of the object.

18. The processor according to claim 15, wherein the one or more control decisions correspond to braking, accelerating, or changing the steering angle of the ego machine.

19. The processor according to claim 15, wherein, The sequence of frames of the sensor data corresponds to a sequence of images of image data, and the neural network is trained using ground truth data generated by correlating training sensor data with training image data representing a training sequence of images, the training sensor data including at least one of LIDAR data from one or more LIDAR sensors, RADAR data from one or more RADAR sensors, SONAR data from one or more SONAR sensors, or ultrasonic data from one or more ultrasonic sensors.

20. The processor according to claim 19, wherein, Use cross-sensor fusion to correlate the training sensor data with the training image data.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Distance space reconstruction method based on deep learning and monocular vision

    CN108108750A

  • Forward collision control method and device, electronic equipment, program and medium

    CN108725440A