Data augmentation including background modification for robust prediction using neural networks

By using data augmentation techniques such as background filtering and hue adjustment, and analyzing inference scores, the system addresses the challenge of creating a robust neural network for hand pose recognition, significantly improving prediction accuracy and generalization.

JP2025087686AInactive Publication Date: 2025-06-10NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025016486
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-09-30
Filing Date
2025-02-04
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing neural networks face challenges in achieving robustness due to the difficulty in constructing a training dataset that effectively covers all potential image features, especially when the network struggles with environmental features, color patterns, and specific hand poses from certain angles.

Method used

The system employs data augmentation techniques, including background filtering and hue adjustment, to generate training images by modifying the background of an object using segmentation masks and integrating the object image with various backgrounds. Additionally, it analyzes inference scores to select appropriate backgrounds and performs early or late fusion using object mask data to enhance neural network performance.

Benefits of technology

This approach enhances the robustness of neural networks by providing a diverse set of training images that improve the network's ability to recognize hand poses and gestures in various environmental conditions, leading to better prediction accuracy and generalization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025087686000001_ABST
    Figure 2025087686000001_ABST
Patent Text Reader

Abstract

To provide systems and methods for realizing data augmentation based on background filtering used to increase the robustness of trained neural networks.SOLUTION: A method for inference using a machine learning model comprises: obtaining at least one neural network trained to perform a predictive task on images using inputs generated from masks that correspond to objects; generating a mask that corresponds to an object in an image, where the object has a background in the image; generating, using the mask, inputs to the at least one neural network; and generating at least one prediction of the predictive task based at least on applying the inputs to the at least one neural network.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] When training a neural network to perform prediction tasks such as object classification, the accuracy of the trained neural network is often limited by the quality of the training dataset. To train to produce a robust neural network, the network should be trained using difficult training images. For example, when training a neural network for hand pose recognition (e.g., thumb up, peace sign, fist, etc.), the network may have difficulty detecting a hand pose in front of certain environmental features. If the pose includes extended fingers, the network may function well when the pose is in front of a plain environment, but may have difficulty when the pose is in front of an environment that includes certain color patterns. As another example, the network may have difficulty with certain poses from certain angles or when the environment and the hand have similar hues.

Background Art

[0002] However, whether a particular training image is difficult for a neural network can depend on a number of factors, such as the prediction task being performed, the architecture of the neural network, and other training images seen by the network. Therefore, it is difficult to construct a training dataset that will result in a well-trained network by anticipating which training images should be used to train the network. It may be possible to estimate which features of training images can be difficult for a neural network. However, even when such an estimate is possible and accurate, it may be impossible or unrealistic to obtain sufficient images showing those features to fully train the network.

Summary of the Invention

Means for Solving the Problem

[0003] Embodiments of the present disclosure relate to data augmentation including background filtering for robust prediction using neural networks. A system and method are disclosed that provide data augmentation techniques, such as those based on background filtering, that can be used to enhance the robustness of a trained neural network.

[0004] In contrast to conventional systems, the present disclosure modifies the background of an object to generate training images. A segmentation mask can be generated and used to generate an object image that includes image data representing the object. The object image can be integrated into different backgrounds and used for data augmentation in the training of a neural network. Other aspects of the present disclosure provide data augmentation that uses (e.g., for an object image) hue adjustment and / or renders three-dimensional capture data corresponding to an object from a selected view direction. The present disclosure also provides an analysis of an inference score for selecting the background of an image that will be included in a training dataset. The background can be selected and training images can be repeatedly added to the training dataset during training (e.g., between epochs). Additionally, the present disclosure performs early or late fusion using object mask data to improve the inference performed by a neural network trained using the object mask data.

[0005] The present system and method for data augmentation including background filtering for robust prediction using a neural network will be described in detail below with reference to the accompanying drawings.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 11C

Figure 11D

DETAILED DESCRIPTION OF THE INVENTION

[0007] Systems and methods are disclosed for data augmentation including background filtering for robust prediction using neural networks. Embodiments of the present disclosure relate to data augmentation including background filtering for robust prediction using neural networks. Systems and methods are disclosed that provide data augmentation techniques, such as those based on background filtering, that can be used to enhance the robustness of trained neural networks.

[0008] The disclosed embodiments may be implemented using a variety of different systems, such as automotive systems, robotics, aerospace systems, intermediate systems, marine systems, smart area monitoring systems, simulation systems, and / or other technical fields. The disclosed techniques may be used for any perception-based or more generally image-based analysis using a machine learning model, for example, for monitoring and / or tracking objects and / or the environment.

[0009] Examples of the application of the disclosed techniques include a multimodal sensor interface that can be applied in a medical context. For example, patients who are intubated or otherwise unable to communicate verbally can use poses or gestures that are interpreted by a computing system. Examples of the application of the disclosed techniques further include autonomous driving and / or vehicle control or interaction. For example, the disclosed techniques can be used to implement the recognition of hand poses or gestures for sensing in the cabin of vehicle 1100 of FIGS. 11A - 11D for controlling convenience features, e.g., control of multimedia options. The recognition of gesture poses can also be applied to the external environment of vehicle 1100 to control any of a variety of autonomous driving control actions, including advanced driver assistance system (ADAS) functions.

[0010] As various examples, the disclosed techniques can be implemented in one or more of a system for performing interactive AI or personal assistant operations, a system for performing simulation operations, a system for performing simulation operations to test or verify autonomous machine applications, a system for performing deep learning operations, a system implemented using edge devices, a system incorporating one or more virtual machines (VMs), a system at least partially implemented in a data center, or a system that can be at least partially implemented using cloud computing resources.

[0011] In contrast to conventional systems, the present disclosure identifies regions within an image corresponding to an object and uses the regions to filter, remove, replace, or otherwise modify the background of the object and / or the object itself to generate training images. According to the present disclosure, a segmentation mask can be generated that identifies one or more segments corresponding to the object and one or more segments corresponding to the background of the object in the source image. The segmentation mask can be applied to the source image to identify regions corresponding to the object, for example, to generate an object image that includes image data representing the object. The object image can be integrated into different backgrounds and used for data augmentation in the training of a neural network. Other aspects of the present disclosure perform data augmentation that uses, for example, hue adjustment of the object image and / or renders three-dimensional capture data corresponding to the object from a selected view direction.

[0012] A further aspect of the present disclosure provides a technique for selecting a background for an object for training a neural network. According to the present disclosure, a machine learning model (MLM) can be at least partially trained, and inference data can be generated by the MLM using images that include different backgrounds. The MLM can include a neural network during training or a different MLM. An inference score corresponding to the inference data can be analyzed to select one or more features of the training image, for example, a specific background or background type of the images that will be included in the training data set. The images can be selected from existing images or generated using an object image and a background using any suitable technique, such as those described herein. In at least one embodiment, one or more features can be selected and one or more corresponding training images can be repeatedly added to the training data set during training.

[0013] The present disclosure further provides a method of using object mask data to improve inferences performed by a neural network trained using the object mask data. Late fusion may be performed, where one set of inference data is generated from a source image and another set of inference data is generated from an image capturing object mask data, e.g., an object image (e.g., using two copies of the neural network). The sets of inference data may be fused and used to update the neural network. In a further example, early fusion may be performed, where the source image and the image capturing the object mask data are combined and inference data is generated from the combined image. The object mask data may be used to stop emphasizing or otherwise modify the background of the source image.

[0014] Referring now to FIG. 1, FIG. 1 is a data flow diagram showing an exemplary process 100 for training one or more machine learning models based at least on integration of an object image with a background, according to some embodiments of the present disclosure. Process 100 is described, by way of example, with respect to a machine learning model (MLM) training system 140. Among a number of potential components, MLM training system 140 may include a background integrator 102, an MLM trainer 104, an MLM postprocessor 106, and a background selector 108.

[0015] At a high level, process 100 may include a background integrator 102 that receives one or more of backgrounds 110 (which may be referred to as background images) and object images 112 corresponding to one or more objects (such as those to be classified, analyzed, and / or detected by MLM 122). The background integrator 102 can integrate the object images 112 with the backgrounds 110 to produce image data that captures (e.g., represents) at least a portion of the backgrounds 110 and the objects within one or more images. The MLM trainer 104 can produce inputs 120 to one or more MLMs 122 from the image data. The MLMs 122 can process the inputs 120 to generate one or more outputs 124. The MLM post-processor 106 can process the outputs 124 to produce prediction data 126 (such as inference scores, object class labels, object bounding boxes or shapes, etc.). The background selector 108 can analyze the prediction data 126 and select one or more of the backgrounds 110 and / or objects for training, at least based on the prediction data 126. In some embodiments, process 100 can be repeated any number of times until one or more of the MLMs 122 are sufficiently trained, or the background selector 108 can be used once or intermittently to select the background 110 for a first training iteration and / or any other iteration.

[0016] For example, and without limitation, the MLM122 described herein can include any type of machine learning model, such as linear regression, logistic regression, decision tree, support vector machine (SVM), naive Bayes, k-nearest neighbor (Knn), K-means clustering, random forest, dimensionality reduction algorithms, gradient boosting algorithms, neural networks (e.g., autoencoders, convolutional, recurrent, perceptron, long / short term memory (LSTM), Hopfield, Boltzmann, deep belief, deconvolutional, adversarial generation, liquid state machines, etc.), and / or other types of machine learning models.

[0017] Process 100 can be used, at least in part, to train one or more of the MLM122s to perform a prediction task. The present disclosure focuses on pose recognition and / or gesture recognition, more specifically hand pose recognition. However, the disclosed techniques are widely applicable for training an MLM to perform various possible prediction tasks, such as image and / or object classification tasks. Examples include object detection, bounding box or shape determination, object classification, pose classification, gesture classification, and many others. For example, FIG. 2 shows an example of an image 246 that can be captured by an input 120 to the MLM122. Process 100 can be used to train the MLM122 to predict the hand pose (e.g., thumb up, thumb down, fist, peace sign, palm, OK sign, etc.) depicted in the image 246.

[0018] In some embodiments, one or more of the iterations of process 100 may not include training one or more of MLMs 122. For example, an iteration may be used by background selector 108 to select one or more of background 110 and / or objects corresponding to object images 112 for inclusion in a training dataset used by MLM trainer 104 for training. Further, in some instances, process 100 can use one MLM 122 to select the background 110 and / or objects of a training dataset in one iteration and (e.g., in a subsequent iteration of process 100 or otherwise) use the training dataset to train the same or a different MLM 122. For example, an MLM 122 used to select from background 110 for training may be partially or fully trained to perform a prediction task. An iteration of process 100 may use a trained or partially trained MLM 122 to select a different MLM 122 from background 110 for training, and the iteration may be used to bootstrap the training of other MLMs 122 by selecting difficult background images and / or combinations of background and object images for training.

[0019] In various examples, one or more iterations of process 100 can form a feedback loop, where background selector 108 uses prediction data of the iteration (e.g., training epoch) to select one or more of background 110 and / or the object corresponding to object image 112 for inclusion in a subsequent training data set. In a subsequent iteration of process 100 (e.g., a subsequent training epoch), background integrator 102 can generate or otherwise prepare or select a corresponding image that MLM trainer 104 can incorporate into the training data set. The training data set can then be applied to MLM 122 that is trained to generate prediction data 126 that is used by background selector 108 to select one or more of background 110 and / or the object corresponding to object image 112 for inclusion in a subsequent training data set. The feedback loop can be used to determine extensions to ensure continuous improvement of the accuracy, generalization ability, and robustness of the trained MLM 122.

[0020] In various examples, based at least on a background selector 108 that selects one or more backgrounds and / or objects, an MLM trainer 104 incorporates one or more images from a background integrator 102 that includes the selected background and / or a combination of the selected background and objects. As an example, the MLM trainer 104 can add one or more images to a training data set used for a previous training iteration and / or epoch. The training data set can grow with each iteration. However, in some cases, the MLM trainer 104 can also remove one or more images from a training data set used for a previous training iteration and / or epoch (e.g., based on selections by the background selector 108 for removal and / or based on a threshold number of images in the training data set). In the example shown, the background integrator 102 can generate one or more images that will be included in the training data set at the beginning of an iteration of process 100 based on selections made by the background selector 108. In other examples, one or more of the images can be pre-generated by the background integrator 102, for example, at least partially, prior to any training that uses process 100 and / or during one or more previous iterations. If an image is pre-generated, the MLM trainer 104 can fetch the pre-generated image from storage based on selections made by the background selector 108.

[0021] The selection of the background described in this specification may refer to the selection of a background for inclusion of at least one image used for training. The selection of the background may also include the selection of an object for inclusion in an image having a background. In at least one example, the background selector 108 can select a background and / or a combination of a background and an object based at least on the reliability of the MLM 122 in one or more predictions made using the MLM 122. In various examples, the reliability can be captured by a set of inference scores corresponding to the prediction of a prediction task performed by one or more of the MLM 122 on one or more images. For example, the prediction can be made in the current iteration and / or one or more previous iterations of the process 100. The inference score can refer to a score that is trained to be provided by the MLM or trained to be provided for a prediction task or a part thereof. In some examples, the inference score can represent the reliability of the MLM with respect to one or more of the corresponding outputs 124 (e.g., tensor data), and / or can be used to determine or calculate the reliability with respect to the prediction task. For example, the inference score can represent the reliability of the MLM 122 in an object detected in an image belonging to a target class (e.g., the probability of the input 120 belonging to the target class).

[0022] The background selector 108 can select one or more of the backgrounds and / or combinations of backgrounds and objects based on inference scores using various possible techniques. In at least one embodiment, the background selector 108 can select one or more specific backgrounds and / or objects based at least on an analysis of the inference scores corresponding to images that include those elements. In at least one embodiment, the background selector 108 is based at least on an analysis of the inference scores corresponding to images that include those elements having one or more specific features of a particular class or type or one or more of them (e.g., a particular background, a particular object, texture, color, lighting conditions, object and / or overlay, hue, viewpoints, orientations, skin color, size, theme, background elements included, etc., each described with respect to FIG. 4) can select one or more backgrounds and / or objects having one or more other specific features. For example, the background selector 108 can select at least one background including a blind and a gesture including an open palm based at least on the inference scores of images that share those features. As another example, the background selector 108 can select a specific background based at least on the inference scores of images that include that background. As an additional example, the background selector 108 can select objects of a specific background and object class (e.g., thumb up, thumb down, etc.) based at least on the inference scores of images that include that background and objects of that object class.

[0023] In some cases, the inference score can be evaluated by the background selector 108 based at least on the calculation of the confusion score. The confusion score can serve as a metric that quantifies the relative network confusion regarding predictions made for one or more images having a particular set of features. If the confusion score exceeds a threshold (e.g., indicating sufficient confusion), the background selector 108 can select at least some elements (e.g., background or a combination of background and object) having a set of features to modify the training dataset. In at least one embodiment, the background confusion can be at least partially based on some accurate and inaccurate predictions made for images including elements having a set of features. For example, the confusion score can be calculated based at least on the ratio of accurate predictions to inaccurate predictions. Additionally or alternatively, the confusion score can be calculated based at least on the difference in the accuracy of the predictions or inference scores of images including elements having a set of features (e.g., indicating that there is a high difference across target classes when a particular background or background type is used).

[0024] The background selector 108 can select one or more backgrounds and / or combinations of background and object for inclusion in at least one image of the training dataset based at least on the selection of one or more different element features. For example, the background selector 108 can select one or more different element features based at least on the corresponding confusion score. The background selector 108 can rank different sets of features and select one or more sets for modification of the training dataset based on the ranking. As a non-limiting example, the background selector 108 can select the top N particular backgrounds or combinations of background and object (or other sets of features), where N is an integer (e.g., for each confusion score exceeding a threshold).

[0025] The background integrator 102 can select, obtain (e.g., from storage), and / or generate one or more images that satisfy the selection made by the background selector 108 to modify the training dataset. When the background integrator 102 generates an image from the selected background, the background integrator 102 may include the entire background image within the image or one or more portions of the background. For example, the background integrator 102 can sample a region (e.g., a rectangle sized based on the input 120 to the MLM 122) from the background using a random or non-random sampling technique. Thus, the image processed by the MLM 122 may include all of the background image or a region of the background image (e.g., the sampled region). Similarly, when the background integrator 102 generates an image from the selected object, the background integrator 102 may include all of the object image within the image or a portion of the object image.

[0026] In at least one embodiment, the background integrator 102 can generate one or more synthetic backgrounds. For example, synthetic backgrounds can be generated for cases where the network bias and sensitivity are understood (experimentally or intuitively from measurements of the network's performance). For example, certain types of synthetic backgrounds (e.g., dots and stripes) can be generated for networks that are sensitive to those particular textures, where the background selector selects that background type. One or more synthetic backgrounds can be generated prior to and / or during the training of the MLM 122 (e.g., between iterations and epochs). For example, various possible techniques can be used to generate synthetic backgrounds, based at least on rendering a 3D virtual environment related to the background type, generating in an algorithm of a texture including the selected pattern, modifying an existing background or image, etc.

[0027] According to aspects of the present disclosure, the object image 112 can be extracted from one or more source images, and the background integrator 102 can use one or more of the backgrounds 110 to replace or modify the initial background of the source image. Referring now to FIG. 2, FIG. 2 is a data flow diagram showing an exemplary process 200 for generating an object image 212 and integrating the object image 212 with one or more of the backgrounds 110, according to some embodiments of the present disclosure.

[0028] Process 200 is described, by way of example, with respect to an object image extraction system 202. Among several potential components, the object image extraction system 202 can include a region identifier 204, a preprocessor 206, and an image data determiner 208.

[0029] Generally speaking, in process 200, the region identifier 204 can be configured to identify regions within the source image. For example, the region identifier 204 can identify a region 212A within the source image 220 that corresponds to an object (e.g., a hand) having a background within the source image 220. The region identifier 204 can further generate a segmentation mask 222 that includes a segment 212B corresponding to the object based on the identification of region 212A. The region identifier 204 can detect the position of the object and define an area 230 of the source image 220 and / or the segmentation mask 222. The preprocessor 206 can process at least a portion of the segment 212B within the area 230 of the segmentation mask 222 to produce an object mask 232. The image data determiner 208 can use the object mask 232 to generate an object image 212 from the source image 220. The object image can then be provided to the background integrator 102 for integration with one or more of the backgrounds 110 (e.g., to super-impose or overlay the object onto a background image).

[0030] According to various embodiments, one or more of the object images 112, e.g., object image 212, may be generated before and / or during the training of the MLM 122. For example, one or more of the object images 112 may be generated (e.g., as shown in FIG. 2), stored, and then retrieved only as needed to generate the input 120 in process 100. As another example, one or more of the object images 112 may be generated during process 100, e.g., on-the-fly or as required by the background integrator 102. In some embodiments, the object images 112 are generated on-the-fly during process 100 and then stored and / or reused for training other MLMs other than the MLM 122 in subsequent iterations of process 100 and / or later.

[0031] As described herein, the region discriminator 204 can identify a region 212A in the source image 220 corresponding to an object (e.g., a hand) having a background in the source image 220. In the example shown, the region discriminator 204 can also identify a region 210A in the source image 220 corresponding to the background of the object. In other examples, the region discriminator 204 can simply identify the region 212A.

[0032] In at least one embodiment, the region discriminator 204 can identify the region 212A and at least determine a segmentation 212B of the source image 220 corresponding to the object based at least on the performance of image segmentation in the source image. The image segmentation can also be used to identify the region 210A and determine a segmentation 210B of the source image 220 corresponding to the background based at least on the performance of image segmentation in the source image 220. In at least one embodiment, the region discriminator 204 can generate data representing a segmentation mask 222 from the source image 220, where the segmentation mask 222 indicates the segmentation 212B corresponding to the object (white pixels in FIG. 2) and / or the segmentation 210B corresponding to the background (black pixels in FIG. 2).

[0033] The region identifier 204 can be implemented in various possible ways, such as using AI-powered background removal. In at least one embodiment, the region identifier 204 includes one or more MLMs trained to classify or label individual or groups of pixels in an image. For example, the MLM can be trained to identify the foreground (e.g., corresponding to an object) and / or background within the image, and the image segmentation can correspond to the foreground and / or background. As a non-limiting example, the region identifier 204 can be implemented using the background removal technology of RTX Greenscreen by NVIDIA Corporation. In some embodiments, the MLM can be trained to identify object types and label pixels accordingly.

[0034] In at least one embodiment, the region identifier 204 can include one or more object detectors, such as object detectors trained to detect objects (e.g., hands) that would be classified by the MLM122. The object detector can be implemented using one or more MLMs trained to detect objects. The object detector can output data indicating the location of the object and can be used to define an area 230 of the source image 220 and / or the segmentation mask 222 that includes the object. For example, the object detector can be trained to provide a bounding box or shape of the object, and the bounding box or shape can be used to define the area 230.

[0035] In the illustrated example, area 230 can be defined by expanding the bounding box, but in other examples, the bounding box can be used as area 230. In this example, region identifier 204 can identify segment 212B corresponding to an object by applying source image 220 to the MLM. In other examples, region identifier 204 can apply area 230 to the MLM instead of (or in addition to in some embodiments) the source image. By applying source image 220 to the MLM, the MLM can have additional context that is not available in area 230 that can improve the accuracy of the MLM.

[0036] In embodiments where area 230 is determined, preprocessor 206 can perform preprocessing based at least on area 230. For example, area 230 of segmentation mask 222 can be preprocessed by preprocessor 206 before being used by image data determiner 208. In at least one embodiment, preprocessor 206 can crop image data corresponding to area 230 from segmentation mask 222 and process the cropped image data to produce object mask 232. Preprocessor 206 can perform various types of preprocessing in area 230 that can improve the capabilities of image data determiner 208.

[0037] Referring now to FIG. 3, FIG. 3 includes an example of preprocessing that can be used to generate an object mask that can be used to generate an object image, according to some embodiments of the present disclosure. As an example, preprocessor 206 can crop segmentation mask 222 that results in object mask 300A. Preprocessor 206 can perform an expansion on object mask 300A that results in object mask 300B. Preprocessor 206 can then blur object mask 300B that results in object mask 232. Object mask 232 can then be used by image data determiner 208 to generate object image 212.

[0038] The preprocessor 206 can use dilation to expand the segment 212B of the object mask 300A corresponding to the object. For example, the segment 212B can be expanded into the segment 210B corresponding to the background. In an embodiment, the preprocessor 206 can perform binary dilation. Other types of dilation, such as grayscale dilation, can be performed. As an example, the preprocessor 206 can first blur the object mask 300A and then perform grayscale dilation. Dilation can be beneficial to enhance the robustness of the object mask 232 against errors in the segment mask 222. For example, when the object includes a hand, the palm of the hand can sometimes be classified as part of the background. Dilation is one approach to correct this potential error. Other mask preprocessing techniques are within the scope of the present disclosure, such as binary or grayscale erosion. As an example, erosion can be performed on the segment 210B corresponding to the background.

[0039] The preprocessor 206 can use blurring (e.g., Gaussian blurring) to assist the background integrator 102 in facilitating the transition between the image data corresponding to the object within the object image 212 and the image data corresponding to the background 110. Without facilitating the transition between the object and the background, the area corresponding to the edge of the object mask 232 can be sharp and artificial. By using blending techniques such as blurring and then applying the object mask, a more natural or realistic transition between the object and the background 110 in the image 246 can be achieved. Mask processing has been described as being performed on the object mask prior to the application of the mask, but in other examples, similar or different image processing operations can be performed by the image data determiner 208 in the application of the object mask (e.g., to the source image 220).

[0040] Returning to FIG. 2, the image data determiner 208 can generally generate an object image 212 using an object mask, e.g., object mask 232. For example, the image data determiner 208 can use the object mask 232 to identify and / or extract a region 242 from the source image 220 corresponding to the object. In other embodiments, the object mask may not be used, and another technique may be used to identify and / or extract the region 242. When using the object mask 232, the image data determiner 208 can multiply the object mask 232 with the source image 220 to obtain an object image 212 that includes image data representing the region 242 (e.g., the foreground of the source image 220) corresponding to the object.

[0041] When integrating the object image 212 with the background 110, the background integrator 102 can use the object image 212 as a mask, and an inversion of the mask can result in the resulting image being blended with the object image and applied to the background 110. For example, the background integrator 102 can perform alpha compositing between the object image 212 and the background 110. The blending of the object image 212 with the background integrator 102 can use various possible blending techniques. In some embodiments, the background integrator 102 can use alpha-blending to integrate the object image 212 with the background 110. Alpha-blending can set the background from the object image 212 to zero when combining the object image 212 with the background 110, and the foreground pixels can be superposed on the background 110 to generate the image 246, or the pixels can be weighted (e.g., from 0 to 1) when combining the image data from the object image 212 and the background 110 by a blur applied by the pre-processor 206 or other means. In at least one embodiment, the background integrator 102 can use one or more seamless blending techniques. The seamless blending technique can aim to create a seamless boundary between the object and the background 110 in the image 246. Examples of seamless blending techniques include gradient domain blending, Laplacian pyramid blending, or Poisson blending.

[0042] Further Examples of Data Augmentation Techniques As described herein, process 200 can be used to augment a training dataset for training an MLM, such as MLM 122 that uses process 100. The present disclosure provides further techniques that can be used to enhance the training dataset. According to at least some embodiments, the hue of an object identified in a source image, such as source image 220, can be modified for data augmentation. As an example, when the object represents at least a portion of a person, the skin color, hair color, and / or other hues can be modified to augment the training dataset. For example, the hue of one or more portions of the region corresponding to the object can be changed (e.g., uniformly or in other ways). In at least one embodiment, the hue can be selected randomly or non-randomly. In some cases, the hue can be selected based on an analysis of prediction data 126. For example, the hue can be used as a feature for selecting or generating one or more training images as described herein (e.g., by background integrator 102).

[0043] Certain areas, such as the background or non-primary or minor sections of regions that tend to have a consistent hue for different realistic variations of an object, can retain the original hue. For example, when the object is a passenger vehicle, the panel can have its hue changed while the hues of the lights, bumper, and tires are maintained. In at least one embodiment, background integrator 102 can perform the hue modification. For example, the hue modification can be performed on one or more portions of the object represented in object image 212. In other examples, object image 212 or object mask 232 can be used to identify image data representing the object and modify the hue in one or more regions of source image 220 (e.g., by image data determiner 208). These examples may not include background integrator 102.

[0044] According to at least some embodiments, the source image 220 can be rendered from various different views of an object in an environment for data augmentation. Referring now to FIG. 4, FIG. 4 is a diagram of how a three-dimensional (3D) capture 402 of an object can be rasterized from multiple views, according to some embodiments of the present disclosure. In at least one embodiment, the object image extraction system 202 can select a view of the object in the environment. For example, the object image extraction system 202 can select from views 406A, 406B, 406C, or any view of the object in the environment 400. The object image extraction system 202 can then generate the source image 220 based at least on the rasterization of the 3D capture of the object within the view in the environment. For example, the three-dimensional (3D) capture 402 can include depth information captured by a physical or virtual depth sensing camera in a physical or virtual environment (which can be different from the environment 400). In one or more embodiments, the 3D capture 402 can include a point cloud that captures at least a portion of the objects and potentially additional elements of the environment 400. For example, if view 406A is selected, the object image extraction system 202 can rasterize the source image 220 from view 406A of the camera 404 using at least the 3D capture 402. In at least one embodiment, the view can be selected randomly or non-randomly. In some cases, the view can be selected based on the analysis of the prediction data 126. For example, the view can be used as a feature for selecting or generating one or more training images (e.g., by the background integrator 102) as described herein (e.g., in combination with the object class). In at least one embodiment, the object can then be rasterized from the view to generate an object image 112 that can be integrated with one or more backgrounds using the techniques described herein. In other examples, the object can be rasterized with the background 110 (a two-dimensional image) or other 3D content of the environment 400 to form the background.

[0045] Example of Inference Using an Object Mask As described herein, the object mask can be used for data augmentation in the training of the MLM 122, for example using process 100. In at least one embodiment, the MLM 122 trained using the mask data can perform inference on an image without leveraging the object mask. For example, the input 120 to the deployed MLM 122 can correspond to one or more images captured by a camera. In such an instance, the object mask can simply be used for data augmentation. In other embodiments, the object mask can also be leveraged for inference. An example of how the object mask can be leveraged for inference will be described with respect to FIGS. 5A and 5B.

[0046] Referring now to FIG. 5A, FIG. 5A is a data flow diagram 500 showing an example of inference using the MLM 122 and early fusion of object mask data, according to some embodiments of the present disclosure. In the example of FIG. 5A, the MLM 122 can be trained to perform inference on the image 506 while leveraging the object image 508 corresponding to the object mask of the image 506. For example, the input 120 can be generated from a combination of the image 506 and the object image 508 and then provided to an MLM 122 (e.g., a neural network) that can generate an output 124 including inference data 510. If the MLM 122 includes a neural network, the inference data 510 can include tensor data from the neural network. Post-processing can be performed on the inference data 510 to generate the prediction data 126.

[0047] The object image 508 is one example of object mask data that can be combined with the image 506 for inference. If early fusion of the object mask data is used for inference, as in the data flow diagram 500, the input 120 to the MLM 122 can be similarly generated during training (e.g., in the process 100). Generally, the object mask data can capture information regarding the manner in which the region discriminator 204 generates a mask from the source image. By leveraging the object mask data during training and inference, the MLM 122 can learn to account for any errors or unnatural artifacts that may be produced by object mask generation. The object mask data can also capture information regarding the manner in which the pre-processor 206 pre-processes the object mask to capture any errors or unnatural artifacts that may be produced or remain after pre-processing.

[0048] The object image 508 can be generated using the object image extraction system 202 similar to the object image 212 (e.g., at inference time). Although the object image 508 is shown in FIGS. 5A and 5B, in other examples, the segmentation mask 222 and / or the object mask 232 can be used in addition to or instead of the object image 508 (either before or after pre-processing by the pre-processor 206).

[0049] Various techniques can be used to generate the input 120 from the combination of the image 506 and the object image 508 (more generally object mask data). In at least one example, the image 506 and the object image 508 are provided as separate inputs 120 to the MLM 122. As a further example, the image 506 and the object image 508 can be combined to form a combined image, and the input 120 can be generated from the combined image. In at least one example, the object mask data can be used to fade, stop emphasizing, mark, indicate, highlight, or otherwise modify one or more portions of the image 506 that represent the background (as captured by the object mask data) with respect to the object or foreground of the image 506. For example, the image 506 can be blended with the object image 508, resulting in the background of the image 506 being faded, blurred, or defocused (e.g., using a depth of field effect). When combining the image 506 and the object image 508, the weights used to determine the resulting pixel color can decrease (e.g., exponentially) with the distance from the object as commanded by the object mask data (e.g., using a concavity effect).

[0050] Referring now to FIG. 5B, FIG. 5B is a data flow diagram 502 showing an example of inference using the MLM 122 and late fusion of object mask data, according to some embodiments of the present disclosure.

[0051] In the example of FIG. 5B, MLM122 can provide separate outputs 124 for the image 506 and the object image 508. The output 124 may include inference data 510A corresponding to the image 506 and inference data 510B corresponding to the object image 508. Further, MLM122 may include separate inputs 120 for the image 506 and the object image 508. For example, MLM122 can include multiple copies of the MLM(122) trained to perform inferences on the image, where one copy performs inferences on the image 506 to generate inference data 510A, and another copy performs inferences on the object image 508 to generate inference data 510B (e.g., in parallel). Post-processing can be performed on the inference data 510A and the inference data 510B to generate prediction data 126 using late fusion. For example, corresponding tensor values across the inference data 510A and the inference data 510B can be combined (e.g., averaged) to fuse the inference data, and then further post-processing can be performed on the fused inference data to generate the prediction data 126. In at least one embodiment, the tensor values of the inference data 510A and the inference data 510B can be combined using weights (e.g., using a weighted average). In at least one embodiment, the weights can be adjusted via a validation dataset.

[0052] Inference using MLM122 can also include temporal filtering of the inference scores to generate prediction data 126 that can improve the temporal stability of the predictions. Additionally, the examples mainly shown relate to the recognition of static poses. However, the disclosed techniques can also be applied to the recognition of dynamic poses, which may be referred to as gestures. To train and use an MLM to predict gestures, in at least one embodiment, multiple images can be provided to MLM122 that capture an object over a period or several or a series of frames. If object mask data is used, the object mask data can be provided for each input image.

[0053] Referring now to FIG. 6, each block of method 600 and other methods described herein can include a computing process that can be executed using any combination of hardware, firmware, and / or software. For example, various functions can be implemented by a processor executing instructions stored in memory. The method can also be implemented as computer-usable instructions stored on a computer storage medium. The method can be provided, for example, as a stand-alone application, service, or host-based service (either stand-alone or in combination with another host-based service), or as a plug-in to another product. Additionally, method 600 is described, by way of example, with respect to system 140 of FIG. 1 and system 202 of FIG. 2. However, the method can be executed additionally or alternatively by any one system, or any combination of systems, including but not limited to those described herein.

[0054] FIG. 6 is a flowchart showing a method 600 for training one or more machine learning models based at least on integrating an object image with at least one background, according to some embodiments of the present disclosure. Method 600 includes, at block B602, identifying a region corresponding to an object having a first background in a first image. For example, region identifier 204 can identify region 212A corresponding to an object having a background in source image 220 in source image 220.

[0055] Method 600 includes, at block B604, determining image data representing the object based at least on the region. For example, image data determiner 208 can determine image data representing the object based at least on region 212A of the object. In at least one embodiment, image data determiner 208 can determine the image data using an object mask 232 or a non-mask-based approach.

[0056] Method 600 includes, at block B606, generating a second image that includes an object having a second background using the image data. The background integrator 102 can generate an image 246 that includes an object having the background 110, at least based on using the image data to integrate the object with the background 110. For example, the image data determiner 208 can incorporate the image data into the object image 212 and provide the object image 212 to the background integrator 102 for integration with the background 110.

[0057] Method 600 includes, at block B608, training at least one neural network to perform a prediction task using the second image. For example, the MLM trainer 104 can train the MLM to classify the objects within the image using the image 246.

[0058] Referring now to FIG. 7, FIG. 7 is a flow diagram showing a method 700 for inference using a machine learning model, according to some embodiments of the present disclosure, where the input corresponds to a mask of an image and at least a portion of the image. Method 700 includes, at block B702, obtaining (or accessing) at least one neural network trained to perform a prediction task on an image using an input generated from a mask corresponding to an object. For example, the MLM 122 of FIG. 5A or 5B can be obtained (accessed) and can be trained according to the process 100 of FIG. 1.

[0059] Method 700 includes, at block B704, generating a mask corresponding to an object within the image, where the object has a background within the image. For example, the region discriminator 204 can generate a segmentation mask 222 corresponding to an object within the source image 220, where the object has a background within the source image 220.

[0060] Method 700 includes, at block B706, generating an input to at least one neural network using a mask. For example, input 120 of FIG. 5A or 5B can be generated using partition mask 222 (or without using object mask data). Input 120 can capture an object having at least a portion of a background.

[0061] Method 700 includes, at block B708, generating at least one prediction of a prediction task based at least on the application of an input to at least one neural network. For example, MLM 122 can be used to generate at least one prediction of a prediction task based at least on the application of input 120 to MLM 122, and prediction data 126 can be determined using output 124 from MLM 122.

[0062] Referring now to FIG. 8, FIG. 8 is a flow diagram showing a method 800 for selecting a background of an object for training one or more machine learning models, according to some embodiments of the present disclosure. Method 800 includes, at block B802, receiving an image of one or more objects having a plurality of backgrounds. For example, MLM training system 140 can receive an image of one or more objects having a plurality of backgrounds 110.

[0063] Method 800 includes, at block B804, generating a set of inference scores corresponding to a prediction task using the image. For example, MLM trainer 104 can provide input 120 to one or more of MLM 122 (or different MLMs) to generate output 124, and MLM postprocessor 106 can process output 124 to produce prediction data 126.

[0064] Method 800 includes, at block B806, selecting a background based at least on one or more of the inference scores. For example, the background selector 108 can select one or more of the backgrounds 110 based at least on the prediction data 126.

[0065] Method 800 includes, at block B808, generating an image based at least on the integration of the object with the background. For example, the background integrator 102 can generate an image based at least on integrating the object with the background (e.g., using the object image 112 and the background 110).

[0066] Method 800 includes, at block B810, training at least one neural network using the image to perform a prediction task. For example, the MLM trainer 104 can train one or more of the MLMs 122 using the image.

[0067] Exemplary computing device FIG. 9 is a block diagram of an exemplary computing device 900 suitable for use in the implementation of some embodiments of the present disclosure. The computing device 900 may include an interconnect system 902 that directly or indirectly couples the following devices: a memory 904, one or more central processing units (CPUs) 906, one or more graphics processing units (GPUs) 908, a communication interface 910, an input / output (I / O) port 912, an input / output component 914, a power supply 916, one or more presentation components 918 (e.g., a display), and one or more logic units 920. In at least one embodiment, the computing device 900 may include one or more virtual machines (VMs), and / or any of its components may include virtual components (e.g., virtual hardware components). By way of non-limiting example, one or more of the GPUs 908 may include one or more vGPUs, one or more of the CPUs 906 may include one or more vCPUs, and / or one or more of the logic units 920 may include one or more virtual logic units. As such, the computing device 900 may include discrete components (e.g., all GPUs dedicated to the computing device 900), virtual components (e.g., a portion of a GPU dedicated to the computing device 900), or a combination thereof.

[0068] The various blocks of FIG. 9 are shown as being connected by lines via an interconnect system 902, but this is not intended to be limiting and is for clarity only. For example, in some embodiments, a presented component 918, such as a display device, may be considered an I / O component 914 (e.g., if the display is a touch screen). As another example, the CPU 906 and / or the GPU 908 may include memory (e.g., memory 904 may represent a storage device in addition to the memory of the GPU 908, CPU 906, and / or other components). In other words, the computing device of FIG. 9 is merely illustrative. Categories such as "workstation", "server", "laptop", "desktop", "tablet", "client device", "mobile device", "handheld device", "gaming console", "electronic control unit (ECU)", "virtual reality system", and / or other device or system types are all intended to be within the scope of the computing device of FIG. 9 and thus are not distinguished.

[0069] The interconnect system 902 can represent one or more links or buses, such as an address bus, a data bus, a control bus, or a combination thereof. The interconnect system 902 can include one or more bus or link types, such as an Industry Standard Architecture (ISA) bus, an Extended Industry Standard Architecture (EISA) bus, a Video Electronics Standards Association (VESA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, and / or another type of bus or link. In some embodiments, there are direct connections between components. As an example, the CPU 906 can be directly connected to the memory 904. Further, the CPU 906 can be directly connected to the GPU 908. If there are direct or point-to-point connections between components, the interconnect system 902 can include a PCIe link for making the connections. In these examples, the PCI bus need not be included in the computing device 900.

[0070] The memory 904 can include any of a variety of computer-readable media. The computer-readable media may be any available media that can be accessed by the computing device 900. The computer-readable media can include both volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, the computer-readable media can comprise computer storage media and communication media.

[0071] A computer storage medium can include both volatile and non-volatile media and / or removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, and / or other data types. For example, memory 904 can store computer-readable instructions (such as for an operating system, for example, representing a program and / or program elements). A computer storage medium can include, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computing device 900. In this specification, a computer storage medium does not include the signal itself.

[0072] A computer storage medium can implement computer-readable instructions, data structures, program modules, and / or other data types in a modulated data signal, such as a carrier wave, or other transport mechanism, and includes any information delivery medium. The term "modulated data signal" can refer to a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, a computer storage medium can include wired media such as a wired network or direct wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Any combination of the foregoing should also be included within the scope of computer-readable media.

[0073] The CPU 906 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to execute one or more of the methods and / or processes described herein. The CPU 906 may each include one or more (e.g., 1, 2, 4, 8, 28, 72, etc.) cores having the ability to process multiple software threads simultaneously. The CPU 906 may include any type of processor and may include different types of processors depending on the type of computing device 900 implemented (e.g., a processor with fewer cores for a mobile device and a processor with more cores for a server). For example, depending on the type of computing device 900, the processor may be an ARM (Advanced RISC Machines) processor implemented using reduced instruction set computing (RISC), or an x86 processor implemented using complex instruction set computing (CISC). The computing device 900 may include one or more CPUs 906 in addition to one or more microprocessors or auxiliary coprocessors, such as a computing coprocessor.

[0074] In addition to or instead of CPU 906, GPU 908 may be configured to execute at least some of the computer-readable instructions to control one or more components of computing device 900 to execute one or more of the methods and / or processes described herein. One or more of GPU 908 may be an integrated GPU (e.g., one that accompanies one or more of CPU 906), and / or one or more of GPU 908 may be a discrete GPU. In an example, one or more of GPU 908 may be one or more coprocessors of CPU 906. GPU 908 may be used by computing device 900 to render graphics (e.g., 3D graphics) or perform general-purpose computing. For example, GPU 908 may be used for general-purpose computing on GPU (GPGPU). GPU 908 may include hundreds or thousands of cores having the ability to process hundreds or thousands of software threads simultaneously. GPU 908 can generate pixel data for an output image in response to rendering commands (e.g., rendering commands from CPU 906 received via a host interface). GPU 908 may include graphics memory, such as display memory, for storing pixel data or any other suitable data, such as GPGPU data. The display memory may be included as part of memory 904. The display memory may be included as part of memory 904. GPU 908 may include two or more GPUs operating in parallel (e.g., via a link). The link can connect directly to the GPUs (e.g., using NVLINK) or connect the GPUs via a switch (e.g., using NVSwitch). When coupled together, each GPU 908 can generate pixel data or GPGPU data for different portions of an output or different outputs (e.g., a first GPU for a first image and a second GPU for a second image). Each GPU can include its own memory or share memory with other GPUs.

[0075] In addition to and / or instead of the CPU 906 and / or the GPU 908, the logic unit 920 may be configured to execute at least some of the computer-readable instructions to control one or more components of the computing device 900 to execute one or more of the methods and / or processes described herein. In an embodiment, the CPU 906, the GPU 908, and / or the logic unit 920 can execute any combination of methods, processes, and / or portions thereof discretely or jointly. One or more of the logic units 920 may be part of and / or integrated with one or more of the CPU 906 and / or the GPU 908, and / or one or more of the logic units 920 may be discrete components with respect to the CPU 906 and / or the GPU 908 or otherwise external to them. In an embodiment, one or more of the logic units 920 may be coprocessors of one or more of the CPU 906 and / or one or more of the GPU 908.

[0076] Examples of the logic unit 920 include one or more processing cores and / or their components, such as tensor cores (TC), tensor processing units (TPU), pixel visual cores (PVC), vision processing units (VPU), graphics processing clusters (GPC), texture processing clusters (TPC), streaming multiprocessors (SM), tree traversal units (TTU), artificial intelligence accelerators (AIA), deep learning accelerators (DLA), arithmetic-logic units (ALU), application-specific integrated circuits (ASIC), floating-point units (FPU), input / output (I / O) elements, peripheral component interconnect (PCI) or peripheral component interconnect express (PCIe) elements, and / or the like.

[0077] The communication interface 910 can include one or more receivers, transmitters, and / or transceivers that enable the computing device 900 to communicate with other computing devices via an electronic communication network, including wired and / or wireless communication. The communication interface 910 can include components and functionality for enabling communication via any of several different networks, such as wireless networks (e.g., Wi-Fi, Z-Wave, Bluetooth®, Bluetooth® LE, ZigBee, etc.), wired networks (e.g., communicating via Ethernet® or InfiniBand), low-power wide-area networks (e.g., LoRaWAN, SigFox, etc.), and / or the Internet.

[0078] The I / O port 912 can enable the computing device 900 to be logically connected to other devices, including some of which can be built in (e.g., integrated) into the computing device 900, such as I / O components 914, presentation components 918, and / or other components. Exemplary I / O components 914 include microphones, mice, keyboards, joysticks, game pads, game controllers, satellite dishes, scanners, printers, wireless devices, etc. The I / O components 914 can provide a natural user interface (NUI) that processes air gestures, voice, or other physiological inputs generated by the user. In some cases, the input can be sent to an appropriate network element for further processing. The NUI can implement any combination of voice recognition, stylus recognition, face recognition, biometric recognition, gesture recognition on and adjacent to the screen, air gestures, head and gaze tracking, and touch recognition related to the display of the computing device 900 (as will be described in more detail later). The computing device 900 can include a depth camera for gesture detection and recognition, such as a stereoscopic camera system, an infrared camera system, an RGB camera system, touch screen technology, and combinations thereof. Additionally, the computing device 900 can include an accelerometer or gyroscope that enables motion detection (e.g., as part of an inertia measurement unit (IMU)). In some examples, the output of the accelerometer or gyroscope can be used by the computing device 900 to render immersive augmented reality or virtual reality.

[0079] The power supply device 916 can include a hard-wired power supply device, a battery power supply device, or a combination thereof. The power supply device 916 can provide power to the computing device 900 to enable the components of the computing device 900 to operate.

[0080] The presentation component 918 may include a display (e.g., a monitor, a touch screen, a television screen, a heads-up display (HUD), other display types, or a combination thereof), a speaker, and / or other presentation components. The presentation component 918 can receive data from other components (e.g., the GPU 908, the CPU 906, etc.) and output the data (e.g., as an image, a video, an audio, etc.).

[0081] Exemplary data center FIG. 10 shows an exemplary data center 1000 that may be used in at least one embodiment of the present disclosure. The data center 1000 may include a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and / or an application layer 1040.

[0082] As shown in FIG. 10, the data center infrastructure layer 1010 may include a resource orchestrator 1012, grouped computing resources 1014, and node computing resources (“node C.R.”) 1016(1) to 1016(N), where “N” represents a natural number of any integer. In at least one embodiment, the node C.R. 1016(1) to 1016(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors or graphics processing units (GPU), etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid state or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and / or cooling modules, etc., but are not limited thereto. In some embodiments, one or more of the node C.R. 1016(1) to 1016(N) may correspond to a server having one or more of the aforementioned computing resources. Additionally, in some embodiments, the node C.R. 1016(1) to 1016(N) may include one or more virtual components, such as vGPU, vCPU, and / or the like, and / or one or more of the node C.R. 1016(1) to 1016(N) may correspond to a virtual machine (VM).

[0083] In at least one embodiment, the grouped computing resources 1014 can include a separate group of node C.R.s 1016 housed within one or more racks (not shown), or multiple racks housed in data centers at various geographical locations (also not shown). Separate groups of node C.R.s 1016 within the grouped computing resources 1014 can include grouped computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s 1016 that include CPUs, GPUs, and / or other processors can be grouped within one or more racks to provide computing resources for supporting one or more workloads. One or more racks can also include any number of power modules, cooling modules, and / or network switches in any combination.

[0084] The resource orchestrator 1022 can configure or otherwise control one or more node C.R.s 1016(1)-1016(N) and / or the grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1022 can include a software design infrastructure ("SDI") management entity of the data center 1000. The resource orchestrator 1022 can include hardware, software, or some combination thereof.

[0085] In at least one embodiment, as shown in FIG. 10, the framework layer 1020 may include a job scheduler 1032, a configuration manager 1034, a resource manager 1036, and / or a distributed file system 1038. The framework layer 1020 may include a framework to support software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. The software 1032 or the application 1042 may each include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. The framework layer 1020 may be of a type of free and open-source software web application framework, such as Apache Spark (trademark) (hereinafter “Spark”), which may use the distributed file system 1038 for large-scale data processing (e.g., “big data”), but is not limited thereto. In at least one embodiment, the job scheduler 1032 may include a Spark driver to facilitate scheduling of workloads supported by various layers of the data center 1000. The configuration manager 1034 may have the ability to configure different layers, such as the software layer 1030 and the framework layer 1020 including Spark and the distributed file system 1038 to support large-scale data processing. The resource manager 1036 may have the ability to manage the clustered or grouped computing resources mapped or assigned for the support of the distributed file system 1038 and the job scheduler 1032. In at least one embodiment, the clustered or grouped computing resources may include the grouped computing resources 1014 in the data center infrastructure layer 1010. The resource manager 1036 may coordinate with the resource orchestrator 1012 to manage these mapped or assigned computing resources.

[0086] In at least one embodiment, the software 1032 included in the software layer 1030 may include at least a portion of nodes C.R. 1016(1) through 1016(N), the grouped computing resources 1014, and / or software used by the distributed file system 1038 of the framework layer 1020. The one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.

[0087] In at least one embodiment, the application 1042 included in the application layer 1040 may include at least a portion of nodes C.R. 1016(1) through 1016(N), the grouped computing resources 1014, and / or one or more types of applications used by the distributed file system 1038 of the framework layer 1020. The one or more types of applications may include, but are not limited to, machine learning applications including any number of genomics applications, cognitive computing, and training or inference software, machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), and / or other machine learning applications used in conjunction with one or more embodiments.

[0088] In at least one embodiment, any one of the configuration manager 1034, the resource manager 1036, and the resource orchestrator 1012 can implement any number and type of self-rewriting actions based on any amount and type of data obtained in any technically possible manner. The self-rewriting actions may free the data center operator of the data center 1000 by making potentially bad configuration decisions and perhaps avoiding underutilized and / or poorly performing portions of the data center.

[0089] Data center 1000 may include tools, services, software, or other resources for training one or more machine learning models or predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, a machine learning model may be trained by calculating weight parameters by a neural network architecture that uses the software and / or computing resources described above with respect to data center 1000. In at least one embodiment, a trained or deployed machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 1000 by using weight parameters calculated via one or more training techniques, which are not limited to, for example, those described herein.

[0090] In at least one embodiment, data center 1000 can use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, and / or other hardware (or corresponding virtual computing resources) for training and / or performing inferences using the aforementioned resources. Further, the aforementioned one or more software and / or hardware resources may be configured as services that enable a user to train or perform inferences on information, such as image recognition, speech recognition, or other artificial intelligence services.

[0091] Exemplary network environment A network environment suitable for use in implementing embodiments of the present disclosure may include one or more client devices, servers, network attached storage (NAS), other backend devices, and / or other device types. Client devices, servers, and / or other device types (e.g., each device) may be implemented as one or more instances of the computing device 900 of FIG. 9, and for example, each device may include similar components, features, and / or functionality of the computing device 900. Additionally, when backend devices (e.g., servers, NAS, etc.) are implemented, the backend devices may be included as part of the data center 1000, an example of which is further detailed herein with respect to FIG. 10.

[0092] The components of the network environment may communicate with each other via a network, which may be wired, wireless, or both. The network may include multiple networks, or a network of networks. By way of example, the network may include one or more wide area networks (WANs), one or more local area networks (LANs), one or more public networks, such as the Internet and / or the public switched telephone network (PSTN), and / or one or more private networks. When the network includes a wireless telecommunications network, components, such as base stations, communication towers, or access points (and other components), may provide a wireless connection.

[0093] A compatible network environment may include one or more peer-to-peer network environments (in which case, a server may or may not be included in the network environment) and one or more client-server network environments (in which case, one or more servers may be included in the network environment). In a peer-to-peer network environment, the functionality described herein with respect to a server may be implemented on any number of client devices.

[0094] In at least one embodiment, the network environment may include one or more cloud-based network environments, a distributed computing environment, a combination thereof, etc. The cloud-based network environment may include a framework layer, a job scheduler, a resource manager, and a distributed file system implemented on one or more of the servers, which may include one or more core network servers and / or edge servers. The framework layer may include a framework to support software in the software layer and / or one or more applications in the application layer. The software or application may each include web-based service software or an application. In an embodiment, one or more of the client devices may use web-based service software or an application (e.g., by accessing the service software and / or application via one or more application programming interfaces (APIs)). The framework layer may be of the type of free and open-source software web application frameworks that may use a distributed file system for large-scale data processing (e.g., "big data"), but is not limited thereto.

[0095] A cloud-based network environment may provide cloud computing and / or cloud storage that implements any combination of the computing and / or data storage functions (or one or more portions thereof) described herein. Any of these various functions may be distributed to multiple locations from a central or core server (such as one or more data centers that may be distributed across a state, region, country, the world, etc.). When the connection to a user (such as a client device) is relatively close to an edge server, the core server may delegate at least a portion of the functionality to the edge server. The cloud-based network environment may be private (e.g., restricted to a single organization), public (e.g., available to multiple organizations), and / or a combination thereof (e.g., a hybrid cloud environment).

[0096] The client device may include at least some of the components, features, and functionality of the exemplary computing device 900 described herein with respect to FIG. 9. By way of example, and not limitation, the client device may be implemented as a personal computer (PC), laptop computer, mobile device, smartphone, tablet computer, smart watch, wearable computer, personal digital assistant (PDA), MP3 player, virtual reality headset, global positioning system (GPS) or device, video player, video camera, surveillance device or system, vehicle, boat, aircraft, virtual machine, drone, robot, handheld communication device, hospital device, gaming device or system, entertainment system, vehicle computer system, embedded system controller, remote control, appliance, consumer electronics device, workstation, edge device, any combination of these described devices, or any other suitable device.

[0097] Exemplary autonomous vehicle FIG. 11A is a diagram of an exemplary autonomous vehicle 1100 according to some embodiments of the present disclosure. The autonomous vehicle 1100 (or referred to herein as "vehicle 1100") can include, but is not limited to, a passenger vehicle, such as a car, truck, bus, first responder vehicle, shuttle, electric or motorized bicycle, motorcycle, fire truck, police vehicle, ambulance, boat, construction vehicle, submarine, drone, and / or another type of vehicle (e.g., unmanned and / or carrying one or more passengers). Autonomous vehicles are generally described in terms of automation levels defined by the National Highway Traffic Safety Administration (NHTSA), a department of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE) "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicle" (Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of this standard). Vehicle 1100 can have the ability to function according to one or more of automation levels 3 to 5 of the autonomous driving level. For example, vehicle 1100 can have the ability of conditional automation (level 3), high automation (level 4), and / or full automation (level 5) according to the embodiment.

[0098] Vehicle 1100 can include components such as the vehicle's chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components. Vehicle 1100 can include a propulsion system 1150, such as an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. Propulsion system 1150 can be connected to the drive train of vehicle 1100 and can include a transmission to enable the propulsion force of vehicle 1100. Propulsion system 1150 can be controlled in response to receiving a signal from throttle / acceleration device 1152.

[0099] The steering system 1154, which may include a steering wheel, can be used to steer the vehicle 1100 (e.g., along a desired path or route) when the propulsion system 1150 is operating (e.g., when the vehicle is in motion). The steering system 1154 can receive a signal from a steering actuator 1156. The steering wheel may be an option for a fully automated (Level 5) function.

[0100] The brake sensor system 1146 can be used to operate vehicle brakes in response to receiving signals from a brake actuator 1148 and / or a brake sensor.

[0101] The controller 1136, which may include one or more system-on-chips (SoCs) 1104 (FIG. 11C) and / or GPUs, can provide signals (e.g., representations of commands) to one or more components and / or systems of the vehicle 1100. For example, the controller can send signals to operate the vehicle brakes via one or more brake actuators 1148, to operate the steering system 1154 via one or more steering actuators 1156, and to operate the propulsion system 1150 via one or more throttle / acceleration devices 1152. The controller 1136 can include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or to assist the driver in operating the vehicle 1100. The controller 1136 can include a first controller 1136 for autonomous driving functions, a second controller 1136 for functional safety functions, a third controller 1136 for artificial intelligence functions (e.g., computer vision), a fourth controller 1136 for infotainment functions, a fifth controller 1136 for redundancy in emergency situations, and / or other controllers. In some examples, a single controller 1136 can process two or more of the aforementioned functions, and two or more controllers 1136 can process a single function and / or any combination thereof.

[0102] Controller 1136 can provide signals for controlling one or more components and / or systems of vehicle 1100 in response to sensor data (e.g., sensor inputs) received from one or more sensors. The sensor data can be received, for example and without limitation, from a global navigation satellite system sensor 1158 (e.g., a global positioning system sensor), a RADAR sensor 1160, an ultrasonic sensor 1162, a LIDAR sensor 1164, an inertial measurement unit (IMU) sensor 1166 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 1196, a stereo camera 1168, a wide view camera 1170 (e.g., a fish-eye camera), an infrared camera 1172, a surround camera 1174 (e.g., a 360-degree camera), a long range and / or mid-range camera 1198, a speed sensor 1144 (e.g., for measuring the speed of vehicle 1100), a vibration sensor 1142, a steering sensor 1140, a brake sensor (e.g., as part of a brake sensor system 1146), and / or other sensor types.

[0103] One or more of the controllers 1136 of the vehicle 1100 may receive an input (e.g., represented by input data) from the instrument cluster 1132 of the vehicle 1100 and provide an output (e.g., represented by output data, display data, etc.) via a human-machine interface (HMI) display 1134, an audible annunciator, a loudspeaker, and / or other components of the vehicle 1100. The output may include information such as vehicle velocity, speed, time, map data (e.g., the HD map 1122 of FIG. 11C), position data (e.g., the position of the vehicle 1100 on a map, etc.), direction, the positions of other vehicles (e.g., occupancy grids), information regarding objects and the situation of objects as perceived by the controller 1136, etc. For example, the HMI display 1134 may display information regarding the presence of one or more objects (e.g., road signs, warning signs, changes in traffic signals, etc.) and / or driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B within 3.22 km (2 miles), etc.).

[0104] The vehicle 1100 further includes a network interface 1124 that can communicate via one or more networks using one or more wireless antennas 1126 and / or a modem. For example, the network interface 1124 may have the ability to communicate via LTE, WCDMA (registered trademark), UMTS, GSM, CDMA2000, etc. The wireless antenna 1126 may also use local area networks such as Bluetooth (registered trademark), Bluetooth (registered trademark) LE, Z-Wave, ZigBee, and / or low power wide-area networks (LPWAN) such as LoRaWAN, SigFox to enable communication between objects (e.g., vehicles, mobile devices, etc.) in the environment.

[0105] FIG. 11B is an example of the camera positions and fields of view of the exemplary autonomous vehicle 1100 of FIG. 11A according to some embodiments of the present disclosure. The cameras and their respective fields of view are one exemplary embodiment and are not intended to be limiting. For example, additional and / or alternative cameras may be included and / or the cameras may be placed at different positions on the vehicle 1100.

[0106] The camera type of the camera may include, but is not limited to, a digital camera adapted to be used with the components and / or systems of the vehicle 1100. The camera can operate at automotive safety integrity level (ASIL) B and / or at another ASIL. The camera type may have the ability to capture images at any rate, such as 60 frames per second (fps), 120 fps, 240 fps, etc., depending on the embodiment. The camera may have the ability to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In some examples, the color filter array may include an RCCC (red clear clear clear) color filter array, an RCCB (red clear clear blue) color filter array, an RBGC (red blue green clear) color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In some embodiments, clear pixel cameras, such as cameras with RCCC, RCCB, and / or RBGC color filter arrays, may be used in efforts to increase light sensitivity.

[0107] In some examples, one or more of the cameras can be used to perform advanced driver assistance system (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, a multi-functional mono-camera can be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlamp control. One or more of the cameras (e.g., all of the cameras) can record and provide image data (e.g., video) simultaneously.

[0108] One or more of the cameras can be mounted on mounting parts such as custom-designed (3D printed) parts to remove stray light and reflections from inside the vehicle that may interfere with the camera's image data capture ability (e.g., reflections from the dashboard reflected in the front windshield mirror). Referring to the side mirror mounting part, the side mirror part can be custom 3D printed so that the camera mounting plate conforms to the shape of the side mirror. In some examples, the camera can be integrated into the side mirror. For side view cameras, the camera can also be integrated into four struts at each corner of the cabin.

[0109] A camera having a field of view that includes a portion of the environment in front of vehicle 1100 (e.g., a forward-facing camera) can be used for surround view to help identify the forward path and obstacles and, with the help of one or more controllers 1136 and / or a control SoC, to help provide information essential for the generation of an occupancy grid and / or the determination of a preferred vehicle path. The forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including emergency braking, pedestrian detection, and collision avoidance. The forward-facing camera can also be used for ADAS functions and systems including other functions such as lane departure warning ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.

[0110] Various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform that includes a CMOS (complementary metal oxide semiconductor) color imaging device. Another example may be a wide-view camera 1170 that can be used to perceive objects entering from a view of the surroundings (e.g., pedestrians, intersecting traffic, or bicycles). Although only one wide-view camera is shown in FIG. 11B, any number of wide-view cameras 1170 may be present on the vehicle 1100. Additionally, a long-range camera 1198 (e.g., a long-view stereo camera pair) can be used for depth-based object detection, particularly for objects for which the neural network has not yet been trained. The long-range camera 1198 can also be used for object detection and classification, as well as for basic object tracking.

[0111] One or more stereo cameras 1168 can also be included in the forward-facing configuration. The stereo camera 1168 can include an integrated control unit with an expandable processing unit that can provide a programmable logic (FPGA) and a multi-core microprocessor with a CAN or Ethernet® interface integrated on a single chip. Such a unit can be used to generate a 3D map of the vehicle's environment that includes distance estimates for all points within the image. An alternative stereo camera 1168 can include a compact stereo vision sensor that includes two camera lenses (one each for left and right) and an image processing chip that can measure the distance from the vehicle to a target object and activate autonomous emergency braking and lane departure warning functions using the generated information (e.g., metadata). Other types of stereo cameras 1168 may be used in addition to, or in place of, those described herein.

[0112] A camera (e.g., a side view camera) having a field of view that includes a portion of the environment relative to the side of vehicle 1100 can be used for surround view to provide information for creating and updating an occupancy grid and for generating a side impact collision warning. For example, surround cameras 1174 (e.g., four surround cameras 1174 as shown in FIG. 11B) can be positioned on vehicle 1100. Surround cameras 1174 can include wide view cameras 1170, fisheye cameras, 360-degree cameras, and / or the like. For example, four fisheye cameras can be disposed at the front, rear, and sides of the vehicle. In an alternative arrangement, the vehicle may use three surround cameras 1174 (e.g., left, right, and rear), and one or more other cameras (e.g., a forward-facing camera) can be utilized as a fourth surround view camera.

[0113] A camera (e.g., a rearview camera) having a field of view that includes a portion of the environment relative to the rear of vehicle 1100 can be used for parking assistance, surround view, rear collision warning, and for creating and updating an occupancy grid. A wide variety of cameras can be used, including but not limited to cameras suitable as forward-facing cameras (e.g., long range and / or midrange cameras 1198, stereo cameras 1168), infrared cameras 1172, etc., as described herein.

[0114] FIG. 11C is a block diagram of an exemplary system architecture of the exemplary autonomous vehicle 1100 of FIG. 11A, according to some embodiments of the present disclosure. It should be understood that this and other arrangements described herein are presented by way of example only. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) may be used in addition to or instead of those shown, and some elements may be excluded altogether. Further, many of the elements described herein are functional entities that may be implemented as individual or distributed components or in combination with other components, and in any suitable combination and location. The various functions described herein as being performed by entities may be implemented by hardware, firmware, and / or software. For example, the various functions may be implemented by a processor executing instructions stored in memory.

[0115] Each of the components, features, and systems of the vehicle 1100 of FIG. 11C is shown as being connected via a bus 1102. The bus 1102 may include a Controller Area Network (CAN) data interface (or referred to as a “CAN bus”). CAN may be a network within the vehicle 1100 used to assist in the control of various features and functions of the vehicle 1100, such as the operation of brakes, acceleration, brakes, steering, windshield wipers, etc. The CAN bus may be configured to have dozens or hundreds of nodes, each having its own unique identifier (e.g., CAN ID). The CAN bus may be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle status indicators. The CAN bus may be ASIL B compliant.

[0116] Bus 1102 is described herein as being a CAN bus, but this is not intended to be limiting. For example, in addition to, or as an alternative to, the CAN bus, FlexRay and / or Ethernet® may be used. Additionally, a single line is used to represent bus 1102, but this is not intended to be limiting. For example, any number of buses 1102 may exist that include one or more CAN buses, one or more FlexRay buses, one or more Ethernet® buses, and / or one or more other types of buses that use different protocols. In some examples, two or more buses 1102 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 1102 may be used for a collision avoidance function and a second bus 1102 may be used for actuation control. In any example, each bus 1102 may communicate with any of the components of vehicle 1100, and two or more buses 1102 may communicate with the same component. In some examples, each SoC 1104, each controller 1136, and / or each computer within the vehicle may have access to the same input data (e.g., an input from a sensor of vehicle 1100) and may be connected to a common bus such as a CAN bus.

[0117] Vehicle 1100 may include one or more controllers 1136, such as those described herein with respect to FIG. 11A. Controller 1136 may be used for various functions. Controller 1136 may be coupled to any of the various other components and systems of vehicle 1100 and may be used for the control of vehicle 1100, the artificial intelligence of vehicle 1100, the infotainment for vehicle 1100, and / or the like.

[0118] Vehicle 1100 may include a system-on-chip (SoC) 1104. The SoC 1104 may include a CPU 1106, a GPU 1108, a processor 1110, a cache 1112, an accelerator 1114, a data store 1116, and / or other components and features not shown. The SoC 1104 may be used to control the vehicle 1100 within various platforms and systems. For example, the SoC 1104 may be coupled in a system (such as the system of the vehicle 1100) having an HD map 1122 that can obtain map refreshes and / or updates via a network interface 1124 from one or more servers (such as the server 1178 of FIG. 11D).

[0119] The CPU 1106 may include a CPU cluster or CPU complex (or also referred to as a "CCPLEX"). The CPU 1106 may include a plurality of cores and / or an L2 cache. For example, in some embodiments, the CPU 1106 may include 8 cores within a coherent multiprocessor configuration. In some embodiments, the CPU 1106 may include 4 dual-core clusters, each cluster having a dedicated L2 cache (such as a 2MB L2 cache). The CPU 1106 (such as a CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of the clusters of the CPU 1106 to become active at any given time.

[0120] CPU 1106 can implement power management capabilities that include one or more of the following features: individual hardware blocks can be automatically clock-gated when in an idle state to conserve dynamic power, each core clock can be gated when the core is not actively executing instructions by the execution of WFI / WFE instructions, each core can be independently power-gated, each core cluster can be independently clock-gated when all cores are clock-gated or power-gated, and / or each core cluster can be independently power-gated when all cores are power-gated. CPU 1106 can further implement enhanced algorithms for managing power states, where the allowed power states and expected wake-up times are specified and the hardware / microcode determines the best power state to input to the cores, clusters, and CCPLEX. The processing cores can support a simplified power state input sequence in software where the work is offloaded to microcode.

[0121] The GPU 1108 may include an integrated GPU (or referred to herein as "iGPU"). The GPU 1108 can be programmable and can be efficient for parallel workloads. In some examples, the GPU 1108 can use an enhanced tensor instruction set. The GPU 1108 may include one or more streaming microprocessors, where each streaming microprocessor may include an L1 cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more of the streaming microprocessors may share a cache (e.g., an L2 cache having a storage capacity of 512 KB). In some embodiments, the GPU 1108 may include at least eight streaming microprocessors. The GPU 1108 can use a compute application programming interface (API). Additionally, the GPU 1108 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0122] The GPU 1108 can be power-optimized for the best performance in automotive and embedded use cases. For example, the GPU 1108 can be manufactured on FinFET (Fin field-effect transistor). However, this is not intended to be limiting, and the GPU 1108 can be manufactured using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate several mixed-precision processing cores partitioned into multiple blocks. By way of example and not limitation, for instance, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In such an example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. Additionally, the streaming microprocessor can include independent parallel integer and floating-point data paths for efficient execution of workloads having a mix of compute and addressing operations. The streaming microprocessor can include independent thread scheduling capabilities to enable higher fine-grained synchronization and cooperation among parallel threads. The streaming microprocessor can include a combined L1 data cache and shared memory unit to simplify programming while improving performance.

[0123] In some examples, GPU 1108 may include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide for a peak memory bandwidth of 900 GB / second. In some examples, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as graphics double data rate type five synchronous random-access memory (GDDR5), may be used.

[0124] GPU 1108 can include unified memory technology that includes access counters to enable more accurate movement of those memory pages to the processor that most frequently accesses the memory pages, thereby improving the efficiency of the memory ranges shared between processors. In some examples, address translation service (ATS) support may be used to enable the GPU 1108 to directly access the CPU 1106 page table. In such examples, when the GPU 1108 memory management unit (MMU) experiences a miss, an address translation request may be sent to the CPU 1106. In response, the CPU 1106 can examine its page table for the virtual-to-physical mapping of the address and send the translation back to the GPU 1108. As such, unified memory technology can enable a single unified virtual address space for the memory of both the CPU 1106 and the GPU 1108, thereby simplifying the GPU 1108 programming and porting of applications to the GPU 1108.

[0125] In addition, GPU 1108 may include an access counter that can record the frequency of access of GPU 1108 to the memory of other processors. The access counter can help ensure that memory pages are moved to the physical memory of the processor that most frequently accesses that page.

[0126] SoC 1104 may include any number of caches 1112, including those described herein. For example, cache 1112 may include an L3 cache that is available to both CPU 1106 and GPU 1108 (e.g., connected to both CPU 1106 and GPU 1108). Cache 1112 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). The L3 cache may have a smaller cache size, but may include 4 MB or more depending on the embodiment.

[0127] SoC 1104 may include an arithmetic logic unit (ALU) that can be utilized when executing processing for any of the various tasks or operations of vehicle 1100 (e.g., processing DNN). In addition, SoC 1104 may include a floating point unit (FPU) (or other mass coprocessor or numeric coprocessor type) for performing mathematical operations within the system. For example, SoC 114 may include one or more FPUs integrated as execution units within CPU 1106 and / or GPU 1108.

[0128] SoC 1104 may include one or more accelerators 1114 (e.g., hardware accelerators, software accelerators, or combinations thereof). For example, SoC 1104 may include a hardware acceleration cluster that may include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other operations. The hardware acceleration cluster may be used to complement the GPU 1108 and to offload some of the tasks of the GPU 1108 (e.g., to free up more cycles of the GPU 1108 for executing other tasks). As an example, the accelerator 1114 may be used for target workloads that are stable enough to be suitable for acceleration (e.g., perception, convolutional neural networks (CNNs)). As used herein, the term "CNN" may include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).

[0129] The acceleration device 1114 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include one or more tensor processing units (TPUs) configured to provide an additional 10 trillion operations per second for deep learning applications and inferences. The TPU may be an acceleration device configured and optimized to execute image processing functions (e.g., those of CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating-point operations, as well as for inferences. The design of the DLA can provide more performance per millimeter than a general-purpose GPU and greatly exceed the performance of the CPU. The TPU can execute several functions, including, for example, a single-instance convolution function that supports INT8, INT16, and FP16 data types for both features and weights, as well as a post-processor function.

[0130] The DLA can rapidly and efficiently execute neural networks, particularly CNNs, with processed or unprocessed data for any of a variety of functions, including but not limited to: CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, CNNs for emergency vehicle detection and identification and detection using data from a microphone, CNNs for face recognition and vehicle owner identification using data from a camera sensor, and / or CNNs for security and / or safety-related events.

[0131] The DLA can execute any function of the GPU 1108, and by using the inference accelerator, for example, a designer can target either the DLA or the GPU 1108 for any function. For example, a designer can focus on processing CNNs and floating-point operations on the DLA and leave other functions to the GPU 1108 and / or other acceleration devices 1114.

[0132] The acceleration device 1114 (e.g., a hardware acceleration cluster) may include, or may be referred to herein as, a programmable vision accelerator (PVA) and a computer vision acceleration device. The PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA may include, but is not limited to, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0133] The RISC core can interact with an image sensor (e.g., the image sensor of any of the cameras described herein), an image signal processor, and / or the like. Each RISC core may include any amount of memory. The RISC core can use any of several protocols, depending on the embodiment. In some examples, the RISC core can execute a real-time operating system (RTOS). The RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (ASICs), and / or memory devices. For example, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0134] The DMA can enable the components of the PVA to access the system memory independent of the CPU1106. The DMA can support any number of features used to provide optimization for the PVA, including but not limited to supporting multidimensional addressing and / or circular addressing. In some examples, the DMA can support addressing up to six dimensions or more, which may include block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0135] The vector processor may be a programmable processor designed to efficiently and flexibly execute the programming of computer vision algorithms and provide signal processing capabilities. In some examples, the PVA may include a PVA core and two vector processing subsystem partitions. The PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. The vector processing subsystem can operate as the primary processing engine of the PVA and may include a vector processing unit (VPU), an instruction cache, and / or vector memory (e.g., VMEM). The VPU core may include a digital signal processor, such as a single instruction, multiple data (SIMD), very long instruction word (VLIW) digital signal processor. The combination of SIMD and VLIW can increase throughput and speed.

[0136] Each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in some examples, each vector processor may be configured to execute independently from other vector processors. In other examples, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in some embodiments, multiple vector processors included in a single PVA may execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA may be able to execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms sequentially on an image or portions of an image. In particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. Additionally, the PVA may include additional error correcting code (ECC) memory to enhance overall system safety.

[0137] The accelerator 1114 (e.g., a hardware acceleration cluster) may include a computer vision network on chip and SRAM to provide high bandwidth, low latency SRAM for the accelerator 1114. In some examples, the on-chip memory may include at least 4MB of SRAM consisting of, for example and without limitation, 8 field-configurable memory blocks that may be accessible by both the PVA and the DLA. Each pair of memory blocks may include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory may be used. The PVA and the DLA may access the memory via a backbone that provides high-speed access to the PVA and the DLA to the memory. The backbone may include a computer vision network on chip that interconnects the PVA and the DLA to the memory (e.g., using the APB).

[0138] The computer vision network-on-chip may include an interface that determines that both the PVA and the DLA are operable and provide valid signals before any control signal / address / data transmission. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transfer. This type of interface can comply with the ISO26262 or IEC61508 standard, although other standards and protocols may be used.

[0139] In some examples, the SoC 1104 may include a real-time ray tracing hardware accelerator as described in U.S. Patent Application No. 16 / 101,232, filed on August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine the position and scale of objects (e.g., within a world model) for generating real-time visualization simulations for RADAR signal interpretation, for acoustic propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulations, for comparison against LIDAR data for localization and / or other functions, and / or for other uses. In some embodiments, one or more tree traversal units (TTUs) may be used to perform one or more ray tracing related operations.

[0140] The acceleration device 1114 (e.g., a hardware acceleration device cluster) has various uses for autonomous driving. The PVA may be a programmable vision acceleration device that can be used in extremely important processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are suitable for areas of algorithms that require predictable processing at low power and low latency. In other words, the PVA functions well with low latency and low power and a predictable execution time, even on small data sets, for semi-dense or dense normal calculations. Therefore, since the PVA is efficient in object detection and integer calculations, in the context of a platform for autonomous vehicles, the PVA is designed to execute classic computer vision algorithms.

[0141] For example, according to one embodiment of the present technology, the PVA is used to execute computer stereo vision. A semi-global matching-based algorithm may be used in some examples, but this is not intended to be limiting. Many applications for level 3-5 autonomous driving require motion estimation / stereo matching on the fly (e.g., SFM (structure from motion), pedestrian recognition, lane detection, etc.). The PVA can execute computer stereo vision functions with inputs from two monocular cameras.

[0142] In some examples, the PVA can be used to execute high-density optical flow. By processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used in flight depth processing time, for example, by processing the raw time of flight data to provide the processed time of flight data.

[0143] DLA can be used to execute any type of network to enhance control and driving safety, including, for example, a neural network that outputs a measure of the reliability of each object detection. Such reliability values can be interpreted as probabilities or as providing the relative "weight" of each detection compared to other detections. This reliability value enables the system to make further decisions regarding which detections should be considered true positive detections rather than false positive detections. For example, the system can set a reliability threshold and consider only detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the vehicle to automatically execute emergency braking, which is clearly undesirable. Therefore, only the most confident detections should be considered as triggers for AEB. DLA can execute a neural network that regresses the reliability value. The neural network can receive as its input at least some subset of parameters such as the bounding box dimensions, the ground plane estimate obtained (e.g., from another subsystem), the vehicle 1100 orientation, distance, 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor 1164 or RADAR sensor 1160), the output of the inertial measurement unit (IMU) sensor 1166 that correlates with the above, and so on.

[0144] SoC 1104 may include a data store 1116 (e.g., memory). The data store 1116 may be the on-chip memory of SoC 1104 and can store neural networks to be executed on the GPU and / or DLA. In some examples, the data store 1116 may have a capacity large enough to store multiple instances of the neural network for redundancy and security. The data store 1112 may include an L2 or L3 cache 1112. References to the data store 1116 may include references to memory related to the PVA, DLA, and / or other accelerators 1114 as described herein.

[0145] SoC 1104 may include one or more processors 1110 (e.g., embedded processors). The processor 1110 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management capabilities and related security enforcement. The boot and power management processor may be part of the SoC 1104 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance with system low-power state transitions, management of the SoC 1104 thermal and temperature sensors, and / or management of the SoC 1104 power state. Each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and SoC 1104 can use the ring oscillator to detect the temperature of the CPU 1106, GPU 1108, and / or accelerator 1114. If the temperature is determined to exceed a threshold, the boot and power management processor can enter a temperature fault routine and place SoC 1104 in a lower power state and / or put the vehicle 1100 in a chauffeur safe stop mode (e.g., safely stop the vehicle 1100).

[0146] Processor 1110 may further include a set of embedded processors that can perform the functions of an audio processing engine. The audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio via multiple interfaces and a wide and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.

[0147] Processor 1110 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake use cases. The always-on processor engine may include a processor core, tightly coupled RAM, supporting peripheral devices (such as timers and interrupt controllers), various I / O controller peripheral devices, and routing logic.

[0148] Processor 1110 may further include a safety cluster engine that includes a dedicated processor subsystem for processing the safety management of automotive applications. The safety cluster engine may include two or more processor cores, tightly coupled RAM, supporting peripheral devices (such as timers, interrupt controllers, etc.), and / or routing logic. In the safety mode, the two or more cores can operate in a lockstep mode and function as a single core having comparison logic for detecting any differences between their operations.

[0149] Processor 1110 may further include a real-time camera engine that may include a dedicated processor subsystem for processing real-time camera management.

[0150] Processor 1110 may further include a high dynamic range signal processor that includes an image signal processor, which is a hardware engine that is part of the camera processing pipeline.

[0151] The processor 1110 may include a video image synthesizer, which may be a processing block (e.g., implemented in a microprocessor) that implements the post-video processing functions required by a video playback application to produce the final image for the player window. The video image synthesizer can perform lens distortion correction with the wide-view camera 1170, with the surround camera 1174, and / or with the in-cabin monitoring camera sensor. The in-cabin monitoring camera sensor is preferably monitored by a neural network running on another instance of a high-level SoC configured to identify in-cabin events and respond appropriately. The in-cabin system can perform lip reading to enable cellular service and make calls, compose emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. Certain functions are available to the driver only when operating in autonomous mode and are otherwise disabled.

[0152] The video image synthesizer may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, if motion occurs in the video, the noise reduction reduces the weight of the information provided by adjacent frames and appropriately weights the spatial information. If the image or a portion of the image does not contain motion, the temporal noise reduction performed by the video image synthesizer can use information from the previous image to reduce the noise in the current image.

[0153] The video image synthesizer may also be configured to perform stereo rectification on the input stereo lens frame. The video image synthesizer can further be used for user interface synthesis when the operating system desktop is in use, and the GPU 1108 is not required to continuously render new surfaces. Even when the GPU 1108 is powered on and actively performing 3D rendering, the video image synthesizer can be used to offload the GPU 1108 to improve performance and responsiveness.

[0154] SoC 1104 may further include a mobile industry processor interface (MIPI) camera serial interface, a high-speed interface, and / or a video input block that can be used for the camera and related pixel input functions to receive video and input from the camera. SoC 1104 may further include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.

[0155] SoC 1104 may further include a wide range of peripheral interfaces to enable communication with peripheral devices, audio codecs, power management, and / or other devices. SoC 1104 may be used to process data from cameras (connected via, for example, Gigabit Multimedia Serial Link and Ethernet®), sensors (such as LIDAR sensor 1164, RADAR sensor 1160, etc., which may be connected via Ethernet®), data from bus 1102 (such as the speed, steering wheel position, etc., of vehicle 1100), and data from GNSS sensor 1158 (connected via, for example, Ethernet® or CAN bus). SoC 1104 may include its own DMA engine and may further include a dedicated high-performance large-capacity storage controller that can be used to free the CPU 1106 from routine data management tasks.

[0156] SoC 1104 may be an end-to-end platform with a flexible architecture that scales to automation levels 3 - 5, thereby leveraging and efficiently using computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack together with deep learning tools, providing an integrated functional safety architecture. SoC 1104 can be faster, more reliable, more energy-efficient, and more space-efficient than conventional systems. For example, when accelerator 1114 is coupled with CPU 1106, GPU 1108, and data store 1016 can provide a fast and efficient platform for level 3 - 5 autonomous vehicles.

[0157] Accordingly, the present technology provides capabilities and functionality that cannot be achieved by conventional systems. For example, computer vision algorithms can be implemented to execute various processing algorithms over a wide variety of visual data and can be executed on a CPU, configured using a high-level programming language such as the C programming language. However, a CPU often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. Specifically, many CPUs cannot execute real-time multi-object detection algorithms, which are requirements for in-vehicle ADAS applications and actual level 3-5 autonomous vehicle requirements.

[0158] In contrast to conventional systems, by providing a CPU complex, a GPU complex, and a hardware acceleration cluster, the technology described herein enables multiple neural networks to be executed simultaneously and / or sequentially and enables the results to be combined to enable level 3-5 autonomous driving functionality. For example, a CNN executed on a DLA or a dGPU (e.g., GPU 1120) can include text and word recognition that enables a supercomputer to read and understand traffic signs, including signs that the neural network has not been specifically trained for. The DLA can further include a neural network that can identify, interpret, and provide a semantic understanding of the signs and pass the semantic understanding to a route planning module executed on the CPU complex.

[0159] As another example, multiple neural networks can be executed simultaneously as required for level 3, 4, or 5 operation. For example, along with the electro-optical, a warning sign consisting of "Attention: Flashing light indicates a frozen state" can be interpreted independently or collectively by some neural networks. The sign itself can be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing light indicates a frozen state" can be interpreted by a second deployed neural network that informs the vehicle's route planning software (preferably running on a CPU complex) that a frozen state exists when a flashing light is detected. The flashing light can be identified by informing the vehicle's route planning software of the presence (or absence) of the flashing light and operating a third deployed neural network through multiple frames. All three neural networks can be executed simultaneously, such as within the DLA and / or on the GPU 1108.

[0160] In some examples, a CNN for face recognition and vehicle owner identification can use data from a camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 1100. A always-on sensor processing engine can be used to unlock and turn on the vehicle when the owner approaches the driver's side door, and to stop the vehicle's operation when the owner leaves the vehicle in security mode. In this way, the SoC 1104 provides security against theft and / or carjacking.

[0161] In another example, the CNN for emergency vehicle detection and identification can detect and identify emergency vehicle sirens using data from microphone 1196. In contrast to conventional systems that use general classifiers to detect sirens and manually extract features, SoC 1104 uses a CNN for classification of environmental and urban sounds, as well as for classification of visual data. In a preferred embodiment, the CNN running on the DLA is trained to identify the relative end velocity of an emergency vehicle (e.g., by using the Doppler effect). The CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by GNSS sensor 1158. Thus, for example, when operating in Europe, the CNN will attempt to detect European sirens, and when in the United States, the CNN will attempt to identify only North American sirens. After an emergency vehicle is detected, a control program can be used to execute an emergency vehicle safety routine that decelerates the vehicle, stops it at the side of the road, parks the vehicle, and / or idles the vehicle, with the assistance of ultrasonic sensor 1162 until the emergency vehicle has passed.

[0162] The vehicle may include a CPU 1118 (e.g., a discrete CPU, or dCPU) that can be coupled to SoC 1104 via a high-speed interconnect (e.g., PCIe). The CPU 1118 may include, for example, an X86 processor. The CPU 1118 can be used to perform any of a variety of functions, including, for example, mediating potential inconsistencies between the ADAS sensors and SoC 1104, and / or monitoring the status and health of controller 1136 and / or infotainment SoC 1130.

[0163] Vehicle 1100 may include a GPU 1120 (e.g., an individual GPU, or dGPU) that can be connected to SoC 1104 via a high-speed interconnect (e.g., NVIDIA's NVLINK). The GPU 1120 can provide additional artificial intelligence capabilities, such as by running redundant and / or different neural networks, and can be used to train and / or update neural networks based on inputs (e.g., sensor data) from the sensors of vehicle 1100.

[0164] Vehicle 1100 may further include a network interface 1124 that may include one or more wireless antennas 1126 (e.g., one or more wireless antennas for different communication protocols such as cellular antennas, Bluetooth® antennas, etc.). The network interface 1124 can be used to enable wireless connections with the cloud via the Internet (e.g., with server 1178 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a passenger's client device). To communicate with other vehicles, a direct link can be established between two vehicles and / or an indirect link can be established (e.g., through a network and via the Internet). The direct link can use and provide a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link can provide vehicle 1100 information regarding vehicles in proximity to vehicle 1100 (e.g., vehicles in front of, beside, and / or behind vehicle 1100). This functionality may be part of the cooperative adaptive cruise control function of vehicle 1100.

[0165] Network interface 1124 may include a SoC that provides modulation and demodulation functions and enables the controller 1136 to communicate via a wireless network. The network interface 1124 may include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. The frequency conversion can be performed through well-known processes and / or using a superheterodyne process. In some examples, the radio frequency front end functions may be provided by a separate chip. The network interface may include wireless capabilities for communicating via LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth®, Bluetooth® LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0166] Vehicle 1100 may further include a data store 1128 that may include storage outside the chip (e.g., outside SoC 1104). The data store 1128 may include one or more storage elements including RAM, SRAM, DRAM, VRAM, flash, hard disk, and / or other components and / or devices capable of storing at least 1 bit of data.

[0167] Vehicle 1100 may further include a GNSS sensor 1158. The GNSS sensor 1158 (e.g., GPS, assisted GPS sensor, differential GPS (DGPS) sensor, etc.) supports mapping, perception, occupancy grid generation, and / or route planning functions. For example, any number of GNSS sensors 1158 may be used, including but not limited to GPS with an Ethernet® to serial (RS-232) bridge USB connector.

[0168] Vehicle 1100 may further include a RADAR sensor 1160. The RADAR sensor 1160 can be used by the vehicle 1100 for long-range vehicle detection even in darkness and / or severe weather conditions. The RADAR functional safety level may be ASIL B. In some examples, the RADAR sensor 1160 can use CAN and / or bus 1102 for control and to access object tracking data (e.g., to transmit data generated by the RADAR sensor 1160) using access to Ethernet (registered trademark) to access raw data. A wide variety of RADAR sensor types can be used. For example, and without limitation, the RADAR sensor 1160 may be suitable for front, rear, and side RADAR use. In some examples, a pulse Doppler RADAR sensor is used.

[0169] The RADAR sensor 1160 can include different configurations, such as long range with a narrow field of view, short range with a wide field of view, short-range side coverage, etc. In some examples, the long-range RADAR can be used for an adaptive cruise control function. The long-range RADAR system can provide a wide field of view realized by two or more independent scans, such as within a range of 250 m. The RADAR sensor 1160 can help distinguish between static and moving objects and can be used by an ADAS system for emergency brake assist and forward collision warning. The long-range RADAR sensor can include a monostatic multi-modal RADAR with a plurality (e.g., six or more) of fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In one example having six antennas, the four central antennas can create a focused beam pattern designed to record the surroundings of the vehicle 1100 at high speed while minimizing interference from traffic in adjacent lanes. The other two antennas can widen the field of view and enable rapid detection of vehicles entering or leaving the lane of the vehicle 1100.

[0170] As an example, a mid-range RADAR system can include a range up to 1160 m (front) or 80 m (rear), and a field of view up to 42 degrees (front) or 1150 degrees (rear). A short-range RADAR system can include, but is not limited to, RADAR sensors designed to be installed at both ends of the rear bumper. When installed at both ends of the rear bumper, such a RADAR sensor system can create two beams that constantly monitor the blind spots behind and adjacent to the vehicle.

[0171] The short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.

[0172] Vehicle 1100 can further include ultrasonic sensors 1162. The ultrasonic sensors 1162, which can be positioned at the front, rear, and / or sides of vehicle 1100, can be used for parking assist and / or for creating and updating occupancy grids. A variety of ultrasonic sensors 1162 can be used, and different ultrasonic sensors 1162 can be used for different ranges of detection (e.g., 2.5 m, 4 m). The ultrasonic sensors 1162 can operate at a functional safety level of ASIL B.

[0173] Vehicle 1100 can include a LIDAR sensor 1164. The LIDAR sensor 1164 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 1164 can also be at a functional safety level of ASIL B. In some examples, vehicle 1100 can include multiple (e.g., 2, 4, 6, etc.) LIDAR sensors 1164 that can use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).

[0174] In some examples, the LIDAR sensor 1164 may have the ability to provide a list of objects and their distances in a 360-degree field of view. Commercially available LIDAR sensors 1164 may have, for example, an accuracy of 2 cm to 3 cm, support for 1100 Mbps Ethernet® connections, and a advertised range of about 1100 m. In some examples, one or more non-protruding LIDAR sensors 1164 may be used. In such examples, the LIDAR sensor 1164 may be implemented as a small device that can be incorporated into the front, rear, sides, and / or corners of the vehicle 1100. In such examples, the LIDAR sensor 1164 may have a range of 200 m even for low-reflectivity objects and can provide up to a 120-degree horizontal and 35-degree vertical field of view. The LIDAR sensor 1164 attached to the front may be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0175] In some examples, LIDAR technologies such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmitter to illuminate the area around the vehicle up to about 200 m. The flash LIDAR unit includes a receptor that records the laser pulse travel time and the reflected light on each pixel, corresponding in sequence to the range from the vehicle to the object. Flash LIDAR can enable a high-precision and distortion-free image of the surroundings to be generated with every laser flash. In some examples, four flash LIDAR sensors can be deployed, one on each side of vehicle 1100. Available 3D flash LIDAR systems include a solid-state 3D steering array LIDAR camera (e.g., a non-scanning LIDAR device) that has no moving parts other than a blower. The flash LIDAR device can use class I (eye-safe) laser pulses of 5 nanoseconds per frame and can capture the reflected laser light in the form of a 3D range point cloud and co-recorded intensity data. By using flash LIDAR, also, since the flash LIDAR is a solid-state device with no moving parts, LIDAR sensor 1164 can be less susceptible to the effects of motion blur, vibration, and / or shock.

[0176] The vehicle may further include an IMU sensor 1166. In some examples, IMU sensor 1166 can be positioned at the center of the rear axle of vehicle 1100. IMU sensor 1166 can include, for example, but is not limited to, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other sensor types. In some examples, in a 6-axis application, for example, IMU sensor 1166 can include an accelerometer and a gyroscope, while in a 9-axis application, IMU sensor 1166 can include an accelerometer, a gyroscope, and a magnetometer.

[0177] In some embodiments, the IMU sensor 1166 may be implemented as a miniature, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical system (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. As such, in some examples, the IMU sensor 1166 may enable the vehicle 1100 to estimate its heading without requiring input from a magnetic sensor by directly observing and correlating changes in velocity from the GPS to the IMU sensor 1166. In some examples, the IMU sensor 1166 and the GNSS sensor 1158 may be combined in a single integrated unit.

[0178] The vehicle may include a microphone 1196 disposed within and / or around the vehicle 1100. The microphone 1196 may be used, among other things, for emergency vehicle detection and identification.

[0179] The vehicle may further include any number of camera types, including a stereo camera 1168, a wide view camera 1170, an infrared camera 1172, a surround camera 1174, a long range and / or medium range camera 1198, and / or other camera types. The cameras may be used to capture image data around the entire outer surface of the vehicle 1100. The type of cameras used is determined according to the embodiments and requirements of the vehicle 1100, and any combination of camera types may be used to achieve the required coverage around the vehicle 1100. Additionally, the number of cameras may vary according to the embodiments. For example, the vehicle may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, and / or another number of cameras. As an example, the cameras may support, but are not limited to, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet (registered trademark). Each camera is described in more detail herein with reference to FIGS. 11A and 11B.

[0180] Vehicle 1100 may further include a vibration sensor 1142. The vibration sensor 1142 can measure the vibration of components of the vehicle, such as an axle. For example, a change in vibration may indicate a change in the road surface. In another example, when two or more vibration sensors 1142 are used, the difference in vibration may be used to determine the friction or slipperiness of the road surface (e.g., when the difference in vibration is between a power-driven axle and a freely rotating axle).

[0181] Vehicle 1100 may include an ADAS system 1138. In some examples, the ADAS system 1138 may include a SoC. The ADAS system 1138 may include autonomous / adaptive / automatic cruise control (ACC), cooperative adaptive cruise control (CACC), forward crash warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functions.

[0182] The ACC system may use a RADAR sensor 1160, a LIDAR sensor 1164, and / or a camera. The ACC system may include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle directly in front of vehicle 1100 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle ahead. Lateral ACC performs distance holding and advises vehicle 1100 to change lanes when necessary. Lateral ACC is related to other ADAS applications such as LCA and CWS.

[0183] CACC can use information from other vehicles that can be received via a wireless link from other vehicles through network interface 1124 and / or wireless antenna 1126, or indirectly through a network connection (e.g., via the Internet). The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (I2V) communication link. In general, the V2V communication concept provides information about the vehicle immediately ahead (e.g., the vehicle immediately in front of vehicle 1100 in the same lane as vehicle 1100), while the I2V communication concept provides information about traffic further ahead. The CACC system can include either or both of an I2V information source and a V2V information source. Given information about the vehicle ahead of vehicle 1100, CACC can be made more reliable and has the potential to make traffic flow smoother and reduce road congestion.

[0184] The FCW system is designed to warn the driver of a hazard so that the driver can take corrective action. The FCW system uses a forward-facing camera and / or RADAR sensor 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibrating component. The FCW system can provide an alert in the form of an acoustic, visual alert, vibration, and / or quick brake pulse.

[0185] The AEB system can detect an impending forward collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. The AEB system can use a forward-facing camera and / or RADAR sensor 1160 connected to a dedicated processor, DSP, FPGA, and / or ASIC. When the AEB system detects a danger, it usually first warns the driver to take corrective action to avoid the collision. If the driver does not take corrective action, the AEB system can automatically apply the brakes as part of an effort to prevent or at least mitigate the impact of the predicted collision. The AEB system may include techniques such as dynamic brake support and / or imminent collision braking.

[0186] The LDW system provides visual, audible, and / or tactile warnings, such as vibrations of the steering wheel or seat, to warn the driver when vehicle 1100 crosses a lane dividing line. The LDW system does not activate when the driver indicates an intentional lane departure by activating the turn indicator. The LDW system can use a front-facing camera connected to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically connected to driver feedback, such as a display, speaker, and / or vibrating component.

[0187] The LKA system is a modified form of the LDW system. The LKA system provides steering input or brakes to correct vehicle 1100 if vehicle 1100 begins to veer out of the lane.

[0188] The BSW system detects and warns the vehicle driver in the blind spots of the vehicle. The BSW system can provide visual, audible, and / or tactile warnings to indicate that a merge or lane change is not safe. The system can provide additional warnings when the driver uses the turn indicator. The BSW system can use a rear-facing camera and / or RADAR sensor coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.

[0189] The RCTW system can provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 1100 is backing up. Some RCTW systems include AEB to ensure that vehicle brakes are applied to avoid a collision. The RCTW system can use one or more rear-facing RADAR sensors 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically connected to driver feedback, such as a display, speaker, and / or vibrating component.

[0190] Conventional ADAS systems enable warning the driver and allowing the driver to determine whether a safety condition actually exists and act accordingly. Thus, conventional ADAS systems, while not usually catastrophic, tend to produce false judgment results that can trouble and distract the driver. However, in the autonomous vehicle 1100, when the results conflict, the vehicle 1100 itself must decide whether to accept the results from the primary computer or the secondary computer (e.g., the first controller 1136 or the second controller 1136). For example, in some embodiments, the ADAS system 1138 may be a backup and / or secondary computer for providing perception information to a backup computer rationality module. The backup computer rationality monitor can execute diverse software that is redundant in hardware components to detect malfunctions in perception and dynamic driving tasks. The output from the ADAS system 1138 can be provided to the supervisory MCU. When the outputs from the primary computer and the secondary computer conflict, the supervisory MCU needs to decide how to reconcile the conflict to ensure safe operation.

[0191] In some examples, the primary computer can be configured to provide a reliability score to the supervisory MCU that indicates the reliability of the primary computer in the selected result. If the reliability score exceeds a threshold, the supervisory MCU can follow the instructions of the primary computer regardless of whether the secondary computer gives conflicting or inconsistent results. If the reliability score does not meet the threshold and the primary and secondary computers indicate different (e.g., conflicting) results, the supervisory MCU can mediate between the computers to determine an appropriate result.

[0192] The supervisory MCU may be configured to execute a neural network trained and configured to determine, based on the outputs from the primary computer and the secondary computer, a state in which the secondary computer provides a false alarm. Thus, the neural network within the supervisory MCU can learn when the output of the secondary computer can be trusted and when it cannot. For example, when the secondary computer is a RADAR-based FCW system, the neural network within the supervisory MCU can learn when the FCW identifies metallic objects such as manhole covers or gratings in a drain that are not actually dangerous and trigger an alarm. Similarly, when the secondary computer is a camera-based LDW system, the neural network within the supervisory MCU can learn to ignore the LDW when a person on a bicycle or a pedestrian is present and lane departure is actually the safest operation. In an embodiment that includes a neural network running on the supervisory MCU, the supervisory MCU may include at least one of a DLA or a GPU suitable for executing the neural network with associated memory. In a preferred embodiment, the supervisory MCU may comprise components of the SoC1104 and / or may be included as components of the SoC1104.

[0193] In other examples, the ADAS system 1138 may include a secondary computer that performs ADAS functions using conventional rules of computer vision. As such, the secondary computer can use classical computer vision rules (if-then), and the presence of the neural network within the supervisory MCU can improve reliability, safety, and performance. For example, the various implementation forms and intentional non-identities make the overall system more fault-tolerant, especially against faults caused by software (or software-hardware interface) functions. For example, if there is a software bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have a greater confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.

[0194] In some examples, the output of the ADAS system 1138 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, if the ADAS system 1138 indicates a forward collision warning due to an object immediately ahead, the perception block can use this information when identifying the object. In other examples, the secondary computer may have its own neural network that is trained as described herein and thus reduces the risk of misjudgment.

[0195] Vehicle 1100 may further include an infotainment SoC 1130 (e.g., an in-vehicle infotainment system (IVI) within the vehicle). Although illustrated and described as an SoC, the infotainment system may not be an SoC and may include two or more individual components. The infotainment SoC 1130 may include a combination of hardware and software used to provide the vehicle 1100 with audio (e.g., music, mobile devices, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connections (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., vehicle-related information such as a navigation system, rear parking assistance, wireless data systems, fuel level, total mileage, brake fuel level, oil level, open / close doors, air filter information, etc.). For example, the infotainment SoC 1130 may include radio, disk player, navigation system, video player, USB and Bluetooth® connections, car computer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, heads-up display (HUD), HMI display 1134, telematics device, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. The infotainment SoC 1130 may be further used to provide information (e.g., visual and / or audible) to the user of the vehicle, such as information from the ADAS system 1138, autonomous driving information such as planned vehicle operations, trajectories, ambient environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.

[0196] The infotainment SoC 1130 may include GPU functionality. The infotainment SoC 1130 can communicate with other devices, systems, and / or components of the vehicle 1100 via a bus 1102 (e.g., CAN bus, Ethernet®, etc.). In some examples, the GPU of the infotainment system can execute some self-driving functions in the event that the primary controller 1136 (e.g., the primary and / or backup computer of the vehicle 1100) fails, and the infotainment SoC 1130 can be coupled to a supervisory MCU. In such examples, the infotainment SoC 1130 can put the vehicle 1100 into the chauffeur's safe stop mode as described herein.

[0197] The vehicle 1100 may further include an instrument cluster 1132 (e.g., digital dash, electronic instrument cluster, digital instrument panel, etc.). The instrument cluster 1132 may include a controller and / or a supercomputer (e.g., an individual controller or supercomputer). The instrument cluster 1132 may include a set of instruments such as a speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, gear shift position indicator, seat belt warning light, parking brake warning light, engine malfunction light, airbag (SRS) system information, lighting control device, safety system control device, navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1130 and the instrument cluster 1132. In other words, the instrument cluster 1132 may be included as part of the infotainment SoC 1130, and vice versa.

[0198] FIG. 11D is a system diagram of communication between a cloud-based server of FIG. 11A and an exemplary autonomous vehicle 1100 according to some embodiments of the present disclosure. System 1176 may include a server 1178, a network 1190, and vehicles including vehicle 1100. Server 1178 may include a plurality of GPUs 1184(A)-1184(H) (collectively referred to herein as GPU 1184), PCIe switches 1182(A)-1182(H) (collectively referred to herein as PCIe switch 1182), and / or CPUs 1180(A)-1180(B) (collectively referred to herein as CPU 1180). The GPUs 1184, CPUs 1180, and PCIe switches may be interconnected with each other by high-speed interconnections such as, but not limited to, an NVLink interface 1188 and / or a PCIe connection 1186 developed by NVIDIA, for example. In some examples, the GPUs 1184 are connected via NVLink and / or an NVSwitch SoC, and the GPUs 1184 and PCIe switches 1182 are connected via a PCIe interconnect. Eight GPUs 1184, two CPUs 1180, and two PCIe switches are illustrated, but this is not intended to be limiting. Depending on the embodiment, each server 1178 may include any number of GPUs 1184, CPUs 1180, and / or PCIe switches. For example, server 1178 may include eight, sixteen, thirty-two, and / or more GPUs 1184, respectively.

[0199] Server 1178 can receive, via network 1190, image data representing an image indicating an unexpected or changed road condition, such as recently started road construction, from a vehicle. Server 1178 can transmit, via network 1190, neural network 1192, updated neural network 1192, and / or map information 1194 including information on traffic and road conditions to the vehicle. The update of map information 1194 may include the update of HD map 1122, such as information on construction sites, potholes, detours, floods, and / or other obstacles. In some examples, neural network 1192, updated neural network 1192, and / or map information 1194 may result from new training and / or experiences represented in data received from any number of vehicles in the environment and / or based on training performed in a data center (e.g., using server 1178 and / or other servers).

[0200] Server 1178 can be used to train a machine learning model (e.g., a neural network) based on training data. The training data can be generated by a vehicle and / or (e.g., using a game engine) generated in a simulation. In some examples, the training data is tagged (e.g., when the neural network benefits from supervised learning) and / or undergoes other preprocessing, while in other examples, the training data is not tagged and / or preprocessed (e.g., when the neural network does not require supervised learning). The training can be performed according to any one or more machine learning techniques, including but not limited to, for example, the following classes: supervised training, semi-supervised training, unsupervised training, self-learning, reinforcement learning, associative learning, transfer learning, feature learning (including principal component and cluster analysis), multi-linear subspace learning, manifold learning, representation learning (including pre-dictionary learning), rule-based machine learning, anomaly detection, and their variants or combinations. After the machine learning model is trained, the machine learning model can be used by the vehicle (e.g., transmitted to the vehicle via network 1190), and / or the machine learning model can be used by server 1178 to remotely monitor the vehicle.

[0201] In some examples, server 1178 can receive data from a vehicle and apply the data to a latest real-time neural network for real-time intelligent inference. Server 1178 can include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1184, such as DGX and DGX Station machines developed by NVIDIA. However, in some examples, server 1178 can include a deep learning infrastructure that uses only CPU-powered data centers.

[0202] The deep learning infrastructure of server 1178 can have the ability of high-speed real-time inference, and use that ability to evaluate and verify the condition of the processors, software, and / or related hardware within vehicle 1100. For example, the deep learning infrastructure can receive periodic updates from vehicle 1100, such as sequences of images and / or objects in which vehicle 1100 is located within those sequences of images (e.g., via computer vision and / or other machine learning object classification techniques). The deep learning infrastructure can run its own neural network to identify objects and compare them with those identified by vehicle 1100. If the results do not match and the infrastructure concludes that the AI within vehicle 1100 is not functioning properly, server 1178 can send a signal to vehicle 1100 that infers control, notifies the passengers, and commands the fail-safe computer of vehicle 1100 to complete a safe parking operation.

[0203] For inference, server 1178 can include GPUs 1184 and one or more programmable inference acceleration devices (e.g., NVIDIA's TensorRT). The combination of GPU-powered servers and inference acceleration can enable real-time responsiveness. In other examples, such as when less performance is required, servers powered by CPUs, FPGAs, and other processors can be used for inference.

[0204] The present disclosure may be described in the general context of computer-executable instructions, such as program modules, being executed by a computer or other machine, such as a mobile information terminal or other handheld device. Generally, program modules include routines, programs, objects, components, data structures, etc., and refer to code that performs particular tasks or implements particular abstract data types. The present disclosure may be implemented in a variety of configurations, including handheld devices, household appliances, general-purpose computers, more specialized computing devices, etc. The present disclosure may also be implemented in a distributed computing environment where tasks are performed by remote processing devices linked through a communications network.

[0205] As used herein, a description of "and / or" with respect to two or more elements should be construed to mean either only one of the elements, or a combination of elements. For example, "element A, element B, and / or element C" may include only element A, only element B, only element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Additionally, "at least one of element A or element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further, "at least one of element A and element B" may include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.

[0206] The subject matter of this disclosure has been described with specificity to meet statutory requirements. However, the description itself is not intended to limit the scope of the disclosure. Rather, the inventors intend that the claimed subject matter may be otherwise implemented, including in combination with other current or future technologies, in different steps or combinations of steps similar to those described in this document. Further, the terms "step" and / or "block" may be used herein to imply different elements of a method being used, but these terms should not be construed as implying any particular order among the various steps disclosed herein except where the order of individual steps is explicitly recited and so described.

Claims

1. receiving an image of one or more objects having a plurality of backgrounds; generating a set of inference scores corresponding to one or more predictions of a prediction task performed on the one or more objects using the images; selecting a background based at least on one or more of the set of inference scores; generating an image based at least on integrating the object with the background based at least on the selection of the background; applying said images during training of at least one neural network to perform said prediction task; A method comprising:

2. The method of claim 1 , wherein the one or more of the inference scores comprises a plurality of inference scores for the group of images including the background, and the selection is based at least on an analysis of the plurality of inference scores.

3. 2. The method of claim 1, wherein the reasoning score is generated using the at least one neural network in a first epoch of training the at least one neural network, and the use of the image is during a second epoch of the training.

4. The method of claim 1 , wherein the generating the image comprises generating the mask based at least on identifying a region of the object in the image and applying a mask to the image.

5. The method of claim 1 , wherein the selection of the background is based at least on determining that at least one reasoning score of the one or more reasoning scores is below a threshold.

6. 2. The method of claim 1 , wherein the one or more of the inference scores correspond to a first cropped region of the background and the integration of the object is with a second cropped region of the background that is different from the first cropped region.

7. The method of claim 1 , wherein the background is a background type, the method further comprising synthetically generating the background based at least on the background type.

8. one or more processors; and one or more memory devices that store instructions, which when executed by the one or more processors: obtaining at least one neural network trained to perform a prediction task on the image using inputs generated from the mask corresponding to the object in the image based at least on a modification of a background of the object using the mask; generating a mask corresponding to an object in an image, the object having a background in the image; generating an input to the at least one neural network using the mask, the input capturing the object having at least a portion of the background; and generating at least one prediction for the prediction task based at least on application of the input to the at least one neural network. The system further comprises:

9. The system of claim 8 , wherein the inputs include a first input representing at least a portion of the image including the object having at least the portion of the background, and a second input representing at least a portion of the mask.

10. The system of claim 8 , wherein the generating of the input comprises generating image data based at least on modifying the background using the mask and at least a portion of the input that corresponds to the image data.

11. 9. The system of claim 8, wherein the generating is based at least on fusing a first set of inference scores and a second set of inference scores, the first set of inference scores corresponding to a first portion of the input representing at least a portion of the image including the object having the at least the portion of the background, and the second set of inference scores corresponding to a second portion of the input representing at least a portion of the mask.

12. The operation, Control systems for autonomous or semi-autonomous machines; Autonomous or semi-autonomous machine perception systems; A system for performing a simulation operation; A system for performing deep learning operations, A system implemented using edge devices; A system implemented using a robot, A system incorporating one or more virtual machines (VMs); A system implemented at least in part in a data center; or System implemented at least in part using cloud computing resources - Patents.com The system of claim 8 , wherein the system is executed by at least one of:

Citation Information

Patent Citations

  • Three-dimensional object recognition device

    JP2005346297A

  • Training a neural network using augmented training datasets

    US20190130218A1

  • Training one-shot instance segmenters using synthesized images

    WO2020042004A1