Training and inference using a neural network to predict an orientation of an object in an image
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2026-03-30
AI Technical Summary
Training neural networks for predicting object orientations in images requires significant memory, time, and computing resources, and is made more challenging when ground truth annotations are unavailable or difficult to obtain.
Training neural networks in a self-supervised manner using loss functions that evaluate image characteristics without ground truth annotations, employing techniques such as generative adversarial networks and variational autoencoders to generate synthetic images for orientation prediction.
Reduces the resource requirements and improves the efficiency of training neural networks for object orientation prediction by leveraging self-supervised learning methods, enabling accurate orientation inference without relying on ground truth data.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical Field]
[0001] This application claims priority to U.S. Patent Application No. 16 / 690,015, filed November 20, 2019, entitled "TRAINING AND INFERENCING USING A NEURAL NETWORK TO PREDICT ORIENTATIONS OF OBJECTS IN IMAGES," the entire contents of which are incorporated herein by reference in their entirety for all purposes.
[0002] At least one embodiment relates to processing resources used to train a neural network for predicting viewpoints of objects in an image. For example, at least one embodiment relates to a processor or computing system used to train a neural network according to various novel techniques described herein. [Background technology]
[0003] Training a neural network can use significant memory, time, or computing resources. Training a neural network that requires ground truth annotations can be more difficult than training a neural network that does not require ground truth annotating of some or all of the training data, at least because ground truth annotations may not always be available and / or may be difficult to obtain. It is possible to improve the amount of memory, time, and / or computing resources used to train a neural network. [Brief explanation of the drawings]
[0004] [Figure 1]FIG. 1 illustrates predicting viewpoints of an object using a neural network trained in a self-supervised manner, according to at least one embodiment. [Figure 2] FIG. 1 illustrates a loss function, according to at least one embodiment. [Figure 3] FIG. 1 illustrates a generative adversarial network, according to at least one embodiment. [Figure 4] FIG. 1 illustrates updating a classifier, according to at least one embodiment. [Figure 5] FIG. 1 illustrates updating a classifier, according to at least one embodiment. [Figure 6] FIG. 10 depicts generator updates, according to at least one embodiment. [Figure 7] FIG. 10 depicts generator updates, according to at least one embodiment. [Figure 8] FIG. 10 depicts symmetry loss, according to at least one embodiment. [Figure 9] FIG. 10 depicts nearest neighbor and farthest neighbor loss, according to at least one embodiment. [Figure 10] FIG. 10 depicts disentanglement loss, according to at least one embodiment. [Figure 11] FIG. 10 illustrates calibration of a neural network, according to at least one embodiment. [Figure 12] FIG. 1 illustrates inference, according to at least one embodiment. [Figure 13] FIG. 1 illustrates an exemplary illustration of a process for training a neural network to predict viewpoints of objects in an image, according to at least one embodiment. [Figure 14] FIG. 10 is a diagram of an illustrative example of a process for training a neural network to predict viewpoints of objects in an image, according to at least one embodiment. [Figure 15A] FIG. 1 illustrates an exemplary illustration of a process for calculating a generation consistency loss, according to at least one embodiment. [Figure 15B]FIG. 1 illustrates an exemplary illustration of a process for computing viewpoint consistency loss, according to at least one embodiment. [Figure 16A] FIG. 1 illustrates an exemplary illustration of a process for calculating symmetry loss, according to at least one embodiment. [Figure 16B] FIG. 1 illustrates an exemplary illustration of a process for calculating symmetry loss, according to at least one embodiment. [Figure 17] FIG. 1 illustrates an exemplary illustration of a process for calculating nearest neighbor and farthest neighbor losses, according to at least one embodiment. [Figure 18A] FIG. 1 illustrates inference and / or training logic, according to at least one embodiment. [Figure 18B] FIG. 1 illustrates inference and / or training logic, according to at least one embodiment. [Figure 19] FIG. 1 illustrates training and deployment of a neural network, according to at least one embodiment. [Figure 20] FIG. 1 illustrates an exemplary data center system, according to at least one embodiment. [Figure 21A] FIG. 1 illustrates an example of an autonomous vehicle, according to at least one embodiment. [Figure 21B] FIG. 21B illustrates example camera locations and fields of view for the autonomous vehicle of FIG. 21A, according to at least one embodiment. [Figure 21C] FIG. 21B is a block diagram illustrating an example system architecture of the autonomous vehicle of FIG. 21A, according to at least one embodiment. [Figure 21D] FIG. 21B illustrates a system for communication between a cloud-based server and the autonomous vehicle of FIG. 21A, according to at least one embodiment. [Figure 22] FIG. 1 is a block diagram illustrating a computer system according to at least one embodiment. [Figure 23] FIG. 1 is a block diagram illustrating a computer system according to at least one embodiment. [Figure 24] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 25] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 26A] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 26B] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 26C] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 26D] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 26E] FIG. 1 illustrates a shared programming model, according to at least one embodiment. [Figure 26F] FIG. 1 illustrates a shared programming model, according to at least one embodiment. [Figure 27] FIG. 1 illustrates an exemplary integrated circuit and associated graphics processor, according to at least one embodiment. [Figure 28A] FIG. 1 illustrates an exemplary integrated circuit and associated graphics processor, according to at least one embodiment. [Figure 28B] FIG. 1 illustrates an exemplary integrated circuit and associated graphics processor, according to at least one embodiment. [Figure 29A] FIG. 10 illustrates additional exemplary graphics processor logic, according to at least one embodiment. [Figure 29B] FIG. 10 illustrates additional exemplary graphics processor logic, according to at least one embodiment. [Figure 30] FIG. 1 illustrates a computer system according to at least one embodiment. [Figure 31A] FIG. 1 illustrates a parallel processor, according to at least one embodiment. [Figure 31B] FIG. 1 illustrates a partition unit, according to at least one embodiment. [Figure 31C] FIG. 1 illustrates a processing cluster, according to at least one embodiment. [Figure 31D]FIG. 1 illustrates a graphics multiprocessor according to at least one embodiment. [Figure 32] FIG. 1 illustrates a multi-graphics processing unit (GPU) system according to at least one embodiment. [Figure 33] FIG. 1 illustrates a graphics processor according to at least one embodiment. [Figure 34] FIG. 1 is a block diagram illustrating a processor micro-architecture for a processor, according to at least one embodiment. [Figure 35] FIG. 1 illustrates a deep learning application processor, according to at least one embodiment. [Figure 36] FIG. 1 is a block diagram illustrating an exemplary neuromorphic processor, according to at least one embodiment. [Figure 37] FIG. 1 illustrates at least a portion of a graphics processor according to one or more embodiments. [Figure 38] FIG. 1 illustrates at least a portion of a graphics processor according to one or more embodiments. [Figure 39] FIG. 1 illustrates at least a portion of a graphics processor according to one or more embodiments. [Figure 40] FIG. 4 is a block diagram of a graphics processing engine 4010 of a graphics processor, according to at least one embodiment. [Figure 41] FIG. 1 is a block diagram of at least a portion of a graphics processor core, according to at least one embodiment. [Figure 42A] FIG. 42 illustrates thread execution logic 4200 including an array of processing elements of a graphics processor core, according to at least one embodiment. [Figure 42B] FIG. 42 illustrates thread execution logic 4200 including an array of processing elements of a graphics processor core, according to at least one embodiment. [Figure 43]FIG. 1 illustrates a parallel processing unit ("PPU"), according to at least one embodiment. [Figure 44] FIG. 1 illustrates a general processing cluster (“GPC”), according to at least one embodiment. [Figure 45] FIG. 1 illustrates a memory partition unit of a parallel processing unit (“PPU”), according to at least one embodiment. [Figure 46] FIG. 1 illustrates a streaming multiprocessor, according to at least one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0005] In at least one embodiment, a neural network is trained to identify the orientation of objects in an image in a self-supervised manner on a set of images as described elsewhere in this disclosure. In at least one embodiment, the neural network is trained to identify the orientation of objects in an image in a self-supervised manner by, at least as part of the training, computing one or more loss functions that evaluate one or more characteristics of the images in the training set (e.g., the set of images). In at least one embodiment, the neural network is trained on a set of images that does not have ground truth annotations or for which ground truth annotations are not available (e.g., such data is not provided to the neural network during training). In at least one embodiment, the neural network is trained to generate a second image having a predicted orientation from an object in a first image having the same orientation. In at least one embodiment, the predicted orientation or viewpoint is encoded as azimuth, elevation, and tilt parameters.
[0006] In at least one embodiment, one or more neural networks are trained in a self-supervised manner on a set of images of different objects of the same category as the object in the image to be inferred. In at least one embodiment, the different objects of the same category may refer to different images, which may be one or more images of a first car in one or more orientations, one or more images of a second car in one or more different orientations, etc. In at least one embodiment, the images of the object to be inferred are included in a set of images used to train the one or more neural networks for inferring orientation. In at least one embodiment, the one or more neural networks are trained in a self-supervised manner by using at least a set of loss functions to evaluate one or more properties of the object in the image. In at least one embodiment, the one or more properties of the object refer to properties of the object that can be used to infer orientation. In at least one embodiment, the neural network is trained in a self-supervised manner to generate synthetic images of the object in a particular orientation, which may be the same as the expected orientation of the input image. In at least one embodiment, the synthetic image is created using or via a deep generative model such as a variational autoencoder (VAE), a differentiable renderer, or a generative adversarial network (GAN). In at least one embodiment, the object whose orientation is inferred may be a vehicle, an aircraft, a drone, a human, a face (e.g., human or animal), etc.
[0007] In at least one embodiment, self-supervised learning (e.g., training) refers to a form of learning in which a neural network is trained on a training set, where the training set data does not include ground truth annotations, but the training set data is partially labeled (e.g., semi-supervised learning). In at least one embodiment, a neural network trained in a self-supervised manner to identify object orientations in images utilizes a training set of images for training, where the images in the training set do not include ground truth annotations indicating the orientations of objects in the images, but do include labels or other information identifying various objects in the images (e.g., one of the images includes labels or other information identifying objects in the image, but does not include annotations indicating the orientation of the object).
[0008] In at least one embodiment, semi-supervised learning refers to a form of learning in which a neural network is trained on a training set, where only a portion of the data in the training set includes ground truth annotations. In at least one embodiment, fully supervised learning refers to a form of learning in which a neural network is trained on a training set, where all of the data in the training set includes ground truth annotations. In at least one embodiment, unsupervised learning refers to a form of learning in which a neural network is trained on a training set, where none of the data in the training set includes ground truth annotations.
[0009] FIG. 1 illustrates a diagram 100 illustrating predicting viewpoints of objects using neural networks trained in a self-supervised manner, according to at least one embodiment. In at least one embodiment, diagram 100 is implemented by one or more systems, such as those described in FIGS. 18-46. In at least one embodiment, diagram 100 includes one or more neural networks associated with a classifier 106 that are trained using self-supervised learning on a set of images of a certain category to infer viewpoints of objects in other images of that category. In at least one embodiment, images are provided as input to the neural network to detect the orientation of objects of a certain category. In at least one embodiment, input images are provided to multiple neural networks trained using the self-supervised learning techniques described herein to identify the orientations or viewpoints of different objects in the input images.
[0010] In at least one embodiment, the viewpoint of an image refers to the orientation of an object in an image, which refers to the three-dimensional orientation of an object captured in a two-dimensional image. In at least one embodiment, a camera is used to capture a two-dimensional image of a real-world object, such as a car, at a particular orientation relative to the camera. In at least one embodiment, the orientation of an object (e.g., viewpoint) is encoded with respect to a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the orientation of an object in an image is encoded as a set of three vectors that define the object's direction relative to the canonical x-, y-, and z-axes.
[0011] In at least one embodiment, an image set 102 is acquired. In at least one embodiment, image set 102 is a collection of one or more images of a certain type of object. In at least one embodiment, image set 102 is used to train one or more neural networks to identify the orientation of objects in images. In at least one embodiment, image set 102 is categorized or labeled as each displaying the same type or category of object. In at least one embodiment, image set 102 is a collection of images of cars, which may include different types of cars in different orientations, in different weather conditions, under different lighting, etc. In at least one embodiment, image set 102 includes images of the same car or the same type of car in different orientations. In at least one embodiment, at least some of image set 102 do not have ground truth annotations specifying the orientation of objects in such training images. In at least one embodiment, all images in image set 102 do not have ground truth annotations specifying the azimuth, elevation, and tilt of objects in the images in the set. In at least one embodiment, ground truth annotations refer to annotations that images in a training set of images may contain when one or more neural networks are configured to determine one or more characteristics of the images, where the one or more neural networks are trained on the training set of images, and that indicate one or more expected characteristics of the images. In at least one embodiment, image set 102 includes one or more synthetic images, such as images created from a variational autoencoder (VAE), a generative adversarial network (GAN), or a renderer. In at least one embodiment, all images in image set 102 are real images, as opposed to images synthesized or created from a generative model, such as a variational autoencoder (VAE), a renderer, or a generative adversarial network. In at least one embodiment, image set 102 is collected and aggregated from a website that sorts images by category.
[0012] In at least one embodiment, the classifier 106 is trained to identify the orientation of an object in the image 104 based, at least in part, on one or more characteristics of the object other than the object's orientation. In at least one embodiment, the classifier 106 is a classifier within one or more neural networks. In at least one embodiment, the classifier 106 is a component of one or more neural networks, including other neural networks, classifiers, and various other machine learning components. In at least one embodiment, the classifier 106 is a generative adversarial network (AGN) classification network. In at least one embodiment, the classifier 106 is part of one or more neural networks and is trained to infer a viewpoint and a set of appearance attributes from input images. In at least one embodiment, the classifier 106 is trained on a set of images of a category (e.g., cars) to infer the orientation of other objects of the same category captured in other images. In at least one embodiment, the classifier 106 is trained on the image set 102 in a self-supervised manner. In at least one embodiment, the classifier 106 is trained to identify object orientations in the images 104 in a self-supervised manner by computing, at least as part of its training, one or more loss functions that evaluate one or more characteristics of a training set of images (e.g., image set 102). In at least one embodiment, a neural network associated with the classifier 106 is trained, at least in part, based on computing a generative consistency loss, a symmetry loss, a nearest neighbor and a farthest neighbor loss, and a disentanglement loss. In at least one embodiment, a neural network for identifying object orientations may be trained according to techniques described in conjunction with FIGS. 2-10. In at least one embodiment, the classifier 106 is trained on a set of images that does not have ground truth annotations or for which ground truth annotations are not available (e.g., such data is not provided during training).
[0013] In at least one embodiment, images 104 are acquired for use with classifier 106. In at least one embodiment, objects in images 104 are of the same type as objects in images in image set 102 used to train one or more neural networks. In at least one embodiment, images 104 are provided to a neural network for inference to predict orientation. In at least one embodiment, a first system trains one or more neural networks, and a second, different system uses these one or more neural networks to perform inference and identify the orientation of objects in the images. In at least one embodiment, classifier 106 is trained in a self-supervised manner on image set 102 of objects of a particular category to infer the orientation of other objects in that category (e.g., objects in image 104). In at least one embodiment, one or more neural networks associated with classifier 106 are trained on a set of images of cars and used to infer the orientation of cars captured in real time by a camera or other suitable video / image capture device mounted on the vehicle. In at least one embodiment, the classifier 106 is trained in a self-supervised manner on the image set 102 to determine the orientation 108 of an object depicted in the image 104. In at least one embodiment, the classifier 106 determines the orientation 108 of a car depicted in the image 104.
[0014] FIG. 2 shows a diagram 200 illustrating a loss function, according to at least one embodiment. In at least one embodiment, diagram 200 is implemented by one or more systems, such as the systems described in FIGS. 18-46. In at least one embodiment, classifier 204 is associated with one or more neural networks and trained using at least one of real image generation consistency loss 208, nearest & farthest neighbor loss 210, symmetry loss 212, and real / fake classification loss 214. In at least one embodiment, classifier 204 is part of one or more neural networks trained to infer viewpoints from input images, the one or more neural networks including various parameters associated with one or more processes of the one or more neural networks, and updated at least in part based on real image generation consistency loss 208, nearest & farthest neighbor loss 210, and symmetry loss 212.
[0015] In at least one embodiment, a set of object images 202 of a certain type of object is obtained for the classifier 204. In at least one embodiment, the set of object images 202 includes images that all include the same type of object. In at least one embodiment, the set of object images 202 includes images that include cars in different orientations, in different weather conditions, under different lighting, etc. In at least one embodiment, the system obtains the set of object images 202 according to techniques described elsewhere in this disclosure, such as in FIG. 13 .
[0016] In at least one embodiment, images from the object image set 202 are selected as input images to the classifier 204. In at least one embodiment, the images from the set may be selected for learning in any suitable manner and may be randomly or pseudo-randomly sampled from a training set. In at least one embodiment, the classifier 204 predicts a viewpoint 206 for the input image. In at least one embodiment, the viewpoint 206 for the input image is inferred by the classifier 204 through one or more processes involving one or more neural networks, the one or more neural networks including one or more input parameters that define the one or more processes involving the one or more neural networks. In at least one embodiment, the viewpoint 206 is determined based on ground truth annotations provided as part of training on at least a portion of the object image set 202. In at least one embodiment, the viewpoint 206 corresponds to a prediction of the orientation of the object in the image input to the classifier 204. In at least one embodiment, the orientation (eg, viewpoint) of an object is encoded with respect to a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter.
[0017] In at least one embodiment, a generative consistency loss 208 is computed for the classifier 204. In at least one embodiment, the generative consistency loss 208 is computed, at least in part, based on an image consistency loss, which compares a selected image to an image generated by a deep generative model, and a viewpoint consistency loss, which compares an input viewpoint to a generative model value predicted by the generative model and the classifier. In at least one embodiment, the generative consistency loss 208 is computed using techniques described elsewhere in this disclosure, such as the techniques discussed in conjunction with FIGS. 4-7. In at least one embodiment, the generative consistency loss 208 includes at least two components: a synthetic image viewpoint consistency loss and a real image consistency loss. In at least one embodiment, the viewpoint consistency loss can be referred to as an orientation consistency loss. In at least one embodiment, the generative consistency loss is applied to real images (e.g., images from the object image set 202), as opposed to synthetic images created by a generator. In at least one embodiment, the viewpoint consistency loss and the image consistency loss are utilized to determine the generative consistency loss. In at least one embodiment, the generative consistency loss is a combination of viewpoint consistency loss and image consistency loss. In at least one embodiment, the generative consistency loss is determined by the following symbolic formula: L gc =L vc +L ic However, L gc corresponds to the production consistency loss, and L vc corresponds to the viewpoint consistency loss, and L ic corresponds to the image consistency loss.
[0018] In at least one embodiment, the image consistency loss is calculated based at least in part on images from the object image set 202 that are input to the classifier 204, which determines at least two properties from the input images: a viewpoint 206 and a set of appearance parameters. In at least one embodiment, the viewpoint 206 and the set of appearance parameters are provided to a generator to create a synthetic image. In at least one embodiment, a generative adversarial network (GAN) receives the viewpoint 206 and the set of appearance parameters and generates a synthetic (e.g., fake) image according to the viewpoint 206 and the set of appearance parameters. In at least one embodiment, the synthetic image and the input image are compared to determine the image consistency loss. In at least one embodiment, the cosine distance between the input image and the synthetic image is compared to determine feature similarity, with closer similarity resulting in a lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss between two images.
[0019] In at least one embodiment, the viewpoint consistency loss is calculated based at least in part on a viewpoint (e.g., viewpoint 206) of the input image. In at least one embodiment, a generator is used to create a synthetic image from viewpoint 206 predicted from the input image by classifier 204. In at least one embodiment, the synthetic image generated from viewpoint 206 is provided to classifier 204, which determines a second viewpoint for the synthetic image. In at least one embodiment, viewpoint 206 is compared to a second viewpoint of a synthetic image generated at least in part based on viewpoint 206. In at least one embodiment, the distance between viewpoint 206 and the second viewpoint of the synthetic image is used to calculate a viewpoint consistency loss, with closer viewpoints resulting in a lower loss. In at least one embodiment, the generative consistency loss is calculated according to the technique described in conjunction with FIG. 15.
[0020] In at least one embodiment, the real / fake classification loss 214 is calculated based on whether the classifier 204 can accurately predict whether an input image to the classifier 204 is real or synthetic. In at least one embodiment, the real / fake classification loss 214 is calculated based on whether the classifier 204 can accurately predict whether a set of input images is real or fake, where the classifier 204 can be provided with either real or fake (e.g., synthetic) images and predicts whether the images are real or fake. As part of training the classifier 204, ground truth about whether the images provided to the classifier 204 are real or fake is available as part of the training (e.g., for calculating the loss).
[0021] In at least one embodiment, the symmetry loss 212 is calculated by at least comparing an input image to a transformed version of the input image. In at least one embodiment, the input image is selected from the object image set 202. In at least one embodiment, a transformation is applied to the input image to generate the transformed image. In at least one embodiment, the input image is horizontally flipped to generate the transformed image. In at least one embodiment, the classifier 204 is used to predict a viewpoint 206 of the input image and a second viewpoint of the transformed image. In at least one embodiment, the viewpoint 206 is predicted for an input image, and the second viewpoint is predicted for a horizontally flipped version of the input image. In at least one embodiment, the loss is calculated based on whether certain properties hold true. In at least one embodiment, a transformation or its inverse is applied to the predicted viewpoint of the transformed version of an input image. In at least one embodiment, if the input image is rotated by (φ, θ, ψ) angles to generate a transformed image, the inferred viewpoint of the transformed image may be rotated back by (-φ, -θ, -ψ) angles. In at least one embodiment, the loss is calculated by comparing the azimuth, elevation, and tilt magnitudes of a viewpoint 206 of an input image with a second viewpoint of the transformed image, where equal magnitudes of the respective orientation parameters result in zero loss. In at least one embodiment, the symmetry loss is calculated according to techniques described elsewhere in this disclosure, such as those discussed in conjunction with Figures 8 and 16. In at least one embodiment, the loss is calculated by determining how closely a first set of appearance parameters predicted by the classifier 204 for an image matches a second set of appearance parameters predicted by the classifier 204 for a transformed version of that image.
[0022] In at least one embodiment, nearest and farthest neighbor losses 210 are calculated by comparing an input image of the object image set 202 to its nearest and farthest neighbors based, at least in part, on a viewpoint graph of the object image set 202. In at least one embodiment, nearest and farthest neighbor losses 210 are calculated according to techniques described in conjunction with FIGS. 4, 9, and 17. In at least one embodiment, the object image set 202 is used to generate a viewpoint graph, where nodes of such a graph correspond to images and edges correspond to their viewpoint covariance distances (e.g., cosine distances). In at least one embodiment, the cosine distances are calculated based on feature similarities between pairs of images using a convolutional neural network (CNN). In at least one embodiment, an anchor image is selected from the object image set 202. In at least one embodiment, the anchor image is located from the viewpoint graph, and the nearest and farthest neighbors are selected based on edge weights. In at least one embodiment, the nearest neighbor has the shortest edge leading to the anchor image. In at least one embodiment, the farthest neighbor has the farthest edge leading to the anchor image. In at least one embodiment, the classifier 204 predicts a first viewpoint for the anchor image (e.g., the predicted viewpoint 206 for the anchor image) and a second viewpoint for the nearest neighbor image (e.g., the predicted viewpoint 206 for the nearest neighbor image), with a loss calculated such that closer distances between these viewpoints correspond to lower losses. In at least one embodiment, the neural network of the classifier 204 predicts a first viewpoint for the anchor image and a third viewpoint for the farthest neighbor image (e.g., the viewpoint 206 for the farthest neighbor image), with a loss calculated such that greater distances between these viewpoints correspond to lower losses.
[0023] In at least one embodiment, the calculated losses (e.g., generative consistency loss 208, nearest and farthest neighbor loss 210, symmetry loss 212, and real / fake classification loss 214) are utilized to update parameters of one or more neural networks associated with classifier 204 that are trained on object image set 202. In at least one embodiment, a system implementing diagram 200 includes executable code for continually updating parameters of one or more neural networks associated with classifier 204, such that the one or more neural networks and classifier 204 are trained to infer viewpoints and other characteristics of input images. In at least one embodiment, training is performed according to any suitable technique and may include selecting and utilizing various additional images of object image set 202 to calculate losses and refine parameters of one or more neural networks trained to infer viewpoints. In at least one embodiment, once training is complete, the trained neural network is made available for inference (e.g., the neural network or its parameters are transferred to a different system).
[0024] FIG. 3 shows a diagram 300 illustrating a generative adversarial network, according to at least one embodiment. In at least one embodiment, diagram 300 is implemented by one or more systems, such as the systems described in FIGS. 18-46. In at least one embodiment, diagram 300 includes a generator 306 that utilizes an input viewpoint 302 and an input set of appearance parameters 304, a synthesized image 308. In at least one embodiment, diagram 300 shows a classifier 310 that utilizes an input image 318 to output an output viewpoint 312, an output decision 314, and an output set of appearance parameters 316. In at least one embodiment, parameters for generator 306 and / or classifier 310 are selected using techniques described in conjunction with FIGS. 4-7.
[0025] In at least one embodiment, the input viewpoint 302 corresponds to the orientation of an object in an image and refers to the three-dimensional orientation of the object captured in the two-dimensional image. In at least one embodiment, the orientation of the object (e.g., viewpoint) is encoded with respect to a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the input viewpoint 302 corresponds to a particular orientation of the object and includes particular values of the set of parameters including the azimuth parameter, the elevation parameter, and the tilt parameter. In at least one embodiment, the input viewpoint 302 indicates a 3D rotation of the object (e.g., the input viewpoint 302 can specify a rotation of the object by a specified number of degrees about a specified axis and its transformation). In at least one embodiment, the input set of appearance parameters 304 are parameters that define the appearance of the object. In at least one embodiment, the object includes a vehicle, an aircraft, a drone, a human, a face (e.g., a human or animal), etc. In at least one embodiment, the input set of appearance parameters 304 correspond to the appearance parameters of a vehicle, such as color, size, wheel type, and various other parameters that define the appearance of the vehicle.
[0026] In at least one embodiment, an input viewpoint 302 and an input set of appearance parameters 304 are provided to a generator 306 to create an image 308. In at least one embodiment, the generator 306 and the classifier 310 are part of a generative adversarial network (GAN). In at least one embodiment, the generator 306 is a generative network within a generative adversarial network. In at least one embodiment, the generator 306 is part of one or more neural networks and is trained to generate images based on the input viewpoint and the input set of appearance parameters. In at least one embodiment, the generator 306 receives the input viewpoint 302 and the input set of appearance parameters 304 and generates an image 308, which is a synthetic (e.g., fake) image according to the input viewpoint 302 and the input set of appearance parameters 304. In at least one embodiment, the generator accepts two separate (e.g., independent) parameters used to create the image 308: an input viewpoint 302 (e.g., encoded azimuth, elevation, and tilt parameters) that indicates a particular viewpoint used in generating the image 308, and appearance parameters 304 that encode appearance properties of the image 308 (e.g., in the case of a car, such properties may include color, make, model, year of manufacture, etc.). In at least one embodiment, the generator 306 generates the image 308, which includes an object generated according to the input set of appearance parameters 304 and oriented according to the input viewpoint 302. In at least one embodiment, the image 308 is a synthetic image including a car, the appearance of the car corresponding to the input set of appearance parameters 304 and the orientation of the car corresponding to the input viewpoint 302.
[0027] In at least one embodiment, the generative adversarial network (GAN) includes a classifier 310. In at least one embodiment, the classifier 310 accepts an input image 318 and generates an output viewpoint 312, an output decision 314, and an output set of appearance parameters 316. In at least one embodiment, the input image 318 can be a real image or a synthetic image. In at least one embodiment, the input image 318 is retrieved from one or more other sources, such as an image database, one or more cameras, and / or variations thereof. In at least one embodiment, the classifier 310 processes the image 308. In at least one embodiment, the classifier 310 includes various neural networks and machine learning processes. In at least one embodiment, the classifier 310 is implemented according to classifiers described elsewhere in this disclosure, such as the classifier discussed in conjunction with FIG. 2. In at least one embodiment, the classifier 310 is associated with one or more neural networks that are trained to infer viewpoints and other characteristics of the input image. In at least one embodiment, the classifier 310 is improved through various processes that involve the calculation of various loss functions, which are used to update various parameters associated with the classifier 310.
[0028] In at least one embodiment, the classifier 310 receives an input image 318 and generates an output viewpoint 312, an output decision 314, and an output set of appearance parameters 316. In at least one embodiment, the output viewpoint 312 is a predicted viewpoint of the input image 318 generated by one or more processes of the classifier 310. In at least one embodiment, the output decision 314 is a decision generated by one or more processes of the classifier 310 indicating whether the input image 318 is a real image or a synthetic (e.g., fake) image. In at least one embodiment, the decision 314 is a binary output (e.g., a TRUE / FALSE indicator as to whether the classifier 310 believes the input image 318 is a real image or a synthetic image). In at least one embodiment, decision 314 is a number between 0 and 1 (inclusive or exclusive of one or both endpoints) that encodes a confidence value for whether classifier 310 considers input image 318 to be real versus fake (e.g., 0.5 indicates the image is equally likely to be real or fake, while 0 indicates the image is highly likely to be fake). In at least one embodiment, output set of appearance parameters 316 is a predicted set of appearance parameters for input image 318 generated by one or more processes of classifier 310.
[0029] In at least one embodiment, if the classifier 310 is accurately calibrated (e.g., the classifier 310 is trained to a desired degree of accuracy or to a desired degree of acceptable loss) and the image 308 is generated by the generator 306, the output decision 314 indicates that the image 308 is fake and that the output viewpoint 312 and output set of appearance parameters 316 are identical to the input viewpoint 302 and input set of appearance parameters 304, respectively. In at least one embodiment, if the classifier 310 is not accurately calibrated (e.g., the classifier 310 is not fully trained to a desired degree of accuracy or to a desired degree of acceptable loss) and the image 308 is generated by the generator 306, the output decision 314 indicates an incorrect decision (e.g., if the image 308 is synthetic, the output decision 314 indicates that the image 308 is real) and that the output viewpoint 312 and output set of appearance parameters 316 are different from the input viewpoint 302 and input set of appearance parameters 304, respectively. In at least one embodiment, a comparison of the output viewpoint 312 and output set of appearance parameters 316 with the input viewpoint 302 and input set of appearance parameters 304 is used to evaluate and further process, train, and / or calibrate the classifier 310 and generator 306, respectively.
[0030] FIG. 4 shows a diagram 400 illustrating classifier updating, according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update parameters of a classifier 404, which is used to predict various outputs from an input image. In at least one embodiment, FIG. 4 shows an input image 402, a classifier 404, a predicted viewpoint 406, a predicted decision 408 as to whether the input image 402 is real or fake, a set of appearance parameters 410, a generator 412, a generated image 414, a real / fake classification loss 416, an image consistency loss 418, nearest and furthest neighbor losses 420, and a symmetry loss 422. In at least one embodiment, FIG. 4 illustrates classifier updating using real images (e.g., images not synthesized by a generator). In at least one embodiment, the techniques described in conjunction with FIG. 4 are coextensive with the techniques described in conjunction with FIGS. 5-7 for training a generator and / or classifier.
[0031] In at least one embodiment, a classifier 404 processes an input image 402. In at least one embodiment, the classifier 404 is associated with one or more neural networks that are trained to infer viewpoints and other characteristics of the input image. In at least one embodiment, the classifier 404 receives the input image 402 and generates a predicted viewpoint 406, a decision 408 as to whether the input image 402 is real or fake, and a set of appearance parameters 410. In at least one embodiment, the viewpoint 406 is a predicted viewpoint of the input image 402 determined by one or more processes of the classifier 404. In at least one embodiment, the viewpoint 406 corresponds to a particular predicted orientation of an object in the image and includes particular values for a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the viewpoint 406 indicates a 3D rotation of the object (e.g., the viewpoint 406 can specify a rotation of the object by a specified number of degrees about a specified axis and its transformation). In at least one embodiment, viewpoint 406 includes a predicted orientation of a vehicle depicted in input image 402. In at least one embodiment, decision 408 is a decision of whether input image 402 is real or fake. In at least one embodiment, a fake image refers to a synthetic image created by a generative adversarial network. In at least one embodiment, decision 408 is a binary value (e.g., a TRUE / FALSE value indicating a prediction of whether input image 402 is real or fake). In at least one embodiment, decision 408 is a non-binary value indicating a confidence level as to whether input image 402 is real or fake. In at least one embodiment, set of appearance parameters 410 is a predicted set of appearance parameters for input image 402 generated by one or more processes of classifier 404. In at least one embodiment, set of appearance parameters 410 are predicted parameters that define the appearance of an object depicted in input image 402.In at least one embodiment, the set of appearance parameters 410 corresponds to a set of predicted appearance parameters for a car depicted in the input image 402, such as predicted color, size, wheel type, and various other parameters that define the appearance of the car depicted in the input image 402.
[0032] In at least one embodiment, viewpoint 406 and set of appearance parameters 410 are provided to generator 412 to generate generated image 414. In at least one embodiment, generator 412 creates a synthetic image. In at least one embodiment, generator 412 is part of a generative adversarial network (GAN). In at least one embodiment, generator 412 receives viewpoint 406 and set of appearance parameters 410 and generates generated image 414, which is a synthetic (e.g., fake) image according to viewpoint 406 and set of appearance parameters 410. In at least one embodiment, generator 412 generates generated image 414, which includes an object generated according to set of appearance parameters 410 and oriented according to viewpoint 406. In at least one embodiment, generated image 414 is a synthetic image including a car, the appearance of the car corresponding to set of appearance parameters 410 and the orientation of the car corresponding to viewpoint 406.
[0033] In at least one embodiment, decision 408 is used to determine a classification loss, such as real / fake classification loss 416. In at least one embodiment, real / fake classification loss 416 is calculated based on whether classifier 404 can accurately predict whether an input image to classifier 404 is a real image or a synthetic image. In at least one embodiment, real / fake classification loss 416 is calculated based on whether classifier 404 can accurately predict whether a set of input images is real or fake, where classifier 404 can be provided with either real or fake (e.g., synthetic) images and predicts whether the images are real or fake. As part of training classifier 404, ground truth about whether the images provided to classifier 404 are real or fake is available as part of the training (e.g., for calculating the loss).
[0034] In at least one embodiment, the generated image 414 and the input image 402 are compared to determine an image consistency loss 418. In at least one embodiment, the cosine distance between the input image 402 and the generated image 414 is compared to determine feature similarity, with closer similarity resulting in a lower loss. In at least one embodiment, at least one of L1, L2, or cosine distance is used to determine the image consistency loss 418 between the input image 402 and the generated image 414. In at least one embodiment, the L1 distance is determined by the following symbolic formula: L1 distance=I in -I gen However, I in corresponds to the representation of the input image, and I gen corresponds to the representation of the generated image.
[0035] In at least one embodiment, the L2 distance is determined by the following symbolic formula:
number
[0036] In at least one embodiment, the cosine distance is determined by the following symbolic formula: Cosine distance = f in -f gen However, f in corresponds to the representation of the features of the input image, and f gen corresponds to the representation of the features of the generated image.
[0037] In at least one embodiment, additional loss functions are calculated as part of updating the classifier as described in conjunction with FIG. 4. In at least one embodiment, nearest and farthest neighbor losses 420 are calculated. In at least one embodiment, symmetry loss 422 is calculated. In at least one embodiment, nearest and farthest neighbor losses 420 and / or symmetry loss 422 are calculated according to techniques described elsewhere, such as the techniques discussed in conjunction with FIGS. 8-10. In at least one embodiment, the calculated losses (e.g., the losses shown in FIG. 4) are used to calculate gradients and update parameters for classifier 410 while holding parameters of generator 412 constant, using any suitable technique, such as gradient descent.
[0038] FIG. 5 shows a diagram 500 illustrating classifier updating, according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update parameters of a classifier 504 used to predict various outputs from input images. In at least one embodiment, FIG. 5 illustrates a viewpoint 502, a set of appearance parameters 504, a generator 506, a generated image 508, a classifier 510, a predicted viewpoint 512, a predicted decision 514 of whether the generated image 508 is real or fake, a predicted set of appearance parameters 516, a viewpoint consistency loss 518, a Z-reconstruction loss 520, and a real / fake classification loss 522. In at least one embodiment, FIG. 5 illustrates classifier updating using synthetic images (e.g., images created by a generative adversarial network). In at least one embodiment, the techniques described in conjunction with FIG. 5 are coextensive with the techniques described in conjunction with FIGS. 4, 6, and 7 for training a generator and / or classifier.
[0039] In at least one embodiment, the viewpoint 502 and the set of appearance parameters 504 are selected in any suitable manner, which may include random selection of parameter values, weighted random selection, etc. In at least one embodiment, the viewpoint 502 and the set of appearance parameters 504 are disentangled parameters that may be independently selected. In at least one embodiment, the generator 506 accepts the viewpoint 502 and the set of appearance parameters 504 as input and creates a generated image 508. In at least one embodiment, the generated image 508 is a synthetic image that is generated based on the set of appearance parameters 504 and appears oriented according to the viewpoint 502.
[0040] In at least one embodiment, a set of images (e.g., generated images 508) is provided as input to a classifier 510, which predicts various properties of the set of images. In at least one embodiment, the classifier 510 receives the generated images 508 and produces a predicted viewpoint 512, a predicted decision 514 as to whether the generated images 508 are real or fake, and a predicted set of appearance parameters 516. In at least one embodiment, the classifier does not have access to the viewpoint 502 and set of appearance parameters 504 used to create the generated images 508 (e.g., such information is not provided to the classifier 510 during prediction). In at least one embodiment, the output of the classifier 510 is used to calculate a loss. In at least one embodiment, the loss function is used to calculate gradients (e.g., using gradient descent) and update the parameters for the classifier 510 while holding the parameters of the generator 506 constant.
[0041] In at least one embodiment, a viewpoint consistency loss 518 is calculated. In at least one embodiment, viewpoint consistency loss refers to a loss function calculated based on how accurate the classifier 510 is in predicting the viewpoint. In at least one embodiment, viewpoint consistency loss 518 is calculated as the difference or distance between the input viewpoint 502 and the predicted viewpoint 512. In at least one embodiment, viewpoint consistency loss is a component of the generative consistency loss. In at least one embodiment, the distance (e.g., L1 distance, L2 distance, cosine distance, and / or variations thereof) between the input viewpoint 502 and the predicted viewpoint 512 is used to calculate viewpoint consistency loss 518, with the loss being lower for closer viewpoints (e.g., viewpoints closer to each other).
[0042] In at least one embodiment, a Z-reconstruction loss 520 is calculated. In at least one embodiment, the Z-reconstruction loss refers to the difference or distance between the input set of appearance parameters 504 and the predicted set of appearance parameters 516. In at least one embodiment, the Z-reconstruction loss refers to a loss function that is calculated based on how accurate the classifier 510 is in predicting the appearance parameters or appearance properties of an image.
[0043] In at least one embodiment, decision 514 is used to determine a classification loss, such as real / fake classification loss 522. In at least one embodiment, real / fake classification loss 522 is calculated based on whether classifier 514 can accurately predict whether generated images 508 submitted to classifier 510 are real or synthetic. In at least one embodiment, real / fake classification loss 522 is calculated based on whether classifier 510 can accurately predict whether a set of input images are real or fake, where classifier 510 can be provided with either real or fake (e.g., synthetic) images and predicts whether the images are real or fake.
[0044] FIG. 6 shows a diagram 600 illustrating generator updating, according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update parameters of a generator used to create synthetic images. In at least one embodiment, FIG. 6 shows an input image 602, a classifier 604, a predicted viewpoint 606, a predicted decision 608 as to whether the input image 602 is real or fake, a predicted set of appearance parameters 610, a generator 612, a generated image 614, an image consistency loss 616, and a real / fake classification loss 618. In at least one embodiment, FIG. 6 illustrates updating a generator using real images (e.g., images not generated by a generative adversarial network). In at least one embodiment, the techniques described in conjunction with FIG. 6 are coextensive with the techniques described in conjunction with FIGS. 4, 5, and 7 for training a generator and / or classifier.
[0045] In at least one embodiment, an input image 602 is input to a classifier 604. In at least one embodiment, the input image 602 is a real image selected from a set of training data. In at least one embodiment, the input image 602 is selected from a set of object images, such as the set of object images described in conjunction with FIG. 2. In at least one embodiment, the classifier 604 processes the input image 602. In at least one embodiment, the classifier 604 receives the input image 602 and generates a predicted viewpoint 606, a predicted decision whether the input image 602 is real or fake, and a predicted set of appearance parameters 608. In at least one embodiment, the viewpoint 606 is a predicted viewpoint of the input image 602 determined by one or more processes of the classifier 604. In at least one embodiment, the decision 608 is a prediction of whether the input image 602 is a real image or a fake image (e.g., a synthetic image created by a generative adversarial network). In at least one embodiment, decision 608 is a binary value or a non-binary range of confidence values (e.g., 0 to 100). In at least one embodiment, set of appearance parameters 610 is a predicted set of appearance parameters for input image 602 generated by one or more processes of classifier 604. In at least one embodiment, set of appearance parameters 610 are predicted parameters that define the appearance of objects depicted in input image 602.
[0046] In at least one embodiment, viewpoint 606 and set of appearance parameters 610 are provided as input to generator 612 to produce generated image 614. In at least one embodiment, generated image 614 is a synthetic (e.g., fake) image created by generator 612, having an orientation and appearance according to viewpoint 606 and set of appearance parameters 610, which are also the predicted viewpoint and appearance of input image 602, as shown in FIG.
[0047] In at least one embodiment, the image consistency loss 616 is calculated based, at least in part, on an input image 602 provided to a classifier 604, which is used to isolate at least two properties from the input image 602: a predicted viewpoint 606 and a predicted set of appearance parameters 610. In at least one embodiment, the predicted viewpoint 606 and the predicted set of appearance parameters 610 are provided to a generator 612 to create a generated image 614. In at least one embodiment, a generative adversarial network (GAN) receives the set of viewpoints and appearance parameters and generates a synthetic (e.g., fake) image according to the provided set of viewpoints and appearance parameters. In at least one embodiment, the input image 602 and the generated image 614 are compared to determine the image consistency loss 616. In at least one embodiment, the cosine distance between the input image and the synthetic image is compared to determine feature similarity, with closer similarity resulting in a lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss 616 between two images.
[0048] In at least one embodiment, decision 608 is used to determine a classification loss, such as a real / fake classification loss 618. In at least one embodiment, the real / fake classification loss 618 is calculated based on whether the classifier 604 can accurately predict whether the input images 602 submitted to the classifier 604 are real or synthetic. In at least one embodiment, the real / fake classification loss 618 is calculated based on whether the classifier 604 can accurately predict whether a set of input images are real or fake, where the classifier 604 can be presented with either real or fake (e.g., synthetic) images and predicts whether the images are real or fake. In at least one embodiment, one or more loss functions are calculated. In at least one embodiment, the parameters of the generator 612 are updated by at least computing gradients (e.g., performing stochastic gradient descent) while holding the classifier parameters fixed.
[0049] FIG. 7 shows a diagram 700 illustrating generator updates, according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update the parameters of the generator used to create the synthetic image. In at least one embodiment, FIG. 7 shows a first viewpoint 702, a set of appearance parameters 704, a generator 706, a first generated image 708, a classifier 710, a predicted first viewpoint 712, a predicted decision 714 of whether the generated image 708 is real or fake, a predicted first set of appearance parameters 716, a viewpoint consistency loss 718, a Z-reconstruction loss 720, a real / fake classification loss 722, a second viewpoint 724, a second generated image 726, a predicted second viewpoint 728, a predicted second set of appearance parameters 730, a viewpoint consistency loss 732, a Z-reconstruction loss 734, and a symmetry loss 736. In at least one embodiment, Figure 6 illustrates updating a generator using fake images (e.g., synthetic or generated images created by a generative adversarial network). In at least one embodiment, the techniques described in conjunction with Figure 7 are coextensive with the techniques described in conjunction with Figures 4-6 for training a generator and / or classifier.
[0050] In at least one embodiment, first viewpoint 702 and set of appearance parameters 704 are independently selectable disentangled parameters. In at least one embodiment, generator 706 accepts first viewpoint 702 and set of appearance parameters 704 as input and creates first generated image 708. In at least one embodiment, first generated image 708 is a synthetic image that is generated based on set of appearance parameters 704 and appears oriented according to first viewpoint 702. Generator 706 may be according to generators described elsewhere in this disclosure, such as the generator discussed in conjunction with FIG. 2.
[0051] In at least one embodiment, a set of images (e.g., first generated images 708) is provided as input to a classifier 710, which predicts various properties of the set of images. In at least one embodiment, the classifier 710 receives the first generated images 708 and produces a predicted first viewpoint 712, a predicted decision 714 of whether the generated images 708 are real or fake, and a predicted first set of appearance parameters 716. In at least one embodiment, the classifier 710 does not have access to the first viewpoint 702 and set of appearance parameters 704 used to create the first generated images 708 (e.g., such information is not provided to the classifier 710 during prediction). In at least one embodiment, the output of the classifier 710 is used to calculate a loss. In at least one embodiment, the loss function is used to calculate gradients (e.g., using gradient descent) and update the parameters for the classifier 710 while holding the parameters of the generator 706 constant.
[0052] In at least one embodiment, a viewpoint consistency loss 718 is calculated. In at least one embodiment, viewpoint consistency loss refers to a loss function calculated based on how accurate the classifier 710 is in predicting the viewpoint. In at least one embodiment, viewpoint consistency loss 718 is calculated as the difference or distance between the first viewpoint 702 and the predicted first viewpoint 712. In at least one embodiment, viewpoint consistency loss is a component of generative consistency loss. In at least one embodiment, the distance (e.g., L1 distance, L2 distance, cosine distance, and / or variations thereof) between the first viewpoint 702 and the predicted first viewpoint 712 is used to calculate viewpoint consistency loss 718, with the loss being lower for closer viewpoints (e.g., viewpoints closer to each other).
[0053] In at least one embodiment, a Z-reconstruction loss 720 is calculated. In at least one embodiment, the Z-reconstruction loss refers to the difference or distance between the set of appearance parameters 704 and the predicted first set of appearance parameters 716. In at least one embodiment, the Z-reconstruction loss refers to a loss function that is calculated based on how accurate the classifier 710 is in predicting the appearance parameters or appearance properties of an image.
[0054] In at least one embodiment, decision 714 is used to determine a classification loss, such as a real / fake classification loss 722. In at least one embodiment, real / fake classification loss 722 is calculated based on whether classifier 714 can accurately predict whether a first generated image 708 submitted to classifier 710 is a real image or a synthetic image. In at least one embodiment, real / fake classification loss 722 is calculated based on whether classifier 710 can accurately predict whether a set of input images is real or fake, where classifier 710 can be provided with either real or fake (e.g., synthetic) images and predicts whether the images are real or fake.
[0055] In at least one embodiment, the second viewpoint 724 is a transformation of the first viewpoint 702. In at least one embodiment, the first viewpoint 702 is flipped horizontally to create the second viewpoint 724. In at least one embodiment, any suitable transformation of azimuth, tilt, and elevation parameters relative to the first viewpoint 702 creates the second viewpoint 724. In at least one embodiment, the second viewpoint 724 is determined according to techniques described elsewhere in this disclosure, such as the technique discussed in conjunction with FIG. 8. In at least one embodiment, if the first viewpoint 702 has azimuth, elevation, and tilt parameters as φ, θ, and ψ, respectively, the second viewpoint 724 has azimuth, elevation, and tilt parameters as −φ, θ, and −ψ.
[0056] In at least one embodiment, the set of appearance parameters 704 is used to generate a second image. In at least one embodiment, the second viewpoint 724 and the set of appearance parameters 704 are provided as inputs to a generator 706 to create a second generated image 726. In at least one embodiment, when the second generated image 726 is flipped horizontally, it produces the first generated image 708.
[0057] In at least one embodiment, symmetry loss 736 is calculated based on the first generated image 708 and the second generated image 726. In at least one embodiment, symmetry loss 736 is calculated by comparing the azimuth, elevation, and tilt magnitudes of the predicted first viewpoint 712 of the generated image 708 with the predicted second viewpoint 728 of the second generated image 726, where equal magnitudes of the parameters for each viewpoint result in zero loss and larger differences in the magnitudes of the parameters for each viewpoint result in larger losses. In at least one embodiment, symmetry loss 736 is calculated according to techniques described elsewhere in this disclosure, such as the techniques discussed in conjunction with FIG. 8. In at least one embodiment, disentanglement loss is applied within the context of FIG. 7. In at least one embodiment, a viewpoint V1 and a set of appearance parameters Z1 are selected and provided to a generator to create a first synthetic image I1. In at least one embodiment, the second image I2 is generated by holding the viewpoint constant (e.g., using V1) and perturbing the appearance parameters by using a second set of appearance parameters Z2 that differs from Z1 used to generate the first image I1. In at least one embodiment, the third image I3 is generated by holding the appearance parameters constant (e.g., using Z1) for image I1 and perturbing the viewpoint by using a second viewpoint V2 that differs from viewpoint V1 to generate the third image I3. In at least one embodiment, a synthetic image I1 generated using a particular viewpoint V1 and set of appearance parameters Z1 is compared to a synthetic image I2 generated using viewpoint V1 and second set of appearance parameters Z2 and / or to a synthetic image I3 generated using a second viewpoint V2 and set of appearance parameters Z1. The technique described in conjunction with FIG. 10 may be applicable to the entanglement loss described in conjunction with FIG. 7, according to at least one embodiment.
[0058] 8 shows a diagram 800 illustrating the calculation of symmetry loss, according to at least one embodiment. In at least one embodiment, symmetry loss is utilized to train one or more neural networks associated with a classifier, such as classifier 806. In at least one embodiment, symmetry loss is utilized in conjunction with one or more other loss functions to refine parameters associated with the classifier.
[0059] In at least one embodiment, diagram 800 includes input image 802. In at least one embodiment, input image 802 is part of a collection of one or more images of a certain type of object. In at least one embodiment, input image 802 is an image depicting a car in a particular orientation and with particular appearance characteristics. In at least one embodiment, input image 802 is selected from one or more collections of images. In at least one embodiment, the images in the collection are selected for learning in any suitable manner and may be randomly or pseudo-randomly sampled from a training set.
[0060] In at least one embodiment, a transform is applied to input image 802 to generate transformed image 804. In at least one embodiment, input image 802 is flipped horizontally to generate transformed image 804. In at least one embodiment, the transform is applied by one or more systems associated with classifier 806. In at least one embodiment, classifier 806 applies one or more image processing techniques to input image 802 to generate transformed image 804. In at least one embodiment, the azimuth and tilt angles (denoted as "az" and "ti" in diagram 800) of the viewpoint of an object in input image 802 are reversed when input image 802 is flipped and / or translated to generate transformed image 804, while the elevation angle (denoted as "el" in diagram 800) remains the same for input image 802 and transformed image 804. In at least one embodiment, transformed image 804 is generated by applying one or more image transforms to input image 802. In at least one embodiment, the transformed image 804 is generated by at least horizontally flipping the input image 802, horizontally flipping the input image 802, flipping the input image 802 about a specified axis, rotating the input image 802 by a specified number of degrees, and / or various other 2D transformations applied to the input image 802.
[0061] In at least one embodiment, a classifier 806 processes an input image 802. In at least one embodiment, the classifier 806 is associated with one or more neural networks that are trained to infer viewpoints and other characteristics of the input image. In at least one embodiment, the classifier 806 receives the input image 802 and generates a first prediction 808. In at least one embodiment, the first prediction 808 corresponds to a particular predicted orientation of an object in the image and includes particular values for a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the first prediction 808 includes a first predicted viewpoint V1 of the input image 802 and a predicted first set of appearance parameters Z1. In at least one embodiment, the first prediction 808 includes a predicted orientation of a vehicle depicted in the input image 802. In at least one embodiment, the classifier 806 processes a transformed image 804. In at least one embodiment, the classifier 806 receives the transformed image 804 and generates a second prediction 810. In at least one embodiment, the second prediction 810 corresponds to a particular predicted orientation of the object in the image and includes particular values for a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the second prediction 810 includes a second predicted viewpoint V2 and a second predicted set of appearance parameters Z2 of the transformed image 804. In at least one embodiment, the second prediction 810 includes a predicted orientation of the car depicted in the transformed image 804.
[0062] In at least one embodiment, a first prediction 808 is predicted relative to an input image 802, and a second prediction 810 is predicted relative to a transformed image 804, which is a horizontally flipped version of the input image 802. In at least one embodiment, a loss is calculated based on whether certain properties hold true. In at least one embodiment, a transformation or its inverse is applied to the second prediction 810. In at least one embodiment, if the input image 802 is rotated by angles (φ, θ, ψ) to produce the transformed image 804, the second prediction 810 may be rotated inversely by angles (-φ, -θ, -ψ). In at least one embodiment, the symmetry loss is calculated by comparing the azimuth, elevation, and tilt magnitudes of a first prediction 808 of the input image 802 with a second prediction 810 of the transformed image 804, where zero loss occurs when the magnitudes of the parameters for each viewpoint are equal and / or the appearance parameters predicted for the input image 802 match the appearance parameters predicted for the transformed image 804. In at least one embodiment, the greater the difference between the magnitudes of the parameters for each viewpoint, the greater the loss. In at least one embodiment, the symmetry loss is calculated according to techniques described elsewhere in this disclosure, such as the technique discussed in conjunction with FIG. 16 . In at least one embodiment, the symmetry loss is calculated, at least in part, based on how closely the appearance parameters for the input image 802 match the appearance parameters predicted for the transformed image 804. In at least one embodiment, the weights and parameters of the classifier 806 are trained to predict whether predicted appearance parameters for the input image 802 will match predicted appearance parameters for a transformed version of the input image 802 (e.g., the transformed image 804).
[0063] In at least one embodiment, the classifier 806 predicts a first set of appearance parameters for the input image 802 and a second set of appearance parameters for the transformed image 804. In at least one embodiment, a loss function is calculated based on how similar the first set of appearance parameters predicted for the input image 802 are to the second set of appearance parameters predicted for the transformed image 804. In at least one embodiment, the classifier 806 is trained to predict whether the parameters for the input image 802 and the transformed image 804 are equivalent.
[0064] FIG. 9 illustrates a diagram 900 depicting a viewpoint graph, according to at least one embodiment. In at least one embodiment, the viewpoint graph 908 is utilized to determine nearest and farthest neighbor losses, which are utilized along with one or more other loss functions to refine parameters for a classifier. In at least one embodiment, the viewpoint graph 908 is constructed through one or more processes and systems associated with the classifier. In at least one embodiment, the viewpoint graph 908 is generated based on a collection of one or more images of a certain type of object. In at least one embodiment, the first image 902 and the second image 904 are part of a collection of one or more images of a certain type of object. In at least one embodiment, the first image 902 and the second image 904 are part of a collection of one or more images including a car. In at least one embodiment, the first image 902 and the second image 904 are images depicting a car in a particular orientation.
[0065] In at least one embodiment, the viewpoint graph 908 is generated based on viewpoint covariance distances (e.g., cosine distances) between images in a collection of images. In at least one embodiment, cosine distance is a mathematical component of cosine similarity (e.g., cosine distance = 1 - cosine similarity). In at least one embodiment, cosine similarity is a measure of similarity between two vectors, which may represent images, text, data, and / or variations thereof, and is based on the cosine of the angle between the vectors. In at least one embodiment, if two images contain features corresponding to objects with similar viewpoints, the cosine distance calculated for the images is small. In at least one embodiment, if two images contain features corresponding to objects with different viewpoints, the cosine distance calculated for the images is large.
[0066] In at least one embodiment, a collection of images is used to generate a viewpoint graph 908, where nodes in the viewpoint graph 908 correspond to images and edges correspond to their cosine distances. In at least one embodiment, the cosine distance is calculated based on feature similarities between pairs of images using a convolutional neural network. In at least one embodiment, the edges in the viewpoint graph 908 are weighted so that a thick edge between two images corresponds to a high degree of similarity between the two images and a thin edge between two images corresponds to a low degree of similarity between the two images. In at least one embodiment, a cosine distance 910 is calculated between a first image 902 and a second image 904. In at least one embodiment, the first image 902 and the second image 904 are used as part of the viewpoint graph 908 and are connected by edges corresponding to the calculated cosine distance 910. In at least one embodiment, the first image 902 and the second image 904 include cars facing similar orientations and / or viewpoints, and the thick edges used between the first image 902 and the second image 904 in the viewpoint graph 908 reflect such similarities.
[0067] FIG. 9 shows a diagram 900 illustrating the calculation of nearest and farthest neighbor losses, according to at least one embodiment. In at least one embodiment, the nearest and farthest neighbor losses are utilized to train one or more neural networks associated with a classifier, such as classifier 912. In at least one embodiment, the nearest and farthest neighbor losses are utilized along with one or more other loss functions to refine parameters associated with the classifier. In at least one embodiment, the nearest and farthest neighbor losses can be thought of as a type of viewpoint-covariant loss. In at least one embodiment, techniques described herein that apply to nearest and farthest neighbor losses (e.g., techniques described in conjunction with FIG. 9 and elsewhere) can also be applied to viewpoint-covariant losses.
[0068] In at least one embodiment, viewpoint graph 908 is constructed through one or more processes and systems associated with classifier 912. In at least one embodiment, viewpoint graph 908 is generated based on a collection of one or more images of a certain type of object. In at least one embodiment, images 902-906 are part of a collection of one or more images of a certain type of object. In at least one embodiment, images 902-906 are part of a collection of one or more images that include cars. In at least one embodiment, images 902-906 are images depicting cars in a particular orientation.
[0069] In at least one embodiment, a classifier 912 processes images 902-906. In at least one embodiment, classifier 912 is associated with one or more neural networks that are trained to infer viewpoints and other characteristics of the input images. In at least one embodiment, classifier 912 receives image 902 and generates viewpoint 914. In at least one embodiment, classifier 912 receives image 904 and generates viewpoint 916. In at least one embodiment, classifier 912 receives image 906 and generates viewpoint 918. In at least one embodiment, viewpoints 914-918 correspond to particular predicted orientations of objects depicted in images 902-906 and include particular values for a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, viewpoints 914-918 include predicted orientations of vehicles depicted in images 902-906, respectively.
[0070] In at least one embodiment, nearest and farthest neighbor losses are calculated by comparing the selected image to its nearest and farthest neighbors based, at least in part, on a viewpoint graph 908 of the set of images. In at least one embodiment, the nearest and farthest neighbor losses include nearest neighbor losses and farthest neighbor losses. In at least one embodiment, an anchor image is selected from a set of training images. In at least one embodiment, image 902 is selected as the anchor image. In at least one embodiment, image 902 is located from the viewpoint graph 908, and the nearest and farthest neighbors are selected based on edge weights. In at least one embodiment, a nearest neighbor, such as image 904, has the shortest edge leading to image 902. In at least one embodiment, a nearest neighbor, such as image 906, has the farthest edge leading to image 902.
[0071] In at least one embodiment, image 904 is determined as the nearest neighbor to image 902. In at least one embodiment, a nearest neighbor loss is calculated between viewpoint 914 and viewpoint 916. In at least one embodiment, the nearest neighbor loss is calculated such that a higher similarity between viewpoint 914 and viewpoint 916 corresponds to a lower loss. In at least one embodiment, the nearest neighbor loss is calculated based on the following symbolic formula:
number
[0072] However, L nn is the nearest neighbor loss,
number
number
number
[0073] In at least one embodiment, image 906 is determined as the farthest neighbor to image 902. In at least one embodiment, a farthest neighbor loss is calculated between viewpoint 914 and viewpoint 916. In at least one embodiment, the farthest neighbor loss is calculated such that a lower similarity between viewpoint 914 and viewpoint 918 corresponds to a lower loss. In at least one embodiment, the farthest neighbor loss is calculated based on the following symbolic formula:
number
[0074] However, L fn is the farthest neighbor loss,
number
number
number
[0075] In at least one embodiment, the nearest and / or farthest neighbors are selected non-deterministically. In at least one embodiment, a probability is assigned to each edge selected as the nearest and / or farthest neighbor of the anchor image. In at least one embodiment, the probability of a nearest neighbor is inversely proportional to the edge weight (e.g., the node with the lowest edge weight leading to the anchor image has the highest probability of being selected). In at least one embodiment, the probability of a farthest neighbor is directly proportional to the edge weight (e.g., the node with the highest edge weight leading to the anchor image has the highest probability of being selected).
[0076] FIG. 10 shows a diagram 1000 illustrating disentanglement loss, according to at least one embodiment. In at least one embodiment, the disentanglement loss is utilized to train one or more neural networks associated with a generator, such as generator 1002. In at least one embodiment, the disentanglement loss is utilized along with one or more other loss functions to refine parameters associated with generator 1002. In at least one embodiment, the disentanglement loss operates when generator parameters are updated. In at least one embodiment, the disentanglement loss operates during the generator update process.
[0077] In at least one embodiment, generator 1002 is associated with a generative model such as a variational autoencoder (VAE), a differentiable renderer, a generative adversarial network, or a renderer. In at least one embodiment, generator 1002 is part of one or more neural networks trained to generate images based on an input set of input viewpoint and appearance parameters, the one or more neural networks including various parameters associated with one or more processes of the one or more neural networks. In at least one embodiment, classifier 1004 is part of one or more neural networks trained to infer a set of viewpoint and appearance attributes from an input image, the one or more neural networks including various parameters associated with one or more processes of the one or more neural networks.
[0078] In at least one embodiment, a first viewpoint 1006A and a first set of appearance attributes 1008A are obtained. In at least one embodiment, the first viewpoint 1006A and the first set of appearance attributes 1008A are randomly generated. In at least one embodiment, the first viewpoint 1006A and the first set of appearance attributes 1008A are obtained from one or more processes associated with the generator 1002 and the classifier 1004. In at least one embodiment, the generator 1002 generates an image 1010A based on the first viewpoint 1006A and the first set of appearance attributes 1008A. In at least one embodiment, the image 1010A is input to the classifier 1004. In at least one embodiment, the classifier 1004 performs one or more processes to determine a first predicted viewpoint 1012A and a first predicted set of appearance attributes 1012B based on the input image 1010A.
[0079] In at least one embodiment, a second set of appearance attributes 1008B is obtained. In at least one embodiment, the second set of appearance attributes 1008B is randomly generated. In at least one embodiment, the second set of appearance attributes 1008B is obtained from one or more processes associated with the generator 1002 and the classifier 1004. In at least one embodiment, the generator 1002 generates an image 1010B based on the first viewpoint 1006A and the second set of appearance attributes 1008B. In at least one embodiment, the image 1010B is input to the classifier 1004. In at least one embodiment, the classifier 1004 performs one or more processes to determine a second predicted viewpoint 1014A and a second predicted set of appearance attributes 1014B based on the input image 1010B.
[0080] In at least one embodiment, a second viewpoint 1006B is obtained. In at least one embodiment, the second viewpoint 1006B is randomly generated. In at least one embodiment, the second viewpoint 1006B is obtained from one or more processes associated with the generator 1002 and the classifier 1004. In at least one embodiment, the generator 1002 generates an image 1010C based on the second viewpoint 1006B and the first set of appearance attributes 1008A. In at least one embodiment, the image 1010C is input to the classifier 1004. In at least one embodiment, the classifier 1004 performs one or more processes to determine a third predicted viewpoint 1016A and a third predicted set of appearance attributes 1016B based on the input image 1010C.
[0081] In at least one embodiment, disentanglement loss includes a z-reconstruction loss and a viewpoint reconstruction loss. In at least one embodiment, the z-reconstruction loss is calculated by comparing a prediction of a set of appearance parameters for an image, generated by a classifier such as classifier 1004, with an input set of appearance parameters used to generate the image generated by a generator such as generator 1002, where the more similar the prediction of a set of appearance parameters is to the input set of appearance parameters, the smaller the loss, and the more dissimilar the prediction of a set of appearance parameters is to the input set of appearance parameters. In at least one embodiment, the viewpoint reconstruction loss is calculated by comparing a prediction of a viewpoint for an image, generated by a classifier such as classifier 1004, with an input viewpoint used to generate the image generated by a generator such as generator 1002, where the more similar the prediction of a viewpoint is to the input viewpoint, the smaller the loss, and the more dissimilar the prediction of a viewpoint is to the input viewpoint.
[0082] In at least one embodiment, a z reconstruction loss and a viewpoint reconstruction loss are calculated for each set of predicted viewpoints and appearance attributes (e.g., a first predicted viewpoint 1012A and a first predicted set of appearance attributes 1012B based on input image 1010A, a second predicted viewpoint 1014A and a second predicted set of appearance attributes 1014B based on input image 1010B, and a third predicted viewpoint 1016A and a third predicted set of appearance attributes 1016B based on input image 1010C). In at least one embodiment, the z reconstruction loss and the viewpoint reconstruction loss are utilized to determine a disentanglement loss. In at least one embodiment, an additional loss function is also calculated as part of the disentanglement loss. In at least one embodiment, the disentanglement loss is calculated based on the following symbolic formula: disentanglement loss = Σ viewpoint reconstruction loss + z reconstruction loss + other loss functions
[0083] In at least one embodiment, the disentanglement loss is utilized to refine one or more parameters associated with generator 1002. In at least one embodiment, the parameters of generator 1002 are updated such that at least the loss from the disentanglement loss is minimized.
[0084] 11 shows a diagram 1100 illustrating the calibration of a neural network, according to at least one embodiment. In at least one embodiment, a classifier is trained on a set of images and calibrated on a portion of the set of images that includes ground truth annotations. In at least one embodiment, classifier 1104 is part of one or more neural networks trained to identify the orientation of an object in an image based, at least in part, on one or more characteristics of the object other than the object's orientation. In at least one embodiment, classifier 1106 is trained on a set of images to infer the orientation of other objects of the same category captured in other images.
[0085] In at least one embodiment, the classifier 1104 is trained on a set of images as described elsewhere in this disclosure to identify viewpoints of objects in the images in a self-supervised manner. In at least one embodiment, the classifier 1104 is trained to identify orientations of objects in the images in a self-supervised manner by, at least as part of training, computing one or more loss functions that evaluate one or more characteristics of the images in the training set. In at least one embodiment, the classifier 1106 is trained at least in part based on computing losses such as one or more of the following: generative consistency loss, symmetry loss, nearest and furthest neighbor loss, and disentanglement loss, which may be in accordance with those described in conjunction with FIGS. 4-10 . In at least one embodiment, the classifier 1104 is trained on a set of images that does not have ground truth annotations or for which ground truth annotations are not available.
[0086] In at least one embodiment, an object image set 1102 is acquired to calibrate a classifier 1104. In at least one embodiment, the object image set 1102 includes images with ground truth annotations. In at least one embodiment, the object image set 1102 includes images depicting an object for which the classifier 1104 is trained to analyze and determine a viewpoint. In at least one embodiment, the object image set 1102 is part of a set of images used to train the classifier 1104. In at least one embodiment, the classifier 1104 is trained on a set of images different from the object image set 1102. In at least one embodiment, the classifier 1104 acquires images from the object image set 1102 and determines a viewpoint 1106 for the images. In at least one embodiment, the classifier 1104 is trained but the viewpoint 1106 does not match the ground truth viewpoint of the images in the object image set 1102. In at least one embodiment, viewpoint 1106 is the correct viewpoint for the images of object image set 1102 that are translated in one or more dimensions.
[0087] In at least one embodiment, a linear model 1108 is determined that translates viewpoints determined by the classifier 1104 for images in the set of object images 1102 to their respective correct positions as indicated by the ground truth annotations. In at least one embodiment, the linear model 1108 is a linear function determined by comparing viewpoints generated by the classifier 1104 for images in the set of object images 1102 to ground truth annotations of viewpoints for the images in the set of object images 1102. In at least one embodiment, the linear model 1108 is determined through one or more processes, such as various regression algorithms, mathematical processes, and / or variations thereof. In at least one embodiment, the linear model 1108 is determined such that viewpoints determined by the classifier 1104 for images can be translated and / or corrected to match the ground truth viewpoints for the images. In at least one embodiment, the linear model 1108 calibrates the viewpoint determined by the classifier 1104 for an image to the coordinate system of the ground truth viewpoint for the image. In at least one embodiment, the linear model 1108 translates the zero position of the viewpoint determined by the classifier 1104 for an image to the zero position of the ground truth viewpoint for the image.
[0088] In at least one embodiment, a linear model 1108 is used to calibrate or correct viewpoint 1106 to generate a corrected viewpoint 1110. In at least one embodiment, corrected viewpoint 1110 is viewpoint 1106 that has been translated to match the ground truth viewpoint. In at least one embodiment, linear model 1108 is used to calibrate or correct the viewpoint determined / generated by classifier 1104 for an image in object image set 1102 to match the ground truth annotation of the viewpoint for that image in object image set 1102.
[0089] FIG. 12 shows a diagram 1200 illustrating inference, according to at least one embodiment. In at least one embodiment, a classifier is trained as part of one or more systems associated with a vehicle's safety system to infer viewpoints of images of other vehicles captured from cameras associated with the vehicle, whereby the inferred viewpoints are utilized to perform one or more actions related to the vehicle, such as braking the vehicle to avoid another vehicle or steering the vehicle to avoid another vehicle. In at least one embodiment, classifier 1204 is trained in a self-supervised manner on a set of images, as described elsewhere in this disclosure, to identify viewpoints of objects in the images. In at least one embodiment, classifier 1204 is trained to identify object orientations in the images in a self-supervised manner, at least by computing, as part of training, one or more loss functions that evaluate one or more characteristics of the images in the training set. In at least one embodiment, the classifier 1204 is trained at least in part based on calculating a generative consistency loss, a symmetry loss, a nearest neighbor and a farthest neighbor loss, and a disentanglement loss, which may be in accordance with those described in conjunction with Figures 4-10. In at least one embodiment, the classifier 1204 is trained on a set of images of cars. In at least one embodiment, the classifier 1204 is trained to identify viewpoints of cars within the images.
[0090] In at least one embodiment, the identifier 1204 is part of one or more systems of the vehicle 1210. In at least one embodiment, the vehicle 1210 is an autonomous vehicle. In at least one embodiment, the vehicle 1210 is a vehicle operated by an operator. In at least one embodiment, the vehicle 1210 includes one or more systems that implement the identifier 1204. In at least one embodiment, the vehicle 1210 includes one or more systems associated with the identifier 1204. In at least one embodiment, the vehicle 1210 includes one or more systems that allow the vehicle 1210 to remotely access and utilize the identifier 1204, which may be implemented in various remote and / or local systems.
[0091] In at least one embodiment, automobile 1210 includes multiple cameras. In at least one embodiment, camera 1202 is located on-board vehicle 1210. In at least one embodiment, camera 1202 is a camera remotely accessible from vehicle 1210. In at least one embodiment, camera 1202 captures or acquires image 1202A. In at least one embodiment, vehicle 1210 is a vehicle operating in an environment including other vehicles with which vehicle 1210 needs to perform one or more actions (e.g., allowing a vehicle to pass, passing a vehicle, braking for an oncoming vehicle, steering to avoid an oncoming vehicle, and / or variations thereof). In at least one embodiment, image 1202A is an image of another vehicle with which automobile 1210 needs to interact. In at least one embodiment, image 1202A is an image captured from an on-board camera of automobile 1210 while automobile 1210 is operating in an environment.
[0092] In at least one embodiment, image 1202A is input to classifier 1204, which determines a viewpoint 1206 of image 1202A. In at least one embodiment, viewpoint 1206 is a viewpoint of a car depicted in image 1202A. In at least one embodiment, safety system 1208 is part of automobile 1210. In at least one embodiment, safety system 1208 includes one or more systems configured to provide assistance to automobile 1210. In at least one embodiment, safety system 1208 includes one or more systems configured to operate one or more systems of automobile 1210, such as brake actuator 1208A and steering actuator 1208B, which are components configured to operate the brakes of automobile 1210 and the steering of automobile 1210, respectively. In at least one embodiment, brake actuator 1208A and steering actuator 1208B are the same as brake actuator 2148 and steering actuator 2156, respectively. In at least one embodiment, the vehicle 1210 may be implemented according to techniques described elsewhere, such as in FIG.
[0093] In at least one embodiment, safety system 1208 obtains viewpoint 1206. In at least one embodiment, safety system 1208 determines a direction of movement of the vehicle depicted in image 1202A based on viewpoint 1206. In at least one embodiment, safety system 1208 determines whether to utilize brake actuator 1208A or steering actuator 1208B based on viewpoint 1206. In at least one embodiment, if safety system 1208 determines that the vehicle depicted in image 1202A is moving in a direction relative to automobile 1210 that requires automobile 1210 to brake, safety system 1208 activates brake actuator 1208A to brake automobile 1210 such that automobile 1210 avoids any potential safety hazards resulting from the movement of the vehicle depicted in image 1202A. In at least one embodiment, if the safety system 1208 determines that the car depicted in image 1202A is moving in a direction relative to the automobile 1210 that requires the automobile 1210 to steer, the safety system 1208 activates the steering actuator 1208B to steer the automobile 1210 so that the automobile 1210 avoids any potential safety hazards resulting from the movement of the car depicted in image 1202A.
[0094] FIG. 13 shows an illustrative illustration of a process 1300 for training a neural network to predict viewpoints of objects in an image, according to at least one embodiment. In at least one embodiment, some or all of process 1300 (or any other process described herein, or variations and / or combinations thereof) runs under the control of one or more computer systems configured with computer-executable instructions and may be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions usable to implement process 1300 are not stored using only transitory signals (e.g., propagating transient electrical or electromagnetic transmissions). A non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transient signal. In at least one embodiment, process 1300 is implemented, at least in part, on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, a first computer system trains one or more neural networks, and a second computer system uses the one or more neural networks to infer (e.g., predict viewpoints of objects in an image). In at least one embodiment, the techniques described in conjunction with FIGS. 1, 12, and 14 are applicable to process 1300.
[0095] In at least one embodiment, a system for implementing at least a portion of process 1300 includes executable code 1302 for obtaining one or more image collections of a type of object. In at least one embodiment, the one or more image collections are used to train one or more neural networks to identify the orientation of objects in the images. In at least one embodiment, the image collections are categorized or labeled as each displaying the same type or category of object. In at least one embodiment, the image collection is a collection of car images, which may include different types of cars in different orientations, in different weather conditions, and under different lighting. In at least one embodiment, the car collection includes images of the same car or the same type of car in different orientations. In at least one embodiment, at least some of the training image collections do not have ground truth annotations specifying the orientation of objects in such training images. In at least one embodiment, all of the images in the image collection do not have ground truth annotations specifying the azimuth, elevation, and tilt of objects in the images in the collection. In at least one embodiment, the collection of images includes one or more synthetic images, such as images created from a generative adversarial network (GAN). In at least one embodiment, all images in the collection of images are real images, as opposed to images synthesized or created from a generative model, such as a variational autoencoder (VAE), a differentiable renderer, a generative adversarial network (GAN), or a renderer. In at least one embodiment, the collection of images is collected and aggregated from a website that sorts images by category.
[0096] In at least one embodiment, the orientation of an object in an image refers to the three-dimensional orientation of an object captured in a two-dimensional image. In at least one embodiment, a camera is used to capture a two-dimensional image of a real-world vehicle at a particular orientation relative to the camera. In at least one embodiment, the orientation of the object (e.g., viewpoint) is encoded with respect to a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the orientation of the object is encoded as a set of three vectors that define the object's direction relative to the x-axis, y-axis, and z-axis.
[0097] In at least one embodiment, a system for performing at least a portion of process 1300 includes executable code for training 1304 one or more neural networks to identify the orientation of an object in an image based, at least in part, on one or more characteristics of the object other than the object's orientation. In at least one embodiment, process 1300 is implemented on a processor including one or more circuits that facilitate training one or more neural networks to identify the orientation of an object in an image based, at least in part, on one or more characteristics of the object other than the object's orientation. In at least one embodiment, one or more neural networks are trained on a set of images to infer the orientation of other objects of the same category captured in other images. In at least one embodiment, the neural network is trained on a set of images of aircraft and, once trained, is used to infer the orientation of other aircraft in other images.
[0098] In at least one embodiment, a neural network is trained to identify the orientation of objects in an image in a self-supervised manner on a set of images as described elsewhere in this disclosure. In at least one embodiment, the neural network is trained to identify the orientation of objects in an image in a self-supervised manner by, at least as part of the training, computing one or more loss functions that evaluate one or more characteristics of the images in the training set (e.g., the set of images). In at least one embodiment, the neural network is trained at least in part based on computing a generative consistency loss, a symmetry loss, a nearest neighbor and a furthest neighbor loss, and a disentanglement loss, which may be in accordance with those described in conjunction with FIGS. 4-10. In at least one embodiment, the neural network is trained on a set of images that does not have ground truth annotations or for which ground truth annotations are not available (e.g., such data is not provided to the neural network during training). In at least one embodiment, a neural network is trained to generate a second image having a predicted orientation from an object in a first image having the same orientation.
[0099] In at least one embodiment, a system implementing process 1300 includes one or more processors for calculating parameters that aid in training one or more neural networks to identify the orientation of an object in an image based, at least in part, on one or more characteristics of the object other than the object's orientation, and one or more memories for storing the parameters. In at least one embodiment, the one or more neural networks are trained to identify the orientation of an object in an image using a set of images of different objects in the same category as the object (e.g., a neural network for inferring the viewpoint of a vehicle is trained on a set of images labeled as vehicles).
[0100] In at least one embodiment, one or more neural networks are trained in a self-supervised manner on a set of images of different objects of the same category as the object in the image to be inferred. In at least one embodiment, the different objects of the same category may refer to different images, which may be one or more images of a first car in one or more orientations, one or more images of a second car in one or more different orientations, etc. In at least one embodiment, the images of the object to be inferred are included in the set of images used to train the one or more neural networks for inferring orientation. In at least one embodiment, the one or more neural networks are trained in a self-supervised manner by using at least a set of loss functions to evaluate one or more properties of the object in the image. In at least one embodiment, the one or more properties of the object refer to properties of the object that can be used to infer orientation. In at least one embodiment, the neural network is trained by calculating one or more of the following: generative consistency loss, symmetry loss, nearest and farthest neighbor loss, and disentanglement loss. In at least one embodiment, a neural network trained in a self-supervised manner is trained to generate synthetic images of objects at a specific orientation, which may be the same as the expected orientation of the input image. In at least one embodiment, the synthetic images are created using a deep generative model, such as a variational autoencoder (VAE), a differentiable renderer, a generative adversarial network (GAN), or a renderer. In at least one embodiment, the object whose orientation is to be inferred may be a vehicle, an aircraft, a drone, a human, a face (e.g., human or animal), etc.
[0101] In at least one embodiment, a system for performing at least a portion of process 1300 includes executable code for acquiring 1306 a second image. In at least one embodiment, the second object in the second image is of the same type of object as a set of images used to train one or more neural networks. In at least one embodiment, the second image is provided to a neural network for inference to predict a second orientation. In at least one embodiment, the images are acquired from a camera capturing still images and / or video comprising multiple frames captured at a variable or fixed rate. In at least one embodiment, the video includes multiple frames (e.g., images). In at least one embodiment, a first system trains one or more neural networks, and a second, different system uses these one or more neural networks to perform inference and identify the orientation of the object in the image.
[0102] In at least one embodiment, a system that performs at least a portion of process 1300 includes 1308 executable code for employing one or more neural networks (e.g., trained as described by numeral 1304) to identify a second orientation of the second object in the second image. In at least one embodiment, the system uses a classifier trained in a self-supervised manner on a collection of images of objects of a particular category to infer the orientation of other objects in that category. In at least one embodiment, one or more neural networks are trained on a collection of images of cars and used to infer the orientation of cars captured in real time by a camera or other suitable video / image capture device mounted on the vehicle.
[0103] In at least one embodiment, a first neural network is trained using self-supervised learning on a first set of images of a first category to infer viewpoints for objects in that first category, and a second neural network is trained using a similar / identical self-supervised learning technique on a second set of images of a second category. In at least one embodiment, images are provided as input to a first neural network to detect first orientations of first objects in the first category, and also to a second neural network to detect second orientations of second objects in the second category. In at least one embodiment, input images are provided to multiple neural networks trained using the self-supervised learning techniques described herein to identify orientations of different objects in the input images.
[0104] FIG. 14 shows an illustrative illustration of a process 1400 for training a neural network to predict viewpoints of objects in an image, according to at least one embodiment. In at least one embodiment, some or all of process 1400 (or any other process described herein, or variations and / or combinations thereof) runs under the control of one or more computer systems configured with computer-executable instructions and may be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions usable to implement process 1300 are not stored using only transitory signals (e.g., propagating transient electrical or electromagnetic transmissions). A non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transient signal. In at least one embodiment, process 1400 is implemented, at least in part, on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, a first computer system trains one or more neural networks, and a second computer system uses the one or more neural networks to infer (e.g., predict viewpoints of objects in an image). In at least one embodiment, the techniques described in conjunction with FIGS. 1, 12, and 13 are applicable to process 1400.
[0105] In at least one embodiment, a system that implements at least a portion of process 1400 includes executable code 1402 that acquires a collection of one or more images of a certain type of object. In at least one embodiment, images of a certain type may mean that all images in the collection of images include cars. In at least one embodiment, the system acquires the one or more collections of images according to techniques described elsewhere in this disclosure, such as in FIG. 13 .
[0106] In at least one embodiment, a system that performs at least a portion of process 1400 includes executable code that selects a first image of a set of images 1404. In at least one embodiment, the images of the set are selected in any suitable manner and may be randomly or pseudo-randomly sampled from a training set for learning.
[0107] In at least one embodiment, a system that implements at least a portion of process 1400 includes executable code that calculates 1406 a generative consistency loss based, at least in part, on comparing a selected image to an image generated by a deep generative model. In at least one embodiment, the generative consistency loss is calculated using techniques described elsewhere in this disclosure, such as the techniques discussed in conjunction with FIGS. 4-7. In at least one embodiment, the generative consistency loss includes at least two components: a viewpoint consistency loss and an image consistency loss. In at least one embodiment, the generative consistency loss is calculated according to the technique described in conjunction with FIG. 15.
[0108] In at least one embodiment, the image consistency loss is calculated based at least in part on a selected image (e.g., an input image) provided to a classifier, which is used to isolate at least two properties from the image: a predicted viewpoint and a set of appearance parameters. In at least one embodiment, the predicted viewpoint and set of appearance parameters are provided to a generator to create the synthetic image. In at least one embodiment, a generative adversarial network (GAN) receives the set of viewpoint and appearance parameters and generates a synthetic (e.g., fake) image according to the provided set of viewpoint and appearance parameters. In at least one embodiment, the synthetic image and the input image are compared to determine the image consistency loss. In at least one embodiment, the cosine distance between the input image and the synthetic image is compared to determine feature similarity, with closer similarity resulting in a lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss between two images.
[0109] In at least one embodiment, a viewpoint (e.g., orientation) consistency loss is calculated based at least in part on the viewpoint of the input image. In at least one embodiment, the viewpoint of the input image is inferred by a classifier. In at least one embodiment, the viewpoint of the input image is determined based on ground truth annotations provided as part of training on at least a portion of a set of training images. In at least one embodiment, a generator is used to create a synthetic image with the same viewpoint as the input image. In at least one embodiment, the synthetic image generated from the viewpoint of the input image is provided to a classifier, which determines a second viewpoint for the synthetic image. In at least one embodiment, a first viewpoint of the input image is compared to a second viewpoint of a synthetic image generated at least in part based on the input image. In at least one embodiment, the distance between the first viewpoint of the input image and the second viewpoint of the synthetic image is used to calculate a viewpoint consistency loss, with closer viewpoints resulting in a lower loss.
[0110] In at least one embodiment, a system implementing at least a portion of process 1400 includes executable code 1408 that calculates a symmetry loss by comparing a selected image with a transformed version of the selected image. In at least one embodiment, the symmetry loss is calculated according to the technique described in conjunction with FIG. 15. In at least one embodiment, an input image is selected from a set of training images. In at least one embodiment, a transformation is applied to the input image to generate a transformed image. In at least one embodiment, the input image is horizontally flipped to generate a flipped image. In at least one embodiment, one or more neural networks are used to predict a first orientation of the input image and a second orientation of the transformed image. In at least one embodiment, the first orientation is predicted for an input image and the second orientation is predicted for a horizontally flipped version of the input image. In at least one embodiment, the loss is calculated based on whether a certain property holds true. In at least one embodiment, a transformation or its inverse is applied to the predicted orientation of the transformed version of an input image. In at least one embodiment, if an input image is rotated by (φ, θ, ψ) angles to produce a transformed image, the inferred orientation of the transformed image may be rotated back by (-φ, -θ, -ψ) angles. In at least one embodiment, the loss is calculated by comparing the magnitude of the azimuth, elevation, and tilt of a first orientation of the input image to a second orientation of the transformed image, where equal magnitudes of the respective orientation parameters result in zero loss. In at least one embodiment, zero loss occurs when the predicted appearance parameters for an input image and a transformed version of the input image match. In at least one embodiment, the symmetry loss is calculated according to techniques described elsewhere in this disclosure, such as the technique discussed in conjunction with FIG. 16.
[0111] In at least one embodiment, a system implementing at least a portion of process 1400 includes executable code 1410 for calculating nearest and farthest neighbor losses by comparing a selected image to its nearest and farthest neighbors based, at least in part, on a viewpoint graph of the set of images. In at least one embodiment, the nearest and farthest neighbor losses are calculated according to the techniques described in conjunction with FIG. 17. In at least one embodiment, a set of images is used to generate a viewpoint graph, where nodes of such a graph correspond to images and edges correspond to their viewpoint covariance distances. In at least one embodiment, the viewpoint covariance distance is calculated based on feature similarities between pairs of images using a convolutional neural network (CNN). In at least one embodiment, an anchor image is selected from a set of training images. In at least one embodiment, the anchor image is located from the viewpoint graph, and the nearest and farthest neighbors are selected based on edge weights. In at least one embodiment, the nearest neighbor has the shortest edge leading to the anchor image. In at least one embodiment, the farthest neighbor has the furthest edge leading to the anchor image. In at least one embodiment, the neural network predicts a first viewpoint for the anchor image and a second viewpoint for the nearest neighbor image, with the loss calculated such that closer distances between the viewpoints correspond to lower loss. In at least one embodiment, the neural network predicts a first viewpoint for the anchor image and a third viewpoint for the farthest neighbor image, with the loss calculated such that greater distances between the viewpoints correspond to lower loss.
[0112] In at least one embodiment, the nearest and / or farthest neighbors are selected non-deterministically. In at least one embodiment, a probability is assigned to each edge selected as the nearest and / or farthest neighbor of the anchor image. In at least one embodiment, the probability of a nearest neighbor is inversely proportional to the edge weight (e.g., the node with the lowest edge weight leading to the anchor image has the highest probability of being selected). In at least one embodiment, the probability of a farthest neighbor is directly proportional to the edge weight (e.g., the node with the highest edge weight leading to the anchor image has the highest probability of being selected).
[0113] In at least one embodiment, a system that performs at least a portion of process 1400 includes executable code 1412 that uses the calculated losses (e.g., from figures 1406-1410) to update parameters of one or more neural networks that are being trained on a set of images. In at least one embodiment, the generators are trained for a symmetry loss, a viewpoint consistency loss, a real / fake classification loss, a disentanglement loss, or any combination thereof. In at least one embodiment, the techniques described in conjunction with Figures 4-7 are used to train the networks according to process 1400.
[0114] In at least one embodiment, a system that performs at least a portion of process 1400 includes executable code 1410 that determines whether to perform further training. In at least one embodiment, training is performed according to any suitable technique and may include selecting a second image and performing steps 1406-1412 using the second selected image to calculate losses and refine parameters of one or more neural networks being trained to infer viewpoints. Once training is complete, the trained neural network may be made available for inference (e.g., the neural network or its parameters may be transferred to a different system).
[0115] In at least one embodiment, a system that performs at least a portion of process 1400 includes executable code that receives 1416 images of the same type as a set of images used to train one or more neural networks. In at least one embodiment, the images are received from a camera or other type of capture device that is capturing images of the surroundings or environment of the system. In at least one embodiment, a system that performs at least a portion of process 1400 includes executable code that uses 1418 the trained neural network to infer viewpoints of objects in the images. In at least one embodiment, a vehicle includes a camera that captures images and provides the images to a neural network trained on a set of car images to determine whether the captured images include cars and / or the orientation of any cars included in the captured images.
[0116] FIG. 15A shows an illustrative illustration of a process 1500A for calculating image consistency loss, according to at least one embodiment. In at least one embodiment, some or all of process 1500A (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and may be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions usable to implement process 1500A are not stored using only transitory signals (e.g., propagating transient electrical or electromagnetic transmissions). A non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transient signal. In at least one embodiment, process 1500A is implemented, at least in part, on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, a first computer system calculates a generative consistency loss. In at least one embodiment, the techniques described in conjunction with FIGS. 4-7 are applicable to process 1500A. In at least one embodiment, process 1500A describes a process for calculating one or more losses (e.g., image consistency loss) that can be used to update parameters of a classifier as part of a training process.
[0117] In at least one embodiment, a system for implementing at least a portion of process 1500A includes 1502 executable code for obtaining input images of a collection of one or more images of a certain type of object. In at least one embodiment, an object of a certain type may mean that all images in the collection of images include cars. In at least one embodiment, the input images depict objects in a particular orientation and with particular appearance characteristics (e.g., attributes). In at least one embodiment, the input images are real images (e.g., as opposed to synthetic images) obtained in accordance with what was described in conjunction with FIG. 4.
[0118] In at least one embodiment, a system for performing at least a portion of process 1500A includes executable code 1504 for predicting, for an input image, a set of appearance attributes, a determination of whether the input image is real or fake, and a viewpoint using a classifier. In at least one embodiment, the classifier is associated with one or more neural networks trained to infer viewpoints and other characteristics from the input image. In at least one embodiment, the viewpoint is a predicted viewpoint of the input image, corresponds to a particular predicted orientation of an object depicted in the input image, and includes particular values of a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter. In at least one embodiment, the set of appearance attributes, or parameters, is a predicted set of appearance attributes of the input image that defines the appearance of the object depicted in the input image. In at least one embodiment, the determination of whether the input image is real or fake is a binary value (e.g., true or false) indicating whether the classifier predicted the input image as real or fake. In at least one embodiment, the predicted decision as to whether the input image is real or fake is used to calculate a real / fake classification loss, which may be in accordance with those described elsewhere in this disclosure, including but not limited to those discussed in conjunction with FIG. 4 and / or FIG. 6.
[0119] In at least one embodiment, a system for implementing at least a portion of process 1500A includes executable code for using a generator to create a synthetic image based, at least in part, on a predicted set of appearance attributes and a predicted viewpoint. In at least one embodiment, the predicted set of appearance attributes and the predicted viewpoint are provided to the generator to generate the synthetic image. In at least one embodiment, the generator is part of a generative adversarial network. In at least one embodiment, the predicted set of appearance attributes and the predicted viewpoint are utilized to generate a synthetic image according to the predicted set of appearance attributes and the predicted viewpoint. In at least one embodiment, the generator generates a synthetic image, the image being generated according to the predicted set of appearance attributes and including an object oriented according to a predicted first viewpoint.
[0120] In at least one embodiment, a system for performing at least a portion of process 1500A includes 1508 executable code for calculating an image consistency loss based on an input image and a composite image. In at least one embodiment, the input image and the composite image are compared to determine the image consistency loss. In at least one embodiment, the cosine distance between the input image and the composite image is compared to determine feature similarity, with closer similarity resulting in a lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss between the input image and the composite image.
[0121] In at least one embodiment, the nearest neighbor and farthest neighbor losses are calculated based at least in part on the input image and the synthesized image as described in conjunction with process 1500A. In at least one embodiment, the nearest neighbor and farthest neighbor losses are calculated according to the techniques described in conjunction with Figures 4 and 9. In at least one embodiment, the symmetry loss is calculated based at least in part on the input image.
[0122] In at least one embodiment, FIG. 15B illustrates a process 1500B for calculating a viewpoint consistency loss. Process 1500B may be implemented by any suitable system, such as the system described in conjunction with FIG. 5. In at least one embodiment, processes 1500A and 1500B are performed by a computer system as part of a training process that adjusts parameters of a classifier and generator used to predict an object's viewpoint. In at least one embodiment, a computer system that performs process 1500B includes executable code 1512 that causes the computer system to obtain a set of viewpoints and appearance parameters. In at least one embodiment, the set of viewpoints and / or appearance parameters is selected randomly.
[0123] In at least one embodiment, a system that performs at least a portion of process 1500B includes executable code that uses a generator to create a composite image from a set of viewpoints and appearance parameters 1514. In at least one embodiment, the composite image is created according to the techniques described above in conjunction with FIG.
[0124] In at least one embodiment, a synthetic image is created and the system is configured to use a classifier to predict a viewpoint, a determination of whether the input image is real or fake, and a set of appearance parameters 1516. In at least one embodiment, the classifier is implemented according to the techniques described in conjunction with FIG. 2 to predict the viewpoint and appearance.
[0125] In at least one embodiment, the system is configured to calculate 1518 a viewpoint consistency loss based at least in part on a predicted viewpoint (e.g., obtained from a classifier that predicts the viewpoint of the synthetic image) and an input viewpoint (e.g., the viewpoint used by the generator to create the synthetic image). In at least one embodiment, the viewpoint consistency loss is calculated according to the techniques described in conjunction with Figure 5 and / or Figure 7.
[0126] FIG. 16A shows an illustrative illustration of a process 1600A for calculating symmetry loss, according to at least one embodiment. In at least one embodiment, some or all of process 1600A (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and may be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions usable to implement process 1600A are not stored using only transitory signals (e.g., propagating transient electrical or electromagnetic transmissions). The non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transient signal. In at least one embodiment, process 1600A is implemented, at least in part, on a computer system, such as a computer system described elsewhere in this disclosure. In at least one embodiment, the first computer system calculates the symmetry loss. In at least one embodiment, the techniques described in conjunction with FIG. 8 are applicable to process 1600A.
[0127] In at least one embodiment, a system for performing at least a portion of process 1600A includes 1602 executable code for obtaining an input image. In at least one embodiment, the input image is an image of a collection of one or more images of a certain type of object. In at least one embodiment, an object of a certain type may mean that all images in the collection of images include cars. In at least one embodiment, the system obtains an input image depicting an object in a particular orientation and with particular appearance characteristics.
[0128] In at least one embodiment, a system for performing at least a portion of process 1600A includes executable code 1604 for generating a transformed image by performing a transformation on an input image. In at least one embodiment, a transformation is applied to an input image and the input image is horizontally flipped to generate the transformed image. In at least one embodiment, the transformation is applied by one or more systems associated with a classifier. In at least one embodiment, one or more image processing techniques are applied to the input image to generate the transformed image. In at least one embodiment, the azimuth and tilt angles of a viewpoint of an object in an input image are reversed when the input image is transformed to generate the transformed image, but the elevation angle remains the same for the input image and the transformed image.
[0129] In at least one embodiment, a system for performing at least a portion of process 1600A includes executable code 1606 for predicting at least a viewpoint for a transformed image. In at least one embodiment, a classifier predicts a viewpoint, a set of appearance parameters, a real / fake classification, or any combination thereof, for a transformed image. In at least one embodiment, the classifier is associated with one or more neural networks that are trained to infer viewpoints and other characteristics from input images. In at least one embodiment, the classifier receives the transformed image and predicts a viewpoint. In at least one embodiment, a viewpoint corresponds to a particular orientation of an object depicted in the image and includes particular values for a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter.
[0130] In at least one embodiment, a system that implements at least a portion of process 1600A includes executable code that applies a transformation to a predicted viewpoint 1608. In at least one embodiment, the transformation or its inverse is applied to the predicted viewpoint. In at least one embodiment, if an input image is rotated by (φ, θ, ψ) angles to produce a transformed image, a predicted viewpoint generated based on the transformed image may be rotated inversely by (−φ, −θ, −ψ) angles.
[0131] In at least one embodiment, a system for performing at least a portion of process 1600A includes executable code 1610 for predicting at least a viewpoint for an input image. In at least one embodiment, a classifier predicts a viewpoint, a set of appearance parameters, a real / fake classification, or any combination thereof for an input image. In at least one embodiment, the classifier receives an input image and predicts a viewpoint. In at least one embodiment, a viewpoint corresponds to a particular orientation of an object depicted in the image and includes particular values for a set of parameters including an azimuth parameter, an elevation parameter, and a tilt parameter.
[0132] In at least one embodiment, a system implementing at least a portion of process 1600A includes executable code 1612 for comparing viewpoints and calculating symmetry loss. In at least one embodiment, the system compares a predicted viewpoint for a transformed image to a predicted viewpoint for an input image, where the transformed image was generated from the input image. In at least one embodiment, symmetry loss is calculated by comparing the magnitude of the azimuth, elevation, and tilt of the predicted viewpoint for the input image to the magnitude of the azimuth, elevation, and tilt of the predicted viewpoint for the transformed image, where equal magnitudes of the respective viewpoint parameters result in zero loss. In at least one embodiment, symmetry loss is calculated, at least in part, based on how closely the appearance parameters of an image (e.g., an input image) and a transformed version of that image match each other.
[0133] FIG. 16B illustrates a process 1600B for calculating a symmetry loss, according to at least one embodiment. In at least one embodiment, FIG. 16B is used to update generator parameters as part of training a neural network to predict viewpoints. In at least one embodiment, a system implementing process 1600B includes executable code for obtaining 1614 a viewpoint and a set of appearance attributes. In at least one embodiment, the system is configured to apply a transformation to the obtained viewpoint to determine 1616 a transformed viewpoint. In at least one embodiment, the viewpoint is flipped to obtain the transformed viewpoint. In at least one embodiment, a transformation function T() is applied to the set of parameters x1, y1, and z1 to obtain a set of transformed parameters x2, y2, z2, which may correspond to azimuth, tilt, and elevation parameters.
[0134] In at least one embodiment, a generator is used to generate 1618 a first composite image based at least in part on a set of transformed viewpoints and appearance parameters, which may be calculated or determined using techniques described in conjunction with other steps in FIG. 16B. In at least one embodiment, the system is configured to apply 1620 a transformation to the first composite image generated from the transformed viewpoint and set of appearance attributes. In at least one embodiment, if the transformed viewpoint is created by applying a transformation T(), the composite image is subjected to an inverse transformation T(). -1 () is applied, in this case T(T -1 (x,y,z))=(x,y,z). In at least one embodiment, if the transformation flips the image horizontally, then the inversion flips the image horizontally back to its original orientation.
[0135] In at least one embodiment, the system includes executable instructions for generating 1622 a second composite image based at least in part on the viewpoint and the set of appearance parameters. In at least one embodiment, the same generator is used to create the first composite image based on the transformed viewpoint and the second composite image based on the original viewpoint. In at least one embodiment, the system is configured to compare the first and second composite images and calculate 1624 a symmetry loss, where a cosine distance of zero between such images means zero loss.
[0136] FIG. 17 shows an illustrative illustration of a process 1700 for calculating nearest neighbor and farthest neighbor losses, according to at least one embodiment. In at least one embodiment, some or all of process 1700 (or any other process described herein, or variations and / or combinations thereof) runs under the control of one or more computer systems configured with computer-executable instructions and may be implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some computer-readable instructions usable to implement process 1700 are not stored using only transitory signals (e.g., propagating transient electrical or electromagnetic transmissions). The non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of a transient signal. In at least one embodiment, process 1700 is implemented, at least in part, on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, the first computer system calculates nearest neighbor and furthest neighbor losses. In at least one embodiment, the techniques described in conjunction with FIG. 9 are applicable to process 1700.
[0137] In at least one embodiment, a system for implementing at least a portion of process 1700 includes executable code for obtaining 1702 a collection of images. In at least one embodiment, the collection of images includes one or more images of a certain type of object. In at least one embodiment, a certain type of object may mean that all images in the collection of images include cars. In at least one embodiment, the system obtains the one or more collections of images according to techniques described elsewhere in this disclosure, such as in FIG. 13 .
[0138] In at least one embodiment, a system implementing at least a portion of process 1700 includes executable code to calculate 1704 a cosine distance for comparing feature similarity for pairs of images in a set. In at least one embodiment, cosine distance is a mathematical component of cosine similarity (e.g., cosine distance = 1 - cosine similarity). In at least one embodiment, cosine similarity is a measure of similarity between two vectors, which may represent images, text, data, and / or variations thereof, and is based on the cosine of the angle between the vectors. In at least one embodiment, if two images contain features corresponding to objects with similar viewpoints, the calculated cosine distance for the images is small. In at least one embodiment, if two images contain features corresponding to objects with different viewpoints, the calculated cosine distance for the images is large. In at least one embodiment, the cosine distance is calculated for each image in a set of images relative to each other image in the set.
[0139] In at least one embodiment, a system for performing at least a portion of process 1700 includes executable code for generating 1706 a viewpoint graph, where nodes of the graph correspond to images of the set and edges correspond to their cosine distances. In at least one embodiment, the cosine distances are calculated based on feature similarities between pairs of images using a convolutional neural network. In at least one embodiment, edges of the viewpoint graph are weighted such that a thick edge between two images corresponds to a high degree of similarity between the two images and a thin edge between two images corresponds to a low degree of similarity between the two images.
[0140] In at least one embodiment, a system for performing at least a portion of process 1700 includes executable code 1708 for selecting an anchor image for a graph to predict a first viewpoint. In at least one embodiment, the anchor image is selected from a set of training images used to generate the viewpoint graph. In at least one embodiment, the classifier is associated with one or more neural networks that are trained to infer viewpoints and other characteristics from input images. In at least one embodiment, the classifier receives the anchor image and predicts a first viewpoint. In at least one embodiment, the first viewpoint corresponds to a particular predicted orientation of an object depicted in the anchor image and includes particular values of a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter, corresponding to the orientation of the object.
[0141] In at least one embodiment, a system for performing at least a portion of process 1700 includes executable code 1710 for selecting nearest neighbors of an anchor image using a graph to predict a second viewpoint. In at least one embodiment, the nearest neighbors are determined based on edge weights of the viewpoint graph. In at least one embodiment, the nearest neighbors are the images in the viewpoint graph with the shortest edge connecting to the anchor image. In at least one embodiment, the nearest neighbors to an anchor image are the images in a collection of images that are most similar to the anchor image. In at least one embodiment, a classifier receives the nearest neighbor images and predicts a second viewpoint. In at least one embodiment, the second viewpoint corresponds to a particular predicted orientation of an object depicted in the nearest neighbor image and includes particular values of a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter, corresponding to the orientation of the object.
[0142] In at least one embodiment, a system for performing at least a portion of process 1700 includes executable code 1712 for selecting a farthest neighbor of the anchor image using the graph to predict a third viewpoint. In at least one embodiment, the farthest neighbor is determined based on edge weights in the viewpoint graph. In at least one embodiment, the farthest neighbor is the image in the viewpoint graph with the longest edge connecting to the anchor image. In at least one embodiment, the farthest neighbor to the anchor image is the image in the set of images that is most different from the anchor image. In at least one embodiment, a classifier receives the farthest neighbor image and predicts a third viewpoint. In at least one embodiment, the third viewpoint corresponds to a particular predicted orientation of an object depicted in the farthest neighbor image and includes particular values of a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter, corresponding to the orientation of the object.
[0143] In at least one embodiment, a system implementing at least a portion of process 1700 includes executable code 1714 for calculating nearest and farthest neighbor losses. In at least one embodiment, the nearest neighbor loss is calculated between a first viewpoint predicted from an anchor image and a second viewpoint predicted from a nearest neighbor image relative to the anchor image. In at least one embodiment, the nearest neighbor loss is calculated such that a higher similarity between the first viewpoint and the second viewpoint corresponds to a lower loss. In at least one embodiment, the farthest neighbor loss is calculated between a first viewpoint predicted from the anchor image and a third viewpoint predicted from a farthest neighbor image relative to the anchor image. In at least one embodiment, the farthest neighbor loss is calculated such that a lower similarity between the first viewpoint and the third viewpoint corresponds to a lower loss. In at least one embodiment, the nearest and farthest neighbor losses are calculated based on a combination of the nearest neighbor loss and the farthest neighbor loss.
[0144] Logic of inference and training Figure 18A illustrates inference and / or training logic 1815 used to perform inference and / or training operations for one or more embodiments. More details regarding inference and / or training logic 1815 are provided below in conjunction with Figures 18A and / or 18B.
[0145] In at least one embodiment, the inference and / or training logic 1815 may include, without limitation, code and / or data storage 1801 for storing forward and / or output weights, and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network that are trained and / or used to infer in one or more embodiments. In at least one embodiment, the training logic 1815 may include or be coupled to code and / or data storage 1801 for storing graph code or other software for controlling the timing and / or sequence of logic loaded with weights and / or other parameter information, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weights or other parameter information into processor ALUs based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data storage 1801 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during forward propagation of the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 1801 may be included with other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache, or system memory.
[0146] In at least one embodiment, any portion of code and / or data storage 1801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 1801 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether code and / or data storage 1801 is internal or external to a processor, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, may depend on the available storage on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in neural network inference and / or training, or any combination of these factors.
[0147] In at least one embodiment, the inference and / or training logic 1815 may include, without limitation, code and / or data storage 1805 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used to infer in accordance with one or more aspects of the embodiment. In at least one embodiment, the code and / or data storage 1805 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments while backpropagating input / output data and / or weight parameters during training and / or inference using one or more aspects of the embodiment. In at least one embodiment, training logic 1815 may include or be coupled to code and / or data storage 1805 for storing graph code or other software for timing and / or sequencing control, where code and / or data storage 1801 is loaded with weights and / or other parameter information to configure logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code, such as graph code, loads weights or other parameter information into processor ALUs based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of code and / or data storage 1805 may be included with other on-chip or off-chip data storage, including processor L1, L2, or L3 cache, or system memory. In at least one embodiment, any portion of code and / or data storage 1805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 1805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether code and / or data storage 1805 is internal or external to the processor, for example, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, may depend on the storage available on-chip versus off-chip, the latency requirements of the training and / or inference functions being performed, the batch size of data used in neural network inference and / or training, or any combination of these factors.
[0148] In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be separate storage structures. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be the same storage structure. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 1801 and code and / or data storage 1805 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0149] In at least one embodiment, inference and / or training logic 1815 may include one or more arithmetic logic units (“ALUs”) 1810, including, without limitation, integer and / or floating point units, for performing logical and / or arithmetic operations based at least in part on or indicated by training and / or inference code (e.g., graph code), the results of which may generate activations (e.g., output values from layers or neurons in a neural network) stored in activation storage 1820, which are functions of input / output and / or weight parameter data stored in code and / or data storage 1801 and / or code and / or data storage 1805. In at least one embodiment, the activations stored in activation storage 1820 are generated according to linear algebra and / or matrix-based calculations performed by ALU 1810 in response to executing instructions or other code, where weight values stored in code and / or data storage 1805 and / or data 1801 are used as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 1805, or code and / or data storage 1801, or in separate storage, on-chip or off-chip.
[0150] In at least one embodiment, ALU 1810 is included within one or more processors or other hardware logic devices or circuits, while in other embodiments, ALU 1810 may be external to the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, ALU 1810 may be included within an execution unit of a processor or may otherwise be included within an ALU bank accessible by execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing unit, graphics processing unit, fixed function unit, etc.). In at least one embodiment, data storage 1801, code and / or data storage 1805, and activation storage 1820 may be in the same processor or other hardware logic devices or circuits, while in other embodiments, they may be in different processors or other hardware logic devices or circuits, or some combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 1820 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic.
[0151] In at least one embodiment, active storage 1820 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, active storage 1820 may be completely or partially internal to or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether active storage 1820 is internal or external to a processor, or whether it is comprised of DRAM, SRAM, flash, or some other type of storage, for example, may depend on available on-chip versus off-chip storage, latency requirements of the training and / or inference functions being performed, batch sizes of data used in neural network inference and / or training, or any combination of these factors. In at least one embodiment, the inference and / or training logic 1815 shown in Figure 18A may be used in conjunction with an application-specific integrated circuit ("ASIC"), such as a Tensorflow® processing unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® (e.g., "Lake Crest") processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 1815 shown in Figure 18A may be used in conjunction with other hardware, such as central processing unit ("CPU") hardware, graphics processing unit ("GPU") hardware, or a field programmable gate array ("FPGA").
[0152] FIG. 18B illustrates inference and / or training logic 1815 according to at least one various embodiment. In at least one embodiment, the inference and / or training logic 1815 may include, without limitation, hardware logic in which computational resources are dedicated to, or otherwise used only in conjunction with, weight values or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 1815 illustrated in FIG. 18B may be used in conjunction with an application-specific integrated circuit (ASIC), such as a Tensorflow® processing unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 1815 illustrated in FIG. 18B may be used in conjunction with other hardware, such as central processing unit (CPU) hardware, graphics processing unit (“GPU”) hardware, or a field-programmable gate array (FPGA). In at least one embodiment, inference and / or training logic 1815 includes, without limitation, code and / or data storage 1801 and code and / or data storage 1805, which may be used to store code (e.g., graph code), weight and / or bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment shown in FIG. 18B , each of code and / or data storage 1801 and code and / or data storage 1805 is associated with dedicated computational resources, such as compute hardware 1802 and compute hardware 1806, respectively. In at least one embodiment, compute hardware 1802 and compute hardware 1806 each include one or more ALUs that perform mathematical functions, such as linear algebraic functions, solely on the information stored in code and / or data storage 1801 and code and / or data storage 1805, respectively, with the results stored in activation storage 1820.
[0153] In at least one embodiment, each of code and / or data storage 1801 and 1805 and corresponding computational hardware 1802 and 1806 corresponds to a different layer of a neural network, whereby activations resulting from one "storage / computation pair 1801 / 1802" of code and / or data storage 1801 and computational hardware 1802 are provided as input to the next "storage / computation pair 1805 / 1806" of code and / or data storage 1805 and computational hardware 1806 to reflect the conceptual organization of the neural network. In at least one embodiment, storage / computation pairs 1801 / 1802 and 1805 / 1806 may correspond to two or more layers of the neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in inference and / or training logic 1815 after or in parallel with storage / computation pairs 1801 / 1802 and 1805 / 1806.
[0154] Neural network training and deployment FIG. 19 illustrates training and deployment of a deep neural network, according to at least one embodiment. In at least one embodiment, an untrained neural network 1906 is trained using a training dataset 1902. In at least one embodiment, the training framework 1904 is the PyTorch framework, while in other embodiments, the training framework 1904 is Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 and enables it to be trained using processing resources described herein to generate a trained neural network 1908. In at least one embodiment, the weights may be selected randomly or by pre-training using a deep belief network. In at least one embodiment, the training may be performed in a supervised, semi-supervised, or unsupervised manner.
[0155] In at least one embodiment, the untrained neural network 1906 is trained using supervised learning, where the training dataset 1902 includes inputs paired with desired outputs, or the training dataset 1902 includes inputs with known outputs, and the outputs of the neural network 1906 are manually scored. In at least one embodiment, the untrained neural network 1906 is trained in a supervised manner, processing inputs from the training dataset 1902 and comparing the resulting outputs to a set of expected or desired outputs. In at least one embodiment, errors are then back-propagated through the untrained neural network 1906. In at least one embodiment, the training framework 1904 adjusts the weights that control the untrained neural network 1906. In at least one embodiment, the training framework 1904 includes tools to monitor how well the untrained neural network 1906 is converging toward a model, such as the trained neural network 1908, that is suitable for generating correct answers, such as in the results 1914, based on known input data, such as the new dataset 1912. In at least one embodiment, the training framework 1904 iteratively trains the untrained neural network 1906 while adjusting weights using a loss function and a tuning algorithm, such as stochastic gradient descent, to refine the output of the untrained neural network 1906. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 until the untrained neural network 1906 reaches a desired accuracy. In at least one embodiment, the trained neural network 1908 can then be deployed to implement any number of machine learning operations.
[0156] In at least one embodiment, the untrained neural network 1906 is trained using unsupervised learning, where the untrained neural network 1906 attempts to train itself using unlabeled data. In at least one embodiment, the training dataset 1902 for unsupervised learning includes input data without any associated output data or “ground truth” data. In at least one embodiment, the untrained neural network 1906 can learn groupings within the training dataset 1902 and determine how individual inputs relate to the untrained dataset 1902. In at least one embodiment, unsupervised training can be used to generate self-organizing maps, which are a type of trained neural network 1908 that can perform operations useful for reducing the dimensionality of the new dataset 1912. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for the identification of data points in the new dataset 1912 that deviate from the normal patterns of the new dataset 1912.
[0157] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled and unlabeled data are mixed in the training dataset 1902. In at least one embodiment, the training framework 1904 may be used to perform incremental learning, such as by transfer learning techniques. In at least one embodiment, incremental learning allows the trained neural network 1908 to adapt to a new dataset 1912 without forgetting the knowledge instilled in the network during initial training.
[0158] Data Center 20 illustrates an exemplary data center 2000 in which at least one embodiment may be used. In at least one embodiment, the data center 2000 includes a data center infrastructure layer 2010, a framework layer 2020, a software layer 2030, and an application layer 2040.
[0159] 20, data center infrastructure layer 2010 may include a resource orchestrator 2012, grouped computing resources 2014, and node computing resources (“node CRs”) 2016(1) through 2016(N), where “N” represents any positive integer. In at least one embodiment, node CRs 2016(1) through 2016(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules. In at least one embodiment, one or more of the nodes CR 2016(1) to 2016(N) may be a server having one or more of the computing resources described above.
[0160] In at least one embodiment, grouped computing resources 2014 may include separate groups of node CRs housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). Separate groups of node CRs within grouped computing resources 2014 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide compute resources to support one or more workloads. In at least one embodiment, one or more racks may also include any number of power supply modules, cooling modules, and network switches in any combination.
[0161] In at least one embodiment, resource orchestrator 2012 may configure or otherwise control one or more nodes CR 2016(1)-2016(N) and / or grouped computing resources 2014. In at least one embodiment, resource orchestrator 2012 may include a software design infrastructure (“SDI”) management entity for data center 2000. In at least one embodiment, resource orchestrator may include hardware, software, or some combination thereof.
[0162] In at least one embodiment shown in FIG. 20 , framework layer 2020 includes a job scheduler 2032, a configuration manager 2034, a resource manager 2036, and a distributed file system 2038. In at least one embodiment, framework layer 2020 may include a framework for supporting software 2032 in software layer 2030 and / or one or more applications 2042 in application layer 2040. In at least one embodiment, software 2032 or applications 2042 may each include web-based service software or applications, such as those offered by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 2020 may be a type of free and open-source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter “Spark”), which can utilize distributed file system 2038 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 2032 may include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 2000. In at least one embodiment, configuration manager 2034 may be capable of configuring different tiers, such as software tier 2030 and framework tier 2020, which includes Spark and distributed file system 2038 to support large-scale data processing. In at least one embodiment, resource manager 2036 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support distributed file system 2038 and job scheduler 2032. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 2014 in data center infrastructure tier 2010.In at least one embodiment, resource manager 2036 may manage these mappings or allocated computing resources in conjunction with resource orchestrator 2012.
[0163] In at least one embodiment, software 2032 included in software layer 2030 may include software used by nodes CR 2016(1)-2016(N), grouped computing resources 2014, and / or at least a portion of distributed file system 2038 of framework layer 2020. The one or more types of software may include, but are not limited to, internet web page searching software, email virus scanning software, database software, and streaming video content software.
[0164] In at least one embodiment, the applications 2042 included in the application layer 2040 may include one or more types of applications used by at least a portion of the nodes CR 2016(1)-2016(N), the grouped computing resources 2014, and / or the distributed file system 2038 of the framework layer 2020. The one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive compute, and machine learning applications including training or inference software, machine learning framework software (e.g., PyTorch, Tensorflow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0165] In at least one embodiment, any of configuration manager 2034, resource manager 2036, and resource orchestrator 2012 may implement any number and types of self-correcting actions based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-correcting actions may enable a data center operator of data center 2000 to avoid determining potentially faulty configurations and eliminate underutilized and / or underperforming portions of the data center.
[0166] In at least one embodiment, data center 2000 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, machine learning models may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to data center 2000. In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to data center 2000 by using weight parameters calculated by one or more techniques described herein.
[0167] In at least one embodiment, the data center may use a CPU, application specific integrated circuit (ASIC), GPU, FPGA, or other hardware to perform training and / or inference using the resources described above. Additionally, one or more of the software and / or hardware resources described above may be configured as a service to enable a user to train or perform inference on information, such as image recognition, speech recognition, or other artificial intelligence services.
[0168] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 20 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0169] In at least one embodiment, the system diagram 20 is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system diagram 20 is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, the system diagram 20 is utilized to implement one or more neural networks including the classifier and generator, and the system diagram 20 is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more properties of the images in a training set.
[0170] Autonomous Vehicles 21A illustrates an example of an autonomous vehicle 2100 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 2100 (alternatively referred to herein as "vehicle 2100") may be a passenger vehicle, such as, without limitation, a car, truck, bus, and / or another type of vehicle that accommodates one or more occupants. In at least one embodiment, the vehicle 2100 may be a semi-tractor trailer truck for transporting cargo. In at least one embodiment, the vehicle 2100 may be an aircraft, a robotic vehicle, or other type of vehicle.
[0171] Autonomous vehicles may be described in terms of levels of automation as defined by the National Highway Traffic Safety Administration (“NHTSA”), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (“SAE”) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., Standard No. J3016-201806, issued June 15, 2018, Standard No. J3016-201609, issued September 30, 2016, and previous and new versions of this standard). In one or more embodiments, vehicle 2100 may be capable of functionality according to one or more of Levels 1 through 5 of autonomous driving. For example, in at least one embodiment, vehicle 2100 may be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.
[0172] In at least one embodiment, vehicle 2100 may include components such as, without limitation, a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. In at least one embodiment, vehicle 2100 may include a propulsion system 2150 such as, without limitation, an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, propulsion system 2150 may be coupled to a drive train of vehicle 2100, which may include, without limitation, a transmission to enable propulsion of vehicle 2100. In at least one embodiment, propulsion system 2150 may be controlled in response to receiving a signal from throttle / accelerator 2152.
[0173] In at least one embodiment, a steering system 2154, which may include without limitation a steering wheel, is used to steer the vehicle 2100 (e.g., along a desired path or route) when the propulsion system 2150 is operating (e.g., when the vehicle is moving). In at least one embodiment, the steering system 2154 may receive signals from a steering actuator 2156. The steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, a brake sensor system 2146 may be used to operate the vehicle brakes in response to receiving signals from a brake actuator 2148 and / or brake sensor.
[0174] In at least one embodiment, controller 2136, which may include, without limitation, one or more systems on chip (“SoC”) (not shown in FIG. 21A ) and / or graphics processing units (“GPUs”), provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 2100. For example, in at least one embodiment, controller 2136 may send signals to operate vehicle brakes via brake actuators 2148, to operate steering system 2154 via steering actuators 2156, and to operate propulsion system 2150 via throttle / accelerator 2152. Controller 2136 may include one or more on-board (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 2100. In at least one embodiment, the controllers 2136 may include a first controller 2136 for autonomous driving functions, a second controller 2136 for functional safety functions, a third controller 2136 for artificial intelligence functions (e.g., computer vision), a fourth controller 2136 for infotainment functions, a fifth controller 2136 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 2136 may handle two or more of the above functionalities, two or more controllers 2136 may handle a single functionality, and / or some combination thereof.
[0175] In at least one embodiment, controller 2136 provides signals to control one or more components and / or systems of vehicle 2100 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, sensor data may be received from, for example, without limitation, global navigation satellite system (“GNSS”) sensors 2158 (e.g., global positioning system sensors), RADAR sensors 2160, ultrasonic sensors 2162, LIDAR sensors 2164, inertial measurement units (“IMUs”), or other sensors. 21A ), a speed sensor 2144 (e.g., for measuring the speed of the vehicle 2100), a vibration sensor 2142, a steering sensor 2140, a brake sensor (e.g., as part of a brake sensor system 2146), and / or other types of sensors.
[0176] In at least one embodiment, one or more of the controllers 2136 may receive input (e.g., represented by input data) from the instrument cluster 2132 of the vehicle 2100 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 2134, an audible annunciator, a loudspeaker, and / or via other components of the vehicle 2100. In at least one embodiment, the output may include information such as vehicle speed, speeding, time, map data (e.g., a high definition map (not shown in FIG. 21A )), location data (e.g., the location of the vehicle 2100 on a map, etc.), direction, the location of other vehicles (e.g., an occupancy grid), information about objects and object conditions sensed by the controller 2136, etc. For example, in at least one embodiment, the HMI display 2134 may display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about a driving maneuver that the vehicle has made, is making, or will make (e.g., currently changing lanes, taking exit 34B in 2 miles, etc.).
[0177] In at least one embodiment, vehicle 2100 further includes network interface 2124, which may use a wireless antenna 2126 and / or a modem for communicating over one or more networks. For example, in at least one embodiment, network interface 2124 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. Additionally, in at least one embodiment, wireless antenna 2126 may enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low power wide-area networks ("LPWAN") such as LoRaWAN, SigFox, etc.
[0178] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 21A for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0179] In at least one embodiment, system FIG. 21A is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, system FIG. 21A is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, system FIG. 21A is utilized to implement one or more neural networks that include a classifier and a generator, and system FIG. 21A is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more characteristics of the images in a training set.
[0180] 21B illustrates example camera locations and fields of view for the autonomous vehicle 2100 of FIG. 21A, according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are an example example and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or cameras may be positioned at different locations on the vehicle 2100.
[0181] In at least one embodiment, the camera type may include, but is not limited to, a digital camera that may be adapted for use with components and / or systems of vehicle 2100. The camera may operate at Automotive Safety Integrity Level (“ASIL”) B and / or another ASIL. In at least one embodiment, the camera type may be capable of any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red, clear, clear, clear ("RCCC") color filter array, a red, clear, clear, blue ("RCCB") color filter array, a red, blue, green, clear ("RBGC") color filter array, a Foveon X3 color filter array, a Bayer sensor (RGGB) color filter array, a monochrome sensor color filter array, and / or another type of color filter array. In at least one embodiment, a clear pixel camera may be used, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, to increase light sensitivity.
[0182] In at least one embodiment, one or more of the cameras may be used to perform advanced driver assistance systems ("ADAS") functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).
[0183] In at least one embodiment, one or more of the cameras may be mounted on a mounting assembly, such as a custom-designed (e.g., three-dimensionally (“3D”) printed) assembly, to eliminate stray light and reflections from the interior of the vehicle (e.g., reflections from the dashboard reflected in the rearview mirror) that may interfere with the camera's image data capture capabilities. With reference to door mirror mounting assemblies, in at least one embodiment, the door mirror assembly may be custom 3D printed so that the camera mounting plate matches the shape of the door mirror. In at least one embodiment, the camera may be integral with the door mirror. For side view cameras, in at least one embodiment, the cameras may again be integrated into the four pillars at each corner of the cabin.
[0184] In at least one embodiment, a camera (e.g., a front-facing camera) having a field of view that includes a portion of the environment ahead of the vehicle 2100 may be used for a surroundings view to facilitate identification of the path and obstacles ahead and, in conjunction with the controller 2136 and / or one or more of the control SoCs, may assist in providing information essential for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the front-facing camera may be used to perform many of the same ADAS functions as LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the front-facing camera may also be used for ADAS features and systems, including, without limitation, other features such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.
[0185] In at least one embodiment, various cameras may be used in a front-facing configuration, including, for example, a monocular camera platform including a CMOS (complementary metal oxide semiconductor) color imager. In at least one embodiment, a wide-angle camera 2170 may be used to sense objects (e.g., pedestrians, cross traffic, or bicycles) coming into view from the periphery. While FIG. 21B shows only one wide-angle camera 2170, in other embodiments, there may be any number (including zero) of wide-angle cameras 2170 on the vehicle 2100. In at least one embodiment, any number of long-range cameras 2198 (e.g., a pair of long-view stereo cameras) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the long-range cameras 2198 may also be used for object detection and classification, as well as basic object tracking.
[0186] In at least one embodiment, any number of stereo cameras 2168 may also be included in a front-facing configuration. In at least one embodiment, one or more stereo cameras 2168 may include an integrated control unit with a scalable processing unit, which may provide a programmable gate array ("FPGA") and a multi-core microprocessor with an integrated controller area network ("CAN") or Ethernet interface on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the vehicle's 2100 environment, including distance estimates for all points in the image. In at least one embodiment, one or more of the stereo cameras 2168 may include, without limitation, a compact stereo vision sensor, which may include, without limitation, two camera lenses (one on each side) and an image processing chip that can measure the distance from the vehicle 2100 to target objects and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning features. In at least one embodiment, other types of stereo cameras 2168 may be used in addition to or instead of those described herein.
[0187] In at least one embodiment, cameras having a field of view that includes a portion of the environment to the sides of the vehicle 2100 (e.g., side view cameras) may be used for the surrounding view to provide information used to create and update the occupancy grid and generate side collision warnings. For example, in at least one embodiment, surrounding cameras 2174 (e.g., four surrounding cameras 2174 as shown in FIG. 21B ) may be disposed on the vehicle 2100. The surrounding cameras 2174 may include, without limitation, any number and combination of wide-angle cameras 2170, fisheye cameras, and / or 360-degree cameras. For example, in at least one embodiment, four fisheye cameras may be disposed in front, behind, and on the sides of the vehicle 2100. In at least one embodiment, the vehicle 2100 may use three surrounding cameras 2174 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a front camera) as a fourth surrounding camera.
[0188] In at least one embodiment, a camera having a field of view that includes a portion of the environment behind the vehicle 2100 (e.g., a rear view camera) may be used for parking assistance, surrounding view, rear collision warning, and to create and update the occupancy grid. In at least one embodiment, a variety of cameras may be used, including, but not limited to, cameras also suitable as front cameras described herein (e.g., long-range camera 2198, and / or mid-range camera 2176, stereo camera 2168, infrared camera 2172, etc.).
[0189] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 21B for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0190] In at least one embodiment, system FIG. 21B is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, system FIG. 21B is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, system FIG. 21B is utilized to implement one or more neural networks that include a classifier and a generator, and system FIG. 21B is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more properties of the images in a training set.
[0191] FIG. 21C is a block diagram illustrating an example system architecture for the autonomous vehicle 2100 of FIG. 21A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 2100 of FIG. 21C is shown as connected via a bus 2102. In at least one embodiment, the bus 2102 may include, without limitation, a CAN data interface (alternatively referred to herein as a (CAN bus)). In at least one embodiment, the CAN may be a network internal to the vehicle 2100 used to assist in controlling various features and functions of the vehicle 2100, such as brake application, acceleration, brake control, steering, windshield wipers, etc. In at least one embodiment, the bus 2102 may be configured to have tens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, the bus 2102 may be read to determine steering wheel angle, ground speed, engine revolutions per minute (“RPM”), button position, and / or other vehicle status indicators. In at least one embodiment, bus 2102 may be an ASIL B compliant CAN bus.
[0192] In at least one embodiment, FlexRay and / or Ethernet may be used in addition to or instead of CAN. In at least one embodiment, there may be any number of buses 2102, including, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 2102 may be used to perform different functions and / or to provide redundancy. For example, a first bus 2102 may be used for collision avoidance functions and a second bus 2102 may be used for actuation control. In at least one embodiment, each bus 2102 may communicate with any of the components of the vehicle 2100, or two or more buses 2102 may communicate with the same component. In at least one embodiment, each of any number of systems on a chip ("SoC") 2104, each of the controllers 2136, and / or each computer in the vehicle may have access to the same input data (e.g., input from sensors in the vehicle 2100) and may be connected to a common bus, such as a CAN bus.
[0193] In at least one embodiment, vehicle 2100 may include one or more controllers 2136, such as those described herein with respect to FIG. 21A. Controller 2136 may be used for a variety of functions. In at least one embodiment, controller 2136 may be coupled to any of a variety of other components and systems of vehicle 2100 and may be used to control vehicle 2100, artificial intelligence of vehicle 2100, and / or infotainment of vehicle 2100, etc.
[0194] In at least one embodiment, the vehicle 2100 may include any number of SoCs 2104. Each of the SoCs 2104 may include, without limitation, a central processing unit ("CPU") 2106, a graphics processing unit ("GPU") 2108, a processor 2110, a cache 2112, an accelerator 2114, a data store 2116, and / or other components and features not shown. In at least one embodiment, the SoCs 2104 may be used to control the vehicle 2100 in a variety of platforms and systems. For example, in at least one embodiment, the SoC 2104 may be incorporated into a system (e.g., that of the vehicle 2100) having a high definition ("HD") map 2122 that can obtain map refreshes and / or updates via a network interface 2124 from one or more servers (not shown in FIG. 21C ).
[0195] In at least one embodiment, CPU2106 may include a CPU cluster, or CPU complex (also referred to herein as a "CCPLEX"). In at least one embodiment, CPU2106 may include multiple cores and / or level 2 ("L2") caches. For example, in at least one embodiment, CPU2106 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, CPU2106 may include four dual-core clusters, where each cluster has a dedicated L2 cache (e.g., 2 MB of L2 cache). In at least one embodiment, CPU2106 (e.g., a CCPLEX) may be configured to support simultaneous cluster operation, allowing any combination of CPU2106 clusters to be active at any given time.
[0196] In at least one embodiment, one or more of the CPUs 2106 may implement power management functionality, including, without limitation, one or more of the following features: individual hardware blocks may be automatically clock gated when idle to save dynamic power; each core clock may be gated when the core is not actively executing instructions due to execution of a Wait for Interrupt ("WFI") / Wait for Event ("WFE") instruction; each core may be independently power gated; when all cores are clock gated or power gated, each core cluster may be independently clock gated; and / or when all cores are power gated, each core cluster may be independently power gated. In at least one embodiment, the CPUs 2106 may further implement an advanced algorithm for managing power states, where, given allowed power states and expected wake-up times, hardware / microcode determines the best power state for cores, clusters, and CCPLEXes to enter. In at least one embodiment, a processing core may support in software a simple sequence of entering power states, with work offloaded to microcode.
[0197] In at least one embodiment, GPU2108 may include an integrated GPU (alternatively referred to herein as an “iGPU”). In at least one embodiment, GPU2108 may be programmable and efficient for parallel workloads. In at least one embodiment, GPU2108 may use an extended tensor instruction set. In one embodiment, GPU2108 may include one or more streaming microprocessors, where each streaming microprocessor may include a level 1 (“L1”) cache (e.g., an L1 cache having at least 96 KB of storage capacity) and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache having 512 KB of storage capacity). In at least one embodiment, GPU2108 may include at least eight streaming microprocessors. In at least one embodiment, GPU2108 may use a compute application programming interface (API). In at least one embodiment, GPU 2108 may use one or more parallel computing platforms and / or programming modules (e.g., NVIDIA's CUDA).
[0198] In at least one embodiment, one or more of the GPUs 2108 may be power-optimized for best performance in automotive and embedded use cases. For example, in one embodiment, the GPUs 2108 may be fabricated on fin field-effect transistors ("FinFETs"). In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores partitioned into multiple blocks. For example, without limitation, 64 PF32 cores and 32 PF64 cores may be partitioned into four processing blocks. In at least one embodiment, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR cores for deep learning matrix operations, a level-zero ("L0") instruction cache, a warp scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor includes independent parallel integer and floating-point data paths to achieve efficient execution of workloads by mixing computational and addressing calculations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling to enable finer-grained synchronization and coordination between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance while simplifying programming.
[0199] In at least one embodiment, one or more of the GPUs 2108 may include high bandwidth memory ("HBM") and / or a 16 GB HBM2 memory subsystem, providing, in some instances, a peak memory bandwidth of approximately 900 GB / s. In at least one embodiment, synchronous graphics random-access memory ("SGRAM"), such as graphics double data rate type five ("GDDR5"), may be used in addition to or in place of the HBM memory.
[0200] In at least one embodiment, the GPU 2108 may include unified memory technology. In at least one embodiment, address translation services ("ATS") support may be used to allow the GPU 2108 to directly access the CPU 2106 page tables. In at least one embodiment, when the GPU 2108 memory management unit ("MMU") encounters a miss, an address translation request may be sent to the CPU 2106. In at least one embodiment, in response, the CPU 2106 may look up the virtual-to-physical address mapping in its page table and send the translation back to the GPU 2108. In at least one embodiment, the unified memory technology allows for a single, unified virtual address space for both the CPU 2106 and the GPU 2108 memory, thereby simplifying programming the GPU 2108 and porting applications to the GPU 2108.
[0201] In at least one embodiment, GPU 2108 may include any number of access counters that can record the frequency of GPU 2108's accesses to the memory of other processors. In at least one embodiment, the access counters may help ensure that memory pages are moved to the physical memory of the processor that is accessing the pages most frequently, thereby improving the efficiency of memory ranges shared between processors.
[0202] In at least one embodiment, one or more of the SoCs 2104 may include any number of caches 2112, including those described herein. For example, in at least one embodiment, the caches 2112 may include a level 3 (“L3”) cache available to both the CPU 2106 and the GPU 2108 (e.g., connected to both the CPU 2106 and the GPU 2108). In at least one embodiment, the caches 2112 may include a write-back cache that can record line state by using a cache coherence protocol or the like (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4 MB or more, depending on the embodiment, although smaller cache sizes may also be used.
[0203] In at least one embodiment, one or more of the SoCs 2104 may include one or more accelerators 2114 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoCs 2104 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, the large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other calculations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU 2108 and offload some of the GPU 2108's tasks (e.g., freeing up more cycles for the GPU 2108 to perform other tasks). In at least one embodiment, accelerator 2114 may be used for targeted workloads that are stable enough to accommodate acceleration (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.). In at least one embodiment, CNNs may include region-based, i.e., regional convolutional neural networks (“RCNNs”), and Fast RCNNs (e.g., used for object detection), or other types of CNNs.
[0204] In at least one embodiment, the accelerator 2114 (e.g., a hardware-accelerated cluster) may include a deep learning accelerator (“DLA”). The DLA may include, without limitation, one or more tensor processing units (“TPU”), which may be further configured to provide tens of trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPU may be an accelerator configured and optimized to perform image processing functions (e.g., CNN, RCNN, etc.). The DLA may further be optimized for a specific set of neural network types and floating-point operations, as well as for inference. In at least one embodiment, the DLA's design allows for improved performance per millimeter over typical general-purpose GPUs, typically significantly exceeding the performance of a CPU. In at least one embodiment, the TPU may perform several functions, including, for example, single-instance convolution functions supporting INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions. In at least one embodiment, the DLA may quickly and efficiently run neural networks, particularly CNNs, on processed or unprocessed data for any of a variety of functions, including, for example, without limitation, a CNN for object identification and detection using data from a camera sensor, a CNN for distance estimation using data from a camera sensor, a CNN for emergency vehicle detection and identification using data from microphone 2196, a CNN for face recognition and vehicle owner identification using data from a camera sensor, and / or a CNN for security and / or safety related events.
[0205] In at least one embodiment, the DLA may perform any function of the GPU 2108, and by using, for example, an inference accelerator, a designer may target either the DLA or the GPU 2108 for any function. For example, in at least one embodiment, a designer may centralize CNN and floating-point processing in the DLA and offload other functions to the GPU 2108 and / or other accelerators 2114.
[0206] In at least one embodiment, the accelerator 2114 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (“PVA”), which may alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (“ADAS”) 2138, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. The PVA may balance performance and versatility. For example, in at least one embodiment, each PVA may include, by way of example and without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.
[0207] In at least one embodiment, the RISC core may interact with an image sensor (e.g., an image sensor of any of the cameras described herein), an image signal processor, and / or the like. In at least one embodiment, each of the RISC cores may include any amount of memory. In at least one embodiment, the RISC core may use any of a number of protocols, depending on the embodiment. In at least one embodiment, the RISC core may execute a real-time operating system ("RTOS"). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application specific integrated circuits ("ASICs"), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0208] In at least one embodiment, the DMA may allow components of the PVA to access system memory independently of the CPU 2106. In at least one embodiment, the DMA may support any number of features used to provide optimizations to the PVA, including, but not limited to, multi-dimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0209] In at least one embodiment, the vector processor may be a programmable processor that may be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, the vector processing subsystem may operate as the primary processing engine of the PVA and may include a vector processing unit ("VPU"), an instruction cache, and / or a vector memory (e.g., "VMEM"). In at least one embodiment, the VPU core may include a digital signal processor, such as a single instruction, multiple data ("SIMD"), very long instruction word ("VLIW") digital signal processor. In at least one embodiment, the combination of SIMD and VLIW may improve throughput and speed.
[0210] In at least one embodiment, each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to execute independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on consecutive images or portions of an image. In at least one embodiment, among other things, any number of PVAs may be included in a hardware-accelerated cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, the PVA may include additional error correction code ("ECC") memory to enhance the overall security of the system.
[0211] In at least one embodiment, the accelerator 2114 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and static random access memory (“SRAM”) to provide high-bandwidth, low-latency SRAM for the accelerator 2114. In at least one embodiment, the on-chip memory may include, for example, without limitation, at least 4 MB of SRAM consisting of eight field-configurable memory blocks, which may be accessible from both the PVA and the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory through a backbone that provides the PVA and DLA with high-speed access to the memory. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using the APB).
[0212] In at least one embodiment, the on-chip computer vision network may include an interface that determines whether both the PVA and DLA provide ready and enable signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-based communication for continuous data transfer. In at least one embodiment, the interface may conform to International Organization for Standardization (“ISO”) 26262 or International Electrotechnical Commission (“IEC”) 61508 standards, although other standards and protocols may be used.
[0213] In at least one embodiment, one or more of the SoCs 2104 may include a real-time ray tracing hardware accelerator, which may be used to quickly and efficiently determine the location and range of objects (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, sound propagation synthesis and / or analysis, SONAR system simulation, general waveform propagation simulation, comparison with LIDAR data for localization and / or other functions, and / or other uses.
[0214] In at least one embodiment, accelerator 2114 (e.g., a hardware accelerator cluster) has diverse uses for autonomous driving. In at least one embodiment, the PVA may be a programmable vision accelerator that can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, the performance of the PVA is well suited to algorithm domains that require low-power, low-latency, and predictable processing. In other words, the PVA works well for semi-dense or dense regular computations that require low-latency, low-power, and predictable run-times, even with small data sets. In at least one embodiment, in an autonomous vehicle such as vehicle 2100, the PVA is designed to run traditional computer vision algorithms because they are effective for object detection and integer arithmetic.
[0215] For example, according to at least one embodiment of the technology, computer stereo vision may be performed using the PVA. In at least one embodiment, algorithms based on semi-global matching may be used in some instances, but this is not intended to be limiting. In at least one embodiment, applications for Level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, the PVA may perform computer stereo vision functions on input from two monocular cameras.
[0216] In at least one embodiment, the PVA may be used to perform dense optical flow. For example, in at least one embodiment, the PVA may process raw RADAR data (e.g., using a 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, the PVA may be used for time-of-flight depth processing, e.g., by processing raw time-of-flight data to provide processed time-of-flight data.
[0217] In at least one embodiment, the DLA may be used to implement any type of network for enhancing control and driving safety, including, for example, without limitation, a neural network that outputs a confidence measure for each object detection. In at least one embodiment, the confidence may be expressed or interpreted as the probability of each detection compared to other detections or as providing its relative “weight.” In at least one embodiment, the confidence may further enable the system to make decisions regarding which detections should be considered positive rather than false. For example, in at least one embodiment, the system may set a threshold for confidence and consider only detections that exceed the threshold to be positive. In embodiments where an automatic emergency braking (“AEB”) system is used, a false detection may cause the vehicle to automatically apply emergency braking, which is clearly undesirable. In at least one embodiment, a highly confident detection may be considered to trigger AEB. In at least one embodiment, the DLA may implement a neural network to regress the confidence value. In at least one embodiment, the neural network may take as its input at least some subset of parameters, such as, among others, the bounding box dimensions, a ground surface estimate obtained (e.g., from another subsystem), an output from the IMU sensor 2166 that correlates with the orientation of the vehicle 2100, distance, and a 3D location estimate of the object obtained from the neural network and / or other sensors (e.g., the LIDAR sensor 2164 or the RADAR sensor 2160).
[0218] In at least one embodiment, one or more of the SoCs 2104 may include a data store 2116 (e.g., memory). In at least one embodiment, the data store 2116 may be on-chip memory of the SoC 2104, which may store neural networks running on the GPU 2108 and / or DLA. In at least one embodiment, the capacity of the data store 2116 may be large enough to store multiple instances of the neural network for redundancy and safety. In at least one embodiment, the data store 2116 may comprise an L2 or L3 cache.
[0219] In at least one embodiment, one or more of the SoCs 2104 may include any number of processors 2110 (e.g., embedded processors). The processors 2110 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and associated security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC 2104 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assist in transitioning the system to a low power state, manage the thermal and temperature sensors of the SoC 2104, and / or manage the power state of the SoC 2104. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 2104 may use the ring oscillator to detect the temperature of the CPU 2106, GPU 2108, and / or accelerator 2114. In at least one embodiment, if the temperature is determined to exceed a threshold, the boot and power management processor may enter a temperature fault routine, place the SoC 2104 in a low power state, and / or place the vehicle 2100 in a driver-safety shutdown mode (e.g., bring the vehicle 2100 to a safety shutdown).
[0220] In at least one embodiment, processor 2110 may further include a set of embedded processors that can act as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem that enables full hardware support for multi-channel audio over multiple interfaces and a wide variety of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core that includes a digital signal processor with dedicated RAM.
[0221] In at least one embodiment, the processor 2110 may further include an always-on processor engine capable of providing the hardware features necessary to support low-power sensor management and bring-up use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0222] In at least one embodiment, the processor 2110 may further include a safety cluster engine, which may include, without limitation, a processor subsystem dedicated to handling safety management for automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In safety mode, in at least one embodiment, two or more cores may operate in lockstep mode and function as a single core with comparison logic to detect any differences between their operation. In at least one embodiment, the processor 2110 may further include a real-time camera engine, which may include, without limitation, a processor subsystem dedicated to handling real-time camera management. In at least one embodiment, the processor 2110 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor, which is a hardware engine that is part of a camera processing pipeline.
[0223] In at least one embodiment, the processor 2110 may include a video image composer, which may be a processing block (e.g., implemented in a microprocessor) that implements video post-processing functions required by a video playback application to generate a final image in a playback device window. In at least one embodiment, the video image composer may perform lens distortion correction for the wide-angle camera 2170, the surrounding camera 2174, and / or the in-cabin surveillance camera sensor. In at least one embodiment, the in-cabin surveillance camera sensor is preferably monitored by a neural network running on a separate instance of the SoC 2104 that is configured to identify in-cabin events and respond accordingly. In at least one embodiment, the in-cabin system may perform lip reading, without limitation, to activate cellular service, make phone calls, write emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, and provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and are unavailable at other times.
[0224] In at least one embodiment, the video image combiner may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when motion occurs in the video, the noise reduction appropriately weights spatial information and downweights information provided by adjacent frames. In at least one embodiment, when an image or portion of an image does not contain motion, the temporal noise reduction performed by the video image combiner may use information from previous images to reduce noise in the current image.
[0225] In at least one embodiment, the video image combiner may also be configured to perform stereo rectification on the input stereo lens frames. In at least one embodiment, the video image combiner may also be used to combine the user interface when the operating system desktop is in use, eliminating the need for the GPU 2108 to continually render new surfaces. In at least one embodiment, the video image combiner may be used to offload the GPU 2108 when it is powered on and actively performing 3D rendering, improving performance and responsiveness.
[0226] In at least one embodiment, one or more of the SoCs 2104 may further include a mobile industry processor interface ("MIPI") camera serial interface for receiving input from video and cameras, a high-speed interface, and / or a video input block that may be used for camera and associated pixel input functions. In at least one embodiment, one or more of the SoCs 2104 may further include an input / output controller, which may be controlled by software and may be used to receive I / O signals that are not tied to a specific role.
[0227] In at least one embodiment, one or more of the SoCs 2104 may further include peripherals, audio encoders / decoders (“codecs”), power management, and / or a wide range of peripheral interfaces to enable communication with other devices. The SoCs 2104 may be used to process data from cameras (e.g., connected via gigabit multimedia serial links and Ethernet), data from sensors (e.g., LIDAR sensors 2164, RADAR sensors 2160, etc., which may be connected via Ethernet), data from bus 2102 (e.g., vehicle 2100 speed, steering wheel position, etc.), data from GNSS sensors 2158 (e.g., connected via Ethernet or CAN bus), etc. In at least one embodiment, one or more of the SoCs 2104 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to offload routine data management tasks from the CPU 2106.
[0228] In at least one embodiment, the SoC2104 may be an end-to-end platform with a flexible architecture spanning levels 3-5 of automation, providing a comprehensive functional safety architecture that leverages and efficiently utilizes computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible and reliable driving software stack, along with deep learning tools. In at least one embodiment, the SoC2104 is faster, more reliable, and more energy- and space-efficient than conventional systems. For example, in at least one embodiment, the accelerator 2114, when combined with the CPU 2106, GPU 2108, and data store 2116, can provide a fast and efficient platform for levels 3-5 of autonomous vehicles.
[0229] In at least one embodiment, computer vision algorithms may run on a CPU, which may be configured using a high-level programming language, such as the C programming language, to perform various processing algorithms across various visual data. However, in at least one embodiment, CPUs often cannot meet the performance requirements of many computer vision applications, such as those related to execution time and power consumption. In at least one embodiment, many CPUs are unable to run the complex object detection algorithms used in in-vehicle ADAS applications and realistic Level 3-5 autonomous vehicles in real time.
[0230] Embodiments described herein enable multiple neural networks to run simultaneously and / or sequentially, and the results can be combined to enable Level 3-5 autonomous driving capabilities. For example, in at least one embodiment, the CNN running on the DLA or a separate GPU (e.g., GPU2120) may include text and word recognition, enabling the supercomputer to read and understand traffic signs, including signs for which the neural network was not specifically trained. In at least one embodiment, the DLA may further include a neural network that can identify and interpret signs and provide a semantic understanding of the signs, which can then be passed to a route planning module running on the CPU complex.
[0231] In at least one embodiment, for Level 3, 4, or 5 driving, multiple neural networks may be run simultaneously. For example, in at least one embodiment, a warning sign displaying "Caution: Flashing Icy Conditions" in conjunction with an electric light may be interpreted separately or collectively by several neural networks. In at least one embodiment, the sign itself may be identified as a traffic sign by a first deployed neural network (e.g., a neural network that has been trained), and the words "Flashing Icy Conditions" may be interpreted by a second deployed neural network, which, if the flashing light is detected, notifies the vehicle's route planning software (preferably running on the CPU complex) that an icy condition exists. In at least one embodiment, the flashing light may be identified by running a third deployed neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may be run simultaneously, such as within the DLA and / or on the GPU 2108.
[0232] In at least one embodiment, a CNN for facial recognition and vehicle owner identification may use data from the camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 2100. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and to disable the vehicle in security mode when the owner leaves the vehicle. In this way, the SoC 2104 provides security against theft and / or carjacking.
[0233] In at least one embodiment, a CNN for emergency vehicle detection and identification may use data from microphone 2196 to detect and identify emergency vehicle sirens. In at least one embodiment, SoC 2104 uses a CNN to classify environmental and urban sounds as well as visual data. In at least one embodiment, the CNN running on the DLA is trained to identify the relative speed at which an emergency vehicle is approaching (e.g., by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the region in which the vehicle is operating, as identified by GNSS sensor 2158. In at least one embodiment, when operating in Europe, the CNN attempts to detect European sirens, and when in the United States, it attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program to execute an emergency vehicle safety routine may be used to slow the vehicle, pull over, stop the vehicle, and / or idle the vehicle in conjunction with ultrasonic sensor 2162 until the emergency vehicle has passed.
[0234] In at least one embodiment, the vehicle 2100 may include a CPU 2118 (e.g., a discrete CPU or dCPU), which may be coupled to the SoC 2104 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, the CPU 2118 may include, for example, an X86 processor. The CPU 2118 may be used to perform any of a variety of functions, including, for example, reconciling potentially inconsistent results between the ADAS sensors and the SoC 2104 and / or monitoring the status and health of the controller 2136 and / or infotainment system on a chip (“infotainment SoC”) 2130.
[0235] In at least one embodiment, vehicle 2100 may include a GPU 2120 (e.g., a discrete GPU or dGPU), which may be coupled to SoC 2104 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, GPU 2120 may provide additional artificial intelligence functionality, such as by running redundant and / or different neural networks, and may be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 2100.
[0236] In at least one embodiment, vehicle 2100 may further include a network interface 2124, which may include, without limitation, a wireless antenna 2126 (e.g., one or more wireless antennas 2126 for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, network interface 2124 may be used to enable wireless connectivity over the Internet with the cloud (e.g., servers and / or other network devices), other vehicles, and / or computing devices (e.g., occupant client devices). In at least one embodiment, to communicate with other vehicles, a direct link may be established between vehicle 210 and the other vehicles and / or an indirect link (e.g., across a network and via the Internet) may be established. In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide vehicle 2100 with information about vehicles in its vicinity (e.g., vehicles in front of, to the sides of, and / or behind vehicle 2100). In at least one embodiment, the above-described functionality may be part of a cooperative adaptive cruise control function of the vehicle 2100.
[0237] In at least one embodiment, the network interface 2124 may include an SoC that provides modulation and demodulation functionality and enables the controller 2136 to communicate over a wireless network. In at least one embodiment, the network interface 2124 may include a radio frequency front end for up-conversion from baseband to radio frequency and down-conversion from radio frequency to baseband. In at least one embodiment, the frequency conversion may be performed in any technically feasible manner. For example, the frequency conversion may be performed by well-known processes and / or using a super-heterodyne process. In at least one embodiment, the radio frequency front end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functionality for communicating via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0238] In at least one embodiment, vehicle 2100 may further include a data store 2128, which may include, without limitation, off-chip (e.g., not on SoC 2104) storage. In at least one embodiment, data store 2128 may include one or more storage elements, including, without limitation, RAM, SRAM, dynamic random access memory (“DRAM”), video random-access memory (“VRAM”), flash, a hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0239] In at least one embodiment, vehicle 2100 may further include GNSS sensors 2158 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or route planning functions. In at least one embodiment, any number of GNSS sensors 2158 may be used, including, for example, without limitation, a GPS using a USB connector with an Ethernet to serial (e.g., RS-232) bridge.
[0240] In at least one embodiment, the vehicle 2100 may further include a RADAR sensor 2160. The RADAR sensor 2160 may be used by the vehicle 2100 for long-range vehicle detection, even in darkness and / or severe weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. The RADAR sensor 2160 may use the CAN and / or bus 2102 for control (e.g., to transmit data generated by the RADAR sensor 2160) and to access object tracking data, and in some instances may have Ethernet access to access raw data. In at least one embodiment, various types of RADAR sensors may be used. For example, without limitation, the RADAR sensor 2160 may be suitable for forward, rearward, and side RADAR use. In at least one embodiment, one or more of the RADAR sensors 2160 are pulse-Doppler RADAR sensors.
[0241] In at least one embodiment, the RADAR sensor 2160 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, and short-range with side coverage. In at least one embodiment, the long-range RADAR may be used for adaptive cruise control functions. In at least one embodiment, the long-range RADAR system may provide a wide field of view, such as within a 250 meter range, achieved by two or more independent scans. In at least one embodiment, the RADAR sensor 2160 may help distinguish between static and moving objects and may be used by the ADAS system 2138 to provide emergency braking assistance and forward collision warning. The sensors 2160 included in a long-range RADAR system may include, without limitation, multiple (e.g., six or more) fixed RADAR antennas, as well as monostatic multi-mode RADAR with high-speed CAN and FlexRay interfaces. In at least one embodiment, where there are six antennas, the center four antennas may generate a focused beam pattern designed to record the surroundings of vehicle 2100 at higher speeds with minimal interference from adjacent lanes. In at least one embodiment, the other two antennas may extend the field of view, allowing for quick detection of vehicles entering or exiting the lane of vehicle 2100.
[0242] In at least one embodiment, the medium-range RADAR system may include, by way of example, a range of up to 160 meters (forward) or 80 meters (rearward) and a field of view of up to 42 degrees (forward) or 150 degrees (rearward). In at least one embodiment, the short-range RADAR system may include, without limitation, any number of RADAR sensors 2160 designed to be mounted on either end of the rear bumper. When mounted on either end of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that constantly monitor blind spots behind and adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used in an ADAS system 2138 to provide blind spot detection and / or lane change assistance.
[0243] In at least one embodiment, the vehicle 2100 may further include ultrasonic sensors 2162. The ultrasonic sensors 2162 may be located at the front, rear, and / or sides of the vehicle 2100 and may be used for parking assistance and / or to generate and update an occupancy grid. In at least one embodiment, multiple ultrasonic sensors 2162 may be used, and different ultrasonic sensors 2162 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensors 2162 may operate at functional safety level ASIL B.
[0244] In at least one embodiment, vehicle 2100 may include a LIDAR sensor 2164. The LIDAR sensor 2164 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensor 2164 may be functional safety level ASIL B. In at least one embodiment, vehicle 2100 may include multiple LIDAR sensors 2164 (e.g., two, four, six, etc.), which may use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).
[0245] In at least one embodiment, the LIDAR sensor 2164 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 2164 may, for example, have an advertised range of approximately 100 meters, an accuracy of 2 cm to 3 cm, and support a 100 Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LIDAR sensors 2164 may be used. In such an embodiment, the LIDAR sensor 2164 may be implemented as a small device that can be integrated into the front, rear, sides, and / or corners of the vehicle 2100. In at least one embodiment, the LIDAR sensor 2164 of such an embodiment may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, with a range of 200 meters, even for low-reflectivity objects. In at least one embodiment, a front-mounted LIDAR sensor 2164 may be configured to provide a horizontal field of view of 45 degrees to 135 degrees.
[0246] In at least one embodiment, LIDAR technology such as 3D flash LIDAR may also be used. 3D flash LIDAR uses a laser flash as a transmission source to illuminate the surroundings of the vehicle 2100 up to approximately 200 meters. In at least one embodiment, the flash LIDAR unit includes, without limitation, a receptor that records the transit time of the laser pulse and the reflected light at each pixel, which corresponds to the range from the vehicle 2100 to the object. In at least one embodiment, flash LIDAR allows a highly accurate, undistorted image of the surroundings to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors may be installed, one on each side of the vehicle 2100. In at least one embodiment, the 3D flash LIDAR system includes, without limitation, a solid-state 3D staring array LIDAR camera (e.g., a non-scanning LIDAR device) with no moving parts other than a fan. In at least one embodiment, the flash LIDAR device may use 5 nanosecond Class I (eye-safe) laser pulses per frame and may capture reflected laser light in the form of a 3D range point cloud and co-registered intensity data.
[0247] In at least one embodiment, the vehicle may further include an IMU sensor 2166. In at least one embodiment, the IMU sensor 2166 may be positioned at the center of the rear axle of the vehicle 2100. In at least one embodiment, the IMU sensor 2166 may include, for example, without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other types of sensors. In at least one embodiment, such as in a 6-axis application, the IMU sensor 2166 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as in a 9-axis application, the IMU sensor 2166 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0248] In at least one embodiment, the IMU sensor 2166 may be implemented as a compact, high-performance GPS-Aided Inertial Navigation System ("GPS / INS") that combines micro-electro-mechanical systems ("MEMS") inertial sensors, a highly sensitive GPS receiver, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor 2166 enables the vehicle 2100 to estimate its orientation by directly observing changes in velocity and correlating them from the GPS to the IMU sensor 2166 without requiring input from a magnetic sensor. In at least one embodiment, the IMU sensor 2166 and the GNSS sensor 2158 may be combined into a single integrated unit.
[0249] In at least one embodiment, vehicle 2100 may include microphones 2196 located in and / or around vehicle 2100. In at least one embodiment, microphones 2196 may be used for, among other things, detection and identification of emergency vehicles.
[0250] In at least one embodiment, vehicle 2100 may further include any number of camera types, including stereo camera 2168, wide-angle camera 2170, infrared camera 2172, perimeter camera 2174, long-range camera 2198, mid-range camera 2176, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of vehicle 2100. In at least one embodiment, the types of cameras used vary depending on vehicle 2100. In at least one embodiment, any combination of camera types may be used to provide the required coverage around vehicle 2100. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, vehicle 2100 may include six cameras, seven cameras, ten cameras, twelve cameras, or another number of cameras. The cameras may support, by way of example and without limitation, Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet. In at least one embodiment, each camera is described in further detail herein above with respect to Figures 21A and 21B.
[0251] In at least one embodiment, vehicle 2100 may further include a vibration sensor 2142. Vibration sensor 2142 may measure vibration of a component of vehicle 2100, such as an axle. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 2142 are used, the difference in vibration may be used to determine the amount of friction or slippage of the road surface (e.g., if there is a vibration difference between a powered axle and a free-spinning axle).
[0252] In at least one embodiment, vehicle 2100 may include an ADAS system 2138. ADAS system 2138 may include, without limitation, an SoC in some instances. In at least one embodiment, the ADAS systems 2138 may include, without limitation, any number and combination of autonomous / adaptive / automatic cruise control ("ACC") systems, cooperative adaptive cruise control ("CACC") systems, forward crash warning ("FCW") systems, automatic emergency braking ("AEB") systems, lane departure warning ("LDW") systems, lane keep assist ("LKA") systems, blind spot warning ("BSW") systems, rear cross-traffic warning ("RCTW") systems, collision warning ("CW") systems, lane centering ("LC") systems, and / or other systems, features, and / or functions.
[0253] In at least one embodiment, the ACC system may use a RADAR sensor 2160, a LIDAR sensor 2164, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to the vehicle immediately preceding the vehicle 2100 and automatically adjusts the speed of the vehicle 2100 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system enforces distance maintenance and notifies the vehicle 2100 to change lanes when necessary. In at least one embodiment, the lateral ACC is related to other ADAS applications, such as LC and CW.
[0254] In at least one embodiment, the CACC system uses information from other vehicles, which may be received by the network interface 2124 and / or wireless antenna 2126 from other vehicles via a wireless link or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, a vehicle-to-vehicle ("V2V") communication link may provide a direct link, while an infrastructure-to-vehicle ("I2V") communication link may provide an indirect link. Generally, the V2V communication concept provides information about the immediate preceding vehicle (e.g., a vehicle immediately in front of the vehicle 2100 and in the same lane), while the I2V communication concept provides information about traffic ahead of that. In at least one embodiment, the CACC system may include either or both I2V and V2V information sources. In at least one embodiment, information about vehicles in front of the vehicle 2100 may make the CACC system more reliable, potentially allowing for smoother traffic flow and reducing congestion on the roads.
[0255] In at least one embodiment, the FCW system is designed to alert the driver to hazards so that the driver can take corrective action. In at least one embodiment, the FCW system uses a front-facing camera and / or RADAR sensor 2160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver, such as a display, speaker, and / or vibration component. In at least one embodiment, the FCW system may provide warnings in the form of an audible, visual warning, vibration, and / or a quick brake pulse.
[0256] In at least one embodiment, the AEB system may detect an imminent frontal collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within specified time or distance parameters. In at least one embodiment, the AEB system may use a front-facing camera and / or RADAR sensor 2160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system typically first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent or at least mitigate the severity of the predicted collision. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-collision braking.
[0257] In at least one embodiment, the LDW system provides visual, audible, and / or tactile warnings, such as vibration of the steering wheel or seat, to advise the driver when the vehicle 2100 crosses a lane marker. In at least one embodiment, the LDW system does not engage if the driver indicates an intentional lane departure by activating a turn signal. In at least one embodiment, the LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC that can be electrically coupled to feedback to the driver, such as a display, speaker, and / or vibration components. In at least one embodiment, the LKA system is a variation of the LDW system. The LKA system provides steering input or brake control to correct the vehicle 2100 if the vehicle 2100 begins to drift out of its lane.
[0258] In at least one embodiment, the BSW system detects vehicles in the vehicle's blind spot and warns the driver. In at least one embodiment, the BSW system may provide visual, audible, and / or haptic alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system may provide an additional warning when the driver uses a turn signal. In at least one embodiment, the BSW system may use a rearview camera and / or RADAR sensor 2160 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs, which are electrically coupled to feedback to the driver, such as a display, speaker, and / or vibration components.
[0259] In at least one embodiment, the RCTW system may provide visual, audible, and / or tactile notifications when an object is detected outside the range of the rear camera when reversing the vehicle 2100. In at least one embodiment, the RCTW system includes an AEB system to ensure vehicle braking is applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear RADAR sensors 2160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to feedback to the driver, such as a display, speaker, and / or vibration components.
[0260] In at least one embodiment, conventional ADAS systems may be prone to false positive results, which can be annoying and distracting to the driver, but are typically not a major concern because conventional ADAS systems advise the driver and allow the driver to determine whether a safety condition truly exists and respond accordingly. In at least one embodiment, in the event of conflicting results, the vehicle 2100 itself determines whether to follow the results from the primary computer (e.g., first controller 2136) or the secondary computer (e.g., second controller 2136). For example, in at least one embodiment, the ADAS system 2138 may be a backup and / or secondary computer to provide perception information to a backup computer rationality module. In at least one embodiment, the backup computer rationality monitor may run redundant software on various hardware components to detect perception errors and dynamic driving tasks. In at least one embodiment, output from the ADAS system 2138 may be provided to a supervisory MCU. In at least one embodiment, if the output from the primary computer and the output from the secondary computer conflict, the supervisory MCU determines how to reconcile the conflict to ensure safe operation.
[0261] In at least one embodiment, the primary computer may be configured to provide the monitor MCU with a reliability score indicating the reliability of the primary computer's selected result. In at least one embodiment, if the reliability score exceeds a threshold, the monitor MCU may follow the primary computer's instructions regardless of whether the secondary computers are providing conflicting or inconsistent results. In at least one embodiment, if the reliability score does not meet the threshold and the primary and secondary computers provide different (e.g., conflicting) results, the monitor MCU may arbitrate between the computers to determine the appropriate result.
[0262] In at least one embodiment, the monitoring MCU may be configured to execute a neural network trained and configured to determine conditions under which the secondary computer will provide a false alarm based at least in part on outputs from the primary and secondary computers. In at least one embodiment, the monitoring MCU's neural network may learn when the secondary computer's output may be trusted and when it may not be trusted. For example, in at least one embodiment, if the secondary computer is a RADAR-based FCW system, the monitoring MCU's neural network may learn when the FCW system identifies a metal object that is not actually a hazard, such as a drain grate or manhole cover, which triggers an alarm. In at least one embodiment, if the secondary computer is a camera-based LDW system, the monitoring MCU's neural network may learn to disable LDW when a bicyclist or pedestrian is present and lane departure is actually the safest maneuver. In at least one embodiment, the monitoring MCU may include at least one of a DLA or a GPU suitable for executing the neural network along with associated memory. In at least one embodiment, the supervisory MCU may comprise and / or be included as a component of the SoC2104.
[0263] In at least one embodiment, the ADAS system 2138 may include a secondary computer that performs ADAS functions using traditional rules of computer vision. In at least one embodiment, the secondary computer may use traditional computer vision rules (if-then rules), and neural networks may reside in the supervisory MCU, improving reliability, safety, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identity may increase the overall system's error tolerance, particularly against errors caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a bug or error in the software running on the primary computer and non-identical software code running on the secondary computer provides the same overall result, the supervisory MCU may have greater confidence that the overall result is correct and that a software or hardware bug on the primary computer did not cause a critical error.
[0264] In at least one embodiment, the output of the ADAS system 2138 may be provided to a perception block of the primary computer and / or a dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 2138 indicates a frontal collision warning due to an immediately preceding object, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have its own neural network that is pre-trained, as described herein, thus reducing the risk of false positives.
[0265] In at least one embodiment, vehicle 2100 may further include an infotainment SoC 2130 (e.g., an in-vehicle infotainment system (IVI)). While the infotainment system 2130 is shown and described as an SoC, in at least one embodiment, it may not be an SoC and may include, without limitation, two or more separate components. In at least one embodiment, the infotainment SoC 2130 may include, without limitation, a combination of hardware and software that may be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear park assist, wireless data system, vehicle-related information such as fuel level, total mileage, brake fuel level, oil level, door opening / closing, air filter information, etc.) to vehicle 2100. For example, infotainment SoC2130 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connectivity, a car computer, in-car entertainment, Wi-Fi, steering wheel audio controls, hands-free voice control, a heads-up display (“HUD”), an HMI display 2134, telematics devices, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC2130 may also be used to provide information (e.g., visual and / or auditory) to a user of the vehicle, such as information from an ADAS system 2138, autonomous driving information such as a vehicle maneuver plan, a trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0266] In at least one embodiment, infotainment SoC 2130 may include any amount and type of GPU functionality. In at least one embodiment, infotainment SoC 2130 may communicate with other devices, systems, and / or components of vehicle 2100 via bus 2102 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, infotainment SoC 2130 may be coupled to a supervisory MCU such that the infotainment system's GPU may perform some self-driving functions when primary controller 2136 (e.g., vehicle's 2100 primary and / or backup computer) fails. In at least one embodiment, infotainment SoC 2130 may place vehicle 2100 in a driver-safety shutdown mode, as described herein.
[0267] In at least one embodiment, the vehicle 2100 may further include an instrument cluster 2132 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 2132 may include, without limitation, a controller and / or a supercomputer (e.g., a separate controller or supercomputer). In at least one embodiment, the instrument cluster 2132 may include any number and combination of instrument sets, such as, without limitation, a speedometer, fuel level, oil pressure, a tachometer, an odometer, turn signals, a shift lever position indicator, a seat belt warning light, a parking brake warning light, an engine malfunction light, supplemental restraint system (e.g., airbag) information, light control, safety system control, navigation information, etc. In some instances, information may be displayed and / or shared between the infotainment SoC 2130 and the instrument cluster 2132. In at least one embodiment, the instrument cluster 2132 may be included as part of the infotainment SoC 2130, or vice versa.
[0268] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 21C for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0269] In at least one embodiment, system FIG. 21C is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, system FIG. 21C is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, system FIG. 21C is utilized to implement one or more neural networks that include a classifier and a generator, and system FIG. 21C is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more properties of the images in a training set.
[0270] 21D is a diagram of a system 2176 for communicating between a cloud-based server and the autonomous vehicle 2100 of FIG. 21A , according to at least one embodiment. In at least one embodiment, the system 2176 may include any number and type of vehicles, including, without limitation, a server 2178, a network 2190, and a vehicle 2100. The server 2178 may include, without limitation, multiple GPUs 2184(A)-2184(H) (collectively referred to herein as GPUs 2184), PCIe switches 2182(A)-2182(H) (collectively referred to herein as PCIe switches 2182), and / or CPUs 2180(A)-2180(B) (collectively referred to herein as CPUs 2180). The GPUs 2184, CPUs 2180, and PCIe switches 2182 may be interconnected by a high-speed interconnect, such as, for example, without limitation, an NVLink interface 2188 developed by NVIDIA and / or a PCIe connection 2186. In at least one embodiment, the GPUs 2184 are connected to each other via an NVLink and / or an NVSwitch SoC, and the GPUs 2184 and PCIe switches 2182 are connected via a PCIe interconnect. In at least one embodiment, eight GPUs 2184, two CPUs 2180, and four PCIe switches 2182 are illustrated, but this is not intended to be limiting. In at least one embodiment, each of the servers 2178 may include any number of GPUs 2184, CPUs 2180, and / or PCIe switches 2182 in any combination, without limitation. For example, in at least one embodiment, the servers 2178 may each include 8, 16, 32, and / or more GPUs 2184.
[0271] In at least one embodiment, server 2178 may receive image data from the vehicle over network 2190 representing images showing unexpected or changed road conditions, such as recently begun road construction. In at least one embodiment, server 2178 may transmit neural network 2192, updated neural network 2192, and / or map information 2194, including, without limitation, information regarding traffic and road conditions, to the vehicle over network 2190. In at least one embodiment, updates to map information 2194 may include, without limitation, updates to HD map 2122, such as information regarding construction sites, potholes, detours, flooding, and / or other obstacles. In at least one embodiment, neural network 2192, updated neural network 2192, and / or map information 2194 may be derived from new training and / or experience represented in data received from any number of vehicles in the environment and / or may be derived based at least in part on training performed at a data center (e.g., using server 2178 and / or other servers).
[0272] In at least one embodiment, server 2178 may be used to train a machine learning model (e.g., a neural network) based at least in part on the training data. The training data may be generated by the vehicle and / or generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of the training data may be tagged and / or otherwise preprocessed (e.g., if the associated neural network benefits from supervised learning). In at least one embodiment, any amount of the training data may not be tagged and / or preprocessed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, it may be used by the vehicle (e.g., transmitted to the vehicle via network 2190) and / or used by server 2178 to remotely monitor the vehicle.
[0273] In at least one embodiment, server 2178 may receive data from vehicles and apply the data to state-of-the-art, real-time neural networks to enable real-time intelligent inference. In at least one embodiment, server 2178 may include a deep learning supercomputer and / or dedicated AI computer powered by GPU 2184, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, server 2178 may also include a deep learning infrastructure using a CPU-powered data center.
[0274] In at least one embodiment, the deep learning infrastructure of server 2178 may be capable of rapid real-time inference and may use that capability to assess and verify the health of the processor, software, and / or associated hardware of vehicle 2100. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 2100, such as a series of images and / or objects that vehicle 2100 has located in that series of images (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them to those identified by vehicle 2100; if the results do not match and the deep learning infrastructure concludes that the AI of vehicle 2100 has failed, server 2178 may send a signal to vehicle 2100 to command the fail-safe computer of vehicle 2100 to assume control, notify the occupants, and complete a safe stopping maneuver.
[0275] In at least one embodiment, server 2178 may include a GPU 2184 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT3). In at least one embodiment, the combination of a GPU-powered server and inference acceleration can enable real-time response. In at least one embodiment, CPU, FPGA, and other processor-powered servers may be used for inference, such as when performance is less critical. In at least one embodiment, a hardware structure 1815 is used to execute one or more embodiments. Details regarding hardware structure (x) 1815 are provided herein in conjunction with FIG. 18A and / or FIG. 18B.
[0276] Computer Systems 22 is a block diagram illustrating an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof 2200 formed with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 2200 may include components such as, without limitation, a processor 2202 to utilize an execution unit that includes logic for executing algorithms for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 2200 may include a processor such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems may be used (including PCs with other microprocessors, engineering workstations, set-top boxes, etc.). In at least one embodiment, computer system 2200 may run a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux), embedded software, and / or graphical user interfaces may also be used.
[0277] Embodiments may be used in other devices, such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and portable PCs. In at least one embodiment, embedded applications may include microcontrollers, digital signal processors ("DSPs"), systems-on-chips, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions according to at least one embodiment.
[0278] In at least one embodiment, computer system 2200 may include, without limitation, a processor 2202, which may include one or more execution units 2208 for performing machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, system 22 is a single-processor desktop or server system, while in other embodiments, system 22 may be a multiprocessor system. In at least one embodiment, processor 2202 may include, without limitation, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 2202 may be coupled to a processor bus 2210, which may transmit data signals between processor 2202 and other components within computer system 2200.
[0279] In at least one embodiment, processor 2202 may include, without limitation, level 1 ("L1") internal cache memory ("cache") 2204. In at least one embodiment, processor 2202 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, cache memory may be external to processor 2202. Other embodiments may include a combination of both internal and external cache, depending on the particular implementation and needs. In at least one embodiment, register file 2206 may store different types of data in various registers, including, without limitation, integer registers, floating-point registers, status registers, and an instruction pointer register.
[0280] In at least one embodiment, processor 2202 also includes an execution unit 2208, including, without limitation, logic for performing integer and floating-point operations. Processor 2202 may also include microcode (“u-code”) read-only memory (“ROM”) that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 2208 may include logic for a packed instruction set 2209. In at least one embodiment, including the packed instruction set 2209, along with associated circuitry for executing the instructions, in the instruction set of general-purpose processor 2202 allows operations used by many multimedia applications to be performed using packed data in general-purpose processor 2202. In one or more embodiments, many multimedia applications can be accelerated and run more efficiently by performing operations on packed data using the full width of the processor's data bus, thereby eliminating the need to transfer smaller units of data between the processor's data bus to perform one or more operations on one data element at a time.
[0281] In at least one embodiment, execution unit 2208 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 2200 may include, without limitation, memory 2220. In at least one embodiment, memory 2220 may be implemented as a dynamic random access memory ("DRAM") device, a static random access memory ("SRAM") device, a flash memory device, or other memory device. Memory 2220 may store instructions 2219 and / or data 2221 represented by data signals that may be executed by processor 2202.
[0282] In at least one embodiment, a system logic chip may be coupled to the processor bus 2210 and the memory 2220. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (“MCH”) 2216, and the processor 2202 may communicate with the MCH 2216 via the processor bus 2210. In at least one embodiment, the MCH 2216 may provide a high-bandwidth memory path 2218 to the memory 2220 for storing instructions and data, and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 2216 may route data signals between the processor 2202, the memory 2220, and other components of the computer system 2200, and may bridge data signals between the processor bus 2210, the memory 2220, and the system I / O 2222. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 2216 may be coupled to memory 2220 via a high-bandwidth memory path 2218, and graphics / video card 2212 may be coupled to MCH 2216 via an Accelerated Graphics Port (“AGP”) interconnect 2214.
[0283] In at least one embodiment, computer system 2200 may use system I / O 2222, a proprietary hub interface bus, to couple MCH 2216 to I / O controller hub (“ICH”) 2230. In at least one embodiment, ICH 2230 may provide direct connectivity to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 2220, a chipset, and processor 2202. Examples may include, without limitation, an audio controller 2229, a firmware hub ("flash BIOS") 2228, a wireless transceiver 2226, data storage 2224, a legacy I / O controller 2223 including user input and keyboard interfaces, a serial expansion port 2227 such as a Universal Serial Bus ("USB"), and a network controller 2234. Data storage 2224 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0284] In at least one embodiment, Figure 22 illustrates a system including interconnected hardware devices or "chips," while in other embodiments, Figure 22 may illustrate an exemplary system-on-a-chip ("SoC"). In at least one embodiment, the devices illustrated in Figure 22 may be interconnected using a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 2200 may be interconnected using a compute express link (CXL) interconnect.
[0285] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 22 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0286] In at least one embodiment, system diagram 22 is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, system diagram 22 is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, system diagram 22 is utilized to implement one or more neural networks that include a classifier and a generator, and system diagram 22 is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more characteristics of the images in a training set.
[0287] 23 is a block diagram illustrating an electronic device 2300 for utilizing a processor 2310, according to at least one embodiment. In at least one embodiment, the electronic device 2300 may be, for example, without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0288] In at least one embodiment, system 2300 may include a processor 2310 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices, including, without limitation, processor 2310 coupled using a bus or interface, such as an I / C bus, a System Management Bus (“SMBus”), a Low Pin Count (“LPC”) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advance Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, and 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, FIG. 23 illustrates a system including interconnected hardware devices or "chips," while in other embodiments, FIG. 23 may illustrate an exemplary system-on-a-chip ("SoC"). In at least one embodiment, the devices illustrated in FIG. 23 may be interconnected with a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 23 may be interconnected using a Compute Express Link (CXL) interconnect.
[0289] In at least one embodiment, FIG. 23 illustrates a display 2324, a touch screen 2325, a touch pad 2330, a Near Field Communications unit ("NFC") 2345, a sensor hub 2340, a thermal sensor 2346, an Express Chipset ("EC") 2335, a Trusted Platform Module ("TPM") 2338, a BIOS / firmware / flash memory ("BIOS,FW flash") 2322, a DSP 2360, a drive ("SSD or HDD") 2320, such as a solid state disk ("SSD") or hard disk drive ("HDD"), a wireless local area network unit ("WLAN") 2340, a wireless local area network ("WLAN") 2340, a wireless local area network ("WLAN") 2340, a wireless local area network ("WLAN") 2340, a wireless local area network ("WLAN") 2340, a wireless local area network ("WLAN") 2340, a wireless local area network ("WLAN") 2340, a wireless local area network ("WLAN") 2360, a wireless local area network ("SSD") 2320, a wireless local area network ("HDD") 232 ... The memory may include a wireless LAN (LAN) 2350, a Bluetooth unit 2352, a wireless wide area network ("WWAN") unit 2356, a global positioning system (GPS) 2355, a camera such as a USB 3.0 camera ("USB 3.0 camera") 2354, or a low power double data rate ("LPDDR") memory unit ("LPDDR3") 2315, implemented, for example, to the LPDDR3 standard. Each of these components may be implemented in any suitable manner.
[0290] In at least one embodiment, other components may be communicatively coupled to the processor 2310 via the components described above. In at least one embodiment, an accelerometer 2341, an ambient light sensor (“ALS”) 2342, a compass 2343, and a gyroscope 2344 may be communicatively coupled to the sensor hub 2340. In at least one embodiment, a thermal sensor 2339, a fan 2337, a keyboard 2346, and a touchpad 2330 may be communicatively coupled to the EC 2335. In at least one embodiment, a speaker 2363, headphones 2364, and a microphone (“mic”) 2365 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 2364, which may be communicatively coupled to the DSP 2360. In at least one embodiment, the audio unit 2364 may include, for example, without limitation, an audio coder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 2357 may be communicatively coupled to the WWAN unit 2356. In at least one embodiment, components such as the WLAN unit 2350 and Bluetooth unit 2352, and the WWAN unit 2356 may be implemented in a Next Generation Form Factor (“NGFF”).
[0291] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 23 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0292] In at least one embodiment, system diagram 23 is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, system diagram 23 is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, system diagram 23 is utilized to implement one or more neural networks that include a classifier and a generator, and system diagram 23 is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more characteristics of the images in a training set.
[0293] 24 illustrates a computer system 2400 according to at least one embodiment. In at least one embodiment, the computer system 2400 is configured to implement the various processes and methods described throughout this disclosure.
[0294] In at least one embodiment, computer system 2400 includes at least one central processing unit ("CPU") 2402 connected to a communication bus 2410 implemented using any suitable protocol, such as, without limitation, PCI (Peripheral Component Interconnect), Peripheral Component Interconnect Express ("PCI-Express"), AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 2400 includes main memory 2404 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in main memory 2404, which may be in the form of random access memory ("RAM"). In at least one embodiment, network interface subsystem (“network interface”) 2422 provides an interface with other computing devices and networks to receive data from other systems and transmit data from computer system 2400 to other systems.
[0295] In at least one embodiment, computer system 2400 includes, without limitation, input device(s) 2408, a parallel processing system 2412, and a display device 2406, which may be implemented using a conventional cathode ray tube ("CRT"), liquid crystal display ("LCD"), light emitting diode ("LED"), plasma display, or other suitable display technology. In at least one embodiment, user input is received from input device(s) 2408, such as a keyboard, mouse, touch pad, microphone, or the like. In at least one embodiment, each of the above modules may be located on a single semiconductor platform to form a processing system.
[0296] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 24 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0297] In at least one embodiment, the system diagram 24 is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system diagram 24 is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, the system diagram 24 is utilized to implement one or more neural networks that include the classifier and generator, and the system diagram 24 is utilized in conjunction with one or more processes that train the one or more neural networks to identify the orientation of objects in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more characteristics of the images in a training set.
[0298] 25 illustrates a computer system 2500 according to at least one embodiment. In at least one embodiment, the computer system 2500 may include, without limitation, a computer 2510 and a USB stick 2520. In at least one embodiment, the computer system 2510 includes, without limitation, any number and type of processor (not shown) and memory (not shown). In at least one embodiment, the computer 2510 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.
[0299] In at least one embodiment, USB stick 2520 includes, without limitation, a processing unit 2530, a USB interface 2540, and USB interface logic 2550. In at least one embodiment, processing unit 2530 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 2530 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, processing core 2530 comprises an application specific integrated circuit ("ASIC") optimized to perform any quantity and type of operations related to machine learning. For example, in at least one embodiment, processing core 2530 is a tensor processing unit ("TPC") optimized to perform machine vision and machine learning inference operations. In at least one embodiment, processing core 2530 is a vision processing unit ("VPU") optimized to perform machine vision and machine learning inference operations.
[0300] In at least one embodiment, USB interface 2540 may be any type of USB connector or socket. For example, in at least one embodiment, USB interface 2540 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, USB interface 2540 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 2550 may include any amount and type of logic that enables processing unit 2530 to interface with a device (e.g., computer 2510) via USB connector 2540.
[0301] Inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 1815 are provided herein in conjunction with Figures 18A and / or 18B. In at least one embodiment, inference and / or training logic 1815 may be used in the system of Figure 25 for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein.
[0302] In at least one embodiment, system diagram 25 is utilized to implement a classifier that is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, system diagram 25 is utilized to implement a generator that is trained to generate images based on an input set of input viewpoint and appearance parameters. In at least one embodiment, system diagram 25 is utilized to implement one or more neural networks that include a classifier and a generator, and system diagram 25 is utilized in conjunction with one or more processes that train the one or more neural networks to identify object orientations in images in a self-supervised manner by computing, at least as part of the training, one or more loss functions that evaluate one or more characteristics of the images in a training set.
[0303] 26A illustrates an exemplary architecture in which multiple GPUs 2610-2613 are communicatively coupled to multiple multi-core processors 2605-2606 via high-speed links 2640-2643 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 2640-2643 support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or more. Various interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0.
[0304] Additionally, in one embodiment, two or more of GPUs 2610-2613 may be interconnected via high-speed links 2629-2630, which may be implemented using the same or different protocol / links as used for high-speed links 2640-2643. Similarly, two or more of multi-core processors 2605-2606 may be connected via high-speed link 2628, which may be a symmetric multi-processor (SMP) bus operating at 20 GB / s, 30 GB / s, 120 GB / s, or more. Alternatively, all communications between the various system components shown in FIG. 26A may be achieved using the same protocol / links (e.g., via a common interconnect fabric).
[0305] In one embodiment, each multi-core processor 2605-2606 is communicatively coupled to processor memory 2601-2602 via a memory interconnect 2626-2627, respectively, and each GPU 2610-2613 is communicatively coupled to GPU memory 2620-2623 via a GPU memory interconnect 2650-2653, respectively. Memory interconnects 2626-2627 and 2650-2653 may utilize the same or different memory access technologies. By way of example, and not limitation, processor memory 2601-2602 and GPU memory 2620-2623 may be volatile memory such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or may be non-volatile memory such as 3D XPoint or Nano-RAM. In one embodiment, some portions of processor memory 2601-2602 may be volatile memory and other portions may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0306] As described herein, the various processors 2605-2606 and GPUs 2610-2613 may each be physically coupled to a particular memory 2601-2602, 2620-2623, but a unified memory architecture may be implemented in which the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, the processor memories 2601-2602 may each have 64 GB of system memory address space, and the GPU memories 2620-2623 may each have 32 GB of system memory address space (resulting in a total of 256 GB of addressable memory in this example).
[0307] 26B shows further details of the interconnection between multi-core processor 2607 and graphics acceleration module 2646 according to one example embodiment. Graphics acceleration module 2646 may include one or more GPU chips integrated on a line card that is coupled to processor 2607 via high-speed link 2640. Alternatively, graphics acceleration module 2646 may be integrated in the same package or chip as processor 2607.
[0308] In at least one embodiment, the illustrated processor 2607 includes multiple cores 2660A-2660D, each having a translation lookaside buffer 2661A-2661D and one or more caches 2662A-2662D. In at least one embodiment, the cores 2660A-2660D may include various other components, not shown, for executing instructions and processing data. The caches 2662A-2662D may comprise level 1 (L1) and level 2 (L2) caches. Additionally, one or more shared caches 2656 may be included in the caches 2662A-2662D and shared by the set of cores 2660A-2660D. For example, one embodiment of the processor 2607 includes 24 cores, each with its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more L2 and L3 caches are shared by two adjacent cores. The processor 2607 and graphics acceleration module 2646 are connected to system memory 2614, which may include processor memories 2601-2602 of FIG. 26A.
[0309] Coherence is maintained for data and instructions stored in the various caches 2662A-2662D, 2656, and system memory 2614 through inter-core communication via coherence bus 2664. For example, each cache may have associated cache coherence logic / circuitry for communicating via coherence bus 2664 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via coherence bus 2664 to monitor cache accesses.
[0310] In one embodiment, proxy circuit 2625 communicatively couples graphics acceleration module 2646 to coherence bus 2664 to enable graphics acceleration module 2646 to participate in cache coherence protocols as a peer of cores 2660A-2660D. In particular, interface 2635 provides a connection to proxy circuit 2625 over high-speed link 2640 (e.g., PCIe bus, NVLink, etc.), and interface 2637 connects graphics acceleration module 2646 to link 2640.
[0311] In one implementation, the accelerator integrated circuit 2636 provides cache management, memory access, content management, and interrupt management services on behalf of the multiple graphics processing engines 2631, 2632, N of the graphics acceleration module 2646. The graphics processing engines 2631, 2632, N may each comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 2631, 2632, N may comprise different types of graphics processing engines within a GPU, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine. In at least one embodiment, the graphics acceleration module 2646 may be a GPU having multiple graphics processing engines 2631-2632, N, or the graphics processing engines 2631-2632, N may be individual GPUs integrated into a common package, line card, or chip.
[0312] In one embodiment, the accelerator integrated circuitry 2636 includes a memory management unit (MMU) 2639 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 2614. The MMU 2639 may also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one implementation, the cache 2638 stores commands and data for efficient access by the graphics processing engines 2631-2632, N. In one embodiment, data stored in the cache 2638 and the graphics memory 2633-2634, M is kept coherent with the core caches 2662A-2662D, 2656, and the system memory 2614. As noted above, this may be achieved via the proxy circuit 2625 on behalf of the cache 2638 and memories 2633-2634, M (e.g., sending updates regarding modifications / accesses of cache lines in the processor caches 2662A-2662D, 2656 to the cache 2638 and receiving updates from the cache 2638).
[0313] A set of registers 2645 stores context data for threads executed by the graphics processing engines 2631-2632, and a context management circuit 2648 manages thread contexts. For example, the context management circuit 2648 may perform save and restore operations to save and restore the context of various threads during a context switch (e.g., where a first thread is saved and a second thread is saved so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 2648 may store current register values in a designated area of memory (e.g., identified by a context pointer). Then, when returning to the context, the context management circuit 2648 may restore the register values. In one embodiment, the interrupt management circuit 2647 receives and processes interrupts received from system devices.
[0314] In one implementation, virtual / effective addresses from the graphics processing engine 2631 are translated to real / physical addresses in the system memory 2614 by the MMU 2639. One embodiment of the accelerator integration circuit 2636 supports multiple (e.g., four, eight, or sixteen) graphics accelerator modules 2646 and / or other accelerator devices. The graphics accelerator modules 2646 may be dedicated to a single application running on the processor 2607 or may be shared among multiple applications. In one embodiment, a virtualized graphics execution environment exists in which the resources of the graphics processing engines 2631-2632, N, are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources may be subdivided into "slices," which are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0315] In at least one embodiment, the accelerator integrated circuitry 2636 acts as a bridge to the system for the graphics acceleration module 2646, providing address translation and system memory caching services. Additionally, the accelerator integrated circuitry 2636 may provide a virtualization facility for the host processor to manage virtualization, interrupts, and memory management for the graphics processing engines 2631-2632.
[0316] The hardware resources of the graphics processing engines 2631-2632, N are explicitly mapped into the real address space seen by the host processor 2607, so that any host processor can directly address these resources using effective address values. In one embodiment, one function of the accelerator integrated circuitry 2636 is to physically separate the graphics processing engines 2631-2632, N so that they appear as independent units to the system.
[0317] In at least one embodiment, one or more graphics memories 2633-2634, M are respectively coupled to each of the graphics processing engines 2631-2632, N. The graphics memories 2633-2634, M store instructions and data that are processed by the respective graphics processing engines 2631-2632, N. The graphics memories 2633-2634, M may be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or may be non-volatile memory such as 3D XPoint or Nano-Ram.
[0318] In one embodiment, to reduce data traffic over link 2640, a biasing technique is used to ensure that the data stored in graphics memory 2633-2634, M is data that will be most frequently used by graphics processing engines 2631-2632, N, and preferably is data that is not used (or at least not frequently used) by cores 2660A-2660D. Similarly, the biasing mechanism attempts to keep data needed by the cores (and therefore preferably not needed by graphics processing engines 2631-2632, N) in the cores' caches 2662A-2662D, 2656 and system memory 2614.
[0319] FIG. 26C shows another exemplary embodiment in which the accelerator integration circuitry 2636 is integrated within the processor 2607. In this embodiment, the graphics processing engines 2631-2632, N communicate directly with the accelerator integration circuitry 2636 via high-speed link 2640 via interface 2637 and interface 2635 (again, any form of bus or interface protocol can be utilized). The accelerator integration circuitry 2636 may perform the same operations as described with respect to FIG. 26B, but may potentially operate at a higher throughput given its proximity to the coherence bus 2664 and caches 2662A-2662D, 2656. One embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuitry 2636 and a programming model controlled by the graphics acceleration module 2646.
[0320] In at least one embodiment, graphics processing engines 2631-2632, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 2631-2632, N, achieving virtualization within a VM / partition.
[0321] In at least one embodiment, the graphics processing engines 2631-2632, N may be shared by multiple VM / application partitions. In at least one embodiment, the sharing model may use a system hypervisor to virtualize the graphics processing engines 2631-2632, N and allow access by each operating system. In a single-partition system without a hypervisor, the graphics processing engines 2631-2632, N are owned by the operating system. In at least one embodiment, the operating system may virtualize the graphics processing engines 2631-2632, N and provide access to each process or application.
[0322] In at least one embodiment, the graphics acceleration module 2646 or the individual graphics processing engines 2631-2632,N selects a process element using a process handle. In one embodiment, the process element is stored in system memory 2614 and is addressable using the effective address to real address translation techniques described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to a host process when registering the host process's context with the graphics processing engines 2631-2632,N (i.e., calling system software to add the process element to the process element linked list). In at least one embodiment, the low-order 16 bits of the process handle may be the offset of the process element within the process element linked list.
[0323] FIG. 26D illustrates an exemplary accelerator integration slice 2690. As used herein, a "slice" comprises a designated portion of the processing resources of the accelerator integration circuitry 2636. An application effective address space 2682 in system memory 2614 stores a process element 2683. In one embodiment, the process element 2683 is stored in response to a GPU call 2681 from an application 2680 executing on the processor 2607. The process element 2683 contains the process state of the corresponding application 2680. A work descriptor (WD) 2684 contained in the process element 2683 can be a single job requested by the application or may contain a pointer to a queue of jobs. In at least one embodiment, the WD 2684 is a pointer to a job request queue in the application's address space 2682.
[0324] The graphics acceleration module 2646 and / or the individual graphics processing engines 2631-2632, N may be shared by all or a subset of the processes in the system. In at least one embodiment, infrastructure may be included for setting process state and sending a WD 2684 to the graphics acceleration module 2646 to start a job in a virtualized environment.
[0325] In at least one embodiment, the dedicated process programming model is implementation specific, in which a single process owns the graphics acceleration module 2646 or an individual graphics processing engine 2631. Because the graphics acceleration module 2646 is owned by a single process, when the graphics acceleration module 2646 is allocated, the hypervisor initializes the accelerator integration circuitry 2636 for the owning partition, and the operating system initializes the accelerator integration circuitry 2636 for the owning process.
[0326] In operation, the WD fetch unit 2691 in the accelerator integrated slice 2690 fetches the next WD 2684, which contains an indication of work to be performed by one or more graphics processing engines of the graphics acceleration module 2646. As shown, data from the WD 2684 is stored in registers 2645 and may be used by the MMU 2639, the interrupt management circuit 2647, and / or the context management circuit 2648. For example, one embodiment of the MMU 2639 includes segment / page walk circuitry for accessing a segment / page table 2686 within the OS virtual address space 2685. The interrupt management circuit 2647 may process interrupt events 2692 received from the graphics acceleration module 2646. When performing graphics operations, effective addresses 2693 generated by the graphics processing engines 2631-2632, N are translated into real addresses by the MMU 2639.
[0327] In one embodiment, the same set of registers 2645 may be replicated for each graphics processing engine 2631-2632, N, and / or graphics acceleration module 2646 and initialized by the hypervisor or operating system. Each of these replicated registers may be included in the accelerator integration slice 2690. Exemplary registers that may be initialized by the hypervisor are shown in Table 1. [Table 1]
[0328] Exemplary registers that may be initialized by the operating system are shown in Table 2. [Table 2]
[0329] In one embodiment, each WD 2684 is specific to a particular graphics acceleration module 2646 and / or graphics processing engine 2631-2632, N. The WD 2684 contain...
Claims
1. One or more processors comprising a circuit for identifying the viewpoint of a first object in an image using one or more neural networks, One or more processors wherein one or more parameters of the one or more neural networks are updated at least in part based on one or more labels corresponding to one or more objects in one or more training images, and the one or more labels indicate one or more characteristics of the one or more objects other than the orientation of the one or more objects.
2. The one or more processors according to claim 1, wherein the one or more neural networks further identify the viewpoint based on at least a portion of a set of images of the same category as the image.
3. One or more processors according to claim 2, wherein ground truth annotation is not included in at least a portion of the set of images.
4. The one or more processors according to claim 1, wherein the one or more characteristics of the one or more objects include symmetrical consistency between the image of the first object and the flipped image of the first object.
5. The one or more processors according to claim 1, wherein the circuit is further configured to use the one or more neural networks to generate a second image depicting the first object in a second orientation, at least partially based on the viewpoint identified for the first object.
6. The one or more processors according to claim 1, wherein the viewpoint of the first object is encoded with respect to a set of parameters including an azimuth parameter, an elevation parameter, and an inclination parameter.
7. A system comprising one or more processors that identify viewpoints of a first object in an image using one or more neural networks, wherein the one or more neural networks are at least A loss value is generated based at least partially on the feature similarity between the characteristics of the first object and the corresponding characteristics of the second object. The one or more neural networks are updated according to the aforementioned loss value. A system that is trained by doing so.
8. The system according to claim 7, wherein the one or more processors are further configured to train the one or more neural networks using an unlabeled training dataset comprising multiple images of objects of the same category.
9. The system according to claim 7, wherein the loss value is further generated at least in part on an image consistency loss calculated based at least on the difference between the viewpoint of the first object and the viewpoint generated by the generative model.
10. The system according to claim 8, wherein ground truth annotation is not included in at least a portion of the training dataset used to train the one or more neural networks.
11. The system according to claim 7, wherein one or more processors are further configured to evaluate the symmetric consistency of the first object by comparing an image of the first object with a transformed version of the same image.
12. The system according to claim 7, wherein one or more neural networks are further trained to infer a composite viewpoint of the first and second objects using a generative adversarial network (GAN).
13. The system according to claim 7, wherein the training includes generating a composite image of the first object in different orientations and comparing the predicted viewpoint of the composite image with the predicted viewpoint of the original image.
14. A method comprising the step of identifying a viewpoint of a first object in an image using one or more neural networks, wherein the one or more neural networks are at least A loss value is generated based at least partially on the feature similarity between the characteristics of the first object and the corresponding characteristics of the second object. The one or more neural networks are updated according to the aforementioned loss value. A method of training by doing so.
15. The method according to claim 14, wherein the one or more neural networks are further trained using an unlabeled training dataset which includes a set of images of objects of the same category as the first object.
16. The method according to claim 15, wherein ground truth annotation is not available in at least a portion of the training dataset.
17. The method according to claim 14, wherein the properties of the first object are evaluated using symmetrical consistency between the image of the first object and a converted version of the image.
18. The method according to claim 14, wherein the training includes using a generator to create a composite image of an object using a plurality of viewpoints, and the composite image is evaluated to calculate the viewpoint consistency loss.
19. The method according to claim 15, wherein the training includes constructing a feature similarity graph across the entire training dataset and calculating nearest neighbor and farthest neighbor losses based on the viewpoints of the objects.
20. The method according to claim 14, wherein the one or more neural networks are further trained to infer a composite viewpoint of the first object and the second object using a generative adversarial network (GAN).