Methods or apparatuses for training and inference using neural networks to predict the direction of objects in images

By training a neural network using self-supervised learning and a deep generative model, and utilizing object features from an image set to generate synthetic images, and training using a loss function, the problem of lack of ground truth annotations during training is solved, and efficient object orientation prediction in images is achieved.

CN114787879BActive Publication Date: 2026-03-31NVIDIA CORP
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-17
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

When training neural networks to predict the viewpoint of objects in an image, the lack of or difficulty in obtaining ground truth annotations leads to a waste of memory, time, and computational resources.

Method used

A self-supervised learning approach is adopted, using deep generative models such as generative adversarial networks and variational autoencoders to train a neural network with object features in an image set, generating synthetic images with specific orientations. The neural network is trained using loss functions such as generative consistency loss, symmetry loss, and neighbor loss to predict the orientation of objects in the image.

Benefits of technology

In the absence of ground truth annotations, it effectively reduces the consumption of memory and computing resources, improves training efficiency and accuracy, and achieves self-supervised prediction of object orientation in images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114787879B_ABST
    Figure CN114787879B_ABST
Patent Text Reader

Abstract

Apparatuses, systems, and techniques for identifying a direction of an object in an image. In at least one embodiment, one or more neural networks are trained to identify a direction of one or more objects based at least in part on one or more features of the object other than the direction of the one or more objects.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application claims priority to U.S. Patent Application No. 16 / 690,015, filed November 20, 2019, entitled “Training and Inferring Using a Neural Network to Predict Orientations of Objects in Images,” the entire contents of which are incorporated herein by reference and used for all purposes. Technical Field

[0003] At least one embodiment relates to processing resources for training a neural network to predict the viewpoint of an object in an image. For example, at least one embodiment relates to a processor or computing system for training a neural network according to the various new techniques described herein. Background Technology

[0004] Training neural networks consumes significant amounts of memory, time, and computational resources. Training a neural network that requires ground truth annotations can be more challenging than training one that doesn't require some or all of its training data to be annotated with ground truth, at least because ground truth annotations may not always be available and / or may be difficult to obtain. The amount of memory, time, and / or computational resources used to train neural networks can be improved. Attached Figure Description

[0005] Figure 1 A diagram illustrating the use of a self-supervised neural network to predict the viewpoint of an object, according to at least one embodiment, is shown.

[0006] Figure 2 A graph depicting the loss function according to at least one embodiment is shown;

[0007] Figure 3 A diagram depicting a generative adversarial network according to at least one embodiment is shown;

[0008] Figure 4 A diagram depicting the discriminator update according to at least one embodiment is shown;

[0009] Figure 5 A diagram depicting the discriminator update according to at least one embodiment is shown;

[0010] Figure 6 A diagram showing the update of the depiction generator according to at least one embodiment is shown;

[0011] Figure 7 A diagram showing the update of the depiction generator according to at least one embodiment is shown;

[0012] Figure 8 A diagram depicting the symmetry loss according to at least one embodiment is shown;

[0013] Figure 9 A diagram depicting nearest neighbor loss and farthest neighbor loss according to at least one embodiment is shown;

[0014] Figure 10 A diagram depicting the decoupling loss according to at least one embodiment is shown;

[0015] Figure 11 A diagram depicting a calibration neural network according to at least one embodiment is shown;

[0016] Figure 12 A diagram depicting reasoning according to at least one embodiment is shown;

[0017] Figure 13 An illustrative example of a process for training a neural network to predict the viewpoint of an object within an image, according to at least one embodiment, is shown.

[0018] Figure 14 An illustrative example of a process for training a neural network to predict the viewpoint of an object within an image, according to at least one embodiment, is shown.

[0019] Figure 15A An illustrative example of a process for calculating generative consistency loss according to at least one embodiment is shown;

[0020] Figure 15B An illustrative example of a process for calculating viewpoint consistency loss according to at least one embodiment is shown;

[0021] Figure 16A An illustrative example of a process for calculating symmetry loss according to at least one embodiment is shown;

[0022] Figure 16B An illustrative example of a process for calculating symmetry loss according to at least one embodiment is shown;

[0023] Figure 17 An illustrative example of the process for calculating the nearest neighbor loss and the farthest neighbor loss according to at least one embodiment is shown;

[0024] Figure 18A The inference and / or training logic according to at least one embodiment is illustrated;

[0025] Figure 18B The inference and / or training logic according to at least one embodiment is illustrated;

[0026] Figure 19The training and deployment of a neural network according to at least one embodiment are illustrated;

[0027] Figure 20 An example data center system according to at least one embodiment is shown;

[0028] Figure 21A An example of an autonomous vehicle according to at least one embodiment is shown;

[0029] Figure 21B The illustration shows an embodiment according to at least one of the embodiments. Figure 21A Examples of camera positions and field of view for autonomous vehicles;

[0030] Figure 21C This is an illustration based on at least one embodiment. Figure 21A A block diagram of an example system architecture for an autonomous vehicle;

[0031] Figure 21D The illustration, according to at least one embodiment, is for one or more cloud-based servers and Figure 21A A diagram of a system for communication between autonomous vehicles;

[0032] Figure 22 This is a block diagram illustrating a computer system according to at least one embodiment;

[0033] Figure 23 This is a block diagram illustrating a computer system according to at least one embodiment;

[0034] Figure 24 A computer system according to at least one embodiment is shown;

[0035] Figure 25 A computer system according to at least one embodiment is shown;

[0036] Figure 26A A computer system according to at least one embodiment is shown;

[0037] Figure 26B A computer system according to at least one embodiment is shown;

[0038] Figure 26C A computer system according to at least one embodiment is shown;

[0039] Figure 26D A computer system according to at least one embodiment is shown;

[0040] Figure 26E and Figure 26F A shared programming model according to at least one embodiment is shown;

[0041] Figure 27An exemplary integrated circuit and a related graphics processor according to at least one embodiment are shown.

[0042] Figure 28A and Figure 28B An exemplary integrated circuit and an associated graphics processor according to at least one embodiment are shown.

[0043] Figure 29A and Figure 29B Additional exemplary graphics processor logic according to at least one embodiment is shown;

[0044] Figure 30 A computer system according to at least one embodiment is shown;

[0045] Figure 31A A parallel processor according to at least one embodiment is shown;

[0046] Figure 31B A partitioning unit according to at least one embodiment is shown;

[0047] Figure 31C A processing cluster according to at least one embodiment is shown;

[0048] Figure 31D A graphics multiprocessor according to at least one embodiment is shown;

[0049] Figure 32 A multi-graphics processing unit (GPU) system according to at least one embodiment is illustrated;

[0050] Figure 33 A graphics processor according to at least one embodiment is shown;

[0051] Figure 34 It is a block diagram illustrating a processor microarchitecture for a processor according to at least one embodiment;

[0052] Figure 35 A deep learning application processor according to at least one embodiment is shown;

[0053] Figure 36 A block diagram of an example neuromorphic processor is shown according to at least one embodiment;

[0054] Figure 37 At least a portion of a graphics processor according to one or more embodiments is shown;

[0055] Figure 38 At least a portion of a graphics processor according to one or more embodiments is shown;

[0056] Figure 39 At least a portion of a graphics processor according to one or more embodiments is shown;

[0057] Figure 40 A block diagram of a graphics processing engine 4010 of a graphics processor is shown according to at least one embodiment;

[0058] Figure 41 This is a block diagram illustrating at least a portion of a graphics processor core according to at least one embodiment;

[0059] Figure 42A and Figure 42B A thread execution logic 4200 according to at least one embodiment is shown, which includes an array of processing elements of a graphics processor core.

[0060] Figure 43 A parallel processing unit (“PPU”) according to at least one embodiment is shown;

[0061] Figure 44 A general-purpose processing cluster (“GPC”) according to at least one embodiment is illustrated;

[0062] Figure 45 A memory partition unit of a parallel processing unit (“PPU”) according to at least one embodiment is shown; and

[0063] Figure 46 A streaming multiprocessor according to at least one embodiment is shown. Detailed Implementation

[0064] In at least one embodiment, the neural network is trained to self-supervisedly identify the orientation of objects within images on a set of images, such as those described elsewhere in this disclosure. In at least one embodiment, the neural network is self-supervised to identify the orientation of objects within images by computing one or more loss functions as part of training to evaluate one or more features of images in a training set (e.g., an image set). In at least one embodiment, the neural network is trained on an image set lacking ground truth annotations or where ground truth annotations are otherwise unavailable (e.g., such data is retained from the neural network during training). In at least one embodiment, the neural network is trained to generate a second image with the same orientation from objects within a first image having a predicted orientation. In at least one embodiment, the predicted orientation or viewpoint is encoded as azimuth parameters, elevation parameters, and tilt parameters.

[0065] In at least one embodiment, one or more neural networks are trained in a self-supervised manner on a set of images of different objects of the same category as the object in the image to be inferred. In at least one embodiment, different objects of the same category can refer to different images, such as one or more images of a first car in one or more directions, one or more images of different second cars in one or more directions, and so on. In at least one embodiment, the image of the object to be inferred is included in the set of images used to train one or more neural networks to inference directions. In at least one embodiment, one or more neural networks are trained in a self-supervised manner by evaluating one or more characteristics of objects within an image using at least a set of loss functions. In at least one embodiment, the one or more characteristics of an object refer to attributes of the object that can be used for inference directions. In at least one embodiment, the neural networks are trained in a self-supervised manner to generate synthetic images of objects with a specific orientation, which may be the same orientation as the predicted orientation of the input image. In at least one embodiment, synthetic images are created using a deep generative model such as a variational autoencoder (VAE), a differentiable renderer, or a generative adversarial network (GAN), or via a renderer. In at least one embodiment, the object whose orientation is to be inferred can be a vehicle, an airplane, a drone, a human, (e.g., a human or animal's) face, etc.

[0066] In at least one embodiment, self-supervised learning (e.g., training) refers to a form of learning in which a neural network is trained on a training set, wherein the data of the training set does not include any ground truth annotations, but the data of the training set is partially labeled (e.g., semi-supervised learning). In at least one embodiment, training a neural network in a self-supervised manner to identify the orientation of objects within an image utilizes a training image set, wherein the images in the training set do not include ground truth annotations representing the orientation of objects within the image, and do not include labels or other information that otherwise identifies the individual objects in the image (e.g., one image in the image includes labels or other information that otherwise identifies the objects in the image, but does not include any annotations representing the orientation of the objects).

[0067] In at least one embodiment, semi-supervised learning refers to a form of learning in which a neural network is trained on a training set, wherein only a portion of the data in the training set includes truth annotations. In at least one embodiment, fully supervised learning refers to a form of learning in which a neural network is trained on a training set, wherein all data in the training set includes truth annotations. In at least one embodiment, unsupervised learning refers to a form of learning in which a neural network is trained on a training set, wherein none of the data in the training set includes truth annotations.

[0068] Figure 1 Figure 100 illustrates the use of a self-supervised neural network to predict the viewpoint of an object according to at least one embodiment. In at least one embodiment, Figure 100 is implemented by one or more systems, such as the systems described in Figures 18-46. In at least one embodiment, Figure 100 includes one or more neural networks associated with a discriminator 106, which is trained using self-supervised learning on a set of images of a category to infer the viewpoint of objects within other images of that category. In at least one embodiment, an image is provided as input to the neural network to detect the orientation of an object of a certain category. In at least one embodiment, the input image is provided to multiple neural networks trained using the self-supervised learning techniques described herein to identify the orientation or viewpoint of different objects in the input image.

[0069] In at least one embodiment, the viewpoint of an image refers to the orientation of an object within the image, specifically the three-dimensional orientation of an object captured within a two-dimensional image. In at least one embodiment, a camera is used to capture a two-dimensional image of a real-world object (e.g., a car) in a specific orientation relative to the camera. In at least one embodiment, the orientation of the object (e.g., viewpoint) is encoded on a set of parameters including azimuth, elevation, and tilt parameters. In at least one embodiment, the orientation of the object within the image is encoded as a set of three vectors that define the orientation of the object relative to the canonical x, y, and z axes.

[0070] In at least one embodiment, an image set 102 is obtained. In at least one embodiment, the image set 102 is a collection of one or more images of an object of one type. In at least one embodiment, the image set 102 is used to train one or more neural networks to identify the orientation of objects within the images. In at least one embodiment, the image set 102 is classified or labeled for each object displaying the same type or category. In at least one embodiment, the image set 102 is a collection of car images, which may include different types of cars in different directions, different weather, different lighting, etc. In at least one embodiment, the image set 102 includes images of the same car or the same type of car in different directions. In at least one embodiment, at least a portion of the image set 102 lacks truth value annotations specifying the orientation of objects within such training images. In at least one embodiment, all images in the image set 102 lack truth value annotations specifying the azimuth, elevation, and tilt of objects within the images in the set. In at least one embodiment, truth value annotations refer to annotations that the images in the training image set may include, for one or more neural networks configured to determine one or more features of an image (where one or more neural networks are trained on a training image set), indicate the expected one or more features of the image. In at least one embodiment, the image set 102 includes one or more synthetic images, such as images created from a variational autoencoder (VAE), a generative adversarial network (GAN), or a renderer. In at least one embodiment, all images in the image set 102 are real images, as opposed to images synthesized or created from generative models such as variational autoencoders (VAEs), renderers, or generative adversarial networks. In at least one embodiment, the image set 102 is collected and aggregated from a website that categorizes images.

[0071] In at least one embodiment, discriminator 106 is trained to identify the orientation of objects within image 104 based at least in part on one or more features of the object rather than the object's orientation. In at least one embodiment, discriminator 106 is a classifier within one or more neural networks. In at least one embodiment, discriminator 106 is a component of one or more neural networks and includes other neural networks, classifiers, and various other machine learning components. In at least one embodiment, discriminator 106 is the discriminator network of a generative adversarial network. In at least one embodiment, discriminator 106 is part of one or more neural networks and is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, discriminator 106 is trained on a set of images of a category (e.g., cars) to infer the orientation of other objects of the same category captured in other images. In at least one embodiment, discriminator 106 is trained in a self-supervised manner on image set 102. In at least one embodiment, discriminator 106 is trained to identify the orientation of objects within image 104 in a self-supervised manner by computing one or more loss functions as part of training to evaluate one or more features of images in the training set (e.g., image set 102). In at least one embodiment, the neural network associated with the discriminator 106 is trained at least in part based on computational generative consistency loss, symmetry loss, nearest neighbor loss, farthest neighbor loss, and decoupling loss. In at least one embodiment, it can be based on a combination of Figure 2-10 The described techniques are used to train a neural network to identify the orientation of objects. In at least one embodiment, the discriminator 106 is trained on a set of images that lack ground truth annotations or whose ground truth annotations are otherwise unavailable (e.g., such data is preserved during training).

[0072] In at least one embodiment, image 104 is obtained for discriminator 106. In at least one embodiment, objects in image 104 belong to the same type as objects in images in image set 102 used to train one or more neural networks. In at least one embodiment, image 104 is provided to the neural network for inference to predict orientation. In at least one embodiment, a first system trains one or more neural networks and a second, different system uses those one or more neural networks to perform inference to identify the orientation of objects in the image. In at least one embodiment, discriminator 106 is trained in a self-supervised manner on image set 102 of objects of a particular category to infer the orientation of other objects of that category (e.g., objects in image 104). In at least one embodiment, one or more neural networks associated with discriminator 106 are trained on a set of car images and used to infer the orientation of cars captured in real time by a camera or other suitable video / image capture device attached to the vehicle. In at least one embodiment, discriminator 106 is trained in a self-supervised manner on image set 102 to determine the orientation 108 of objects depicted in image 104. In at least one embodiment, discriminator 106 determines the orientation 108 of cars depicted in image 104.

[0073] Figure 2 Figure 200 illustrates a depiction of a loss function according to at least one embodiment. In at least one embodiment, Figure 200 is implemented by one or more systems, such as the systems described in Figures 18-46. In at least one embodiment, a discriminator 204 is associated with one or more neural networks and trained using at least one of a true image generative consistency loss 208, a nearest and farthest neighbor loss 210, a symmetry loss 212, and a true / false classification loss 214. In at least one embodiment, the discriminator 204 is part of one or more neural networks trained to infer viewpoints from input images, said one or more neural networks including various parameters associated with one or more processes of said one or more neural networks, and updated at least in part based on the true image generative consistency loss 208, the nearest and farthest neighbor loss 210, and the symmetry loss 212.

[0074] In at least one embodiment, a set 202 of object images of one type of object is obtained for the discriminator 204. In at least one embodiment, the set 202 of object images includes all images comprising objects of the same type. In at least one embodiment, the set 202 of object images includes images of cars in different orientations, different weather conditions, different lighting conditions, etc. In at least one embodiment, the system uses techniques described elsewhere in this disclosure (e.g., Figure 13 ) Obtain the object image collection 202.

[0075] In at least one embodiment, images from the object image set 202 are selected as input images to the discriminator 204. In at least one embodiment, images from the set are selected for learning in any suitable manner, which may be randomly or pseudo-randomly sampled from the training set. In at least one embodiment, the discriminator 204 predicts a viewpoint 206 of the input image. In at least one embodiment, the viewpoint 206 of the input image is inferred by the discriminator 204 through one or more processes involving one or more neural networks, which include one or more input parameters indicating one or more processes involved in said one or more neural networks. In at least one embodiment, the viewpoint 206 is determined based on truth annotations provided as part of training as at least a portion of the object image set 202. In at least one embodiment, the viewpoint 206 corresponds to a prediction of the orientation of an object within an image input to the discriminator 204. In at least one embodiment, the orientation of the object (e.g., viewpoint) is encoded on a set of parameters including azimuth parameters, elevation parameters, and tilt parameters.

[0076] In at least one embodiment, a generative consistency loss 208 is calculated for the discriminator 204. In at least one embodiment, the generative consistency loss 208 is calculated at least in part based on image consistency loss by comparing the selected image with an image generated by the deep generative model and viewpoint consistency loss by comparing the input viewpoint with the generative model and its predicted values ​​by the discriminator. In at least one embodiment, the generative consistency loss 208 is calculated using techniques described elsewhere in this disclosure, such as combining... Figure 4-7 The discussion focuses on those aspects. In at least one embodiment, the generative consistency loss 208 comprises at least two components: a synthetic image viewpoint consistency loss and a grounded image consistency loss. In at least one embodiment, the viewpoint consistency loss can be represented as a orientation consistency loss. In at least one embodiment, the generative consistency loss is applied to grounded images (e.g., images from the object image set 202), rather than synthetic images created by the generator. In at least one embodiment, the viewpoint consistency loss and the image consistency loss are used to determine the generative consistency loss. In at least one embodiment, the generative consistency loss is a combination of the viewpoint consistency loss and the image consistency loss. In at least one embodiment, the generative consistency loss is determined by the following symbolic mathematical equation:

[0077] L gc =L vc +L ic

[0078] Where L gc Corresponding to generative consistency loss, L vc Corresponding to viewpoint consistency loss, L icCorresponding image consistency loss.

[0079] In at least one embodiment, image consistency loss is calculated at least in part based on images in a set of object images 202 input to a discriminator 204, which determines at least two attributes from the input images: a viewpoint 206 and an appearance parameter set. In at least one embodiment, the viewpoint 206 and the appearance parameter set are provided to a generator to create a synthetic image. In at least one embodiment, a generative adversarial network (GAN) receives the viewpoint 206 and the appearance parameter set and generates a synthetic (e.g., fake) image based on the viewpoint 206 and the appearance parameter set. In at least one embodiment, the synthetic image and the input image are compared to determine the image consistency loss. In at least one embodiment, the cosine distance between the input image and the synthetic image is compared to determine feature similarity, where closer similarity corresponds to a lower loss. In at least one embodiment, L1 distance, L2 distance, or cosine distance is used to determine the image consistency loss between two images.

[0080] In at least one embodiment, the viewpoint consistency loss is calculated at least in part based on the viewpoint of the input image (e.g., viewpoint 206). In at least one embodiment, a generator is used to create a synthetic image from viewpoint 206 predicted by discriminator 204 based on the input image. In at least one embodiment, the synthetic image generated from viewpoint 206 is provided to discriminator 204, which determines a second viewpoint of the synthetic image. In at least one embodiment, viewpoint 206 is compared with a second viewpoint of the synthetic image generated at least in part based on viewpoint 206. In at least one embodiment, the distance between viewpoint 206 and the second viewpoint of the synthetic image is used to calculate the viewpoint consistency loss, wherein a closer viewpoint corresponds to a lower loss. In at least one embodiment, the generative consistency loss is calculated according to the technique described in conjunction with FIG15.

[0081] In at least one embodiment, the true / false classification loss 214 is calculated based on whether the discriminator 204 can correctly predict whether the input image to the discriminator 204 is a real image or a synthetic image. In at least one embodiment, the true / false classification loss 214 is calculated based on whether the discriminator 204 can correctly predict whether the input image set is real or fake, wherein the discriminator 204 may be provided with real or fake (e.g., synthetic) images and used to predict whether these images are real or fake. As part of training the discriminator 204, ground truth values ​​regarding whether the images provided to the discriminator 204 are real or fake can be used as part of the training (e.g., for calculating the loss).

[0082] In at least one embodiment, a symmetric loss 212 is calculated at least by comparing an input image with a transformed version of the input image. In at least one embodiment, an input image is selected from a set of object images 202. In at least one embodiment, a transformation is applied to the input image to generate a transformed image. In at least one embodiment, the input image is horizontally flipped to generate a transformed image. In at least one embodiment, a discriminator 204 is used to predict a viewpoint 206 of the input image and a second viewpoint of the transformed image. In at least one embodiment, the viewpoint 206 of the input image is predicted, and a second viewpoint of the horizontally flipped version of the input image is predicted. In at least one embodiment, a loss is calculated based on whether certain properties hold true. In at least one embodiment, a transformation or its inverse transformation is applied to the predicted viewpoint of the transformed version of the input image. In at least one embodiment, if the input image is rotated by an angle (Φ, θ, ψ) to generate a transformed image, the inferred viewpoint of the transformed image can be rotated in the opposite direction by an angle (-Φ, -θ, -ψ). In at least one embodiment, the loss is calculated by comparing the magnitudes of the azimuth, elevation, and tilt of viewpoint 206 of the input image with those of a second viewpoint of the transformed image, wherein zero loss is produced when the magnitudes of each direction parameter are equal. In at least one embodiment, a symmetric loss is calculated according to techniques described elsewhere in this disclosure, such as combining... Figure 8 And those discussed in Figure 16. In at least one embodiment, the loss is calculated by determining how well the first set of appearance parameters predicted by the discriminator 204 for the image matches the second set of appearance parameters predicted by the discriminator 204 for the transformed version of the image.

[0083] In at least one embodiment, the nearest neighbor loss and the farthest neighbor loss 210 are calculated at least partially by comparing an input image in the object image set 202 with its nearest and farthest neighbors based on the viewpoint map of the object image set 202. In at least one embodiment, the nearest neighbor and farthest neighbor losses 210 are based on a combination of... Figure 4 , Figure 9 and Figure 17The techniques described are used for computation. In at least one embodiment, a set of object images 202 is used to generate a viewpoint map, wherein the nodes of such a map correspond to images and the edges correspond to their viewpoint equivariant distances (e.g., cosine distances). In at least one embodiment, the cosine distance is computed using a convolutional neural network (CNN) based on the feature similarity of image pairs. In at least one embodiment, an anchor image is selected from the set of object images 202. In at least one embodiment, the anchor image is located from the viewpoint map and the nearest and farthest neighbors are selected based on edge weights. In at least one embodiment, the nearest neighbor has the shortest edge connected to the anchor image. In at least one embodiment, the farthest neighbor has the farthest edge connected to the anchor image. In at least one embodiment, a discriminator 204 predicts a first viewpoint of the anchor image (e.g., viewpoint 206 predicted for the anchor image) and predicts a second viewpoint of the nearest neighbor image (e.g., viewpoint 206 predicted for the nearest neighbor image), and computes a loss such that closer distances between these viewpoints correspond to smaller losses. In at least one embodiment, the discriminator 204’s neural network predicts a first viewpoint of the anchor image and a third viewpoint of the farthest neighbor image (e.g., viewpoint 206 of the farthest neighbor image) and calculates a loss such that a longer distance between those viewpoints corresponds to a smaller loss.

[0084] In at least one embodiment, the calculated losses (e.g., generative consistency loss 208, nearest and farthest neighbor loss 210, symmetry loss 212, and true / false classification loss 214) are used to update the parameters of one or more neural networks associated with a discriminator 204 trained on the object image set 202. In at least one embodiment, the system implementing FIG200 includes executable code for continuously updating the parameters of one or more neural networks associated with the discriminator 204, such that the one or more neural networks and the discriminator 204 are trained to infer viewpoints and other features of the input images. In at least one embodiment, training is performed according to any suitable technique and may include selecting and utilizing various additional images from the object image set 202 to calculate losses and refine parameters for one or more neural networks trained to infer viewpoints. In at least one embodiment, once training is complete, the trained neural networks are made available (e.g., the neural networks or their parameters are transferred to a different system) for inference.

[0085] Figure 3Figure 300 illustrates a generative adversarial network according to at least one embodiment. In at least one embodiment, Figure 300 is implemented by one or more systems, such as the systems described in Figures 18-46. In at least one embodiment, Figure 300 includes a generator 306 that utilizes an input viewpoint 302 and an input appearance parameter set 304 to synthesize an image 308. In at least one embodiment, Figure 300 illustrates a discriminator 310 that utilizes the input image 318 and outputs an output viewpoint 312, an output determination 314, and an output appearance parameter set 316. In at least one embodiment, a combination of... Figure 4-7 The described techniques are used to select the parameters of generator 306 and / or discriminator 310.

[0086] In at least one embodiment, the input viewpoint 302 corresponds to the orientation of an object within an image, referring to the three-dimensional orientation of the object captured within a two-dimensional image. In at least one embodiment, the orientation of the object (e.g., viewpoint) is encoded on a set of parameters including azimuth, elevation, and tilt parameters. In at least one embodiment, the input viewpoint 302 corresponds to a specific orientation of the object and includes specific values ​​for the parameter set, which includes azimuth, elevation, and tilt parameters. In at least one embodiment, the input viewpoint 302 indicates a 3D rotation of the object (e.g., the input viewpoint 302 may specify that the object is rotated by a specified number of degrees on a specified axis and variations thereof). In at least one embodiment, the input appearance parameter set 304 is a set of parameters that define the appearance of the object. In at least one embodiment, the object includes a vehicle, an aircraft, a drone, a person, (e.g., a human or animal) face, etc. In at least one embodiment, the input appearance parameter set 304 corresponds to the appearance parameters of a car, such as color, size, wheel type, and various other parameters that define the appearance of the car.

[0087] In at least one embodiment, an input viewpoint 302 and an input appearance parameter set 304 are provided to a generator 306 to create an image 308. In at least one embodiment, the generator 306 and the discriminator 310 are part of a generative adversarial network (GAN). In at least one embodiment, the generator 306 is a generative network within a generative adversarial network. In at least one embodiment, the generator 306 is part of one or more neural networks and is trained to generate images based on the input viewpoint and the input appearance parameter set. In at least one embodiment, the generator 306 receives the input viewpoint 302 and the input appearance parameter set 304 and generates an image 308, which is a synthesized (e.g., fake) image based on the input viewpoint 302 and the input appearance parameter set 304. In at least one embodiment, the generator accepts two separate (e.g., independent) parameters for creating image 308: an input viewpoint 302, which indicates the specific viewpoint from which image 308 will be generated (e.g., encoded azimuth, elevation, and tilt parameters), and appearance parameters 304, which encode appearance attributes of image 308 (e.g., for a car, such attributes may include color, brand, model, year of manufacture, etc.). In at least one embodiment, generator 306 generates image 308 comprising objects generated according to the input appearance parameter set 304, oriented relative to the input viewpoint 302. In at least one embodiment, image 308 is a composite image including a car, wherein the car's appearance corresponds to the input appearance parameter set 304, and the car's orientation corresponds to the input viewpoint 302.

[0088] In at least one embodiment, the generative adversarial network (GAN) includes a discriminator 310. In at least one embodiment, the discriminator 310 accepts an input image 318 and generates an output viewpoint 312, an output determination 314, and an output appearance parameter set 316. In at least one embodiment, the input image 318 may be a real image or a synthetic image. In at least one embodiment, the input image 318 is retrieved from one or more other sources, such as an image database, one or more cameras, and / or variations thereof. In at least one embodiment, the discriminator 310 processes an image 308. In at least one embodiment, the discriminator 310 includes various neural networks and machine learning processes. In at least one embodiment, the discriminator 310 is implemented according to those described elsewhere in this disclosure, such as combining... Figure 2 The aforementioned. In at least one embodiment, the discriminator 310 is associated with one or more neural networks trained to infer viewpoints and other features of the input image. In at least one embodiment, the discriminator 310 is refined through various processes involving the computation of various loss functions used to update various parameters associated with the discriminator 310.

[0089] In at least one embodiment, the discriminator 310 receives an input image 318 and generates an output viewpoint 312, an output determination 314, and an output appearance parameter set 316. In at least one embodiment, the output viewpoint 312 is a predicted viewpoint of the input image 318 generated by one or more processes of the discriminator 310. In at least one embodiment, the output determination 314 is a determination generated by one or more processes of the discriminator 310, indicating whether the input image 318 is a real image or a synthetic (e.g., fake) image. In at least one embodiment, the determination 314 is a binary output (e.g., a true / false indicator that the discriminator 310 believes the input image 318 is a real image or a synthetic image). In at least one embodiment, the determination 314 is a value between 0 and 1 (including or excluding one or both endpoints) that encodes a confidence value for whether the discriminator 310 considers the input image 318 to be real or fake (e.g., 0.5 indicates that the image is equally likely to be real or fake; 0 indicates that the image is highly likely to be fake). In at least one embodiment, the output appearance parameter set 316 is a predicted appearance parameter set of the input image 318 generated by one or more processes of the discriminator 310.

[0090] In at least one embodiment, if the discriminator 310 is accurately calibrated (e.g., the discriminator 310 is trained to the desired accuracy or the desired acceptable loss) and image 308 is generated by generator 306, then output determination 314 indicates that image 308 is fake, and output viewpoint 312 and output appearance parameter set 316 are exactly the same as input viewpoint 302 and input appearance parameter set 304, respectively. In at least one embodiment, if the discriminator 310 is not accurately calibrated (e.g., the discriminator 310 is not fully trained to the desired accuracy or the desired acceptable loss) and image 308 is generated by generator 306, then output determination 314 indicates an incorrect determination (e.g., if image 308 is synthetic, output determination 314 would indicate that image 308 is real), and output viewpoint 312 and output appearance parameter set 316 are different from input viewpoint 302 and input appearance parameter set 304, respectively. In at least one embodiment, comparisons between the output viewpoint 312 and the output appearance parameter set 316, and between the input viewpoint 302 and the input appearance parameter set 304, are used to evaluate and further process, train and / or calibrate the discriminator 310 and the generator 306, respectively.

[0091] Figure 4 Figure 400 depicts a discriminator update according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update the parameters of discriminator 404, which is used to predict various outputs from the input image. In at least one embodiment, Figure 4The diagram illustrates an input image 402; a discriminator 404; a predicted viewpoint 406; a prediction determination 408 to determine whether the input image 402 is real or fake; an appearance parameter set 410; a generator 412; a generated image 414; a real / fake classification loss 416; an image consistency loss 418; nearest and farthest neighbor losses 420; and a symmetry loss 422. In at least one embodiment, Figure 4 The discriminator update using a real image (e.g., not an image synthesized by a generator) is illustrated. In at least one embodiment, combined with Figure 4 The described technology and combination Figure 5-7 The techniques described for training generators and / or discriminators have the same scalability.

[0092] In at least one embodiment, discriminator 404 processes input image 402. In at least one embodiment, discriminator 404 is associated with one or more neural networks trained to infer viewpoint and other features of the input image. In at least one embodiment, discriminator 404 receives input image 402 and generates a predicted viewpoint 406, a determination 408 of whether input image 402 is real or fake, and an appearance parameter set 410. In at least one embodiment, viewpoint 406 is a predicted viewpoint of input image 402 determined by one or more processes of discriminator 404. In at least one embodiment, viewpoint 406 corresponds to a predicted specific orientation of an object within the image and includes specific values ​​for a parameter set including azimuth, elevation, and tilt parameters. In at least one embodiment, viewpoint 406 indicates a 3D rotation of an object (e.g., viewpoint 406 may specify that the object is rotated a specified number of degrees on a specified axis and variations thereof). In at least one embodiment, viewpoint 406 includes a predicted orientation of a car depicted in input image 402. In at least one embodiment, determination 408 determines whether the input image 402 is a real image or a fake image. In at least one embodiment, a fake image refers to a synthetic image created by a generative adversarial network. In at least one embodiment, determination 408 is a binary value (e.g., a predicted true / false value indicating whether the input image 402 is real or fake). In at least one embodiment, determination 408 is a non-binary value indicating a confidence level indicating whether the input image 402 is real or fake. In at least one embodiment, appearance parameter set 410 is a predicted set of appearance parameters for the input image 402 generated by one or more processes of discriminator 404. In at least one embodiment, appearance parameter set 410 is a set of predicted parameters defining the appearance of an object depicted in the input image 402. In at least one embodiment, appearance parameter set 410 corresponds to a predicted set of appearance parameters for a car depicted in the input image 402, such as predicted color, size, wheel type, and various other parameters defining the appearance of the car depicted in the input image 402.

[0093] In at least one embodiment, viewpoint 406 and appearance parameter set 410 are provided to generator 412 to generate generated image 414. In at least one embodiment, generator 412 creates synthetic image. In at least one embodiment, generator 412 is part of a generative adversarial network (GAN). In at least one embodiment, generator 412 receives viewpoint 406 and appearance parameter set 410 and generates generated image 414, which is a synthetic (e.g., fake) image based on viewpoint 406 and appearance parameter set 410. In at least one embodiment, generator 412 generates generated image 414 which includes an object generated according to appearance parameter set 410 and oriented according to viewpoint 406. In at least one embodiment, generated image 414 is a synthetic image including a car, wherein the appearance of the car corresponds to appearance parameter set 410 and the orientation of the car corresponds to viewpoint 406.

[0094] In at least one embodiment, determination 408 is used to determine a classification loss, such as a true / false classification loss 416. In at least one embodiment, the true / false classification loss 416 is calculated based on whether the discriminator 404 can correctly predict whether the input image of the discriminator 404 is a real image or a synthetic image. In at least one embodiment, the true / false classification loss 416 is calculated based on whether the discriminator 404 can correctly predict whether the input image set is real or fake, wherein the discriminator 404 may be provided with real or fake (e.g., synthetic) images and used to predict which images are real or fake. As part of training the discriminator 404, ground truth values ​​regarding whether the images provided to the discriminator 404 are real or fake may be used as part of the training (e.g., to calculate the loss).

[0095] In at least one embodiment, the generated image 414 and the input image 402 are compared to determine an image consistency loss 418. In at least one embodiment, the cosine distance between the input image 402 and the generated image 414 is compared to determine feature similarity, wherein closer similarity corresponds to lower loss. In at least one embodiment, at least one of L1, L2, or cosine distance is used to determine the image consistency loss 418 between the input image 402 and the generated image 414. In at least one embodiment, the L1 distance is determined by the following symbolic mathematical equation:

[0096] L1Distance=I in -I gen

[0097] Where I in The representation of the corresponding input image, I gen The representation of the corresponding generated image.

[0098] In at least one embodiment, the L2 distance is determined by the following symbolic mathematical equation:

[0099]

[0100] Where I in The representation of the corresponding input image, I gen The representation of the corresponding generated image.

[0101] In at least one embodiment, the cosine distance is determined by the following symbolic mathematical equation:

[0102] Cosine Distance = f in -f gen

[0103] Where f in The representation of the features of the corresponding input image, f gen The representation of the features of the generated image.

[0104] In at least one embodiment, as a combination Figure 4 An additional loss function is calculated as part of the discriminator update described herein. In at least one embodiment, the nearest and farthest neighbor loss 420 is calculated. In at least one embodiment, a symmetric loss 422 is calculated. In at least one embodiment, the nearest and farthest neighbor loss 420 and / or the symmetric loss 422 are calculated according to techniques described elsewhere, such as combining... Figure 8-10 Those discussed. In at least one embodiment, the calculated loss (e.g., Figure 4 The parameters shown are used to calculate gradients and update the parameters of discriminator 410, while keeping the parameters of generator 412 constant using any suitable technique (e.g., gradient descent).

[0105] Figure 5 Figure 500 depicts a discriminator update according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update the parameters of discriminator 504, which is used to predict various outputs from the input image. In at least one embodiment, Figure 5 Viewpoint 502; appearance parameter set 504; generator 506; generated image 508; discriminator 510; predicted viewpoint 512; prediction determination of whether the generated image 508 is real or fake 514; predicted appearance parameter set 516; viewpoint consistency loss 518; Z reconstruction loss 520; and real / fake classification loss 522 are shown. In at least one embodiment, Figure 5 The illustration depicts a discriminator update using a synthetic image (e.g., an image created by a generative adversarial network). In at least one embodiment, it combines... Figure 5 The described technology and combination Figure 4 , Figure 6 and Figure 7The techniques described for training generators and / or discriminators have the same scalability.

[0106] In at least one embodiment, the viewpoint 502 and the appearance parameter set 504 are selected in any suitable manner, which may include random selection of parameter values, weighted random selection, etc. In at least one embodiment, the viewpoint 502 and the appearance parameter set 504 are decoupled parameters that can be selected independently. In at least one embodiment, the generator 506 accepts the viewpoint 502 and the appearance parameter set 504 as input and creates a generated image 508. In at least one embodiment, the generated image 508 is a synthetic image having an appearance generated based on the appearance parameter set 504 and oriented according to the viewpoint 502.

[0107] In at least one embodiment, an image set (e.g., generated image 508) is provided as input to a discriminator 510, and the discriminator predicts various attributes of the image set. In at least one embodiment, the discriminator 510 receives the generated image 508 and produces a predicted viewpoint 512; a prediction determination 514 of whether the generated image 508 is real or fake; and a predicted set of appearance parameters 516. In at least one embodiment, the discriminator does not have access to the viewpoint 502 and the set of appearance parameters 504 used to create the generated image 508 (e.g., such information is concealed from the discriminator 510 during prediction). In at least one embodiment, the output of the discriminator 510 is used to compute a loss. In at least one embodiment, the loss function is used to compute gradients (e.g., using gradient descent) and update the parameters of the discriminator 510 while fixing the parameters of the generator 506 to a constant.

[0108] In at least one embodiment, a viewpoint consistency loss 518 is calculated. In at least one embodiment, the viewpoint consistency loss refers to a loss function calculated based on the accuracy of the discriminator 510 in predicting the viewpoint. In at least one embodiment, the viewpoint consistency loss 518 is calculated as the difference or distance between the input viewpoint 502 and the predicted viewpoint 512. In at least one embodiment, the viewpoint consistency loss is a component of the generative consistency loss. In at least one embodiment, the distance between the input viewpoint 502 and the predicted viewpoint 512 (e.g., L1 distance, L2 distance, cosine distance, and / or variations thereof) is used to calculate the viewpoint consistency loss 518, wherein closer viewpoints (e.g., viewpoints with shorter distances to each other) correspond to a lower loss.

[0109] In at least one embodiment, a Z-reconstruction loss 520 is calculated. In at least one embodiment, the Z-reconstruction loss refers to the difference or distance between the input appearance parameter set 504 and the predicted appearance parameter set 516. In at least one embodiment, the Z-reconstruction loss is a loss function calculated based on the accuracy of the discriminator 510 in predicting the appearance parameters or appearance characteristics of the image.

[0110] In at least one embodiment, determination 514 is used to determine a classification loss, such as a true / false classification loss 522. In at least one embodiment, the true / false classification loss 522 is calculated based on whether discriminator 514 can correctly predict whether the generated image 508 submitted to discriminator 510 is a real image or a synthetic image. In at least one embodiment, the true / false classification loss 522 is calculated based on whether discriminator 510 can correctly predict whether the input image set is real or fake, wherein discriminator 510 may be provided with real or fake (e.g., synthetic) images and used to predict which images are real or fake.

[0111] Figure 6 Figure 600 illustrates a depiction of generator updates according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update the parameters of the generator used to create the synthetic image. In at least one embodiment, Figure 6 The diagram illustrates an input image 602; a discriminator 604; a predicted viewpoint 606; a prediction determination 608 to determine whether the input image 602 is real or fake; a predicted set of appearance parameters 610; a generator 612; a generated image 614; an image consistency loss 616; and a real / fake classification loss 618. In at least one embodiment, Figure 6 The illustration depicts a generator update using real images (e.g., images not created by a generative adversarial network). In at least one embodiment, it combines... Figure 6 The described technology and combination Figure 4 , Figure 5 and Figure 7 The techniques described for training generators and / or discriminators have the same scalability.

[0112] In at least one embodiment, the input image 602 is input to the discriminator 604. In at least one embodiment, the input image 602 is a real image selected from a training dataset. In at least one embodiment, the input image 602 is selected from a set of object images, such as a combination of... Figure 2The descriptions are as follows. In at least one embodiment, discriminator 604 processes input image 602. In at least one embodiment, discriminator 604 receives input image 602 and generates a predicted viewpoint 606, a predicted determination that input image 602 is a real image or a fake image, and a predicted set of appearance parameters 608. In at least one embodiment, viewpoint 606 is a predicted viewpoint of input image 602 determined by one or more processes of discriminator 604. In at least one embodiment, determination 608 is a prediction that input image 602 is a real image or a fake image (e.g., a synthetic image created by a generative adversarial network). In at least one embodiment, determination 608 is a confidence value in a binary value or a non-binary range (e.g., 0-100). In at least one embodiment, appearance parameter set 610 is a predicted set of appearance parameters of input image 602 generated by one or more processes of discriminator 604. In at least one embodiment, appearance parameter set 610 is a set of predicted parameters that define the appearance of objects depicted in input image 602.

[0113] In at least one embodiment, viewpoint 606 and appearance parameter set 610 are provided as input to generator 612 to produce generated image 614. In at least one embodiment, generated image 614 is a synthetic (e.g., fake) image created by generator 612, having an orientation and appearance based on viewpoint 606 and appearance parameter set 610, such as... Figure 6 As shown, viewpoint 606 and appearance parameter set 610 are also the predicted viewpoint and appearance of input image 602.

[0114] In at least one embodiment, an image consistency loss 616 is calculated at least in part based on an input image 602, which is provided to a discriminator 604 for decomposing at least two attributes from the input image 602: a predicted viewpoint 606 and a predicted set of appearance parameters 610. In at least one embodiment, the predicted viewpoint 606 and the predicted set of appearance parameters 610 are provided to a generator 612 to create a generated image 614. In at least one embodiment, a generative adversarial network (GAN) receives the viewpoint and the set of appearance parameters and generates a synthetic (e.g., fake) image based on any of the provided viewpoint and the set of appearance parameters. In at least one embodiment, the input image 602 and the generated image 614 are compared to determine the image consistency loss 616. In at least one embodiment, the cosine distance between the input image and the synthetic image is compared to determine feature similarity, where closer similarity corresponds to a lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss 616 between the two images.

[0115] In at least one embodiment, determination 608 is used to determine a classification loss, such as a true / false classification loss 618. In at least one embodiment, the true / false classification loss 618 is calculated based on whether the discriminator 604 can correctly predict whether the input image 602 submitted to the discriminator 604 is a real image or a synthetic image. In at least one embodiment, the true / false classification loss 618 is calculated based on whether the discriminator 604 can correctly predict whether the input image set is real or fake, wherein the discriminator 604 may be provided with real or fake (e.g., synthetic) images and used to predict whether these images are real or fake. In at least one embodiment, one or more loss functions are computed. In at least one embodiment, the parameters of the generator 612 are updated at least by computed gradients (e.g., performing stochastic gradient descent) while the discriminator parameters are fixed.

[0116] Figure 7 Figure 700 illustrates a depiction generator update according to at least one embodiment. In at least one embodiment, a loss function is calculated and used to update the parameters of the generator used to create the synthetic image. In at least one embodiment, Figure 7 The diagram illustrates a first viewpoint 702; an appearance parameter set 704; a generator 706; a first generated image 708; a discriminator 710; a predicted first viewpoint 712; a prediction determination 714 of whether the generated image 708 is real or fake; a predicted first appearance parameter set 716; viewpoint consistency loss 718; Z-reconstruction loss 720; real / fake classification loss 722; a second viewpoint 724; a second generated image 726; a predicted second viewpoint 728; a predicted second appearance parameter set 730; viewpoint consistency loss 732; Z-reconstruction loss 734; and symmetry loss 736. In at least one embodiment, Figure 6 A generator update using fake images (e.g., synthetic or generated images created by a generative adversarial network) is illustrated. In at least one embodiment, combined with Figure 7 The described technology and combination Figure 4-6 The techniques described for training generators and / or discriminators have the same scalability.

[0117] In at least one embodiment, the first viewpoint 702 and the appearance parameter set 704 are independently selectable decoupled parameters. In at least one embodiment, the generator 706 accepts the first viewpoint 702 and the appearance parameter set 704 as input and creates a first generated image 708. In at least one embodiment, the first generated image 708 is a synthetic image whose appearance is generated based on the appearance parameter set 704 and oriented according to the first viewpoint 702. The generator 706 may be configured according to those described elsewhere in this disclosure, for example, in combination with... Figure 2 Those that were discussed.

[0118] In at least one embodiment, an image set (e.g., a first generated image 708) is provided as input to a discriminator 710, and the discriminator predicts various attributes of the image set. In at least one embodiment, the discriminator 710 receives the first generated image 708 and generates a predicted first viewpoint 712; a prediction determination 714 of whether the generated image 708 is real or fake; and a predicted first set of appearance parameters 716. In at least one embodiment, the discriminator 710 has no access to the first viewpoint 702 and the set of appearance parameters 704 used to create the first generated image 708 (e.g., such information is concealed from the discriminator 710 during prediction). In at least one embodiment, the output of the discriminator 710 is used to calculate a loss. In at least one embodiment, the loss function is used to calculate a gradient (e.g., using gradient descent) and update the parameters of the discriminator 710 while fixing the parameters of the generator 706 to a constant.

[0119] In at least one embodiment, a viewpoint consistency loss 718 is calculated. In at least one embodiment, the viewpoint consistency loss is a loss function calculated based on the accuracy of the discriminator 710 in predicting viewpoints. In at least one embodiment, the viewpoint consistency loss 718 is calculated as the difference or distance between the first viewpoint 702 and the predicted first viewpoint 712. In at least one embodiment, the viewpoint consistency loss is a component of the generative consistency loss. In at least one embodiment, the distance between the first viewpoint 702 and the predicted first viewpoint 712 (e.g., L1 distance, L2 distance, cosine distance, and / or variations thereof) is used to calculate the viewpoint consistency loss 718, wherein closer viewpoints (e.g., viewpoints with shorter distances to each other) correspond to a lower loss.

[0120] In at least one embodiment, a Z-reconstruction loss 720 is calculated. In at least one embodiment, the Z-reconstruction loss refers to the difference or distance between the appearance parameter set 704 and the predicted first appearance parameter set 716. In at least one embodiment, the Z-reconstruction loss is a loss function calculated based on the accuracy of the discriminator 710 in predicting the appearance parameters or appearance attributes of the image.

[0121] In at least one embodiment, determination 714 is used to determine a classification loss, such as a true / false classification loss 722. In at least one embodiment, the true / false classification loss 722 is calculated based on whether discriminator 714 can correctly predict whether the first generated image 708 submitted to discriminator 710 is a real image or a synthetic image. In at least one embodiment, the true / false classification loss 722 is calculated based on whether discriminator 710 can correctly predict whether the input image set is real or fake, wherein discriminator 710 may be provided with real or fake (e.g., synthetic) images and used to predict which images are real or fake.

[0122] In at least one embodiment, the second viewpoint 724 is a transformation of the first viewpoint 702. In at least one embodiment, the first viewpoint 702 is horizontally flipped to produce the second viewpoint 724. In at least one embodiment, any suitable transformation of the azimuth, tilt, and elevation parameters on the first viewpoint 702 produces the second viewpoint 724. In at least one embodiment, the second viewpoint 724 is determined according to techniques described elsewhere in this disclosure, such as in combination with... Figure 8 Those discussed. In at least one embodiment, if the azimuth, elevation, and tilt parameters of the first viewpoint 702 are Φ, θ, and ψ, respectively, then the azimuth, elevation, and tilt parameters of the second viewpoint 724 are -Φ, θ, and -ψ.

[0123] In at least one embodiment, an appearance parameter set 704 is used to generate a second image. In at least one embodiment, a second viewpoint 724 and the appearance parameter set 704 are provided as input to a generator 706 to create a second generated image 726. In at least one embodiment, if the second generated image 726 is horizontally flipped, it produces a first generated image 708.

[0124] In at least one embodiment, a symmetry loss 736 is calculated based on a first generated image 708 and a second generated image 726. In at least one embodiment, the symmetry loss 736 is calculated by comparing the magnitudes of the azimuth, elevation, and tilt of the predicted first viewpoint 712 from the generated image 708 with the magnitudes of the predicted second viewpoint 728 from the second generated image 726, wherein zero loss occurs when the magnitudes of each viewpoint parameter are equal, and the loss increases as the difference between the magnitudes of each viewpoint parameter increases. In at least one embodiment, the symmetry loss 736 is calculated according to techniques described elsewhere in this disclosure, such as in combination with… Figure 8 Those discussed. In at least one embodiment, in... Figure 7 The decoupling loss is applied in the context of [the specific context]. In at least one embodiment, a viewpoint V1 and an appearance parameter set Z1 are selected and provided to the generator to produce a first synthetic image i1. In at least one embodiment, a second image I2 is generated by keeping the viewpoint constant (e.g., using V1) and perturbing the appearance parameters using a second appearance parameter set Z2 different from the Z1 used to generate the first image i1. In at least one embodiment, a third image I3 is generated by keeping the appearance parameters constant relative to image I1 (e.g., using Z1) and perturbing the viewpoint using a second viewpoint V2 different from the viewpoint V1 to generate the third image I3. In at least one embodiment, the synthetic image I1 generated using a specific viewpoint I1 and appearance parameter set Z1 is compared with the synthetic image I2 generated using viewpoint V1 and the second appearance parameter set Z2 and / or the synthetic image I3 generated using the second viewpoint V2 and appearance parameter set Z1. According to at least one embodiment, [the following is combined with the context of the first image i1]. Figure 10The described techniques can be applied to combining Figure 7 The coupling loss is described.

[0125] Figure 8 Figure 800 illustrates the computation of a symmetric loss according to a description of at least one embodiment. In at least one embodiment, the symmetric loss is used to train one or more neural networks associated with a discriminator, such as discriminator 806. In at least one embodiment, the symmetric loss is used in conjunction with one or more other loss functions to refine the parameters associated with the discriminator.

[0126] In at least one embodiment, FIG800 includes an input image 802. In at least one embodiment, the input image 802 is part of a set of one or more images of a type of object. In at least one embodiment, the input image 802 is an image depicting a car in a specific orientation (including specific appearance features). In at least one embodiment, the input image 802 is selected from a set of one or more images. In at least one embodiment, images from the set are selected for learning in any suitable manner, and may be sampled randomly or pseudo-randomly from a training set.

[0127] In at least one embodiment, a transformation is applied to the input image 802 to generate a transformed image 804. In at least one embodiment, the input image 802 is horizontally flipped to generate the transformed image 804. In at least one embodiment, the transformation is applied by one or more systems associated with the discriminator 806. In at least one embodiment, the discriminator 806 applies one or more image processing techniques to the input image 802 to generate the transformed image 804. In at least one embodiment, when the input image 802 is flipped and / or transformed to generate the transformed image 804, the azimuth and tilt angles (depicted as “az” and “ti” in FIG. 800) of the viewpoints of objects within the input image 802 are reversed, while the elevation angle (depicted as “el” in FIG. 800) remains the same between the input image 802 and the transformed image 804. In at least one embodiment, the transformed image 804 is generated by applying one or more image transformations to the input image 802. In at least one embodiment, the transformed image 804 is generated by at least horizontally flipping the input image 802, horizontally flipping the input image 802, flipping the input image 802 according to a specified axis, rotating the input image 802 by a specified degree, and / or applying various other 2D transformations to the input image 802.

[0128] In at least one embodiment, discriminator 806 processes input image 802. In at least one embodiment, discriminator 806 is associated with one or more neural networks trained to infer viewpoint and other features of the input image. In at least one embodiment, discriminator 806 receives input image 802 and generates a first prediction 808. In at least one embodiment, the first prediction 808 corresponds to a specific direction of a predicted object within the image and includes specific values ​​for a parameter set including azimuth, elevation, and tilt parameters. In at least one embodiment, the first prediction 808 includes a first predicted viewpoint V1 and a first predicted appearance parameter set Z1 for input image 802. In at least one embodiment, the first prediction 808 includes a predicted direction of a car depicted in input image 802. In at least one embodiment, discriminator 806 processes transformed image 804. In at least one embodiment, discriminator 806 receives transformed image 804 and generates a second prediction 810. In at least one embodiment, the second prediction 810 corresponds to a specific direction of the predicted object within the image and includes specific values ​​for a parameter set, which includes azimuth, elevation, and tilt parameters. In at least one embodiment, the second prediction 810 includes a second predicted viewpoint V2 and a second predicted appearance parameter set Z2 of the transformed image 804. In at least one embodiment, the second prediction 810 includes the predicted direction of a car depicted in the transformed image 804.

[0129] In at least one embodiment, a first prediction 808 is predicted for an input image 802 and a second prediction 810 is predicted for a transformed image 804, which is a horizontally flipped version of the input image 802. In at least one embodiment, a loss is calculated based on whether certain properties remain true. In at least one embodiment, a transformation or its inverse transformation is applied to the second prediction 810. In at least one embodiment, if the input image 802 is rotated by an angle (Φ, θ, ψ) to produce the transformed image 804, the second prediction 810 can be rotated in the opposite direction by an angle (-Φ, -θ, -ψ). In at least one embodiment, a symmetric loss is calculated by comparing the magnitudes of the azimuth, elevation, and tilt of the first prediction 808 of the input image 802 with the second prediction 810 of the transformed image 804, wherein zero loss is produced when the magnitudes of each viewpoint parameter are equal and / or the appearance parameter predicted for the input image 802 matches the appearance parameter predicted for the transformed image 804. In at least one embodiment, the loss increases with the difference between the magnitudes of each viewpoint parameter. In at least one embodiment, a symmetry loss is calculated according to techniques described elsewhere in this disclosure, such as those discussed in conjunction with FIG16. In at least one embodiment, the symmetry loss is calculated at least in part based on the degree of matching between the predicted appearance parameters for input image 802 and the predicted appearance parameters for transformed image 804. In at least one embodiment, the weights and parameters of discriminator 806 are trained to predict whether the appearance parameters of input image 802 match the predicted appearance parameters for a transformed version of input image 802 (e.g., transformed image 804).

[0130] In at least one embodiment, the discriminator 806 predicts a first set of appearance parameters for the input image 802 and a second set of appearance parameters for the transformed image 804. In at least one embodiment, a loss function is computed based on the similarity between the first set of appearance parameters predicted for the input image 802 and the second set of appearance parameters predicted for the transformed image 804. In at least one embodiment, the discriminator 806 is trained to predict that the parameters of the input image 802 and the transformed image 804 are equivalent.

[0131] Figure 9Figure 900 illustrates a viewpoint map according to at least one embodiment. In at least one embodiment, the viewpoint map 908 is used to determine nearest neighbor and farthest neighbor losses, which are used in conjunction with one or more other loss functions to refine the parameters associated with the discriminator. In at least one embodiment, the viewpoint map 908 is constructed through one or more processes and systems associated with the discriminator. In at least one embodiment, the viewpoint map 908 is generated based on a set of one or more images of an object of a certain type. In at least one embodiment, the first image 902 and the second image 904 are part of a set of one or more images of an object of a certain type. In at least one embodiment, the first image 902 and the second image 904 are part of a set of one or more images including a car. In at least one embodiment, the first image 902 and the second image 904 are images depicting a car in a specific direction.

[0132] In at least one embodiment, viewpoint map 908 is generated based on viewpoint equivariant distances (e.g., cosine distances) between images in an image set. In at least one embodiment, cosine distance is a mathematical complement to cosine similarity (e.g., cosine distance = 1 - cosine similarity). In at least one embodiment, cosine similarity is a similarity measure between two vectors that can represent images, text, data, and / or variations thereof based on the cosine of the angle between them. In at least one embodiment, the cosine distance calculated for the images is low when the two images include features corresponding to objects with similar viewpoints. In at least one embodiment, the cosine distance calculated for the images is high when the two images include features corresponding to objects with different viewpoints.

[0133] In at least one embodiment, a set of images is used to generate a viewpoint map 908, wherein the nodes of the viewpoint map 908 correspond to images and the edges correspond to their cosine distances. In at least one embodiment, a convolutional neural network is used to compute the cosine distance based on the feature similarity of the image pairs. In at least one embodiment, the edges of the viewpoint map 908 are weighted such that a thicker edge between two images corresponds to a higher similarity between the two images, and a thinner edge between two images corresponds to a lower similarity between the two images. In at least one embodiment, a cosine distance 910 is computed between a first image 902 and a second image 904. In at least one embodiment, the first image 902 and the second image 904 are used as part of the viewpoint map 908 and are connected to edges corresponding to the computed cosine distance 910. In at least one embodiment, the first image 902 and the second image 904 include cars oriented in similar directions and / or viewpoints, and the thicker edge used between the first image 902 and the second image 904 within the viewpoint map 908 reflects this similarity.

[0134] Figure 9 Figure 900 illustrates the computation of nearest neighbor and farthest neighbor losses according to at least one embodiment. In at least one embodiment, the nearest neighbor and farthest neighbor losses are used to train one or more neural networks associated with a discriminator, such as discriminator 912. In at least one embodiment, the nearest neighbor and farthest neighbor losses are used in conjunction with one or more other loss functions to refine the parameters associated with the discriminator. In at least one embodiment, the nearest and farthest neighbor losses can be viewed as a type of viewpoint equivariant loss. In at least one embodiment, the techniques described herein applied to the nearest and farthest neighbor losses (e.g., combined with...) Figure 9 (As described elsewhere) It also applies to viewpoint-variable losses.

[0135] In at least one embodiment, viewpoint diagram 908 is constructed through one or more processes and systems associated with discriminator 912. In at least one embodiment, viewpoint diagram 908 is generated based on a set of one or more images of an object of a certain type. In at least one embodiment, images 902-906 are part of a set of one or more images of an object of a certain type. In at least one embodiment, images 902-906 are part of a set of one or more images including a car. In at least one embodiment, images 902-906 are images depicting a car in a particular orientation.

[0136] In at least one embodiment, discriminator 912 processes images 902-906. In at least one embodiment, discriminator 912 is associated with one or more neural networks trained to infer viewpoints and other features of the input images. In at least one embodiment, discriminator 912 receives image 902 and generates viewpoint 914. In at least one embodiment, discriminator 912 receives image 904 and generates viewpoint 916. In at least one embodiment, discriminator 912 receives image 906 and generates viewpoint 918. In at least one embodiment, viewpoints 914-918 correspond to a predicted specific direction of an object depicted in images 902-906 and include specific values ​​for a set of parameters including azimuth, elevation, and tilt parameters. In at least one embodiment, viewpoints 914-918 respectively include the predicted direction of a car depicted in images 902-906.

[0137] In at least one embodiment, the nearest neighbor and farthest neighbor losses are calculated by comparing the selected image with its nearest and farthest neighbors, at least in part, based on the viewpoint map 908 of the image set. In at least one embodiment, the nearest neighbor and farthest neighbor losses include a nearest neighbor loss and a farthest neighbor loss. In at least one embodiment, an anchor image is selected from the set of training images. In at least one embodiment, image 902 is selected as the anchor image. In at least one embodiment, image 902 is located from the viewpoint map 908, and the nearest and farthest neighbors are selected based on edge weights. In at least one embodiment, the nearest neighbor (such as image 904) has the shortest edge connected to image 902. In at least one embodiment, the farthest neighbor (such as image 906) has the farthest edge connected to image 902.

[0138] In at least one embodiment, image 904 is determined to be the nearest neighbor of image 902. In at least one embodiment, a nearest neighbor loss is calculated between viewpoint 914 and viewpoint 916. In at least one embodiment, the nearest neighbor loss is calculated such that a higher similarity between viewpoint 914 and viewpoint 916 corresponds to a smaller loss. In at least one embodiment, the nearest neighbor loss is calculated based on the following symbolic mathematical equation:

[0139]

[0140] Where L nn It was the nearest neighbor's loss. It is the viewpoint of the anchor image. It is the viewpoint of the nearest neighbor image, and It is a function that determines the distance (e.g., L1 distance, L2 distance, cosine distance and / or its variants) or difference between viewpoints.

[0141] In at least one embodiment, image 906 is determined to be the farthest neighbor of image 902. In at least one embodiment, a farthest neighbor loss is calculated between viewpoints 914 and 916. In at least one embodiment, a farthest neighbor loss is calculated such that a lower similarity between viewpoints 914 and 918 corresponds to a smaller loss. In at least one embodiment, the farthest neighbor loss is calculated based on the following symbolic mathematical equation:

[0142]

[0143] Where L fn It is the loss of the furthest neighbor. It is the viewpoint of the anchor image. The viewpoint is the image of the farthest neighbor. `min()` is a function that determines the minimum value between two values. It is a function that determines the distance (e.g., L1 distance, L2 distance, cosine distance, and / or variations thereof) or difference between viewpoints, and t is a minimum threshold determined by one or more processes. In at least one embodiment, t is determined by a discriminator, one or more systems associated with the discriminator, and / or variations thereof. In at least one embodiment, t is a parameter set by one or more processes before training begins, and is the threshold at which its farthest neighbor loss is considered to have no loss.

[0144] In at least one embodiment, the nearest and / or farthest neighbor is selected nondeterministically. In at least one embodiment, a probability is assigned to each edge to be selected as the nearest and / or farthest neighbor of the anchor image. In at least one embodiment, the probability of the nearest neighbor is inversely proportional to the edge weight (e.g., the node connected to the anchor image with the lowest edge weight has the highest probability of being selected). In at least one embodiment, the probability of the farthest neighbor is directly proportional to the edge weight (e.g., the node connected to the anchor image with the highest edge weight has the highest probability of being selected).

[0145] Figure 10 Figure 1000 illustrates a decoupling loss according to at least one embodiment. In at least one embodiment, the decoupling loss is used to train one or more neural networks associated with a generator, such as generator 1002. In at least one embodiment, the decoupling loss is used in conjunction with one or more other loss functions to refine the parameters associated with generator 1002. In at least one embodiment, the decoupling loss operates when updating generator parameters. In at least one embodiment, the decoupling loss operates during the generator update phase.

[0146] In at least one embodiment, generator 1002 is associated with a generative model, such as a variational autoencoder (VAE), a differentiable renderer, a generative adversarial network, or a renderer. In at least one embodiment, generator 1002 is part of one or more neural networks trained to generate images based on an input viewpoint and an input set of appearance parameters, and the one or more neural networks include various parameters associated with one or more processes of the one or more neural networks. In at least one embodiment, discriminator 1004 is part of one or more neural networks trained to infer viewpoint and appearance attribute sets from an input image, and the one or more neural networks include various parameters associated with one or more processes of the one or more neural networks.

[0147] In at least one embodiment, a first viewpoint 1006A and a first appearance attribute set 1008A are obtained. In at least one embodiment, the first viewpoint 1006A and the first appearance attribute set 1008A are randomly generated. In at least one embodiment, the first viewpoint 1006A and the first appearance attribute set 1008A are obtained from one or more processes associated with generator 1002 and discriminator 1004. In at least one embodiment, generator 1002 generates image 1010A based on the first viewpoint 1006A and the first appearance attribute set 1008A. In at least one embodiment, image 1010A is input to discriminator 1004. In at least one embodiment, discriminator 1004 performs one or more processes and determines a first predicted viewpoint 1012A and a first predicted appearance attribute set 1012B based on the input image 1010A.

[0148] In at least one embodiment, a second appearance attribute set 1008B is obtained. In at least one embodiment, the second appearance attribute set 1008B is randomly generated. In at least one embodiment, the second appearance attribute set 1008B is obtained from one or more processes associated with generator 1002 and discriminator 1004. In at least one embodiment, generator 1002 generates image 1010B based on first viewpoint 1006A and second appearance attribute set 1008B. In at least one embodiment, image 1010B is input to discriminator 1004. In at least one embodiment, discriminator 1004 performs one or more processes and determines a second predicted viewpoint 1014A and a second predicted appearance attribute set 1014B based on input image 1010B.

[0149] In at least one embodiment, a second viewpoint 1006B is obtained. In at least one embodiment, the second viewpoint 1006B is randomly generated. In at least one embodiment, the second viewpoint 1006B is obtained from one or more processes associated with generator 1002 and discriminator 1004. In at least one embodiment, generator 1002 generates image 1010C based on the second viewpoint 1006B and a first appearance attribute set 1008A. In at least one embodiment, image 1010C is input to discriminator 1004. In at least one embodiment, discriminator 1004 performs one or more processes and determines a third predicted viewpoint 1016A and a third predicted appearance attribute set 1016B based on the input image 1010C.

[0150] In at least one embodiment, the decoupling loss includes z-reconstruction loss and viewpoint reconstruction loss.

[0151] In at least one embodiment, the z-reconstruction loss is calculated by comparing a prediction of the appearance parameter set of an image generated by a discriminator such as discriminator 1004 with an input appearance parameter set generated by a generator (e.g., generator 1002) for generating the image, wherein a lower loss is generated when the prediction of the appearance parameter set is more similar to the input appearance parameter set, and a higher loss is generated when the prediction of the appearance parameter set is less similar to the input appearance parameter set. In at least one embodiment, the viewpoint reconstruction loss is calculated by comparing a prediction of the viewpoint of an image (generated by a discriminator such as discriminator 1004) with an input viewpoint generated by a generator such as generator 1002 for generating the image, wherein a lower loss is generated when the prediction of the viewpoint is more similar to the input viewpoint, and a higher loss is generated when the prediction of the viewpoint is less similar to the input viewpoint.

[0152] In at least one embodiment, a z-reconstruction loss and a viewpoint reconstruction loss are calculated for each predicted viewpoint set and appearance attribute (e.g., a first predicted viewpoint 1012A and a first predicted appearance attribute set 1012B based on input image 1010A, a second predicted viewpoint 1014A and a second predicted appearance attribute set 1014B based on input image 1010B, and a third predicted viewpoint 1016A and a third predicted appearance attribute set 1016B based on input image 1010C). In at least one embodiment, the z-reconstruction loss and the viewpoint reconstruction loss are used to determine the decoupling loss. In at least one embodiment, an additional loss function is also calculated as part of the decoupling loss. In at least one embodiment, the decoupling loss is calculated based on the following symbolic mathematical equation:

[0153] Disentanglement Loss

[0154] =∑vlewpoint reconstruction loss+z reconstruction loss

[0155] +other loss functions

[0156] In at least one embodiment, the decoupling loss is used to refine one or more parameters associated with generator 1002. In at least one embodiment, the parameters of generator 1002 are updated such that the loss from at least the decoupling loss is minimized.

[0157] Figure 11Figure 1100 illustrates a depiction of a calibration neural network according to at least one embodiment. In at least one embodiment, a discriminator is trained on a set of images and calibrated on portions of the image set including truth annotations. In at least one embodiment, discriminator 1104 is part of one or more neural networks trained to identify the orientation of objects within an image based at least in part on one or more features of the object rather than the object's orientation. In at least one embodiment, discriminator 1106 is trained on a set of images to infer the orientation of other objects of the same category captured in other images.

[0158] In at least one embodiment, discriminator 1104 is trained to self-supervisedly identify the viewpoint of an object within an image set, such as those described elsewhere in this disclosure. In at least one embodiment, discriminator 1104 is trained to self-supervisedly identify the orientation of an object within an image by computing one or more loss functions as part of training to evaluate one or more features of images in the training set. In at least one embodiment, discriminator 1106 is trained at least in part based on computing losses such as one or more of the following: generative consistency loss, symmetry loss, nearest neighbor and farthest neighbor loss, and decoupling loss, which may be combined with... Figure 4-10 The descriptions are consistent. In at least one embodiment, the discriminator 1104 is trained on a set of images that lack truth annotations or where truth annotations are otherwise unavailable.

[0159] In at least one embodiment, an object image set 1102 is obtained to calibrate the discriminator 1104. In at least one embodiment, the object image set 1102 includes images with truth annotations. In at least one embodiment, the object image set 1102 includes images depicting objects, and the discriminator 1104 is trained to analyze and determine the viewpoint of those images. In at least one embodiment, the object image set 1102 is part of an image set used to train the discriminator 1104. In at least one embodiment, the discriminator 1104 is trained on an image set different from the object image set 1102. In at least one embodiment, the discriminator 1104 obtains images from the object image set 1102 and determines the viewpoint 1106 of those images. In at least one embodiment, although the discriminator 1104 is trained, the viewpoint 1106 does not match the truth viewpoint of the images in the object image set 1102. In at least one embodiment, the viewpoint 1106 is the correct viewpoint of the images in the object image set 1102 translated in one or more dimensions.

[0160] In at least one embodiment, a linear model 1108 is determined that transforms the viewpoints of the images in the object image set 1102, as determined by the discriminator 1104, to their respective correct positions indicated by truth annotations. In at least one embodiment, the linear model 1108 is a linear function determined by comparing the viewpoints generated by the discriminator 1104 for the images in the object image set 1102 with truth annotations for the viewpoints of the images in the object image set 1102. In at least one embodiment, the linear model 1108 is determined by one or more processes, such as various regression algorithms, mathematical processes, and / or variations thereof. In at least one embodiment, the linear model 1108 is determined such that the viewpoints of the images determined by the discriminator 1104 can be transformed and / or corrected to match the truth viewpoints of the images. In at least one embodiment, the linear model 1108 calibrates the viewpoints of the images determined by the discriminator 1104 to the coordinate system of the truth viewpoints of the images. In at least one embodiment, the linear model 1108 converts the zero position of the viewpoint of the image determined by the discriminator 1104 into the zero position of the true viewpoint of the image.

[0161] In at least one embodiment, linear model 1108 is used to calibrate or otherwise correct viewpoint 1106 to generate corrected viewpoint 1110. In at least one embodiment, corrected viewpoint 1110 is viewpoint 1106 that has been transformed to match the truth viewpoint. In at least one embodiment, linear model 1108 is used to calibrate or otherwise correct viewpoints determined / generated by discriminator 1104 for images in object image set 1102 to match the truth annotations of viewpoints in said images in object image set 1102.

[0162] Figure 12Figure 1200 illustrates a descriptive inference according to at least one embodiment. In at least one embodiment, as part of one or more systems associated with a safety system for a motor vehicle, a discriminator is trained to infer viewpoints from images of other motor vehicles captured from a camera associated with the motor vehicle, such that the inferred viewpoints are used to perform one or more operations associated with the motor vehicle, such as braking the motor vehicle to avoid another motor vehicle, or steering the motor vehicle to avoid another motor vehicle. In at least one embodiment, discriminator 1204 is trained in a self-supervised manner on a set of images to identify viewpoints of objects within the images, such as those described elsewhere in this disclosure. In at least one embodiment, discriminator 1204 is trained in a self-supervised manner to identify the orientation of objects within the images by at least computing one or more loss functions as part of training to evaluate one or more features of the images in the training set. In at least one embodiment, discriminator 1204 is trained at least in part based on computing generative consistency loss, symmetry loss, nearest neighbor and farthest neighbor loss, and decoupling loss, which may be based on a combination of Figure 4-10 Those described. In at least one embodiment, the discriminator 1204 is trained on a set of vehicle images. In at least one embodiment, the discriminator 1204 is trained to identify the viewpoint of a vehicle within an image.

[0163] In at least one embodiment, the authenticator 1204 is part of one or more systems of the motor vehicle 1210. In at least one embodiment, the motor vehicle 1210 is an autonomous vehicle. In at least one embodiment, the motor vehicle 1210 is an operator-operated vehicle. In at least one embodiment, the motor vehicle 1210 includes one or more systems implementing the authenticator 1204. In at least one embodiment, the motor vehicle 1210 includes one or more systems associated with the authenticator 1204. In at least one embodiment, the motor vehicle 1210 includes one or more systems enabling the motor vehicle 1210 to remotely access and utilize the authenticator 1204, which may be implemented on various remote and / or local systems.

[0164] In at least one embodiment, the motor vehicle 1210 includes a plurality of cameras. In at least one embodiment, camera 1202 is located on the motor vehicle 1210. In at least one embodiment, camera 1202 is a remotely accessible camera of the motor vehicle 1210. In at least one embodiment, camera 1202 captures or otherwise obtains image 1202A. In at least one embodiment, the motor vehicle 1210 is a vehicle operating in an environment including other motor vehicles, and the motor vehicle 1210 must perform one or more actions related to other motor vehicles (e.g., allowing a motor vehicle to pass, overtaking, braking for an oncoming motor vehicle, steering to avoid an oncoming vehicle, and / or variations thereof). In at least one embodiment, image 1202A is an image of another motor vehicle with which the motor vehicle 1210 must interact. In at least one embodiment, image 1202A is an image captured from an onboard camera of the motor vehicle 1210 while the motor vehicle 1210 is operating in the environment.

[0165] In at least one embodiment, image 1202A is input to discriminator 1204, which determines viewpoint 1206 of image 1202A. In at least one embodiment, viewpoint 1206 is the viewpoint of the vehicle depicted in image 1202A. In at least one embodiment, safety system 1208 is part of motor vehicle 1210. In at least one embodiment, safety system 1208 includes one or more systems configured to provide assistance to motor vehicle 1210. In at least one embodiment, safety system 1208 includes one or more systems configured to operate one or more systems of motor vehicle 1210, such as brake actuator 1208A and steering actuator 1208B, which are components configured to operate the brakes and steering of motor vehicle 1210, respectively. In at least one embodiment, brake actuator 1208A and steering actuator 1208B are identical to brake actuator 2148 and steering actuator 2156, respectively. In at least one embodiment, the motor vehicle 1210 may be implemented according to techniques described elsewhere herein, such as FIG21.

[0166] In at least one embodiment, the safety system 1208 obtains a viewpoint 1206. In at least one embodiment, the safety system 1208 determines the direction of travel of the vehicle depicted in image 1202A based on viewpoint 1206. In at least one embodiment, the safety system 1208 determines whether to use brake actuator 1208A or steering actuator 1208B based on viewpoint 1206. In at least one embodiment, if the safety system 1208 determines that the vehicle depicted in image 1202A is traveling in a direction relative to motor vehicle 1210 in which braking of motor vehicle 1210 is required, the safety system 1208 activates brake actuator 1208A to brake motor vehicle 1210, such that motor vehicle 1210 avoids any potential safety problems caused by the travel of the vehicle depicted in image 1202A. In at least one embodiment, if the safety system 1208 determines that the car depicted in image 1202A is traveling in a direction relative to the motor vehicle 1210 in which the motor vehicle 1210 needs to be steered, the safety system 1208 activates the steering actuator 1208B to steer the motor vehicle 1210 so that the motor vehicle 1210 avoids any potential safety problems caused by the travel of the car depicted in image 1202A.

[0167] Figure 13An illustrative example of a process 1300 for training a neural network to predict the viewpoint of an object within an image, according to at least one embodiment, is shown. In at least one embodiment, some or all of process 1300 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and can be implemented by hardware, software, or a combination thereof as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1300 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). Non-transitory computer-readable media do not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 1300 is performed at least partially on a computer system such as those described elsewhere in this disclosure. In at least one embodiment, a first computer system trains one or more neural networks, and a second computer system uses said one or more neural networks to perform inference (e.g., predicting the viewpoint of objects within an image). In at least one embodiment, combined with Figure 1 , Figure 12 and Figure 14 The described technique is applicable to process 1300.

[0168] In at least one embodiment, the system performing at least a portion of process 1300 includes executable code for obtaining a set of one or more images of an object of a certain type 1302.

[0169] In at least one embodiment, a set of one or more images is used to train one or more neural networks to identify the orientation of objects within the images. In at least one embodiment, the image set is classified or labeled for each object displaying the same type or category. In at least one embodiment, the image set is a set of car images, which may include different types of cars in different directions, in different weather conditions, and under different lighting conditions. In at least one embodiment, the car set includes images of the same car or the same type of car in different directions. In at least one embodiment, at least a portion of the training image set lacks ground truth annotations specifying the orientation of objects within such training images. In at least one embodiment, all images in the image set lack ground truth annotations specifying the azimuth, elevation, and tilt angles of objects within the images in the set. In at least one embodiment, the image set includes one or more synthetic images, such as images created from a generative adversarial network (GAN). In at least one embodiment, all images in the image set are real images, not those synthesized or created from generative models such as variational autoencoders (VAEs), differentiable renderers, generative adversarial networks (GANs), or renderers. In at least one embodiment, the image set is collected and aggregated from a website that classifies images by category.

[0170] In at least one embodiment, the orientation of an object within an image refers to the three-dimensional orientation of the object captured within a two-dimensional image. In at least one embodiment, a camera is used to capture a two-dimensional image of a real-world car, which is oriented relative to the camera in a specific direction. In at least one embodiment, the orientation of the object (e.g., viewpoint) is encoded on a set of parameters including azimuth, elevation, and tilt parameters. In at least one embodiment, the orientation of the object is encoded as a set of three vectors that define the orientation of the object relative to the x, y, and z axes.

[0171] In at least one embodiment, a system performing at least a portion of process 1300 includes executable code for training 1304 one or more neural networks to identify the orientation of objects within an image based at least in part on one or more features of the object rather than the object's orientation. In at least one embodiment, process 1300 is implemented on a processor including one or more circuits for assisting in training one or more neural networks to identify the orientation of objects within an image based at least in part on one or more features of the object rather than the object's orientation. In at least one embodiment, one or more neural networks are trained on a set of images to infer the orientation of other objects of the same category captured in other images. In at least one embodiment, the neural network is trained on a set of aircraft images and, once trained, is used to infer the orientation of other aircraft in other images.

[0172] In at least one embodiment, the neural network is trained to self-supervisedly identify the orientation of objects within an image set, such as those described elsewhere in this disclosure. In at least one embodiment, the neural network is self-supervised to identify the orientation of objects within an image by computing one or more loss functions as part of training to evaluate one or more features of images in a training set (e.g., a set of images). In at least one embodiment, the neural network is trained at least in part based on computing generative consistency loss, symmetry loss, nearest neighbor and farthest neighbor losses, and decoupling loss, which may be based on a combination of... Figure 4-10 Those described.

[0173] In at least one embodiment, the neural network is trained on a set of images that lack truth annotations or where truth annotations are otherwise unavailable (e.g., such data is concealed from the neural network during training).

[0174] In at least one embodiment, the neural network is trained to generate a second image with the same orientation from objects within a first image having a predicted orientation.

[0175] In at least one embodiment, the system implementing process 1300 includes one or more processors for computing parameters to help train one or more neural networks to identify the orientation of objects within an image based at least in part on one or more features of the object rather than the orientation of the object; and one or more memories for storing the parameters. In at least one embodiment, one or more neural networks are trained to identify the orientation of objects within an image using a set of images of different objects of the same category as the object (e.g., a neural network for inferring vehicle viewpoints is trained on a set of images labeled as vehicles).

[0176] In at least one embodiment, one or more neural networks are trained in a self-supervised manner on a set of images of different objects of the same category as the object in the image to be inferred. In at least one embodiment, different objects of the same category can refer to different images, such as one or more images of a first car in one or more directions, one or more images of different second cars in one or more directions, and so on. In at least one embodiment, the image of the object to be inferred is included in the set of images used to train one or more neural networks to infer direction. In at least one embodiment, one or more neural networks are trained in a self-supervised manner by evaluating at least one or more features of the object within the image using a set of loss functions. In at least one embodiment, the one or more features of the object refer to attributes of the object that can be used to infer direction. In at least one embodiment, the neural network is trained by calculating one or more of the following: generative consistency loss; symmetry loss; nearest neighbor and farthest neighbor loss; and decoupling loss. In at least one embodiment, the self-supervised neural network is trained to generate synthetic images of objects with a specific orientation, which may be the same orientation as the predicted orientation of the input image. In at least one embodiment, a deep generative model, such as a variational autoencoder (VAE), a differentiable renderer, a generative adversarial network (GAN), or a renderer, is used to create synthetic images. In at least one embodiment, the object whose orientation is to be inferred can be a vehicle, an aircraft, a drone, a human, (e.g., a human or animal) face, etc.

[0177] In at least one embodiment, the system performing at least a portion of process 1300 includes executable code for obtaining a second image 1306. In at least one embodiment, a second object within the second image is of the same type as a set of images used to train one or more neural networks. In at least one embodiment, the second image is provided to the neural network for inference to predict a second orientation. In at least one embodiment, images are obtained from a camera that captures still images and / or comprises a video consisting of multiple frames captured at a variable or fixed rate. In at least one embodiment, the video comprises multiple frames (e.g., images). In at least one embodiment, a first system trains one or more neural networks and a second, different system uses those one or more neural networks to perform inference to identify the orientation of an object within the image.

[0178] In at least one embodiment, the system performing at least a portion of process 1300 includes executable code for use with 1308 for use on one or more neural networks (e.g., trained as described in numeral 1304) to identify a second orientation of the second object within the second image. In at least one embodiment, the system uses a discriminator trained in a self-supervised manner on a set of images of objects of a particular category to infer the orientation of other objects of said category. In at least one embodiment, one or more neural networks are trained on a set of images of vehicles and used to infer the orientation of vehicles captured in real time by a camera or other suitable video / image capture device attached to the vehicle.

[0179] In at least one embodiment, a first neural network is trained using self-supervised learning on a first set of images of a first category to infer the viewpoint of objects of that first category, while a second neural network is trained using a similar / the same self-supervised learning technique on a second set of images of a second category. In at least one embodiment, an image is provided as input to the first neural network to detect a first orientation of a first object of a first category, and also as input to the second neural network to detect a second orientation of a second object of a second category. In at least one embodiment, an input image is provided to multiple neural networks trained using the self-supervised learning techniques described herein to identify the orientation of different objects in the input image.

[0180] Figure 14An illustrative example of a process 1400 for training a neural network to predict the viewpoint of an object within an image, according to at least one embodiment, is shown. In at least one embodiment, some or all of process 1400 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and can be implemented by hardware, software, or a combination thereof as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1300 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). A non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 1400 is performed at least partially on a computer system such as those described elsewhere in this disclosure. In at least one embodiment, a first computer system trains one or more neural networks, and a second computer system uses said one or more neural networks to perform inference (e.g., predicting the viewpoint of objects within an image). In at least one embodiment, combined with Figure 1 , Figure 12 and Figure 13 The described technique is applicable to process 1400.

[0181] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for obtaining a collection of one or more images of an object of a certain type, 1402. In at least one embodiment, the image type may mean that all images in the image collection include images of cars. In at least one embodiment, the system operates according to techniques described elsewhere in this disclosure (e.g., Figure 13 (To obtain a set of one or more images)

[0182] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for selecting a first image from a set of images 1404. In at least one embodiment, images in the set are selected for learning in any suitable manner, and may be sampled randomly or pseudo-randomly from a training set.

[0183] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for calculating a generative consistency loss 1406 based at least in part on comparing a selected image with an image generated by a deep generative model. In at least one embodiment, techniques described elsewhere in this disclosure are used to calculate the generative consistency loss, such as combining... Figure 4-7 The aforementioned aspects. In at least one embodiment, the generative consistency loss comprises at least two components: viewpoint consistency loss and image consistency loss. In at least one embodiment, the generative consistency loss is calculated according to the technique described in conjunction with Figure 15.

[0184] In at least one embodiment, an image consistency loss is calculated at least in part based on a selected image (e.g., an input image), which is provided to a discriminator for decomposing at least two attributes from the image: a predicted set of viewpoint and appearance parameters. In at least one embodiment, the predicted set of viewpoint and appearance parameters is provided to a generator to create a synthetic image. In at least one embodiment, a generative adversarial network (GAN) receives the set of viewpoint and appearance parameters and generates a synthetic (e.g., fake) image based on any one of the provided set of viewpoint and appearance parameters. In at least one embodiment, the synthetic image and the input image are compared to determine the image consistency loss. In at least one embodiment, the cosine distance between the input image and the synthetic image is compared to determine feature similarity, where closer similarity corresponds to a lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss between the two images.

[0185] In at least one embodiment, a viewpoint (e.g., orientation) consistency loss is calculated at least in part based on the viewpoint of the input image. In at least one embodiment, the viewpoint of the input image is inferred by a discriminator. In at least one embodiment, the viewpoint of the input image is determined based on truth annotations provided as part of training for at least a portion of the training image set. In at least one embodiment, a generator is used to create a synthetic image having the same viewpoint as the input image. In at least one embodiment, the synthetic image generated from the viewpoint of the input image is provided to a discriminator that determines a second viewpoint of the synthetic image. In at least one embodiment, a first viewpoint of the input image is compared with a second viewpoint of the synthetic image generated at least in part based on the input image. In at least one embodiment, the distance between the first viewpoint of the input image and the second viewpoint of the synthetic image is used to calculate the viewpoint consistency loss, wherein a closer viewpoint corresponds to a lower loss.

[0186] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for calculating a symmetry loss 1408 by comparing a selected image with a transformed version of the selected image. In at least one embodiment, the symmetry loss is calculated according to the technique described in conjunction with FIG. 15. In at least one embodiment, the input image is selected from a training image set. In at least one embodiment, a transformation is applied to the input image to generate a transformed image. In at least one embodiment, the input image is horizontally flipped to generate a flipped image. In at least one embodiment, one or more neural networks are used to predict a first orientation of the input image and a second orientation of the transformed image. In at least one embodiment, a first orientation is predicted for the input image and a second orientation is predicted for a horizontally flipped version of the input image. In at least one embodiment, the loss is calculated based on whether certain properties hold true. In at least one embodiment, a transformation or its inverse transformation is applied to the predicted orientation of a transformed version of the input image. In at least one embodiment, if the input image is rotated by an angle (Φ, θ, ψ) to produce a transformed image, the inferred orientation of the transformed image can be rotated in the opposite direction by an angle (-Φ, -θ, -ψ). In at least one embodiment, the loss is calculated by comparing the magnitude of the azimuth, elevation, and tilt of a first direction of the input image with a second direction of the transformed image, wherein zero loss is produced when the magnitudes of each direction parameter are equal. In at least one embodiment, zero loss is produced when the appearance parameters predicted for the input image match the transformed version of the input image. In at least one embodiment, a symmetric loss is calculated according to techniques described elsewhere in this disclosure, such as those discussed in conjunction with Figure 16.

[0187] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for calculating 1410 nearest and farthest neighbor losses by comparing a selected image with its nearest and farthest neighbors, at least partially based on a viewpoint map of the image set. In at least one embodiment, the nearest and farthest neighbor losses are based on a combination of... Figure 17The described technique is computationally oriented. In at least one embodiment, a set of images is used to generate a viewpoint map, wherein the nodes of the map correspond to images and the edges correspond to their viewpoint equivariant distances. In at least one embodiment, a convolutional neural network (CNN) is used to compute the viewpoint equivariant distances based on the feature similarity of image pairs. In at least one embodiment, an anchor image is selected from a training image set. In at least one embodiment, the anchor image is located from the viewpoint map and the nearest and farthest neighbors are selected based on edge weights. In at least one embodiment, the nearest neighbor has the shortest edge connected to the anchor image. In at least one embodiment, the farthest neighbor has the farthest edge connected to the anchor image. In at least one embodiment, the neural network predicts a first viewpoint of the anchor image and predicts a second viewpoint of the nearest neighbor image and computes a loss such that closer distances between those viewpoints correspond to smaller losses. In at least one embodiment, the neural network predicts a first viewpoint of the anchor image and predicts a third viewpoint of the farthest neighbor image and computes a loss such that longer distances between those viewpoints correspond to smaller losses.

[0188] In at least one embodiment, the nearest and / or farthest neighbors are selected nondeterministically. In at least one embodiment, a probability is assigned to each edge to be selected as the nearest and / or farthest neighbor of the anchor image. In at least one embodiment, the probability of the nearest neighbor is inversely proportional to the edge weight (e.g., the node connected to the anchor image with the lowest edge weight has the highest probability of being selected). In at least one embodiment, the probability of the farthest neighbor is directly proportional to the edge weight (e.g., the node connected to the anchor image with the highest edge weight has the highest probability of being selected).

[0189] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for updating the parameters of one or more neural networks trained on the image set using a loss (e.g., from numbers 1406-1410) computed using 1412. In at least one embodiment, the generator is trained with respect to a symmetry loss, viewpoint consistency loss, true / false classification loss, decoupling loss, or any combination thereof. In at least one embodiment, combined with Figure 4-7 The described technique is used to train a network according to process 1400.

[0190] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for determining whether 1410 should perform further training. In at least one embodiment, training is performed according to any suitable technique and may include selecting a second image and using the second selected image to perform steps 1406-1412 to compute a loss and refine the parameters of one or more neural networks trained for inference viewpoints. Once training is complete, the trained neural network (e.g., the neural network or its parameters transferred to a different system) can be used for inference.

[0191] In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for receiving 1416 images of the same type as a set of images used to train one or more neural networks. In at least one embodiment, images are received from a camera or other type of capture device that is capturing images of the system's surroundings or environment. In at least one embodiment, the system performing at least a portion of process 1400 includes executable code for inferring the viewpoint of objects in an image using a neural network trained on 1418. In at least one embodiment, the vehicle includes a camera that captures images and provides those images to a neural network trained on a set of vehicle images to determine whether the captured images include a vehicle and / or the orientation of any vehicles contained in the captured images.

[0192] Figure 15A An illustrative example of a process 1500A for calculating image consistency loss according to at least one embodiment is shown. In at least one embodiment, some or all of process 1500A (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and can be implemented by hardware, software, or a combination thereof as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1500A are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). The non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 1500A is executed at least partially on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, the first computer system calculates the generative consistency loss. In at least one embodiment, in conjunction with Figure 4-7 The described techniques are applicable to process 1500A. In at least one embodiment, process 1500A describes a process for calculating one or more losses (e.g., image consistency loss), which can be used to update the parameters of the discriminator as part of the training process.

[0193] In at least one embodiment, the system performing at least a portion of process 1500A includes executable code for obtaining an input image of a set of one or more images of an object of a certain type. In at least one embodiment, an object of a certain type may mean that all images in the image set include images of cars. In at least one embodiment, the input image depicts an object in a particular orientation, including specific appearance features (e.g., attributes). In at least one embodiment, the input image is based on a combination of... Figure 4 The described real images (e.g., as opposed to synthetic images).

[0194] In at least one embodiment, a system performing at least a portion of process 1500A includes executable code for using a 1504 discriminator to predict a set of appearance attributes of an input image, determine whether the input image is real or fake, and a viewpoint. In at least one embodiment, the discriminator is associated with one or more neural networks trained to infer viewpoints and other features from the input image. In at least one embodiment, the viewpoint is a predicted viewpoint of the input image and corresponds to a predicted specific orientation of an object depicted in the input image, including specific values ​​for a set of parameters including azimuth, elevation, and tilt parameters. In at least one embodiment, the set of appearance attributes or parameters is a predicted set of appearance attributes of the input image and defines the appearance of an object depicted in the input image. In at least one embodiment, the determination of whether the input image is real or fake is a binary value (e.g., real or fake) instructing the discriminator to predict whether the input image is real or fake. In at least one embodiment, the prediction of whether the input image is real or fake is used to calculate a real / fake classification loss, which can be those described elsewhere in this disclosure, including but not limited to combinations of... Figure 4 and / or Figure 6 Those that were discussed.

[0195] In at least one embodiment, a system performing at least a portion of process 1500A includes executable code for creating a synthetic image using a generator 1506 based at least in part on a predicted set of appearance attributes and a predicted viewpoint. In at least one embodiment, the predicted set of appearance attributes and the predicted viewpoint are provided to the generator to generate the synthetic image. In at least one embodiment, the generator is part of a generative adversarial network. In at least one embodiment, the predicted set of appearance attributes and the predicted viewpoint are used to generate a synthetic image consistent with the predicted set of appearance attributes and the predicted viewpoint. In at least one embodiment, the generator generates a synthetic image comprising objects generated according to the predicted set of appearance attributes and oriented according to a predicted first viewpoint.

[0196] In at least one embodiment, the system executing at least a portion of process 1500A includes executable code for calculating image consistency loss 1508 based on the input image and the synthesized image. In at least one embodiment, the input image and the synthesized image are compared to determine the image consistency loss. In at least one embodiment, the cosine distance between the input image and the synthesized image is compared to determine feature similarity, wherein closer similarity corresponds to lower loss. In at least one embodiment, L1, L2, or cosine distance is used to determine the image consistency loss between the input image and the synthesized image.

[0197] In at least one embodiment, the nearest and farthest neighbor losses are calculated based at least in part on the input image and the synthesized image described in the fusion process 1500A. In at least one embodiment, fusion is used. Figure 4 and Figure 9 The described technique is used to compute nearest and farthest neighbor losses. In at least one embodiment, the symmetric loss is computed at least in part based on the input image.

[0198] In at least one embodiment, Figure 15B The procedure 1500B for calculating viewpoint consistency loss is illustrated. Procedure 1500B can be implemented by any suitable system, such as in conjunction with... Figure 5 Those described. In at least one embodiment, processes 1500A and 1500B are executed by a computer system as part of a training process that adjusts the parameters of the discriminator and generator used to predict the viewpoint of an object. In at least one embodiment, the computer system executing process 1500B includes executable code that enables the computer system to obtain a set of 1512 viewpoint and appearance parameters. In at least one embodiment, the viewpoint and / or appearance parameter set is randomly selected.

[0199] In at least one embodiment, the system performing at least a portion of process 1500B includes executable code for creating a composite image from a set of viewpoint and appearance parameters using a generator 1514. In at least one embodiment, the composite image is based on the above-described combination... Figure 2 The technology described was created.

[0200] In at least one embodiment, a synthetic image is created and the system is configured to use a 1516 discriminator to predict the viewpoint, determine whether the input image is real or fake, and a set of appearance parameters. In at least one embodiment, based on the combination Figure 2 The described technique enables a discriminator to predict viewpoint and appearance.

[0201] In at least one embodiment, the system is configured to calculate a 1518 viewpoint consistency loss based at least in part on the predicted viewpoint (e.g., obtained from a discriminator predicting the viewpoint of the synthetic image) and the input viewpoint (e.g., the viewpoint used by the generator to create the synthetic image). In at least one embodiment, based on the combination Figure 5 and / or Figure 7 The described technique is used to calculate viewpoint consistency loss.

[0202] Figure 16A An illustrative example of a process 1600A for calculating symmetry loss according to at least one embodiment is shown. In at least one embodiment, some or all of the process (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and can be implemented by hardware, software, or a combination thereof as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1600A are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). Non-transitory computer-readable media do not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 1600A is executed at least partially on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, the first computer system calculates the symmetry loss. In at least one embodiment, combined with Figure 8 The described technique is applicable to process 1600A.

[0203] In at least one embodiment, the system performing at least a portion of process 1600A includes executable code for obtaining an input image 1602. In at least one embodiment, the input image is an image in a collection of one or more images of an object of a certain type. In at least one embodiment, an object of a certain type can mean that all images in the image collection include images of a car. In at least one embodiment, the system obtains an input image depicting an object in a specific orientation including specific appearance features.

[0204] In at least one embodiment, a system performing at least a portion of process 1600A includes performing a 1604 transformation on an input image to generate executable code for a transformed image. In at least one embodiment, the transformation is applied to the input image, wherein the input image is horizontally flipped to generate the transformed image. In at least one embodiment, the transformation is applied by one or more systems associated with a discriminator. In at least one embodiment, one or more image processing techniques are applied to the input image to generate the transformed image. In at least one embodiment, when transforming the input image to generate the transformed image, the azimuth and tilt angles of the viewpoints of objects within the input image are reversed, and the elevation angle remains the same on both the input image and the transformed image.

[0205] In at least one embodiment, the system performing at least a portion of process 1600A includes executable code for predicting at least a viewpoint of the transformed image 1606. In at least one embodiment, the discriminator predicts the viewpoint, a set of appearance parameters, a true / false classification, or any combination thereof of the transformed image. In at least one embodiment, the discriminator is associated with one or more neural networks trained to infer viewpoints and other features from the input image. In at least one embodiment, the discriminator receives the transformed image and predicts the viewpoint. In at least one embodiment, the viewpoint corresponds to a specific orientation of an object depicted in the image and includes specific values ​​for a set of parameters, including azimuth, elevation, and tilt parameters.

[0206] In at least one embodiment, a system performing at least a portion of process 1600A includes executable code that applies a transformation 1608 to a predicted viewpoint. In at least one embodiment, the transformation or its inverse transformation is applied to the predicted viewpoint. In at least one embodiment, if an input image is rotated by an angle (Φ, θ, ψ) to produce a transformed image, the predicted viewpoint generated based on the transformed image can be rotated in the opposite direction by an angle (-Φ, -θ, -ψ).

[0207] In at least one embodiment, the system performing at least a portion of process 1600A includes executable code for predicting at least a viewpoint for an input image 1610. In at least one embodiment, a discriminator predicts the viewpoint, a set of appearance parameters, a true / false classification, or any combination thereof for the input image. In at least one embodiment, the discriminator receives the input image and predicts the viewpoint. In at least one embodiment, the viewpoint corresponds to a specific orientation of an object depicted in the image and includes specific values ​​for a set of parameters, which includes azimuth parameters, elevation parameters, and tilt parameters.

[0208] In at least one embodiment, the system performing at least a portion of process 1600A includes executable code for comparing 1612 viewpoints to calculate a symmetry loss. In at least one embodiment, the system compares a predicted viewpoint of a transformed image with a predicted viewpoint of an input image, wherein the transformed image is generated from the input image. In at least one embodiment, the symmetry loss is calculated by comparing the magnitudes of the azimuth, elevation, and tilt of the predicted viewpoint of the input image with the magnitudes of the azimuth, elevation, and tilt of the predicted viewpoint of the transformed image, wherein zero loss is produced when the magnitudes of each viewpoint parameter are equal. In at least one embodiment, the symmetry loss is calculated at least in part based on the appearance parameters of an image (e.g., the input image) and how closely the transformed versions of that image match each other.

[0209] Figure 16B The illustration depicts a process 1600B for calculating symmetry loss according to at least one embodiment. In at least one embodiment, Figure 16B The parameters of the generator are used to update the generator as part of training the neural network to predict the viewpoint. In at least one embodiment, the system performing process 1600B includes executable code for obtaining a set of 1614 viewpoints and appearance attributes. In at least one embodiment, the system is configured to apply a 1616 transformation to the obtained viewpoint to determine a transformed viewpoint. In at least some embodiments, the viewpoint is flipped to obtain the transformed viewpoint. In at least one embodiment, a transformation function T() is applied to parameter sets x1, y1, and z1 to obtain a transformed parameter set x2, y2, z2, which may correspond to azimuth, tilt, and elevation parameters.

[0210] In at least one embodiment, the generator is used to generate a first synthetic image 1618 based at least in part on a transformed set of appearance parameters, the set of appearance parameters being able to combine... Figure 16B The other steps described are used to calculate or otherwise determine the transformation. In at least one embodiment, the system is configured to apply transformation 1620 to a first composite image generated from a transformed set of viewpoint and appearance attributes. In at least one embodiment, if a transformed viewpoint is generated by performing transformation T(), then inverse transformation T() is performed. -1 () is applied to synthesized images, where T(T) -1 (x,y,z))=(x,y,z). In at least one embodiment, if the transform horizontally flips the image, the inverse transform horizontally flips the image back to its original orientation.

[0211] In at least one embodiment, the system includes executable instructions for generating a second synthetic image 1622 based at least in part on a set of viewpoint and appearance parameters. In at least one embodiment, the same generator is used to generate a first synthetic image based on a transformed viewpoint and a second synthetic image based on the original viewpoint. In at least one embodiment, the system is configured to compare the first synthetic image 1624 and the second synthetic image to calculate a symmetry loss, wherein the zero cosine distance between the images is associated with zero loss.

[0212] Figure 17 An illustrative example of a process 1700 for calculating nearest and farthest neighbor losses according to at least one embodiment is shown. In at least one embodiment, some or all of process 1700 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and can be implemented by hardware, software, or a combination thereof as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) that executes jointly on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1700 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). Non-transitory computer-readable media do not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 1700 is executed at least partially on a computer system, such as those described elsewhere in this disclosure. In at least one embodiment, the first computer system calculates the nearest and farthest neighbor losses. In at least one embodiment, combined with Figure 9 The described technique is applicable to process 1700.

[0213] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code for obtaining a set of images 1702. In at least one embodiment, the set of images includes one or more images of an object of one type. In at least one embodiment, an object of one type can mean that all images in the set of images include images of cars. In at least one embodiment, the system operates according to techniques described elsewhere in this disclosure (e.g., Figure 13 () Obtain a collection of one or more images.

[0214] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code for calculating a cosine distance for comparing feature similarities for image pairs in the set at 1704. In at least one embodiment, the cosine distance is a mathematical complement to cosine similarity (e.g., cosine distance = 1 - cosine similarity). In at least one embodiment, cosine similarity is a similarity measure between two vectors that may represent images, text, data, and / or variations thereof based on the cosine of the angle between them. In at least one embodiment, a lower cosine distance is calculated for the images when the two images include features corresponding to objects with similar viewpoints. In at least one embodiment, a higher cosine distance is calculated for the images when the two images include features corresponding to objects with different viewpoints. In at least one embodiment, a cosine distance is calculated for each image in the image set relative to each other image in the set.

[0215] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code that generates a viewpoint graph 1706, wherein the nodes of the graph correspond to images in a set and the edges correspond to their cosine distances. In at least one embodiment, a convolutional neural network is used to compute the cosine distance based on the feature similarity of image pairs. In at least one embodiment, the edges of the viewpoint graph are weighted such that a thicker edge between two images corresponds to a higher similarity between the two images, while a thinner edge between two images corresponds to a lower similarity between the two images.

[0216] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code for selecting an anchor image of map 1708 and predicting a first viewpoint. In at least one embodiment, an anchor image is selected from a training image set for generating the viewpoint map. In at least one embodiment, a discriminator is associated with one or more neural networks trained to infer viewpoints and other features from input images. In at least one embodiment, the discriminator receives the anchor image and predicts the first viewpoint. In at least one embodiment, the first viewpoint corresponds to a predicted specific orientation of an object depicted in the anchor image and includes specific values ​​of a set of parameters, including azimuth, elevation, and tilt parameters corresponding to the orientation of the object.

[0217] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code for using graph 1710 to select the nearest neighbor of the anchor image and predict a second viewpoint.

[0218] In at least one embodiment, the nearest neighbor is determined based on the edge weights of the viewpoint graph. In at least one embodiment, the nearest neighbor is the image with the shortest edge connected to the anchor image in the viewpoint graph. In at least one embodiment, the nearest neighbor of the anchor image is the image most similar to the anchor image in the image set. In at least one embodiment, the discriminator receives the nearest neighbor image and predicts a second viewpoint. In at least one embodiment, the second viewpoint corresponds to a predicted specific direction of an object depicted in the nearest neighbor image and includes specific values ​​of a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter corresponding to the direction of the object.

[0219] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code for using graph 1712 to select the farthest neighbor of the anchor image and predict a third viewpoint.

[0220] In at least one embodiment, the farthest neighbor is determined based on the edge weights of the viewpoint graph. In at least one embodiment, the farthest neighbor is the image with the longest edge connected to the anchor image in the viewpoint graph. In at least one embodiment, the farthest neighbor of the anchor image is the image most different from the anchor image in the image set. In at least one embodiment, the discriminator receives the farthest neighbor image and predicts a third viewpoint. In at least one embodiment, the third viewpoint corresponds to a predicted specific orientation of an object depicted in the farthest neighbor image and includes specific values ​​of a set of parameters, including an azimuth parameter, an elevation parameter, and a tilt parameter corresponding to the orientation of the object.

[0221] In at least one embodiment, the system performing at least a portion of process 1700 includes executable code for calculating 1714 nearest and farthest neighbor losses. In at least one embodiment, a nearest neighbor loss is calculated between a first viewpoint predicted from an anchor image and a second viewpoint predicted from a nearest neighbor image of the anchor image. In at least one embodiment, the nearest neighbor loss is calculated such that a higher similarity between the first and second viewpoints corresponds to a smaller loss. In at least one embodiment, a farthest neighbor loss is calculated between a first viewpoint predicted from an anchor image and a third viewpoint predicted from a farthest neighbor image of the anchor image. In at least one embodiment, the farthest neighbor loss is calculated such that a lower similarity between the first and third viewpoints corresponds to a smaller loss. In at least one embodiment, the nearest and farthest neighbor losses are calculated based on a combination of the nearest and farthest neighbor losses.

[0222] Reasoning and training logic

[0223] Figure 18A Inference and / or training logic 1815 for performing inference and / or training operations associated with one or more embodiments is shown. The following is in conjunction with... Figure 18A and / or Figure 18B Provide details regarding reasoning and / or training logic 1815.

[0224] In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, code and / or data storage 1801 for storing forward and / or output weights and / or input / output data, and / or other parameters configuring neurons or layers of a neural network trained for and / or used for inference in one or more embodiments. In at least one embodiment, the training logic 1815 may include or be coupled to code and / or data storage 1801 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic, including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code (e.g., graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, the code and / or data storage 1801 stores weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using one or more embodiments.

[0225] In at least one embodiment, any portion of the code and / or data storage 1801 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0226] In at least one embodiment, any portion of the code and / or data storage 1801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1801 may be a cache memory, dynamic random-addressable memory (“DRAM”), static random-addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or code and / or data storage 1801 is internal or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip and off-chip storage, the latency requirements of the training and / or inference functions being performed, the data batch size used in neural network inference and / or training, or some combination of these factors.

[0227] In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, code and / or data storage 1805 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, the code and / or data storage 1805 stores the weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during backpropagation of input / output data and / or weight parameters during training, and / or inference using one or more embodiments. In at least one embodiment, training logic 1815 may include or be coupled to code and / or data storage 1805 to store graph code or other software to control timing and / or sequence, wherein weights and / or other parameter information are loaded to configure the logic, including integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weights or other parameter information into the processor ALU based on the architecture of the neural network corresponding to the code. In at least one embodiment, any portion of the code and / or data storage 1805 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment Any portion of the code and / or data storage 1805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1805 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or data storage 1805 is internal or external to the processor, or whether it comprises DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip and off-chip storage, the latency requirements of the training and / or inference functions being performed, the data batch size used in neural network inference and / or training, or some combination of these factors.

[0228] In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be separate storage structures. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be the same storage structure. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be partially identical and partially separate storage structures. In at least one embodiment, any portion of code and / or data storage 1801 and code and / or data storage 1805 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.

[0229] In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 1810 (including integer and / or floating-point units) for performing logical and / or mathematical operations at least in part based on or instructed by training and / or inference code (e.g., graph code), the results of which may produce activations (e.g., output values ​​from layers or neurons within a neural network) stored in activation storage 1820, which are functions of input / output and / or weight parameter data stored in code and / or data storage 1801 and / or code and / or data storage 1805. In at least one embodiment, activation is activated in response to execution instructions or other code, and linear algebraic and / or matrix-based mathematical generation performed by ALU 1810 is stored in activation storage 1820. The weight values ​​stored in code and / or data storage 1805 and / or data 1801 are used as operands with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters. Any or all of these can be stored in code and / or data storage 1805 or code and / or data storage 1801 or other on-chip or off-chip storage.

[0230] In at least one embodiment, one or more ALUs 1810 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment, one or more ALUs 1810 may be located outside the processor or other hardware logic device or the circuits using them (e.g., coprocessors). In at least one embodiment, one or more ALUs 1810 may be included within an execution unit of a processor, or otherwise included in a group of ALUs accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., central processing unit, graphics processing unit, fixed-function unit, etc.). In at least one embodiment, data storage 1801, code and / or data storage 1805, and activation storage 1820 may be in the same processor or other hardware logic device or circuit, while in another embodiment, they may be in different processors or other hardware logic devices or circuits, or some combination of the same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 1820 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.

[0231] In at least one embodiment, the active memory 1820 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 1820 may be wholly or partially located inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 1820 is internal to or external to the processor, for example, or composed of DRAM, SRAM, flash memory, or some other memory type, may depend on the available on-chip or off-chip storage, the latency requirements for training and / or inference functions, the batch size of data used in inference and / or training the neural network, or some combination of these factors. In at least one embodiment, Figure 18A The inference and / or training logic 1815 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM The inference processing unit (IPU) or from Intel. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 18A The inference and / or training logic 1815 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”)

[0232] Figure 18B Inference and / or training logic 1815 according to at least one embodiment is illustrated. In at least one embodiment, the inference and / or training logic 1815 may include, but is not limited to, hardware logic, wherein computational resources are dedicated or otherwise uniquely used in conjunction with weight values ​​or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 18B The inference and / or training logic 1815 shown can be used in conjunction with an application-specific integrated circuit (ASIC), such as those from Google. Processing unit, from Graphcore TM The inference processing unit (IPU) or from Intel. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 18BThe inference and / or training logic 1815 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, the inference and / or training logic 1815 includes, but is not limited to, code and / or data storage 1801 and code and / or data storage 1805, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 18B In at least one embodiment shown, each of code and / or data storage 1801 and code and / or data storage 1805 is associated with dedicated computing resources (e.g., computing hardware 1802 and computing hardware 1806), respectively. In at least one embodiment, each of computing hardware 1802 and computing hardware 1806 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) only on the information stored in code and / or data storage 1801 and code and / or data storage 1805, respectively, and the results of the function execution are stored in activation memory 1820.

[0233] In at least one embodiment, each of the code and / or data storage 1801 and 1805 and the corresponding computing hardware 1802 and 1806 corresponds to a different layer of the neural network, such that activation obtained from one “store / computation pair 1801 / 1802” of the code and / or data storage 1801 and computing hardware 1802 provides input as input to the next “store / computation pair 1805 / 1806” of the code and / or data storage 1805 and computing hardware 1806, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each store / computation pair 1801 / 1802 and 1805 / 1806 may correspond to more than one neural network layer. In at least one embodiment, additional store / computation pairs (not shown) may be included in the inference and / or training logic 1815 following or paralleling the store / computation pairs 1801 / 1802 and 1805 / 1806.

[0234] Neural network training and deployment

[0235] Figure 19Training and deployment of a deep neural network according to at least one embodiment are illustrated. In at least one embodiment, an untrained neural network 1906 is trained using a training dataset 1902. In at least one embodiment, the training framework 1904 is the PyTorch framework, while in other embodiments, the training framework 1904 is Tensorflow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 and enables it to be trained using the processing resources described herein to generate a trained neural network 1908. In at least one embodiment, the weights may be randomly selected or pre-trained using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.

[0236] In at least one embodiment, supervised learning is used to train an untrained neural network 1906, wherein the training dataset 1902 includes inputs paired with desired outputs for input, or wherein the training dataset 1902 includes inputs with known outputs and the neural network 1906 is manually graded output. In at least one embodiment, the untrained neural network 1906 is trained in a supervised manner, processing inputs from the training dataset 1902 and comparing the resulting outputs with a set of expected or desired outputs. In at least one embodiment, errors are then propagated back through the untrained neural network 1906. In at least one embodiment, a training framework 1904 adjusts the weights controlling the untrained neural network 1906. In at least one embodiment, the training framework 1904 includes tools for monitoring the degree to which the untrained neural network 1906 converges to a model (e.g., a trained neural network 1908) adapted to generate the correct answer (e.g., result 1914) based on known input data (e.g., a new dataset 1912). In at least one embodiment, the training framework 1904 repeatedly trains the untrained neural network 1906 while adjusting the weights to improve the output of the untrained neural network 1906 using a loss function and tuning algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 1904 trains the untrained neural network 1906 until the untrained neural network 1906 reaches the desired accuracy. In at least one embodiment, the trained neural network 1908 can then be deployed to implement any number of machine learning operations.

[0237] In at least one embodiment, unsupervised learning is used to train an untrained neural network 1906, wherein the untrained neural network 1906 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 1902 will include input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 1906 can learn groupings within the training dataset 1902 and can determine how each input relates to the untrained dataset 1902. In at least one embodiment, unsupervised training can be used to generate a self-organizing graph, which is a type of trained neural network 1908 capable of performing operations useful for reducing the dimensionality of the new dataset 1912. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows the identification of data points in the new dataset 1912 that deviate from the normal patterns of the new dataset 1912.

[0238] In at least one embodiment, semi-supervised learning can be used, which is a technique that includes a mixture of labeled and unlabeled data in the training dataset 1902. In at least one embodiment, the training framework 1904 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 1908 to adapt to a new dataset 1912 without forgetting the knowledge injected into the network during initial training.

[0239] Data Center

[0240] Figure 20 An example data center 2000 that can be used in at least one embodiment is shown. In at least one embodiment, the data center 2000 includes a data center infrastructure layer 2010, a framework layer 2020, a software layer 2030, and an application layer 2040.

[0241] In at least one embodiment, such as Figure 20As shown, the data center infrastructure layer 2010 may include a resource coordinator 2012, packet computing resources 2014, and node computing resources (“nodes CR”) 2016(1)-2016(N), where “N” represents any whole positive integer. In at least one embodiment, nodes CR 2016(1)-2016(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., solid-state drives or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 2016(1)-2016(N) may be servers having one or more of the aforementioned computing resources.

[0242] In at least one embodiment, the grouped computing resource 2014 may include individual groups (not shown) of node CRs housed within one or more racks, or a plurality of racks (also not shown) housed within data centers in various geographical locations. The individual groups of node CRs within the individually grouped computing resource 2014 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches, in any combination.

[0243] In at least one embodiment, the resource coordinator 2012 may be configured or otherwise control one or more nodes CR2016(1)-2016(N) and / or grouped computing resources 2014. In at least one embodiment, the resource coordinator 2012 may include a Software Design Infrastructure (“SDI”) management entity for the data center 2000. In at least one embodiment, the resource coordinator may include hardware, software, or some combination thereof.

[0244] In at least one embodiment, such as Figure 20As shown, the framework layer 2020 includes a job scheduler 2032, a configuration manager 2034, a resource manager 2036, and a distributed file system 2038. In at least one embodiment, the framework layer 2020 may include a framework of software 2032 supporting the software layer 2030 and / or one or more applications 2042 of the application layer 2040. In at least one embodiment, the software 2032 or application 2042 may respectively include web-based service software or applications, such as services or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 2020 may be, but is not limited to, a free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark") which can utilize the distributed file system 2038 for large-scale data processing (e.g., "big data"). In at least one embodiment, the job scheduler 2032 may include a Spark driver to facilitate the scheduling of job loads supported by the various layers of the data center 2000. In at least one embodiment, configuration manager 2034 may be able to configure different layers, such as software layer 2030 and framework layer 2020 including Spark and distributed file system 2038 for supporting large-scale data processing. In at least one embodiment, resource manager 2036 is able to manage cluster or group computing resources mapped to or allocated to support distributed file system 2038 and job scheduler 2032. In at least one embodiment, cluster or group computing resources may include group computing resources 2014 on data center infrastructure layer 2010. In at least one embodiment, resource manager 2036 may coordinate with resource coordinator 2012 to manage these mapped or allocated computing resources.

[0245] In at least one embodiment, the software 2032 included in the software layer 2030 may include software used by at least a portion of the nodes CR2016(1)-2016(N), the grouped computing resources 2014, and / or the distributed file system 2038 of the framework layer 2020. One or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.

[0246] In at least one embodiment, one or more applications 2042 included in application layer 2040 may include one or more types of applications used by at least a portion of nodes CR2016(1)-2016(N), grouped computing resources 2014, and / or the distributed file system 2038 of framework layer 2020. One or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0247] In at least one embodiment, any of the configuration manager 2034, resource manager 2036, and resource coordinator 2012 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by the data center operator of the data center 2000 and can prevent underutilization and / or poor performance of the data center.

[0248] In at least one embodiment, the data center 2000 may include tools, services, software, or other resources to train one or more machine learning models or to use one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to the data center 2000. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to the data center 2000, by using weight parameters calculated through one or more training techniques described herein.

[0249] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.

[0250] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18BDetails regarding the inference and / or training logic 1815 are provided. In at least one embodiment, the inference and / or training logic 1815 can be implemented in the system. Figure 20 Used in this context for inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0251] In at least one embodiment, the system Figure 20 This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 20 This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 20 Used to implement one or more neural networks including discriminators and generators, and the system Figure 20 and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0252] Autonomous vehicles

[0253] Figure 21A An example of an autonomous vehicle 2100 according to at least one embodiment is shown. In at least one embodiment, the autonomous vehicle 2100 (which may alternatively be referred to herein as "vehicle 2100") may be, but is not limited to, a passenger vehicle, such as a car, truck, bus, and / or another type of vehicle capable of accommodating one or more passengers. In at least one embodiment, vehicle 2100 may be a semi-tractor-trailer for hauling goods. In at least one embodiment, vehicle 2100 may be an aircraft, robotic vehicle, or other type of vehicle.

[0254] Autonomous vehicles can be described according to the levels of automation defined by the National Highway Traffic Safety Administration (“NHTSA”) and the Society of Automotive Engineers (“SAE”) of the U.S. Department of Transportation in their standard “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., standard number J3016-201806, published June 15, 2018; standard number J3016-201609, published September 30, 2016; and previous and future versions of this standard). In one or more embodiments, vehicle 2100 may be able to function according to one or more of the levels of autonomous driving from Level 1 to Level 5. For example, in at least one embodiment, vehicle 2100 may be able to perform conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5).

[0255] In at least one embodiment, vehicle 2100 may include, but is not limited to, components such as chassis, body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other vehicle components. In at least one embodiment, vehicle 2100 may include, but is not limited to, propulsion system 2150, such as an internal combustion engine, a hybrid powertrain, an all-electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 2150 may be connected to the drivetrain of vehicle 2100, which may include, but is not limited to, a transmission, to enable propulsion of vehicle 2100. In at least one embodiment, propulsion system 2150 may be controlled in response to receiving a signal from throttle / accelerator 2152.

[0256] In at least one embodiment, when the propulsion system 2150 is operating (e.g., when the vehicle is in motion), the steering system 2154 (which may include, but is not limited to, a steering wheel) is used to steer the vehicle 2100 (e.g., along a desired path or route). In at least one embodiment, the steering system 2154 may receive signals from the steering actuator 2156. The steering wheel may be optional for fully automated (Level 5) functionality. In at least one embodiment, the brake sensor system 2146 may be used to operate the vehicle brakes in response to signals received from the brake actuator 2148 and / or brake sensors.

[0257] In at least one embodiment, controller 2136 may include, but is not limited to, one or more system-on-chips (“SoCs”). Figure 21AA controller 2136 (not shown) and / or a graphics processing unit (“GPU”) provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 2100. For example, in at least one embodiment, controller 2136 may send signals to operate vehicle braking via brake actuator 2148, to operate steering system 2154 via one or more steering actuators 2156, and to operate propulsion system 2150 via one or more throttles / accelerators 2152. One or more controllers 2136 may include one or more onboard (e.g., integrated) computing devices (e.g., supercomputers) that process sensor signals and output operating commands (e.g., signals representing commands) to enable autonomous driving and / or assist a driver in driving vehicle 2100. In at least one embodiment, one or more controllers 2136 may include a first controller 2136 for autonomous driving functions, a second controller 2136 for functional safety functions, a third controller 2136 for artificial intelligence functions (e.g., computer vision), a fourth controller 2136 for infotainment functions, a fifth controller 2136 for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller 2136 may handle two or more of the above functions, and two or more controllers 2136 may handle a single function and / or any combination thereof.

[0258] In at least one embodiment, one or more controllers 2136 provide signals for controlling one or more components and / or systems of vehicle 2100 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, the sensor data can be received from sensors, including, but not limited to, one or more Global Navigation Satellite System (“GNSS”) sensors 2158 (e.g., one or more Global Positioning System sensors), one or more RADAR sensors 2160, one or more ultrasonic sensors 2162, one or more LIDAR sensors 2164, one or more inertial measurement unit (IMU) sensors 2166 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 2196, one or more stereo cameras 2168, one or more wide-angle cameras 2170 (e.g., fisheye cameras), one or more infrared cameras 2172, one or more surround cameras 2174 (e.g., 360-degree cameras), and remote cameras (…). Figure 21A (not shown in the image), medium-range camera ( Figure 21A(Not shown in the diagram) One or more speed sensors 2144 (e.g., for measuring the speed of vehicle 2100), one or more vibration sensors 2142, one or more steering sensors 2140, one or more brake sensors (e.g., as part of brake sensor system 2146) and / or other sensor types are received.

[0259] In at least one embodiment, one or more controllers 2136 may receive input (e.g., represented by input data) from the dashboard 2132 of the vehicle 2100 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 2134, a voice signaler, a speaker, and / or other components of the vehicle 2100. In at least one embodiment, the output may include information such as vehicle speed, velocity, time, map data (e.g., high-definition map). Figure 21A The HMI display 2134 may display information such as (not shown in the image), location data (e.g., the location of vehicle 2100, for example, on a map), direction, the location of other vehicles (e.g., occupancy raster), information about objects, and the state of objects perceived by one or more controllers 2136. For example, in at least one embodiment, the HMI display 2134 may display information about the presence of one or more objects (e.g., road signs, warning signs, traffic light changes, etc.) and / or information about driving operations that the vehicle has already made, is making, or will make (e.g., changing lanes now, exiting exit 34B within two miles, etc.).

[0260] In at least one embodiment, vehicle 2100 further includes a network interface 2124 that can communicate over one or more networks using one or more wireless antennas 2126 and / or one or more modems. For example, in at least one embodiment, network interface 2124 may be able to communicate via Long Term Evolution (“LTE”), Wideband Code Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multicarrier (“CDMA2000”), etc. In at least one embodiment, one or more wireless antennas 2126 may also enable communication between objects in the environment (e.g., vehicles, mobile devices) using one or more local area networks (e.g., Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more low-power wide area networks (hereinafter “LPWAN”) (e.g., LoRaWAN, SigFox, etc.).

[0261] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding the inference and / or training logic 1815 are provided. In at least one embodiment, the inference and / or training logic 1815 can be implemented in the system. Figure 21A The operation is used to infer or predict the operation based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0262] In at least one embodiment, the system Figure 21A This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 21A This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 21A Used to implement one or more neural networks including discriminators and generators, and the system Figure 21A and one One or more processes are used in combination to train one or more neural networks in a self-supervised manner by computing one or more loss functions as part of training to evaluate one or more features of images in the training set, in order to identify the orientation of objects within an image.

[0263] Figure 21B The illustration shows an embodiment according to at least one of the embodiments. Figure 21A Examples of camera positions and fields of view for an autonomous vehicle 2100. In at least one embodiment, the camera and its respective field of view are exemplary embodiments and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located at different positions on the vehicle 2100.

[0264] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera suitable for use with components and / or systems of vehicle 2100. One or more cameras may operate at Automotive Safety Integrity Level (“ASIL”) B and / or other ASILs. In at least one embodiment, the camera type may have any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc. In at least one embodiment, the camera may be able to use a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a red-to-clear (“RCCC”) color filter array, a red-to-clear-blue (“RCCB”) color filter array, a red-blue-green (“RBGC”) color filter array, a Foveon X3 color filter array, a Bayer sensor (“RGGB”) color filter array, a monochrome sensor color filter array, and / or other types of color filter arrays. In at least one embodiment, a transparent pixel camera, such as a camera with RCCC, RCCB, and / or RBGC color filter arrays, may be used to improve photosensitivity.

[0265] In at least one embodiment, one or more cameras may be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundancy or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed to provide functions including lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).

[0266] In at least one embodiment, one or more cameras may be mounted in a mounting assembly, such as a custom-designed (3D-printed) assembly, to cut out stray light and reflections from within the vehicle (e.g., dashboard reflections in the windshield mirror), which may interfere with the camera's image data capture capabilities. Regarding the rearview mirror mounting assembly, in at least one embodiment, the rearview mirror assembly may be 3D-printed custom-made such that the camera mounting plate matches the shape of the rearview mirror. In at least one embodiment, one or more cameras may be integrated into the rearview mirror. In a lesser embodiment, for side-view cameras, one or more cameras may also be integrated within four pillars at each corner of the cabin.

[0267] In at least one embodiment, a camera (e.g., a forward-facing camera) having a field of view including a portion of the environment in front of the vehicle 2100 can be used for surround view and, with the assistance of one or more controllers 2136 and / or control SoCs, to help identify the forward path and obstacles, thereby providing information crucial for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the forward-facing camera can be used to perform many of the same ADAS functions as LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the forward-facing camera can also be used for ADAS functions and systems, including but not limited to lane departure warning (“LDW”), adaptive cruise control (“ACC”), and / or other functions (e.g., traffic sign recognition).

[0268] In at least one embodiment, various cameras can be used in a forward-facing configuration, including, for example, a monocular camera platform including a CMOS (“complementary metal-oxide-semiconductor”) color imager. In at least one embodiment, a wide-angle camera 2170 can be used to sense objects entering from the periphery (e.g., pedestrians, people crossing the street, or bicycles). Although in Figure 21B Only one wide-angle camera 2170 is shown in the illustration; however, in other embodiments, the vehicle 2100 may have any number (including zero) of one or more wide-angle cameras 2170. In at least one embodiment, any number of remote cameras 2198 (e.g., a pair of remote stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not yet been trained. In at least one embodiment, the remote cameras 2198 can also be used for object detection and classification, as well as basic object tracking.

[0269] In at least one embodiment, any number of stereo cameras 2168 may also be included in a forward configuration. In at least one embodiment, one or more stereo cameras 2168 may include an integrated control unit comprising a scalable processing unit that may provide programmable logic (“FPGA”) and a multi-core microprocessor with a controller area network (“CAN”) or Ethernet interface integrated on a single chip. In at least one embodiment, such a unit may be used to generate a 3D map of the environment of vehicle 2100, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 2168 may include, but are not limited to, a compact stereo vision sensor, which may include, but is not limited to, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance from vehicle 2100 to a target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions.

[0270] In at least one embodiment, other types of stereo cameras 2168 may be used in addition to those described herein.

[0271] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view including a portion of the environment on the side of vehicle 2100 can be used for surround viewing, thereby providing information for creating and updating the occupied grid, and generating a side collision warning. For example, in at least one embodiment, a surround camera 2174 (e.g., such as...) Figure 21B The four surround cameras 2174 shown can be positioned on the vehicle 2100. One or more surround cameras 2174 can include, but are not limited to, any number and combination of one or more wide-angle cameras 2170, one or more fisheye lenses, one or more 360-degree cameras, etc. For example, in at least one embodiment, the four fisheye lens cameras can be located at the front, rear, and sides of the vehicle 2100. In at least one embodiment, the vehicle 2100 can use three surround cameras 2174 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., forward-facing cameras) as a fourth surround-view camera.

[0272] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view including a portion of the environment behind vehicle 2100 can be used for parking assistance, surround view, rear collision warning, and creating and updating occupancy raster. In at least one embodiment, a wide variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward-facing cameras (e.g., long-range camera 2198 and / or one or more mid-range cameras 2176, one or more stereo cameras 2168, one or more infrared cameras 2172, etc.), as described herein.

[0273] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. Figure 18A and / or Figure 18B Details regarding inference and / or training logic 1815 are provided herein. In at least one embodiment, inference and / or training logic 1815 may be Figure 21B Used in systems for reasoning or predicting operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.

[0274] In at least one embodiment, the system Figure 21B This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 21BThis is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 21B Used to implement one or more neural networks including discriminators and generators, and the system Figure 21B and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0275] Figure 21C The illustration shows an embodiment according to at least one of the embodiments. Figure 21A A block diagram of an example system architecture for an autonomous vehicle 2100. In at least one embodiment, Figure 21C Each of one or more components, one or more features, and one or more systems of vehicle 2100 is shown connected via bus 2102. In at least one embodiment, bus 2102 may include, but is not limited to, a CAN data interface (which may alternatively be referred to herein as “CAN bus”). In at least one embodiment, CAN may be a network within vehicle 2100 used to help control various features and functions of vehicle 2100, such as brake actuation, acceleration, braking, steering, windshield wipers, etc. In one embodiment, bus 2102 may be configured to have dozens or even hundreds of nodes, each node having its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 2102 can be read to find steering wheel angle, ground speed, engine rotation speed (“RPM”), button position, and / or other vehicle status indicators. In at least one embodiment, bus 2102 may be an ASIL B compliant CAN bus.

[0276] In at least one embodiment, FlexRay and / or Ethernet may be used in addition to or from CAN. In at least one embodiment, there may be any number of buses 2102, which may include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 2102 may be used to perform different functions and / or may be used for redundancy. For example, a first bus 2102 may be used for collision avoidance functions, and a second bus 2102 may be used for actuation control. In at least one embodiment, each bus 2102 may communicate with any component of vehicle 2100, and two or more buses 2102 may communicate with the same component. In at least one embodiment, each of any number of system-on-chip (“SoC”) 2104, each of one or more controllers 2136, and / or each computer within the vehicle may access the same input data (e.g., input from sensors of vehicle 2100) and may be connected to a common bus, such as a CAN bus.

[0277] In at least one embodiment, vehicle 2100 may include one or more controllers 2136, such as those described herein. Figure 21A The described ones. One or more controllers 2136 can be used for a variety of functions. In at least one embodiment, controller 2136 can be coupled to any of various other components and systems of vehicle 2100 and can be used to control vehicle 2100, artificial intelligence of vehicle 2100, infotainment of vehicle 2100, etc.

[0278] In at least one embodiment, vehicle 2100 may include any number of SoCs 2104. Each SoC 2104 may include, but is not limited to, a central processing unit (“one or more CPUs”) 2106, a graphics processing unit (“one or more GPUs”) 2108, one or more processors 2110, one or more caches 2112, one or more accelerators 2114, one or more data storage 2116, and / or other components and features not shown. In at least one embodiment, one or more SoCs 2104 may be used to control vehicle 2100 on various platforms and systems. For example, in at least one embodiment, one or more SoCs 2104 may be combined with a high-definition (“HD”) map 2122 in a system (e.g., the system of vehicle 2100), the high-definition map 2122 being available from one or more servers via a network interface 2124. Figure 21C (Not shown in the image) Get map refresh and / or update.

[0279] In at least one embodiment, one or more CPUs 2106 may include CPU clusters or CPU complexes (which may alternatively be referred to herein as “CCPLEX”). In at least one embodiment, one or more CPUs 2106 may include multiple cores and / or a secondary (“L2”) cache. For example, in at least one embodiment, one or more CPUs 2106 may include eight cores in an intercoupled multiprocessor configuration. In at least one embodiment, one or more CPUs 2106 may include four dual-core clusters, each cluster having a dedicated L2 cache (e.g., 2MB L2 cache). In at least one embodiment, one or more CPUs 2106 (e.g., CCPLEX) may be configured to support simultaneous cluster operation, such that any combination of clusters of one or more CPUs 2106 can be active at any given time.

[0280] In at least one embodiment, one or more CPUs 2106 may implement power management functions, including but not limited to one or more of the following features: automatic clock gating of individual hardware modules to conserve dynamic power when idle; clock gating of each core when the core is not actively executing instructions due to executing Wait for Interrupt (“WFI”) / Event Wait (“WFE”) instructions; independent power supply for each core; independent clock gating for each core cluster when all cores are clock-gated or power-gated; and / or independent power gating for each core cluster when all cores are power-gated. In at least one embodiment, one or more CPUs 2106 may further implement an enhanced algorithm for managing power states, wherein allowed power states and expected wake-up times are specified, and the hardware / microcode determines the optimal power state for cores, clusters, and CCPLEX inputs. In at least one embodiment, the processing core may support a simplified power state input sequence in software, wherein the work is offloaded to the microcode.

[0281] In at least one embodiment, one or more GPUs 2108 may include integrated GPUs (or "iGPUs" herein). In at least one embodiment, one or more GPUs 2108 may be programmable and efficient for parallel workload loading. In at least one embodiment, one or more GPUs 2108 may use an enhanced tensor instruction set. In one embodiment, one or more GPUs 2108 may include one or more streaming microprocessors, wherein each streaming microprocessor may include a Level 1 ("L1") cache (e.g., an L1 cache with at least 96KB of storage capacity), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with 512KB of storage capacity). In at least one embodiment, one or more GPUs 2108 may include at least eight streaming microprocessors. In at least one embodiment, one or more GPUs 2108 may use a computation application programming interface (API). In at least one embodiment, one or more GPUs 2108 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA).

[0282] In at least one embodiment, one or more GPUs 2108 may be power-optimized for optimal performance in automotive and embedded use cases. For example, in one embodiment, one or more GPUs 2108 may be fabricated on a FinFET (“FinFET”). In at least one embodiment, each streaming microprocessor may include multiple mixed-precision processing cores divided into multiple blocks. For example, but not limited to, 64 PF32 cores and 32 PF64 cores may be divided into four processing blocks. In at least one embodiment, each processing block may be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor cores for deep learning matrix arithmetic, a level-zero (“L0”) instruction cache, a thread bundle scheduler, a dispatch unit, and / or a 64KB register file. In at least one embodiment, the streaming microprocessor may include independent parallel integer and floating-point data paths to provide efficient execution of workloads that mix computation and addressing operations. In at least one embodiment, the streaming microprocessor may include independent thread scheduling capabilities to enable finer-grained synchronization and collaboration between parallel threads. In at least one embodiment, the streaming microprocessor may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.

[0283] In at least one embodiment, one or more GPUs 2108 may include high-bandwidth memory (“HBM”) and / or a 16GB HBM2 memory subsystem to provide a peak storage bandwidth of approximately 900GB / s in some examples. In at least one embodiment, in addition to or instead of HBM memory, synchronous graphics random access memory (“SGRAM”) may be used, such as graphics double data rate type five synchronous random access memory (“GDDR5”).

[0284] In at least one embodiment, one or more GPUs 2108 may include unified memory technology. In at least one embodiment, address translation service (“ATS”) support may be used to allow one or more GPUs 2108 to directly access the page tables of one or more CPUs 2106. In at least one embodiment, when one or more GPUs 2108 memory management units (“MMUs”) experience a miss, an address translation request may be sent to one or more CPUs 2106. In response, in at least one embodiment, one or more CPUs 2106 may look up the virtual-physical mapping of the address in their page tables and transfer the translation back to one or more GPUs 2108. In at least one embodiment, unified memory technology may allow a single unified virtual address space to be used for the memory of both one or more CPUs 2106 and one or more GPUs 2108, thereby simplifying the programming of one or more GPUs 2108 and the porting of applications to one or more GPUs 2108.

[0285] In at least one embodiment, one or more GPUs 2108 may include any number of access counters that can track the frequency of memory accesses by one or more GPUs 2108 to other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses the pages most frequently, thereby improving the efficiency of shared memory ranges between processors.

[0286] In at least one embodiment, one or more SoCs 2104 may include any number of caches 2112, including those described herein. For example, in at least one embodiment, one or more caches 2112 may include a Level 3 (“L3”) cache available for one or more CPUs 2106 and one or more GPUs 2108 (e.g., connected to both CPUs 2106 and GPUs 2108). In at least one embodiment, one or more caches 2112 may include a write-back cache that can, for example, track the state of a line using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, although a smaller cache size may be used, according to an embodiment, the L3 cache may include 4 MB or more.

[0287] In at least one embodiment, one or more SoCs 2104 may include one or more accelerators 2114 (e.g., hardware accelerators, software accelerators, or combinations thereof). In at least one embodiment, one or more SoCs 2104 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4MB of SRAM) enables the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to supplement one or more GPUs 2108 and offload some tasks from one or more GPUs 2108 (e.g., freeing up more loops on one or more GPUs 2108 to perform other tasks). In at least one embodiment, one or more accelerators 2114 may be used to load target workloads that are sufficiently stable to withstand acceleration testing (e.g., perceptual, convolutional neural networks (“CNN”), recurrent neural networks (“RNN”), etc.). In at least one embodiment, the CNN may include region-based or region convolutional neural networks (“RCNN”) and fast RCNN (e.g., for object detection) or other types of CNNs.

[0288] In at least one embodiment, one or more accelerators 2114 (e.g., a hardware acceleration cluster) may include one or more deep learning accelerators (“DLAs”). One or more DLAs may include, but are not limited to, one or more Tensor Processing Units (“TPUs”), which may be configured to provide an additional 10 trillion operations per second for deep learning applications and inference. In at least one embodiment, the TPU may be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). One or more DLAs may be further optimized for specific sets of neural network types and floating-point operations and inference. In at least one embodiment, one or more DLAs are designed to provide higher performance per millimeter than typical general-purpose GPUs and typically significantly outperform CPUs. In at least one embodiment, one or more TPUs may perform several functions, including single-instance convolution functions supporting, for example, INT8, INT16, and FP16 data types for features and weights, as well as post-processor functions. In at least one embodiment, one or more DLAs can execute neural networks, particularly CNNs, quickly and efficiently on processed or unprocessed data for any of the various functions, including, but not limited to: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, recognition, and identification using data from microphone 2196; CNNs for face recognition and vehicle owner recognition using data from camera sensors; and / or CNNs for safety and / or safety-related events.

[0289] In at least one embodiment, the DLA can perform any function of one or more GPUs 2108, and by using an inference accelerator, for example, the designer can target one or more DLAs or one or more GPUs 2108 for any function. For example, in at least one embodiment, the designer can concentrate the CNN processing and floating-point operations on one or more DLAs, leaving other functions to one or more GPUs 2108 and / or one or more other accelerators 2114.

[0290] In at least one embodiment, one or more accelerators 2114 (e.g., a hardware acceleration cluster) may include one or more programmable vision accelerators (“PVAs”), which may alternatively be referred to herein as computer vision accelerators. In at least one embodiment, one or more PVAs may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (“ADAS”) 2138, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. One or more PVAs may strike a balance between performance and flexibility. For example, in at least one embodiment, each of one or more PVAs may include, for example, but not limited to, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.

[0291] In at least one embodiment, the RISC core can be associated with an image sensor (e.g., the image sensor of any camera described herein), an image signal processor, and so on. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, the RISC core can use any of a variety of protocols, depending on the embodiment. In at least one embodiment, the RISC core can execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core can be implemented using one or more integrated circuit devices, application-specific integrated circuits (“ASICs”), and / or storage devices. For example, in at least one embodiment, the RISC core can include an instruction cache and / or tightly coupled RAM.

[0292] In at least one embodiment, DMA enables one or more components of a PVA to access system memory independently of one or more CPUs 2106. In at least one embodiment, DMA can support any number of features for providing optimization to the PVA, including but not limited to, support for multidimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more addressing dimensions, which may include, but are not limited to, block width, block height, block depth, horizontal block step, vertical block step, and / or depth step.

[0293] In at least one embodiment, the vector processor may be a programmable processor designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem may serve as the main processing engine of the PVA and may include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”). In at least one embodiment, the VPU core may include a digital signal processor, such as a Single Instruction Multiple Data (“SIMD”) or Very Long Instruction Word (“VLIW”) digital signal processor. In at least one embodiment, the combination of SIMD and VLIW can improve throughput and speed.

[0294] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each vector processor may be configured to execute independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA may be configured to employ data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm, except on different regions of an image. In at least one embodiment, vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on the same image, or even execute different algorithms on a sequence of images or portions of images. In at least one embodiment, among others, any number of PVAs may be included in the hardware-accelerated cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, one or more PVAs may include additional error-correcting code (“ECC”) memory to enhance overall system security.

[0295] In at least one embodiment, one or more accelerators 2114 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and static random access memory (“SRAM”) for providing high-bandwidth, low-latency SRAM to one or more accelerators 2114. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, consisting of, for example, but not limited to, eight field-configurable memory blocks accessible to both the PVA and DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory via a backbone providing high-speed access to the memory for both the PVA and DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using an APB).

[0296] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as bursty communication for continuous data transmission. In at least one embodiment, although other standards and protocols may be used, the interface may conform to the International Organization for Standardization (“ISO”) 26262 or the International Electrotechnical Commission (“IEC”) 61508 standard.

[0297] In at least one embodiment, one or more SoCs 2104 may include a real-time eye-tracking hardware accelerator. In at least one embodiment, the real-time eye-tracking hardware accelerator may be used to quickly and efficiently determine the location and extent of an object (e.g., within a world model) to generate real-time visualization simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison with LIDAR data for localization and / or other functions, and / or for other purposes.

[0298] In at least one embodiment, one or more accelerators 2114 (e.g., a cluster of hardware accelerators) have broad applicability for autonomous driving. In at least one embodiment, the PVA can be a programmable vision accelerator that can be used in key processing stages in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of the PVA at low power and low latency are well-matched to algorithmic domains requiring predictable processing. In other words, the PVA performs well in semi-intensive or intensive conventional computations, even on small datasets that require predictable runtimes with low latency and low power consumption. In at least one embodiment, the PVA in an autonomous vehicle (e.g., vehicle 2100) is designed to run classical computer vision algorithms because they are efficient in object detection and integer mathematical operations.

[0299] For example, according to at least one embodiment of the technology, PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, although this is not intended to be limiting. In at least one embodiment, applications for Level 3-5 autonomous driving use dynamic estimation / stereo matching during operation (e.g., structure recovery from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, PVA can perform computer stereo vision functions on input from two monocular cameras.

[0300] In at least one embodiment, the PVA can be used to perform intensive optical flow. For example, in at least one embodiment, the PVA can process raw RADAR data (e.g., using 4D Fast Fourier Transform) to provide processed RADAR data. In at least one embodiment, the PVA is used for time-of-flight depth processing, for example, by processing raw time-of-flight data to provide processed time-of-flight data.

[0301] In at least one embodiment, the DLA can be used to run any type of network to enhance control and driving safety, including, but not limited to, neural networks whose output is used for a confidence score for each object detection. In at least one embodiment, the confidence score can be represented or interpreted as a probability, or as providing a relative “weight” for each detection relative to other detections. In at least one embodiment, the confidence score enables the system to make further decisions about which detections should be considered true positives rather than false positives. For example, in at least one embodiment, the system can set a threshold for the confidence score and only consider detections exceeding the threshold as true positives. In embodiments using an Automatic Emergency Braking (“AEB”) system, false positives would cause the vehicle to automatically perform emergency braking, which is obviously undesirable. In at least one embodiment, a highly confident detection can be considered a trigger for AEB. In at least one embodiment, the DLA can run a neural network for regressing the confidence score value. In at least one embodiment, the neural network may take at least a subset of parameters as its input, such as bounding box size, obtained ground plane estimate (e.g., from another subsystem), and outputs of one or more IMU sensors 2166 related to the vehicle orientation, distance, and 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., one or more LiDAR sensors 2164 or one or more RADAR sensors 2160).

[0302] In at least one embodiment, one or more SoCs 2104 may include one or more data storage devices 2116 (e.g., memory). In at least one embodiment, one or more data storage devices 2116 may be on-chip memory of one or more SoCs 2104, which may store neural networks to be executed on one or more GPUs 2108 and / or DLAs. In at least one embodiment, one or more data storage devices 2116 may have a sufficiently large capacity to store multiple instances of the neural network for redundancy and security. In at least one embodiment, one or more data storage devices 2112 may include L2 or L3 caches.

[0303] In at least one embodiment, one or more SoCs 2104 may include any number of one or more processors 2110 (e.g., embedded processors). One or more processors 2110 may include a startup and power management processor, which may be a dedicated processor and subsystem for handling startup power and management functions, as well as associated security implementations. In at least one embodiment, the startup and power management processor may be part of a startup sequence of one or more SoCs 2104 and may provide runtime power management services. In at least one embodiment, the startup power and management processor may provide clock and voltage programming, assist system low-power state transitions, thermal and temperature sensor management of one or more SoCs 2104s, and / or power state management of one or more SoCs 2104s. In at least one embodiment, each temperature sensor may be implemented with its output frequency proportional to temperature, and one or more SoCs 2104s may use the ring oscillator to detect the temperature of one or more CPUs 2106, one or more GPUs 2108, and / or one or more accelerators 2114. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the startup and power management processor may enter a temperature fault routine and place one or more SoCs 2104 into a lower power state and / or place the vehicle 2100 into a driver’s safe stopping pattern (e.g., bring the vehicle 2100 to a safe stop).

[0304] In at least one embodiment, one or more processors 2110 may further include a set of embedded processors that can be used as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem capable of providing full hardware support for multi-channel audio through multiple interfaces and a wide and flexible range of audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.

[0305] In at least one embodiment, one or more processors 2110 may further include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. In at least one embodiment, the processor on the always-on processor engine may include, but is not limited to, a processor core, tightly coupled RAM, support for peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.

[0306] In at least one embodiment, one or more processors 2110 may further include a secure cluster engine, which includes, but is not limited to, a dedicated processor subsystem for handling security management of automotive applications. In at least one embodiment, the secure cluster engine may include, but is not limited to, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.) and / or routing logic. In secure mode, in at least one embodiment, the two or more cores may operate in lockstep mode and may be used as a single core with comparison logic for detecting any differences between their operations. In at least one embodiment, one or more processors 2110 may further include a real-time camera engine, which may include, but is not limited to, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 2110 may further include a high dynamic range signal processor, which may include, but is not limited to, an image signal processor, which is a hardware engine as part of the camera processing pipeline.

[0307] In at least one embodiment, one or more processors 2110 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by the video playback application to generate the final video for the player window. In at least one embodiment, the video image synthesizer may perform lens distortion correction on one or more wide-angle cameras 2170, one or more surround cameras 2174, and / or one or more cabin monitoring camera sensors. In at least one embodiment, preferably, the cabin monitoring camera sensors are monitored by a neural network running on another instance of the SoC 2104, the neural network being configured to recognize cabin events and respond accordingly. In at least one embodiment, the cabin system may perform, but is not limited to, lip reading to activate cellular service and make phone calls, instruct emails, change the vehicle's destination, activate or change the vehicle's infotainment system and settings, or provide voice-activated web browsing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode, and are otherwise disabled.

[0308] In at least one embodiment, the video image synthesizer may include enhanced temporal denoising for simultaneous spatial and temporal denoising. For example, in at least one embodiment, when motion occurs in the video, denoising appropriately weights spatial information, thereby reducing the weight of information provided by adjacent frames. In at least one embodiment, when the image or a portion of the image does not contain motion, temporal denoising performed by the video image synthesizer may use information from previous images to reduce noise in the current image.

[0309] In at least one embodiment, the video image compositor can also be configured to perform stereoscopic correction on the input stereo lens frames. In at least one embodiment, when using an operating system desktop, the video image compositor can also be used for user interface compositing and does not require one or more GPUs 2108 to continuously render new surfaces. In at least one embodiment, when one or more GPUs 2108 are powered and actively performing 3D rendering, the video image compositor can be used to offload one or more GPUs 2108 to improve performance and responsiveness.

[0310] In at least one embodiment, one or more SoCs 2104 may further include a Mobile Industrial Processor Interface (“MIPI”) camera serial interface, a high-speed interface, and / or a video input block that can be used for receiving video and input from a camera and associated pixel input functions. In at least one embodiment, one or more SoCs 2104 may further include an input / output controller that can be software controlled and can be used to receive I / O signals not assigned to a specific role.

[0311] In at least one embodiment, one or more SoCs 2104 may further include extensive peripheral interfaces to enable communication with peripheral devices, audio encoders / decoders (“codecs”), power management and / or other devices. One or more SoCs 2104 may be used to process data from (e.g., via gigabit multimedia serial links and Ethernet connections) cameras, sensors (e.g., one or more LiDAR sensors 2164, one or more RADAR sensors 2160, etc., which may be connected via Ethernet), data from bus 2102 (e.g., vehicle 2100 speed, steering wheel position, etc.), data from one or more GNSS sensors 2158 (e.g., via Ethernet or CAN bus connections), etc. In at least one embodiment, one or more SoCs 2104 may further include a dedicated high-performance mass storage controller, which may include its own DMA engine and may be used to free one or more CPUs 2106 from routine data management tasks.

[0312] In at least one embodiment, one or more SoCs 2104 can be an end-to-end platform with a flexible architecture spanning automation levels 3-5, providing a comprehensive functional safety architecture that leverages and effectively utilizes computer vision and ADAS technologies to achieve diversity and redundancy. This provides a platform offering a flexible and reliable driving software stack as well as deep learning tools. In at least one embodiment, one or more SoCs 2104 can be faster, more reliable, and even more energy and space efficient than conventional systems. For example, in at least one embodiment, one or more accelerators 2114, when combined with one or more CPUs 2106, one or more GPUs 2108, and one or more data storage devices 2116, can provide a fast and efficient platform for Level 3-5 autonomous vehicles.

[0313] In at least one embodiment, the computer vision algorithm can be executed on a CPU, which can be configured using a high-level programming language (e.g., C) to execute multiple processing algorithms on a variety of visual data. However, in at least one embodiment, the CPU typically cannot meet the performance requirements of many computer vision applications, such as performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real time, which are used in automotive ADAS applications and practical Level 3-5 autonomous vehicles.

[0314] The embodiments described herein allow multiple neural networks to be executed simultaneously and / or sequentially, and allow the results to be combined to achieve Level 3-5 autonomous driving capabilities. For example, in at least one embodiment, a CNN executed on a DLA or discrete GPU (e.g., one or more GPUs 2120) may include text and word recognition, thereby allowing a supercomputer to read and understand traffic signs, including signs for which the neural network has not yet been specifically trained. In at least one embodiment, the DLA may also include a neural network capable of recognizing, interpreting, and providing semantic understanding of symbols, and passing this semantic understanding to a path planning module running on a CPU Complex.

[0315] In at least one embodiment, for drives of levels 3, 4, or 5, multiple neural networks can run simultaneously. For example, in at least one embodiment, a warning sign consisting of "Caution: flashing lights indicate icy conditions" along with a light can be interpreted independently or jointly by multiple neural networks. In at least one embodiment, the sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "flashing lights indicate icy conditions" can be interpreted by a second deployed neural network, which informs the vehicle's path planning software (preferably executed on a CPU complex) that icing conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be identified by operating a third deployed neural network across multiple frames, informing the vehicle's path planning software of the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks can run simultaneously, for example within a DLA and / or on one or more GPUs 2108.

[0316] In at least one embodiment, the CNN for facial recognition and vehicle owner identification can use data from camera sensors to identify the presence of an authorized driver and / or the owner of vehicle 2100. In at least one embodiment, a normally open sensor processor engine can be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and, in security mode, can be used to disable the vehicle when the owner leaves the vehicle. In this way, one or more SoCs 2104 provide protection against theft and / or carjacking.

[0317] In at least one embodiment, the CNN for emergency vehicle detection and identification can use data from microphone 2196 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 2104 use the CNN to classify environmental and urban sounds, as well as visual data. In at least one embodiment, the CNN running on DLA is trained to identify the relative approach speed of emergency vehicles (e.g., by using the Doppler effect). In at least one embodiment, the CNN can also be trained to identify emergency vehicles in the area where the vehicle is operating, as identified by one or more GNSS sensors 2158. In at least one embodiment, when operating in Europe, the CNN will seek to detect European sirens, while when operating in the United States, the CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used, with the assistance of one or more ultrasonic sensors 2162, to execute emergency vehicle safety routines, slow down the vehicle, pull the vehicle to the side of the road, stop, and / or leave the vehicle idle until one or more emergency vehicles pass.

[0318] In at least one embodiment, vehicle 2100 may include one or more CPUs 2118 (e.g., one or more discrete CPUs or one or more dCPUs) that may be coupled to one or more SoCs 2104 via high-speed interconnects (e.g., PCIe). In at least one embodiment, one or more CPUs 2118 may include x86 processors. For example, one or more CPUs 2118 may be used to perform any of the various functions, such as arbitrating the results of potential inconsistencies between ADAS sensors and one or more SoCs 2104, and / or monitoring the status and health of one or more monitoring controllers 2136 and / or on-chip information systems (“information SoCs”) 2130.

[0319] In at least one embodiment, vehicle 2100 may include one or more GPUs 2120 (e.g., one or more discrete GPUs or one or more dGPUs) that may be coupled to one or more SoCs 2104 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, one or more GPUs 2120 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update the neural networks based at least in part on inputs from sensors of vehicle 2100 (e.g., sensor data).

[0320] In at least one embodiment, vehicle 2100 may further include a network interface 2124, which may include, but is not limited to, one or more wireless antennas 2126 (e.g., one or more wireless antennas 2126 for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 2124 may be used to enable wireless connectivity with other vehicles and / or computing devices (e.g., passenger client devices) via the Internet and the cloud (e.g., using servers and / or other network devices). In at least one embodiment, for communication with other vehicles, a direct link and / or an indirect link (e.g., via a network and the Internet) may be established between vehicle 2100 and other vehicles. In at least one embodiment, a vehicle-to-vehicle communication link may be used to provide a direct link. The vehicle-to-vehicle communication link may provide vehicle 2100 with information about vehicles near vehicle 2100 (e.g., vehicles in front, to the side, and / or behind vehicle 2100). In at least one embodiment, the foregoing functionality may be part of a cooperative adaptive cruise control function of vehicle 2100.

[0321] In at least one embodiment, network interface 2124 may include a System-on-Chip (SoC) that provides modulation and demodulation functions and enables one or more controllers 2136 to communicate over a wireless network. In at least one embodiment, network interface 2124 may include a radio frequency (RF) front-end for up-conversion from baseband to RF and down-conversion from RF to baseband. In at least one embodiment, frequency conversion may be performed in any technically feasible manner. For example, frequency conversion may be performed using known processes and / or using a superheterodyne process. In at least one embodiment, the RF front-end functionality may be provided by a separate chip. In at least one embodiment, the network interface may include wireless functions for communication over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0322] In at least one embodiment, vehicle 2100 may further include one or more data storage units 2128, which may include, but are not limited to, off-chip (e.g., one or more SoC 2104) storage. In at least one embodiment, one or more data storage units 2128 may include, but are not limited to, one or more storage elements, including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disk and / or other components and / or devices capable of storing at least one bit of data.

[0323] In at least one embodiment, vehicle 2100 may further include one or more GNSS sensors 2158 (e.g., GPS and / or auxiliary GPS sensors) to assist in map creation, perception, occupancy raster generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 2158 may be used, including, for example, but not limited to, GPS sensors connected to a serial interface (e.g., RS-232) bridge using a USB connector with Ethernet.

[0324] In at least one embodiment, vehicle 2100 may further include one or more RADAR sensors 2160. One or more RADAR sensors 2160 can be used by vehicle 2100 for remote vehicle detection, even in dark and / or inclement weather conditions. In at least one embodiment, the RADAR functional safety level may be ASIL B. One or more RADAR sensors 2160 may use CAN and / or bus 2102 (e.g., to transmit data generated by one or more RADAR sensors 2160) for control and access to object tracking data, and in some examples, may access Ethernet to access raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example, but not limited to, one or more of the RADAR sensors 2160 may be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more RADAR sensors 2160 are one or more pulse Doppler RADAR sensors.

[0325] In at least one embodiment, one or more RADAR sensors 2160 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In at least one embodiment, the long-range RADAR can be used for adaptive cruise control functions. In at least one embodiment, the long-range RADAR system can provide a wide field of view achieved through two or more independent scans (e.g., within a 250m range). In at least one embodiment, one or more RADAR sensors 2160 can help distinguish between stationary and moving objects and can be used by the ADAS system 2138 for emergency braking assistance and forward collision warning. One or more sensors 2160 included in the long-range RADAR system may include, but are not limited to, a monostatic multimode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In at least one embodiment, having six antennas, with the four central antennas, can create a focused beammap designed to record the surrounding environment of the vehicle 2100 at a high speed while minimizing traffic interference from adjacent lanes. In at least one embodiment, the other two antennas can expand the field of view, thereby enabling rapid detection of lanes entering or leaving vehicle 2100.

[0326] In at least one embodiment, as an example, a mid-range RADAR system may include, for example, a range of up to 160m (front) or 80m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system may include, but is not limited to, any number of RADAR sensors 2160 designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that continuously monitor the blind spots at and near the rear of the vehicle. In at least one embodiment, the short-range RADAR system may be used in ADAS system 2138 for blind spot detection and / or lane change assistance.

[0327] In at least one embodiment, vehicle 2100 may further include one or more ultrasonic sensors 2162. One or more ultrasonic sensors 2162, which may be positioned at the front, rear, and / or sides of vehicle 2100, can be used for parking assistance and / or creating and updating occupancy detectors. In at least one embodiment, a wide variety of ultrasonic sensors 2162 can be used, and different ultrasonic sensors 2162 can be used for different detection ranges (e.g., 2.5m, 4m). In at least one embodiment, the ultrasonic sensors 2162 can operate at the ASIL B functional safety level.

[0328] In at least one embodiment, vehicle 2100 may include one or more LiDAR sensors 2164. The one or more LiDAR sensors 2164 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the one or more LiDAR sensors 2164 may be of functional safety level ASIL B. In at least one embodiment, vehicle 2100 may include multiple (e.g., two, four, six, etc.) LiDAR sensors 2164 that can use Ethernet (e.g., providing data to a Gigabit Ethernet switch).

[0329] In at least one embodiment, one or more LiDAR sensors 2164 may be able to provide a list of objects and their distances for a 360-degree field of view. In at least one embodiment, one or more commercially available LiDAR sensors 2164 may, for example, have an advertising range of approximately 100m, an accuracy of 2cm-3cm, and, for example, support a 100Mbps Ethernet connection. In at least one embodiment, one or more non-protruding LiDAR sensors 2164 may be used. In such embodiments, one or more LiDAR sensors 2164 may be implemented as small devices that can be embedded in the front, rear, sides, and / or corners of a vehicle 2100. In at least one embodiment, one or more LiDAR sensors 2164, in such embodiments, can provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for objects with low reflectivity, and have a range of 200m. In at least one embodiment, one or more forward-facing LiDAR sensors 2164 may be configured for a horizontal field of view between 45 degrees and 215 degrees.

[0330] In at least one embodiment, LIDAR technology (e.g., 3D flash LIDAR) may also be used. 3D flash LIDAR uses a laser flash as a transmission source to illuminate approximately 200m around the vehicle 2100. In at least one embodiment, the flash LIDAR unit includes, but is not limited to, a receiver that records the laser pulse propagation time and reflected light on each pixel, which in turn corresponds to the range from the vehicle 2100 to the object. In at least one embodiment, flash LIDAR can allow the generation of highly accurate and distortion-free images of the surrounding environment using each laser flash. In at least one embodiment, four flash LIDAR sensors may be deployed, one on each side of the vehicle 2100. In at least one embodiment, the 3D flash LIDAR system includes, but is not limited to, a solid-state 3D line-of-sight array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, the flash LIDAR device can use a 5-nanosecond Class I (eye-safe) laser pulse per frame and can capture reflected laser light in the form of a 3D ranging point cloud and co-registered intensity data.

[0331] In at least one embodiment, the vehicle may further include one or more IMU sensors 2166. In at least one embodiment, one or more IMU sensors 2166 may be located at the center of the rear axle of the vehicle 2100. In at least one embodiment, one or more IMU sensors 2166 may include, for example, but not limited to, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In at least one embodiment, for example in a six-axis application, one or more IMU sensors 2166 may include, but are not limited to, accelerometers and gyroscopes. In at least one embodiment, for example in a nine-axis application, one or more IMU sensors 2166 may include, but are not limited to, accelerometers, gyroscopes, and magnetometers.

[0332] In at least one embodiment, one or more IMU sensors 2166 may be implemented as a miniature, high-performance GPS-assisted inertial navigation system (“GPS / INS”) combining a microelectromechanical system (“MEMS”) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide position, velocity, and attitude estimates; in at least one embodiment, one or more IMU sensors 2166 may enable vehicle 2100 to estimate heading without input from a magnetic sensor obtained by directly observing and correlating velocity changes from GPS to one or more IMU sensors 2166. In at least one embodiment, one or more IMU sensors 2166 and one or more GNSS sensors 2158 may be combined in a single integrated unit.

[0333] In at least one embodiment, vehicle 2100 may include one or more microphones 2196 placed inside and / or around vehicle 2100. In at least one embodiment, in addition, one or more microphones 2196 may be used for emergency vehicle detection and identification.

[0334] In at least one embodiment, vehicle 2100 may further include any number of camera types, including one or more stereo cameras 2168, one or more wide-angle cameras 2170, one or more infrared cameras 2172, one or more surround cameras 2174, one or more long-range cameras 2198, one or more mid-range cameras 2176, and / or other camera types. In at least one embodiment, the cameras can be used to capture image data around the entire perimeter of vehicle 2100. In at least one embodiment, the type of camera used depends on vehicle 2100. In at least one embodiment, any combination of camera types can be used to provide the necessary coverage around vehicle 2100. In at least one embodiment, the number of cameras can vary depending on the embodiment. For example, in at least one embodiment, vehicle 2100 may include six cameras, seven cameras, ten cameras, twelve cameras, or other numbers of cameras. The cameras may be examples, but are not limited to, supporting Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet. In at least one embodiment, references previously made herein... Figure 21A and Figure 21B Each camera can be described in more detail, whether it is one or more cameras.

[0335] In at least one embodiment, vehicle 2100 may further include one or more vibration sensors 2142. One or more vibration sensors 2142 can measure vibrations of components of vehicle 2100 (e.g., axles). For example, in at least one embodiment, changes in vibration can indicate changes in road surface conditions. In at least one embodiment, when two or more vibration sensors 2142 are used, differences between vibrations can be used to determine road surface friction or slippage (e.g., when there is a vibration difference between a power drive axle and a free-rotating axle).

[0336] In at least one embodiment, vehicle 2100 may include ADAS system 2138. ADAS system 2138 may include, but is not limited to, SoC. In at least one embodiment, ADAS system 2138 may include, but is not limited to, any number of autonomous / adaptive / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keeping assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functions, and combinations thereof.

[0337] In at least one embodiment, the ACC system may use one or more RADAR sensors 2160, one or more LIDAR sensors 2164, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to vehicles adjacent to vehicle 2100 and automatically adjusts the speed of vehicle 2100 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance holding and suggests that vehicle 2100 change lanes when necessary. In at least one embodiment, lateral ACC is associated with other ADAS applications, such as LC and CW.

[0338] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles via network interface 2124 and / or one or more wireless antennas 2126 via a wireless link or indirectly via a network connection (e.g., via the Internet). In at least one embodiment, the direct link may be provided by a vehicle-to-vehicle (“V2V”) communication link, while the indirect link may be provided by an infrastructure-to-vehicle (“I2V”) communication link. Typically, the V2V communication concept provides information about the vehicle immediately preceding it (e.g., a vehicle immediately in front of vehicle 2100 and in the same lane as it), while the I2V communication concept provides information about traffic further ahead. In at least one embodiment, the CACC system may include one or both of the I2V and V2V information sources. In at least one embodiment, given information about vehicles preceding vehicle 2100, the CACC system can be more reliable and has the potential to improve traffic flow smoothness and reduce road congestion.

[0339] In at least one embodiment, the FCW system is designed to warn the driver of danger so that the driver can take corrective action. In at least one embodiment, the FCW system uses a forward-facing camera and / or one or more RADAR sensors 2160, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration components. In at least one embodiment, the FCW system can provide warnings, for example, in the form of audible, visual warnings, vibrations, and / or rapid braking pulses.

[0340] In at least one embodiment, the AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use one or more forward-facing cameras and / or one or more RADAR sensors 2160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, it typically first warns the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system may automatically apply brakes to attempt to prevent or at least mitigate the effects of the predicted collision. In at least one embodiment, the AEB system may include techniques such as dynamic braking to support and / or brakes in the event of an impending collision.

[0341] In at least one embodiment, when vehicle 2100 crosses lane markings, the LDW system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver. In at least one embodiment, the LDW system is inactive when the driver indicates intentional lane departure by activating turn signals. In at least one embodiment, the LDW system may use a front-facing camera coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which is electrically coupled to driver feedback such as a display, speaker, and / or vibration components. In at least one embodiment, the LKA system is a variant of the LDW system. If vehicle 2100 begins to leave the lane, the LKA system provides steering input or braking to correct vehicle 2100.

[0342] In at least one embodiment, the BSW system detects and warns the driver of a vehicle in the blind spot. In at least one embodiment, the BSW system can provide visual, auditory, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, the BSW system can provide additional warnings when the driver uses the turn signal. In at least one embodiment, the BSW system can use one or more rear-facing cameras and / or one or more RADAR sensors 2160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, electrically coupled to driver feedback, such as a display, speaker, and / or vibration assembly.

[0343] In at least one embodiment, the RCTW system can provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera while the vehicle 2100 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure the application of the vehicle brakes to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear-facing RADAR sensors 2160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to driver feedback such as a display, speaker, and / or vibration assembly.

[0344] In at least one embodiment, conventional ADAS systems may be prone to generating false alarms, which can be annoying and distracting to the driver, but are generally not catastrophic because conventional ADAS systems warn the driver and allow the driver to determine whether a safe situation truly exists and take appropriate action. In at least one embodiment, in the event of conflicting results, vehicle 2100 itself decides whether to follow the result of the primary computer or the secondary computer (e.g., the first controller 2136 or the second controller 2136). For example, in at least one embodiment, ADAS system 2138 may be a backup and / or auxiliary computer for providing perception information to a backup computer rationality module. In at least one embodiment, the backup computer rationality monitor may run redundant software on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, the output from ADAS system 2138 may be provided to a monitoring MCU. In at least one embodiment, if the outputs from the primary computer and the auxiliary computer conflict, the monitoring MCU decides how to reconcile the conflict to ensure safe operation.

[0345] In at least one embodiment, the master computer may be configured to provide a confidence score to the supervisory MCU to indicate the master computer's confidence in the selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU may follow the master computer's instructions regardless of whether the auxiliary computer provides conflicting or inconsistent results. In at least one embodiment, if the confidence score does not meet the threshold, and if the master computer and the auxiliary computer indicate different results (e.g., conflict), the supervisory MCU may arbitrate between the computers to determine the appropriate result.

[0346] In at least one embodiment, the supervisory MCU may be configured to run a neural network trained and configured to determine, at least in part, the conditions under which the auxiliary computer provides a false alarm based on outputs from both the host computer and the auxiliary computer. In at least one embodiment, the neural network in the supervisory MCU may learn when the outputs of the auxiliary computer can be trusted and when they cannot. For example, in at least one embodiment, when the auxiliary computer is a RADAR-based FCW system, the neural network in the supervisory MCU may learn when the FCW system recognizes a metallic object that is not actually dangerous, such as a drain grating or manhole cover that would trigger an alarm. In at least one embodiment, when the auxiliary computer is a camera-based LDW system, the neural network in the supervisory MCU may learn to override the LDW when a cyclist or pedestrian is present and lane departure is actually the safest operation. In at least one embodiment, the supervisory MCU may include at least one of a DLA or GPU suitable for running a neural network with associated memory. In at least one embodiment, the supervisory MCU may include and / or be included as a component of one or more SoC 2104s.

[0347] In at least one embodiment, the ADAS system 2138 may include an auxiliary computer that performs ADAS functions using conventional computer vision rules. In at least one embodiment, the auxiliary computer may use classic computer vision rules (if-then), and the presence of a neural network in the supervisory MCU can improve reliability, security, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identity make the entire system more fault-tolerant, especially for failures caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if a software vulnerability or bug exists in the software running on the host computer, and different software code running on the auxiliary computer provides the same overall result, the supervisory MCU can more confidently assume that the overall result is correct and that the vulnerability in the software or hardware on the host computer will not lead to a significant error.

[0348] In at least one embodiment, the output of the ADAS system 2138 may be input to the perception module and / or the dynamic driving task module of the host computer. For example, in at least one embodiment, if the ADAS system 2138 indicates a forward collision warning due to an object directly ahead, the perception block may use this information when identifying the object. In at least one embodiment, as described herein, the assistance computer may have its own neural network trained to reduce the risk of false alarms.

[0349] In at least one embodiment, vehicle 2100 may further include an infotainment SoC 2130 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, in at least one embodiment, the infotainment system 2130 may not be an SoC and may include, but is not limited to, two or more discrete components. In at least one embodiment, the infotainment SoC 2130 may include, but is not limited to, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephone (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.) and / or information services (e.g., navigation system, rear parking assist, radio data system, vehicle-related information such as fuel level, total coverage distance, brake fuel level, fuel level, door opening / closing, air filter information, etc.) to vehicle 2100. For example, the infotainment SoC 2130 may include a radio, disk player, navigation system, video player, USB and Bluetooth connectivity, automobile, in-vehicle entertainment system, WiFi, steering wheel audio controls, hands-free voice control, head-up display (“HUD”), HMI display 2134, telematics device, control panel (e.g., for controlling and / or interacting with various components, features and / or systems) and / or other components. In at least one embodiment, the infotainment SoC 2130 may be further used to provide information (e.g., visual and / or auditory) to users of the vehicle, such as information from ADAS system 2138, autonomous driving information (e.g., planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.) and / or other information.

[0350] In at least one embodiment, the infotainment SoC 2130 may include any number and type of GPU functionality. In at least one embodiment, the infotainment SoC 2130 may communicate with other devices, systems, and / or components of the vehicle 2100 via bus 2102 (e.g., CAN bus, Ethernet, etc.). In at least one embodiment, the infotainment SoC 2130 may be coupled to a monitoring MCU, enabling the GPU of the infotainment system to perform some autonomous driving functions in the event of a failure of the main controller 2136 (e.g., the main computer and / or backup computer of the vehicle 2100). In at least one embodiment, the infotainment SoC 2130 may cause the vehicle 2100 to enter a driver-to-safe-stop mode, as described herein.

[0351] In at least one embodiment, vehicle 2100 may further include instrument panel 2132 (e.g., digital instrument panel, electronic instrument panel, digital instrument control panel, etc.). Instrument panel 2132 may include, but is not limited to, controllers and / or supercomputers (e.g., discrete controllers or supercomputers). In at least one embodiment, instrument panel 2132 may include, but is not limited to, any number and combination of a set of instruments, such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information may be displayed and / or shared between infotainment SoC 2130 and instrument panel 2132. In at least one embodiment, instrument panel 2132 may be included as part of infotainment SoC 2130, or vice versa.

[0352] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding the inference and / or training logic 1815 are provided. In at least one embodiment, the inference and / or training logic 1815 can be implemented in the system. Figure 21C The operation is used to infer or predict the operation based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0353] In at least one embodiment, the system Figure 21C This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 21C This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 21C Used to implement one or more neural networks including discriminators and generators, and the system Figure 21C and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0354] Figure 21D It is based on at least one embodiment in a cloud-based server and Figure 21A A diagram of a system 2176 for communication between autonomous vehicles 2100. In at least one embodiment, system 2176 may include, but is not limited to, one or more servers 2178, one or more networks 2190, and any number and type of vehicles, including vehicle 2100. One or more servers 2178 may include, but are not limited to, multiple GPUs 2184(A)-2184(H) (collectively referred to herein as GPU 2184), PCIe switches 2182(A)-2182(H) (collectively referred to herein as PCIe switch 2182), and / or CPUs 2180(A)-2180(B) (collectively referred to herein as CPU 2180). GPU 2184, CPU 2180, and PCIe switch 2182 may be interconnected with high-speed interconnects, such as, but not limited to, NVLink interface 2188 developed by NVIDIA and / or PCIe connection 2186. The GPU 2184 is connected via NVLink and / or NVSwitchSoC, and the GPU 2184 and PCIe switch 2182 are connected via PCIe interconnect. In at least one embodiment, although eight GPUs 2184, two CPUs 2180, and four PCIe switches 2182 are shown, this is not intended to be limiting. In at least one embodiment, each of one or more servers 2178 may include, but is not limited to, any combination of any number of GPUs 2184, CPUs 2180, and / or PCIe switches 2182. For example, in at least one embodiment, one or more servers 2178 may each include eight, sixteen, thirty-two, and / or more GPUs 2184.

[0355] In at least one embodiment, one or more servers 2178 may receive image data representing images from vehicles via one or more networks 2190, the images showing unexpected or changed road conditions, such as recently commenced roadworks. In at least one embodiment, one or more servers 2178 may transmit neural network 2192, updated neural network 2192, and / or map information 2194, including but not limited to information about traffic and road conditions, to vehicles via one or more networks 2190. In at least one embodiment, updates to map information 2194 may include, but are not limited to, updates to HD map 2122, such as information about construction sites, potholes, sidewalks, floods, and / or other obstacles. In at least one embodiment, neural network 2192, updated neural network 2192, and / or map information 2194 may be generated from new training and / or experience represented by data received from any number of vehicles in the environment, and / or at least based on training performed in a data center (e.g., using one or more servers 2178 and / or other servers).

[0356] In at least one embodiment, one or more servers 2178 may be used to train a machine learning model (e.g., a neural network) at least in part based on training data. The training data may be generated by the vehicle and / or may be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., where the associated neural network benefits from supervised learning) and / or undergoes other preprocessing. In at least one embodiment, no amount of training data is labeled and / or preprocessed (e.g., where the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by the vehicle (e.g., transmitted to the vehicle via one or more networks 2190), and / or the machine learning model may be used by one or more servers 2178 to remotely monitor the vehicle.

[0357] In at least one embodiment, one or more servers 2178 may receive data from the vehicle and apply the data to state-of-the-art real-time neural networks for real-time intelligent inference. In at least one embodiment, one or more servers 2178 may include a deep learning supercomputer and / or a dedicated AI computer powered by one or more GPUs 2184, such as the DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 2178 may include a deep learning infrastructure in a data center using CPU power.

[0358] In at least one embodiment, the deep learning infrastructure of one or more servers 2178 may be capable of fast, real-time inference and may use this capability to assess and verify the health of the processor, software, and / or associated hardware in vehicle 2100. For example, in at least one embodiment, the deep learning infrastructure may receive cyclical updates from vehicle 2100, such as image sequences and / or objects located by vehicle 2100 in the image sequence (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them with objects identified by vehicle 2100, and if the results do not match and the deep learning infrastructure determines that the AI ​​in vehicle 2100 is malfunctioning, one or more servers 2178 may signal to vehicle 2100 to instruct the fail-safe computer of vehicle 2100 to take control, notify passengers, and complete a safe stopping operation.

[0359] In at least one embodiment, one or more servers 2178 may include one or more GPUs 2184 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3). In at least one embodiment, the combination of GPU-driven servers and inference acceleration enables real-time response. In at least one embodiment, for example, where performance is less critical, servers driven by CPUs, FPGAs, and other processors may be used for inference. In at least one embodiment, hardware architecture 1815 is used to execute one or more embodiments. This document incorporates... Figure 18A and / or Figure 18B Provides details about the hardware architecture 1815.

[0360] Computer System

[0361] Figure 22 This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor 2200, which may include an execution unit to execute instructions. In at least one embodiment, according to this disclosure, such as the embodiments described herein, the computer system 2200 may include, but is not limited to, components such as processor 2202, whose execution unit includes logic to execute algorithms for process data. In at least one embodiment, the computer system 2200 may include a processor, such as those available from Intel Corporation of Santa Clara, California. Processor family, Xeon™ XScale™ and / or StrongARM™ Core TM or Nervana TM A microprocessor may be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors may also be used. In at least one embodiment, computer system 2200 may execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces may also be used.

[0362] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip (SoC), a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system that can execute one or more instructions according to at least one embodiment.

[0363] In at least one embodiment, computer system 2200 may include, but is not limited to, processor 2202, which may include, but is not limited to, one or more execution units 2208, to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, system 2200 is a single-processor desktop or server system, but in another embodiment, system 2200 may be a multiprocessor system. In at least one embodiment, processor 2202 may include, but is not limited to, a Complex Instruction Set Computer (“CISC”) microprocessor, a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 2202 may be coupled to processor bus 2210, which can transmit data signals between processor 2202 and other components in computer system 2200.

[0364] In at least one embodiment, processor 2202 may include, but is not limited to, a Level 1 (“L1”) internal cache memory (“cache”) 2204. In at least one embodiment, processor 2202 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache memory may reside external to processor 2202. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 2206 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.

[0365] In at least one embodiment, an execution unit 2208, including but not limited to logic for performing integer and floating-point operations, is also located within processor 2202. Processor 2202 may also include a microcode (“ucode”) read-only memory (“ROM”) for storing microcode of certain macro instructions. In at least one embodiment, execution unit 2208 may include logic for processing a packaged instruction set 2209. In at least one embodiment, by including the packaged instruction set 2209 in the instruction set of general-purpose processor 2202, along with associated circuitry for executing the instructions, packaged data in general-purpose processor 2202 can be used to perform operations used by numerous multimedia applications. In one or more embodiments, many multimedia applications can be executed more quickly and efficiently by using the full width of the processor’s data bus to perform operations on the packaged data, which may eliminate the need to transfer smaller data units on the processor’s data bus to perform one or more operations on one data element at a time.

[0366] In at least one embodiment, execution unit 2208 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, computer system 2200 may include, but is not limited to, memory 2220. In at least one embodiment, memory 2220 may be implemented as a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other memory device. Memory 2220 may store instructions 2219 and / or data 2221 represented by data signals that can be executed by processor 2202.

[0367] In at least one embodiment, the system logic chip may be coupled to processor bus 2210 and memory 2220. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 2216, and processor 2202 may communicate with MCH 2216 via processor bus 2210. In at least one embodiment, MCH 2216 may provide a high-bandwidth memory path 2218 to memory 2220 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, MCH 2216 may initiate data signals between processor 2202, memory 2220, and other components in computer system 2200, and bridge data signals between processor bus 2210, memory 2220, and system I / O 2222. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 2216 may be coupled to memory 2220 via high-bandwidth memory path 2218, and graphics / video card 2212 may be coupled to MCH 2216 via Accelerated Graphics Port (“AGP”) interconnect 2214.

[0368] In at least one embodiment, computer system 2200 may use system I / O 2222, a proprietary hub interface bus, to couple MCH 2216 to I / O controller hub (“ICH”) 2230. In at least one embodiment, ICH 2230 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to memory 2220, chipset, and processor 2202. Examples may include, but are not limited to, an audio controller 2229, a firmware hub (“Flash BIOS”) 2228, a wireless transceiver 2226, data storage 2224, a conventional I / O controller 2223 including a user input and keyboard interface, a serial expansion port 2227 (e.g., Universal Serial Bus (USB), and a network controller 2234). Data storage 2224 may include a hard disk drive, floppy disk drive, CD-ROM device, flash memory device, or other mass storage device.

[0369] In at least one embodiment, Figure 22 A system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 22An exemplary system-on-chip (“SoC”) may be illustrated. In at least one embodiment, the device shown in Figure cc may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of system 2200 are interconnected using compute fast link (CXL) interconnects.

[0370] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details are provided regarding the inference and / or training logic 1815. In at least one embodiment, the inference and / or training logic 1815 can... Figure 22 Used in systems for reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0371] In at least one embodiment, the system Figure 22 This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 22 This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 22 Used to implement one or more neural networks including discriminators and generators, and the system Figure 22 and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0372] Figure 23 This is a block diagram illustrating an electronic device 2300 for utilizing a processor 2310 according to at least one embodiment. In at least one embodiment, the electronic device 2300 may be, for example, but not limited to, a laptop computer, tower server, rack server, blade server, desktop computer, tablet computer, mobile device, telephone, embedded computer, or any other suitable electronic device.

[0373] In at least one embodiment, system 2300 may include, but is not limited to, processor 2310 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 2310 uses a bus or interface coupling, such as an I2C bus, system management bus (“SMBus”), low pin count (LPC) bus, serial peripheral interface (“SPI”), high-definition audio (“HDA”) bus, serial advanced technology accessory (“SATA”) bus, universal serial bus (“USB”) (versions 1, 2, 3), or universal asynchronous receiver / transmitter (“UART”) bus. In at least one embodiment, Figure 23 The system shown includes interconnected hardware devices or "chips," while in other embodiments, Figure 23 An exemplary system-on-chip (“SoC”) may be illustrated. In at least one embodiment, Figure 23 The device shown can be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, Figure 23 One or more components are interconnected using Computational Fast Link (CXL) interconnects.

[0374] In at least one embodiment, Figure 23 It may include a display 2324, a touch screen 2325, a touchpad 2330, a near field communication unit (“NFC”) 2345, a sensor hub 2340, a thermal sensor 2346, a fast chipset (“EC”) 2335, a trusted platform module (“TPM”) 2338, a BIOS / firmware / flash (“BIOS, FW Flash”) 2322, a DSP 2360, a drive (“SSD or HDD”) 2320 (e.g., a solid-state drive (“SSD”) or a hard disk drive (“HDD”) 2350), a wireless local area network unit (“WLAN”) 2350, a Bluetooth unit 2352, a wireless wide area network unit (“WWAN”) 2356, a global positioning system (GPS) 2355, a camera (“USB 3.0 camera”) 2354 (e.g., a USB 3.0 camera), or a low-power double data rate (“LPDDR”) memory unit (“LPDDR3”) 2315 implemented in, for example, the LPDDR3 standard. These components can each be implemented in any suitable way.

[0375] In at least one embodiment, other components may be communicatively coupled to processor 2310 via the components described above. In at least one embodiment, accelerometer 2341, ambient light sensor (“ALS”) 2342, compass 2343, and gyroscope 2344 may be communicatively coupled to sensor hub 2340. In at least one embodiment, thermal sensor 2339, fan 2337, keyboard 2346, and touchpad 2330 may be communicatively coupled to EC 2335. In at least one embodiment, speaker 2363, earphone 2364, and microphone (“mic”) 2365 may be communicatively coupled to audio unit (“audio codec and Class D amplifier”) 2364, which in turn may be communicatively coupled to DSP 2360. In at least one embodiment, audio unit 2364 may include, for example, but not limited to, audio encoder / decoder (“codec”) and Class D amplifier. In at least one embodiment, SIM card (“SIM”) 2357 may be communicatively coupled to WWAN unit 2356. In at least one embodiment, the components (e.g., WLAN unit 2350, Bluetooth unit 2352, and WWAN unit 2356) may be implemented as next-generation form factor (NGFF).

[0376] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding the inference and / or training logic 1815 are provided. In at least one embodiment, the inference and / or training logic 1815 can be implemented in the system. Figure 23 It is used in the context of reasoning or predicting operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures or neural network use cases described herein.

[0377] In at least one embodiment, the system Figure 23 This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 23 This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 23 Used to implement one or more neural networks including discriminators and generators, and the system Figure 23 and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0378] Figure 24A computer system 2400 according to at least one embodiment is shown. In at least one embodiment, the computer system 2400 is configured to implement various processes and methods described throughout this disclosure.

[0379] In at least one embodiment, the computer system 2400 includes, but is not limited to, at least one central processing unit (“CPU”) 2402 connected to a communication bus 2410 implemented using any suitable protocol, such as PCI (“Peripheral Interconnect”), Peripheral Component Interconnect Express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, the computer system 2400 includes, but is not limited to, main memory 2404 and control logic (e.g., implemented in hardware, software, or a combination thereof), and data may be stored in main memory 2404 in the form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“network interface”) 2422 provides an interface to other computing devices and networks for receiving data from the computer system 2400 and transferring data to other systems.

[0380] In at least one embodiment, the computer system 2400 includes, but is not limited to, an input device 2408, a parallel processing system 2412, and a display device 2406, which may be implemented using conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light-emitting diode (“LED”), plasma display, or other suitable display technologies. In at least one embodiment, user input is received from the input device 2408 (e.g., a keyboard, mouse, touchpad, microphone, etc.). In at least one embodiment, each of the foregoing modules may reside on a single semiconductor platform to form the processing system.

[0381] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding the inference and / or training logic 1815 are provided. In at least one embodiment, the inference and / or training logic 1815 can be implemented in the system. Figure 24 It is used to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architecture or neural network use cases described herein.

[0382] In at least one embodiment, the system Figure 24 This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 24This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 24 Used to implement one or more neural networks including discriminators and generators, and the system Figure 24 and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0383] Figure 25 A computer system 2500 according to at least one embodiment is illustrated. In at least one embodiment, the computer system 2500 includes, but is not limited to, a computer 2510 and a USB flash drive 2520. In at least one embodiment, the computer 2510 may include, but is not limited to, any number and type of processors (not shown) and memory (not shown). In at least one embodiment, the computer 2510 includes, but is not limited to, a server, a cloud instance, a laptop computer, and a desktop computer.

[0384] In at least one embodiment, the USB flash drive 2520 includes, but is not limited to, a processing unit 2530, a USB interface 2540, and USB interface logic 2550. In at least one embodiment, the processing unit 2530 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing core 2530 can include, but is not limited to, any number and type of processing cores (not shown). In at least one embodiment, the processing core 2530 includes an application-specific integrated circuit (“ASIC”) optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, the processing core 2530 is a tensor processing unit (“TPC”) optimized to perform machine learning inference operations. In at least one embodiment, the processing core 2530 is a vision processing unit (“VPU”) optimized to perform machine vision and machine learning inference operations.

[0385] In at least one embodiment, the USB interface 2540 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, the USB interface 2540 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, the USB interface 2540 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 2550 may include any number and type of logic enabling the processing unit 2530 to connect to an OR device (e.g., computer 2510) via the USB connector 2540.

[0386] The inference and / or training logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding the inference and / or training logic 1815 are provided. In at least one embodiment, the inference and / or training logic 1815 can be implemented in the system. Figure 25 In use, at least in part, the operation is based on weight parameters, neural network functions and / or architectures computed using neural network training operations, or neural network use cases described herein to infer or predict operations.

[0387] In at least one embodiment, the system Figure 25 This is used to implement a discriminator, which is trained to infer a set of viewpoint and appearance attributes from an input image. In at least one embodiment, the system... Figure 25 This is used to implement a generator, which is trained to generate images based on an input viewpoint and an input set of appearance parameters. In at least one embodiment, the system... Figure 25 Used to implement one or more neural networks including discriminators and generators, and the system Figure 25 and one One or more processes are used in combination to identify the orientation of objects within an image by training one or more neural networks in a self-supervised manner, at least by computing one or more loss functions as part of training to evaluate one or more features of images in a training set.

[0388] Figure 26A An exemplary architecture is illustrated, in which multiple GPUs 2610-2613 are communicatively coupled to multiple multi-core processors 2605-2606 via high-speed links 2640-2643 (e.g., bus / point-to-point interconnect, etc.). In one embodiment, the high-speed links 2640-2643 support communication throughput of 4GB / s, 30GB / s, 80GB / s, or higher. Various interconnect protocols can be used, including but not limited to PCIe 4.0 or 5.0 and NVLink 2.0.

[0389] Furthermore, in one embodiment, two or more GPUs 2610-2613 are interconnected via high-speed links 2629-2630, which may use the same or different protocols / links as those used for high-speed links 2640-2643. Similarly, two or more multi-core processors 2605-2606 may be connected via high-speed link 2628, which may be a symmetric multiprocessor (SMP) bus operating at speeds of 20GB / s, 30GB / s, 120GB / s, or higher. Alternatively, the same protocol / link may be used (e.g., via a common interconnect structure). Figure 26A This shows all communication between the various system components.

[0390] In one embodiment, each multi-core processor 2605-2606 is communicatively coupled to processor memories 2601-2602 via memory interconnects 2626-2627, and each GPU 2610-2613 is communicatively coupled to GPU memories 2620-2623 via GPU memory interconnects 2650-2653. Memory interconnects 2626-2627 and 2650-2653 may utilize the same or different memory access technologies. By way of example and not limitation, processor memories 2601-2602 and GPU memories 2620-2623 may be volatile memories, such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or may be non-volatile memories, such as 3D XPoint or Nano-RAM. In one embodiment, some portions of the processor memories 2601-2602 may be volatile memory, while other portions may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).

[0391] As described in this article, although the various processors 2605-2606 and GPUs 2610-2613 can be physically coupled to specific memories 2601-2602 and 2620-2623 respectively, a unified memory architecture can be implemented, in which the same virtual system address space (also known as the “effective address” space) is distributed among the various physical memories. For example, processor memories 2601-2602 can each contain 64GB of system memory address space, and GPU memories 2620-2623 can each contain 32GB of system memory address space (resulting in a total of 256GB of addressable memory in this example).

[0392] Figure 26B Additional details are shown regarding the interconnection between a multi-core processor 2607 and a graphics acceleration module 2646, according to an exemplary embodiment. The graphics acceleration module 2646 may include one or more GPU chips integrated on a line card coupled to the processor 2607 via a high-speed link 2640. Alternatively, the graphics acceleration module 2646 may be integrated on the same package or chip as the processor 2607.

[0393] In at least one embodiment, the processor 2607 shown includes multiple cores 2660A-2660D, each core having a translation back buffer 2661A-2661D and one or more caches 2662A-2662D. In at least one embodiment, cores 2660A-2660D may include various other components (not shown) for executing instructions and processing data. Caches 2662A-2662D may include level 1 (L1) and level 2 (L2) caches. Furthermore, one or more shared caches 2656 may be included in caches 2662A-2662D and shared by the respective groups of cores 2660A-2660D. For example, one embodiment of the processor 2607 includes 24 cores, each core having its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches. In this embodiment, two adjacent cores share one or more L2 and L3 caches. Processor 2607 and graphics acceleration module 2646 are connected to system memory 2614, which may include Figure 26A The processor memory 2601-2602 in the memory.

[0394] The consistency bus 2664 maintains consistency for data and instructions stored in the various caches 2662A-2662D, 2656 and system memory 2614 via inter-core communication. For example, each cache may have associated cache consistency logic / circuitology to communicate via the consistency bus 2664 in response to the detection of a read or write to a specific cache line. In one implementation, a cache snooping protocol is implemented via the consistency bus 2664 to snoop on cache accesses.

[0395] In one embodiment, proxy circuitry 2625 communicatively couples graphics acceleration module 2646 to coherence bus 2664, thereby allowing graphics acceleration module 2646 to participate in cache coherence protocols as a peer of cores 2660A-2660D. Specifically, interface 2635 provides connectivity to proxy circuitry 2625 via high-speed link 2640 (e.g., PCIe bus, NVLink, etc.), and interface 2637 connects graphics acceleration module 2646 to link 2640.

[0396] In one implementation, the accelerator integrated circuit 2636 represents multiple graphics processing engines 2631, 2632, N of the graphics acceleration module, providing cache management, memory access, context management, and interrupt management services. The graphics processing engines 2631, 2632, N may each include a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 2631, 2632, N may include different types of graphics processing engines within the GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module 2646 may be a GPU having multiple graphics processing engines 2631-2632, N, or the graphics processing engines 2631-2632, N may be individual GPUs integrated on a general-purpose package, line card, or chip.

[0397] In one embodiment, the accelerator integrated circuit 2636 includes a memory management unit (MMU) 2639 for performing various memory management functions, such as virtual-to-physical memory translation (also known as effective-to-real memory translation), and a memory access protocol for accessing system memory 2614. The MMU 2639 may also include a translation back buffer (“TLB”) (not shown) for caching virtual / effective-to-physical / real address translations. In one implementation, cache 2638 stores commands and data for efficient access by graphics processing engines 2631-2632,N. In one embodiment, data stored in cache 2638 and graphics memories 2633-2634,M is kept consistent with core caches 2662A-2662D, 2656 and system memory 2614. As previously mentioned, this task can be accomplished via proxy circuitry 2625 representing cache 2638 and graphics memory 2633-2634,M (e.g., sending updates related to the modification / access of cache lines on processor caches 2662A-2662D, 2656 to cache 2638 and receiving updates from cache 2638).

[0398] A set of registers 2645 stores the context data of the threads executed by graphics processing engines 2631-2632,N, and context management circuitry 2648 manages the thread context. For example, context management circuitry 2648 can perform save and restore operations to save and restore the context of individual threads during context switching (e.g., saving the first thread and storing the second thread so that the second thread can be executed by the graphics processing engine). For example, during context switching, context management circuitry 2648 can store the current register value into a designated area in memory (e.g., identified by a context pointer). The register value can then be restored when returning to the context. In one embodiment, interrupt management circuitry 2647 receives and processes interrupts received from system devices.

[0399] In one implementation, MMU 2639 translates virtual / effective addresses from graphics processing engine 2631 into real / physical addresses in system memory 2614. One embodiment of accelerator integrated circuit 2636 supports multiple (e.g., 4, 8, 16) graphics accelerator modules 2646 and / or other accelerator devices. Graphics accelerator module 2646 may be dedicated to a single application executing on processor 2607, or may be shared among multiple applications. In one embodiment, a virtualized graphics execution environment is presented, where resources of graphics processing engines 2631-2632,N are shared with multiple applications or virtual machines (VMs). In at least one embodiment, resources may be subdivided into “slices” based on processing requirements and priorities associated with VMs and / or applications, which are then allocated to different VMs and / or applications.

[0400] In at least one embodiment, the accelerator integrated circuit 2636 acts as a bridge to the system of the graphics acceleration module 2646, providing address translation and system memory caching services. Additionally, the accelerator integrated circuit 2636 can provide virtualization facilities for the host processor to manage the virtualization, interrupt, and memory management of the graphics processing engines 2631-2632,N.

[0401] Because the hardware resources of the graphics processing engines 2631-2632,N are explicitly mapped to the actual address space seen by the host processor 2607, any host processor can directly address these resources using valid address values. In one embodiment, one function of the accelerator integrated circuit 2636 is to physically separate the graphics processing engines 2631-2632,N, making them appear as independent units to the system.

[0402] In at least one embodiment, one or more graphics memories 2633-2634,M are coupled to each graphics processing engine 2631-2632,N. The graphics memories 2633-2634,M store instructions and data processed by each graphics processing engine 2631-2632,N. The graphics memories 2633-2634,M can be volatile memory, such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory, such as 3D XPoint or Nano-RAM.

[0403] In one embodiment, to reduce data traffic on link 2640, a biasing technique is used to ensure that the data stored in graphics memories 2633-2634,M is the data most frequently used by graphics processing engines 2631-2632,N, and preferably not used (or at least infrequently used) by cores 2660A-2660D. Similarly, the biasing mechanism attempts to keep data needed by the cores (and preferably not by graphics processing engines 2631-2632,N) in the core caches 2662A-2662D, 2656 and system memory 2614.

[0404] Figure 26C Another exemplary embodiment is shown, in which the accelerator integrated circuit 2636 is integrated within the processor 2607. In this embodiment, the graphics processing engines 2631-2632,N communicate directly with the accelerator integrated circuit 2636 via a high-speed link 2640 through interfaces 2637 and 2635 (which may also utilize any form of bus or interface protocol). The accelerator integrated circuit 2636 can perform operations related to... Figure 26B The operations described are the same. However, due to its close proximity to the coherence bus 2664 and caches 2662A-2662D, 2656, it may have higher throughput. One embodiment supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by accelerator integrated circuit 2636 and a programming model controlled by graphics acceleration module 2646.

[0405] In at least one embodiment, graphics processing engines 2631-2632,N are dedicated to a single application or process within a single operating system. In at least one embodiment, a single application can funnel requests from other applications to graphics processing engines 2631-2632,N, thereby providing virtualization within a VM / partition.

[0406] In at least one embodiment, graphics processing engines 2631-2632,N can be shared by multiple VM / application partitions. In at least one embodiment, the shared model can use a hypervisor to virtualize graphics processing engines 2631-2632,N to allow each operating system to access them. For a single-partition system without a hypervisor, the operating system owns graphics processing engines 2631-2632,N. In at least one embodiment, the operating system can virtualize graphics processing engines 2631-2632,N to provide access to each process or application.

[0407] In at least one embodiment, the graphics acceleration module 2646 or the individual graphics processing engines 2631-2632,N uses a process handle to select a process element. In one embodiment, the process element is stored in system memory 2614 and can be addressed using the effective address to physical address translation techniques described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering its context with the graphics processing engines 2631-2632,N (i.e., invoking system software to add the process element to the process element linked list). In at least one embodiment, the lower 16 bits of the process handle may be an offset of the process element in the process element linked list.

[0408] Figure 26D An exemplary accelerator integration slice 2690 is shown. As used herein, a “slice” includes a designated portion of the processing resources of the accelerator integrated circuit 2636. The application’s effective address space 2682 in system memory 2614 stores process element 2683. In one embodiment, process element 2683 is stored in response to a GPU call 2681 from an application 2680 executing on processor 2607. Process element 2683 contains the proc...

Claims

1. A processor comprising: one or more circuits to facilitate training of one or more neural networks to identify a direction of an object within an image based at least in part on one or more labels indicative of one or more features of the object other than the direction of the object, wherein the one or more circuits are to train the one or more neural networks at least by: obtaining an input image; determining, using a discriminator, at least a predicted viewpoint and a predicted appearance parameter set; creating, using a generator, a synthetic image based at least in part on the predicted viewpoint and the predicted appearance parameter set; and computing a viewpoint consistency loss based at least in part on the input image and the synthetic image.

2. The processor of claim 1, wherein the one or more circuits are to facilitate training of the one or more neural networks on a set of images of a same class as the image.

3. The processor of claim 2, wherein ground truth annotations are not available for at least a portion of the set of images.

4. The processor of claim 1, wherein the one or more features of the object include a symmetry consistency between the image of the object and an inverted image of the object.

5. The processor of claim 1, wherein one or more circuits are to facilitate training of the one or more neural networks to generate a second image of the object having a second direction.

6. The processor of claim 1, wherein the direction of the object is encoded on a parameter set comprising an azimuth parameter, an elevation parameter, and a tilt parameter.

7. A system comprising: one or more processors to compute parameters to facilitate training of one or more neural networks to identify a direction of an object within an image based at least in part on one or more labels indicative of one or more features of the object other than the direction of the object; and one or more memories to store the parameters, wherein the one or more processors are to train the one or more neural networks at least by: obtaining an input image; determining, using a discriminator, at least a predicted viewpoint and a predicted appearance parameter set; creating, using a generator, a synthetic image based at least in part on the predicted viewpoint and the predicted appearance parameter set; and computing a viewpoint consistency loss based at least in part on the input image and the synthetic image.

8. The system of claim 7, wherein the one or more processors to compute the parameters to facilitate training of the one or more neural networks are to facilitate training of the one or more neural networks on a set of images of different objects of a same class as the object.

9. The system of claim 7, wherein the input image is a real image.

10. The system of claim 8, wherein the one or more processors are to train the one or more neural networks at least by: obtaining a first viewpoint and a first appearance parameter set; creating, using a generator, a synthetic image based at least in part on the first viewpoint and the first appearance parameter set; predicting, using a discriminator, a second view and a second set of appearance parameters based on the synthetic image; computing, based at least in part on the first view and the second view, a view consistency loss; and computing, based at least in part on the first set of appearance parameters and the second set of appearance parameters, a reconstruction loss.

11. The system of claim 8, wherein the one or more processors are to train the one or more neural networks at least by: creating, using a generator, a first synthetic image based at least in part on a first view and a set of appearance parameters; performing a transformation on the first view to obtain a second view; creating, using the generator, a second synthetic image based at least in part on the second view and the set of appearance parameters; and computing, based at least in part on the first synthetic image and the second synthetic image, a symmetry loss.

12. The system of claim 11, wherein the transformation horizontally flips the first view to obtain the second view.

13. A method comprising: training one or more neural networks to identify a direction of an object within an image based at least in part on one or more labels indicative of one or more features of the object other than the direction of the object, wherein the one or more neural networks are trained at least by: obtaining an input image; determining, using a discriminator, a predicted view and a predicted set of appearance parameters; creating, using a generator, a synthetic image based at least in part on the predicted view and the predicted set of appearance parameters; and computing, based at least in part on the input image and the synthetic image, a view consistency loss.

14. The method of claim 13, wherein training the one or more neural networks comprises: training the one or more neural networks in a self-supervised manner on a set of images of different objects of a same class as the object within the image.

15. The method of claim 14, wherein training the one or more neural networks in the self-supervised manner comprises: evaluating the one or more features of the object other than the direction of the object using a set of loss functions.

16. The method of claim 14, wherein the subject belongs to a first category, and the method further comprises: training the one or more neural networks to identify a second direction of a second object using a second set of images, wherein: the second object belongs to a second class different from the first class; and the second set of images belongs to objects of a second class different from the second object.

17. The method of claim 14, wherein training the one or more neural networks in the self-supervised manner comprises: training the one or more neural networks at least to: predict, using the discriminator, a view and a set of parameters from the input image; create, using the generator, a synthetic image based at least in part on the view and the set of parameters; and compute one or more gradients based at least in part on the synthetic image and update parameters of the discriminator.

18. The method of claim 17, wherein the generator is a deep generative model.

19. The method of claim 18, wherein the deep generative model is a renderer, a variational autoencoder, or a generative adversarial network (GAN).

20. The method of claim 13, wherein the object is a vehicle. one or more circuits to train one or more neural networks to identify one or more directions of an object within an image based at least in part on one or more labels indicative of one or more features of the object other than the one or more directions of the object, 21. A processor comprising: ​ wherein the one or more neural networks comprise: a generator to create a synthetic image based at least in part on a specified viewpoint and a specified appearance parameter set; and a discriminator to determine, from one or more images, a predicted viewpoint and a predicted appearance parameter set.

22. The processor of claim 21, wherein the one or more neural networks are trained on a set of images of different objects of a same category as the object.

23. The processor of claim 21, wherein ground truth annotations are not available for the set of images.

24. The processor of claim 21, wherein the one or more features of the object include symmetry consistency between the image of the object and an inverted image of the object.

25. The processor of claim 21, wherein the one or more orientations of the object are encoded on a parameter set comprising an azimuth parameter, an elevation parameter, and a tilt parameter.

26. A system comprising: one or more memories; and one or more processors to train one or more neural networks to identify one or more orientations of an object within an image based at least in part on one or more labels indicative of one or more features of the object other than the one or more orientations of the object, wherein the one or more neural networks comprise: a generator to create a synthetic image based at least in part on a specified viewpoint and a specified appearance parameter set; and a discriminator to determine, from one or more images, a predicted viewpoint and a predicted appearance parameter set.

27. The system of claim 26, wherein the one or more processors to train the one or more neural networks facilitate training the one or more neural networks on a set of images having different objects, wherein the different objects belong to a same category as the object.

28. The system of claim 26, wherein the one or more processors to train the one or more neural networks are to train the one or more neural networks at least by: computing a first set of gradients to update a first set of parameters of the generator; and computing a second set of gradients to update a second set of parameters of the discriminator.

29. The system of claim 26, wherein the one or more processors to train the one or more neural networks are to train the one or more neural networks at least by computing a decoupling loss, the decoupling loss computed at least by: generating a first synthetic image using a first viewpoint and a first appearance parameter set; generating a second synthetic image using the first viewpoint and a second appearance parameter set; and generating a third synthetic image using a second viewpoint and the first appearance parameter set.

30. The system of claim 26, wherein the one or more orientations are relative to a canonical orientation.

31. The system of claim 26, wherein the one or more orientations each comprise an azimuth parameter, an elevation parameter, and a tilt parameter. ​ 32. A method comprising: training one or more neural networks to identify one or more directions of an object within an image based at least in part on one or more labels indicative of one or more features of the object other than the one or more directions of the object, wherein the one or more neural networks comprise: a generator to create a synthetic image based at least in part on a specified viewpoint and a specified set of appearance parameters; and a discriminator to determine, from one or more images, a predicted viewpoint and a predicted set of appearance parameters.

33. The method of claim 32, wherein the one or more neural networks are trained in a self-supervised manner on a set of images that share the same labels as the image, the labels being indicative of a feature of one or more features of the object other than the one or more directions of the object.

34. The method of claim 33, wherein the one or more neural networks are trained in a self-supervised manner to identify directions of the set of images based on labels of the set of images other than the directions.

35. The method of claim 32, wherein the generator is a deep generative model.

36. The method of claim 33, wherein the object is a human.

37. The method of claim 33, wherein the one or more directions of the object are encoded on a parameter set comprising an azimuth parameter, an elevation parameter, and a tilt parameter.

38. An automobile comprising: one or more cameras to capture images of one or more objects and one or more neural networks to identify one or more directions of the one or more objects based at least in part on one or more labels indicative of one or more features of the one or more objects other than the one or more directions of the one or more objects, wherein the one or more neural networks are trained by: obtaining an input image; determining, using a discriminator, at least a predicted viewpoint and a predicted set of appearance parameters; creating, using a generator, a synthetic image based at least in part on the predicted viewpoint and the predicted set of appearance parameters; and computing a viewpoint consistency loss based at least in part on the input image and the synthetic image.

39. The automobile of claim 38, wherein the one or more neural networks are trained in a self-supervised manner on a set of images that share the same labels as the image, the labels being indicative of a feature of one or more features of the one or more objects other than the one or more directions of the one or more objects.

40. The automobile of claim 38, wherein the one or more features of the one or more objects comprise a symmetry consistency between an image of the one or more objects and an inverted image of the one or more objects.

41. The automobile of claim 38, wherein the one or more neural networks are trained to generate a second image having the one or more directions of the one or more objects.

42. The automobile of claim 38, wherein the one or more directions of the one or more objects are three-dimensional directions.

43. The automobile of claim 38, wherein the one or more objects are people.

44. The automobile of claim 38, wherein the one or more objects are vehicles other than the automobile.

Citation Information

Patent Citations

  • Face direction estimation apparatus and program thereof

    JP2018022416A

  • Target recognizing device, target recognizing method, and program

    JP2019152543A

  • Calibration target detection apparatus, calibration target detecting method for detecting calibration target, and program for calibration target detection apparatus

    US20120033087A1

  • Subcategory-aware convolutional neural networks for object detection

    US20170124415A1

  • Method and system for performing segmentation of image having a sparsely distributed object

    US20190080456A1