Optimization of Parameter Estimation for Training Neural Networks
By constructing a proxy dataset and using similarity scores to select unique images, the method efficiently estimates neural network parameters, addressing the computational challenges of training on high-dimensional medical images and reducing time and resources needed.
Patent Information
- Application Number
- JP2022082067
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-05-28
- Filing Date
- 2022-05-19
- Publication Date
- 2025-07-17
- Estimated Expiration
- 2042-05-19
AI Technical Summary
Training neural networks for medical image segmentation is computationally costly due to the high dimensionality of medical images, requiring significant memory and time, and existing methods like AutoML do not efficiently determine optimal hyperparameters.
Construct a proxy dataset that is a smaller subset of the training data, using similarity scores to select unique images, and train a proxy neural network model to estimate parameters efficiently, reducing computational burden while maintaining performance.
The proposed method significantly reduces the computational time for hyperparameter optimization from days to hours while achieving performance comparable to training on the full dataset, making it suitable for high-dimensional medical imaging tasks.
Smart Images

Figure 0007709938000009 
Figure 0007709938000010 
Figure 0007709938000011
Abstract
Description
Technical Field
[0001] At least one embodiment relates to processing resources to train one or more neural networks based on the uniqueness of data used to train one or more neural networks. For example, at least one embodiment relates to a processor or computing system used to select a subset of training data for estimating parameters used to train one or more neural networks according to various novel techniques described herein, based on the uniqueness of the training data.
Summary of the Invention
[0002] In various contexts, it can be difficult to determine training data utilized to estimate parameters for training a neural network model. Often, the training data is high-dimensional and thus includes medical images that require significant computational cost to estimate parameters. When the training data includes large-sized images, the memory capacity, time, or computational resources used to estimate parameters for training a neural network become important. There is room for improvement in the memory capacity, time, or computational resources used to estimate parameters for training a neural network from the training data.
Brief Description of the Drawings
[0003]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11A
Figure 11B
Figure 11C
Figure 11D
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16A
Figure 16B
Figure 16C
Figure 16D
Figure 16E
Figure 16F
Figure 17
Figure 18A
Figure 18B
Figure 19A
Figure 19B
Figure 20
Figure 21A
Figure 21B
Figure 21C
Figure 21D
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32A
Figure 32B
Figure 33
Figure 34
Figure 35
Figure 36
Figure 37
Figure 38
Figure 39
Figure 40A
Figure 40B
Figure 41A
Figure 41B
DETAILED DESCRIPTION
[0004] In at least one embodiment, a deep learning model for medical image segmentation is mainly data-driven. In at least one embodiment, a model trained with more data leads to improved performance and generalization ability. In at least one embodiment, a neural network is trained to learn various patterns corresponding to various medical events (e.g., diseases, injuries, etc.). In at least one embodiment, medical images are supplied to the neural network for training. In at least one embodiment, parameters (e.g., learning rate, size of the neural network, topology of the neural network) are estimated and used when training one or more neural networks. In at least one embodiment, it is possible to determine the parameters that are most suitable for use in training the neural network.
[0005] Neural networks are often trained to infer information from medical images. In many cases, medical images are high-dimensional and are obtained with a large amount of information (e.g., labels) to assist training. Data-driven methods have become the main approach for medical image segmentation-based tasks in most imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI). However, the performance of deep learning methods depends greatly on parameters (e.g., learning rate, optimizer, augmentation probability, etc.), which are referred to herein as hyperparameters or values, that are utilized when training the model. Hyperparameter optimization (HPO) can be performed by using automated machine learning (AutoML) or another system, but using AutoML and other systems is generally a very computationally costly process because the computational cost becomes severe for medical imaging tasks where the data is high-dimensional, i.e., often a three-dimensional (3D) volume.
[0006] In at least one embodiment, the parameters used to train a neural network can be efficiently determined by first constructing a proxy dataset that is relatively much smaller in size, even though it represents the entire training dataset. The proxy dataset may also be referred to herein as proxy data, a portion of the training data, or a subset of the training data. In at least one embodiment, the parameters used to train a neural network can be estimated using one or more proxy models that represent a larger / complete network structure, despite the proxy dataset and a substantially smaller network. In at least one embodiment, both the proxy dataset and the proxy network can be utilized to significantly reduce the computational burden of AutoML (e.g., from days to hours), and yet it may be possible to estimate parameters that can lead to performance improvements. In at least one embodiment, by using the proxy dataset and the proxy model, the parameters can be calculated more efficiently and accurately. In at least one embodiment, the techniques described herein construct the proxy dataset via conventional metrics of mutual information and normalized cross-correlation. In at least one embodiment, the proxy network (e.g., proxy neural network) can be systematically reconstructed by reducing the residual blocks, channels, and / or number of levels of a convolutional neural network (e.g., U-net). In at least one embodiment, the techniques described herein can generalize across multiple different imaging modalities.
[0007] In at least one embodiment, the processor comprises one or more circuits for performing parameter estimation to train a neural network based on the uniqueness of the training data (e.g., the similarity between the training data). In at least one embodiment, the processor performs parameter estimation by executing instructions stored in a computer-readable storage medium, and the processor may be part of a system as described below. In at least one embodiment, the uniqueness of the data is determined based at least in part on the similarity between various data within the training data. In at least one embodiment, the similarity is first determined by identifying regions of interest within training images from the training data. In at least one embodiment, the region of interest is compared to corresponding regions of interest in other training images from the training data. In at least one embodiment, a score indicating uniqueness (e.g., an importance score, a similarity score, an indicator) is generated based on the comparison. In at least one embodiment, the score is added to a list of scores (e.g., an index, a table). In at least one embodiment, the list comprises scores arranged in rank order. In at least one embodiment, the lower the similarity of the region of interest to the regions of interest of other images, the less redundant the less similar image is to other images, and thus the region of interest will have a higher score. In at least one embodiment, all training images from the training data are analyzed and scores are assigned based on how high or low the similarity of the regions of interest within the training image are to the regions of interest in other training images. In at least one embodiment, a processor comprising one or more circuits performs parameter estimation to train a neural network by using a subset of the training images based on the scores within the list. In at least one embodiment, the subset of training images is selected based on the scores associated with each training image.In at least one embodiment, a processor comprising one or more circuits generates a proxy model (e.g., a machine learning model different from the neural network to be trained) and estimates (e.g., calculates, determines) parameters for training the neural network using a selected subset of training images. In at least one embodiment, by scoring each image, the data used to estimate the parameters is more representative of the entire data set, and as a result, the parameters can be calculated more efficiently and accurately when training the neural network.
[0008] FIG. 1 shows an exemplary framework 100 in which one or more neural networks 116 are trained based on the uniqueness of training data 102 according to at least one embodiment. In at least one embodiment, framework 100 includes one or more graphics processing units (GPUs) configured to train model 116 through training data 102 obtained through network 104. In at least one embodiment, GPU 106 selects a portion (e.g., proxy data) 110 of training data 102 for estimating parameters 114 that can be used to train model 116 based on the uniqueness of training data 102.
[0009] In at least one embodiment, GPU 106 includes one or more graphics processing systems. In at least one embodiment, GPU 106 is a parallel processing unit (PPU). In at least one embodiment, GPU 106 is one or more processors comprising one or more circuits for implementing and training various models such as models 112, 116 and / or neural networks. In at least one embodiment, models 112, 116 are classification models that identify one or more categories to which one or more features of the input data belong. In at least one embodiment, models 112, 116 are multi-label classification models. In at least one embodiment, models 112, 116 include one or more neural models. In at least one embodiment, models 112, 116 include one or more convolutional neural network (CNN) architectures. In at least one embodiment, models 112, 116 are U-nets. In at least one embodiment, models 112, 116 include one or more neural networks that classify one or more aspects of the input data based on the input data. In at least one embodiment, models 112, 116 include one or more neural networks that classify one or more features of medical imaging data. In at least one embodiment, models 112, 116 include various neural networks, image feature detection, and embedding generation components.
[0010] In at least one embodiment, the training data 102 is obtained by the GPU 106 through the network 104. In at least one embodiment, the network 104 represents any suitable communication path between the GPU 106 and one or more other systems. In at least one embodiment, the network 104 includes one or more networks, such as the Internet, a local area network, a wide area network, and / or variations thereof. In at least one embodiment, different types of networks 104 are described in connection with FIG. 20 below. In at least one embodiment, the training data 102 includes image data and associated text data. In at least one embodiment, the image data includes two-dimensional (2D), three-dimensional (3D), and / or four-dimensional (4D) medical images. In at least one embodiment, the training data 102 includes medical images obtained from an external service or storage device not shown in FIG. 1. In at least one embodiment, the training data 102 includes data captured by one or more cameras, medical imaging equipment (e.g., X-ray), scanners, or image recording devices. In at least one embodiment, the training data 102 includes video, audio, and / or speech data.
[0011] In at least one embodiment, GPU 106 processes training data 102 (including training images) to estimate parameters (e.g., hyperparameters, values) used to control how a neural network is trained. In at least one embodiment, the GPU processes training data 102 to estimate parameters before the neural network 116 to be trained uses the training data for training. In at least one embodiment, the parameters are variables that determine the network structure (e.g., the number of hidden layers) and / or how the network is trained (e.g., the learning rate). In at least one embodiment, the parameters include values of activation functions, dropout normalization, and the number of neurons. In at least one embodiment, the parameters are set and adjustable before training the neural network (before optimizing weights and biases). In at least one embodiment, GPU 106 determines the region of interest in each training image and calculates a score 108. In at least one embodiment, the region of interest in the training image is determined using a bounding box around a particular object that is to be recognized by the neural network 116 being trained. In at least one embodiment, FIG. 1 shows that each training image includes a bounding box around a product similar to a shipping box (although other products, or in the case of medical images, organs may also be used in addition to the shipping box), and the region of interest is compared to all other corresponding regions of interest in all other training images. In at least one embodiment, the box in the first image is compared to other boxes in the corresponding image. In at least one embodiment, GPU 106 calculates a score (e.g., a similarity score, importance score, indicator) such as a value of 1 out of 10, and based on the comparison, assigns the score to the region of interest in the first image. Assigning scores on a scale of 0 to 10 is only one exemplary embodiment, and the scoring method may vary (e.g., assign scores on a scale of 1 to 100) and is not limited to scores based on a scale of 0 to 10. In at least one embodiment, the scores are stored and added to a list.In at least one embodiment, the list is stored in a storage device, buffer, cache, or any storage medium configured to store data. In at least one embodiment, a score of 1 out of 10 indicates that the object in the region of interest is very similar to other shipping box objects in the corresponding image. In at least one embodiment, the score and the training image are then stored with a lower ranking than other scores in the list. In at least one embodiment, if the similarity of a region of interest to other regions of interest in the corresponding image is lower, the less similar image is not redundant with respect to other images, and thus a higher importance score will be assigned. Storing scores in a list based on a comparison between regions of interest and ranking the scores higher based on the dissimilarity between the corresponding regions of interest is only one exemplary embodiment, and other ways of storing and ranking scores (e.g., ranking scores with higher similarity higher instead of lower) can also be utilized. In at least one embodiment, the list is configured to store scores in an order sorted either in ascending or descending order.
[0012] In at least one embodiment, the GPU 106 then uses the list to identify a subset of the training data (e.g., proxy data) 110. In at least one embodiment, the subset of training images 110 is selected based on the importance scores from the list. In at least one embodiment, the GPU 106 uses the proxy data 110 to estimate what the parameters for training should be 114. In at least one embodiment, the proxy data 110 is about 15% of the total training data 102 (although other percentages are applicable). In at least one embodiment, the GPU 106 uses a neural network different from the neural network 116 to be trained (e.g., a proxy network) 112 to estimate the parameters 114 based on the proxy data 110. In at least one embodiment, the proxy network 112 is a U-net having a reduced number of levels, channels, and / or residual blocks compared to the neural network 116 to be trained. In at least one embodiment, the GPU 106 uses the proxy network 112 to estimate the parameters 114 based on the proxy data 110. In at least one embodiment, the GPU 106 then uses the estimated parameters 114 and the training data 102 to train the model 116 to yield a fully trained machine learning model (e.g., a neural network model).
[0013] FIG. 2 shows an exemplary framework 200 for selecting a subset of training data for training a neural network based on the uniqueness of training data, according to at least one embodiment. The top view of framework 200 shows ways (e.g., using all data or a random subset of data) that can be used to perform parameter estimation. In at least one embodiment, a processor comprising one or more circuits executes the bottom row of framework 200 to train one or more neural networks based on the uniqueness of the data. In at least one embodiment, a processor comprising one or more circuits executes the bottom row of framework 200 to estimate parameters for training one or more neural networks based on the uniqueness of the data (e.g., the similarity between input training data). In at least one embodiment, a GPU acquires data (e.g., training data including 3D medical images) used to train one or more neural networks. In at least one embodiment, proxy data (e.g., a subset of data determined based on the uniqueness of all data) is selected and trained using a proxy network (e.g., a smaller representation of the one or more neural networks to be trained). In at least one embodiment, a GPU estimates (e.g., calculates) parameters using the proxy network and proxy data. In at least one embodiment, the GPU then trains the one or more neural networks using all the data and the estimated parameters.
[0014] In at least one embodiment, a proxy data selection strategy is executed. In at least one embodiment, to select proxy data, a dataset D is determined that includes a set of data points {x1, x2, … x n}, where x i is a single data sample. Data points may also be referred to herein as regions of interest within the training data. In at least one embodiment, a single data point x iTo estimate the importance of each data point, the utility of each data point is estimated in relation to other data points x j and as a result, a set of paired measures is brought about. In at least one embodiment, the pair of x1 is {(x1,x1),(x1,x2),(x1,x3)..(x1,x n )}. In at least one embodiment, the average of the above measures is utilized as an indicator of the importance of each data point (for example, an indicator representing the similarity between each data point among other data points). In at least one embodiment, the mutual information (MI) shown in Equation 1 below is measured with respect to the flattened vector of the 3D image and the normalized local cross-correlation (NCC) within the (9,9,9) local window size of each data pair (x i ,x j ). [Number]
[0015] In at least one embodiment, P(x i ) and P(x j ) are marginal probability distributions, while P(x i ,x j ) is a joint probability distribution. [Number]
[0016] In at least one embodiment, p i is the 3D voxel position within the window around p, [Number] and [Number] correspondingly are the voxel positions p i in x j , p iis the local average within the window surrounding it. In at least one embodiment, Ω is a voxel coordinate space.
[0017] In at least one embodiment, the comparison of training images to determine the similarity between the above images focuses on task-specific regions of interest. In at least one embodiment, the acquisition parameters (e.g., number of slices, resolution, etc.) of different 3D volume scans can vary. In at least one embodiment, when considering the pair (x i , x j ), even if x i is resampled to the x j image size, there are inconsistencies in the region of interest (ROI) (e.g., the organ to be annotated by the model). In at least one embodiment, task-specific ROIs are utilized by analyzing information from existing labels. In at least one embodiment, the selected volume is trimmed using the ROI and resampled to a cube patch size. In at least one embodiment, the data points are ranked by importance, and the data points with the lowest mutual information or the lowest correlation are selected within a given budget B (e.g., the amount of available computational resources).
[0018] In at least one embodiment, it is a version of U-net that is smaller than the neural network to be trained. In at least one embodiment, the U-net is used for medical image segmentation tasks. In at least one embodiment, a 5-level U-net with two residual blocks per level and skip connections between the encoder block and the decoder block is used as the neural network to be trained. In at least one embodiment, the rough hyperparameters of the U-net are the number of channels in the first encoder block (subsequent encoder blocks are multiples of 2), the number of residual blocks, and the number of levels. In at least one embodiment, to create a proxy network, the number of channels can be reduced to 4 and the number of residual blocks is reduced from 2 to 1. In at least one embodiment, a variant of the proxy network is created by reducing the number of levels to 5, 4, and 3 (for example, reducing the number of encoding and decoding blocks). In at least one embodiment, a recurrent neural network (RNN) which is part of overarching reinforcement learning (RL) for estimating hyperparameters is used.
[0019] Figure 3 shows an exemplary visualization of results 300 when using a framework for training a neural network based on the uniqueness of training data according to at least one embodiment. In at least one embodiment, a subset of the training data based on the uniqueness of the training data can be selected to efficiently estimate hyperparameters. In at least one embodiment, the exemplary illustration 300 shows splits by different seeds. In at least one embodiment, image “A)” is the “upper bound” Dice of the Beyond the Cranial Vault (BTCV) external dataset and is reported when trained on all Medical Segmentation Decathlon (MSD) data. In at least one embodiment, this is compared to the Dice on BTCV when trained on a proxy dataset across the hyperparameter space. In at least one embodiment, image “B)” shows that random data is used instead of proxy data. In at least one embodiment, 24% of the total data was used for A) and B) respectively. In at least one embodiment, image “C)” shows the “upper bound” Dice on an external dataset provided by PROSTATEx that is reported when trained on all MSD data. In at least one embodiment, this is compared to the Dice on PROSTATEx when trained on a proxy dataset across the hyperparameter space. In at least one embodiment, image “D)” shows that random data is used instead of proxy data. In at least one embodiment, 31% of all data was used for C) and D) respectively. In at least one embodiment, example 300 shows that 24% of the data by the proxy data selection method can achieve up to 90% of the performance of the entire dataset.
[0020] Figure 4 shows an exemplary visualization 400 of validation scores when using a framework for training a neural network based on the uniqueness of training data according to at least one embodiment. In at least one embodiment, the images in the upper row (e.g., using the spleen task dataset from MSD) show the Dice score correlation from a proxy network of the Dice score from the proxy network for the use of all data regarding the internal validation of the ground truth. In at least one embodiment, the images show, from left to right, that the number of levels of the U-net is reduced from 5 to 3 (the corresponding levels from left to right are 5, 4, and 3). In at least one embodiment, all proxy U-nets have 4 channels and 1 residual block per level. In at least one embodiment, the images in the lower row (e.g., using the prostate task dataset from MSD) show the varying number of U-net levels, and correlations are compared regarding the internal validation of the prostate data, similar to the spleen.
[0021] Figure 5 shows an exemplary visualization 500 of results regarding using a framework for training a neural network based on the uniqueness of training data according to at least one embodiment. In at least one embodiment, image "A)" shows the results of the spleen BTCV dataset, with the Dice score plotted against the data usage. In at least one embodiment, image "B)" shows the results of the PROSTATEx dataset and the Dice score plotted against the data usage. In at least one embodiment, image "C)" shows the results of the spleen dataset, showing the estimated hyperparameters of the probability of intensity shift by the learning rate and the relative distance to the ground truth. In at least one embodiment, image "D)" is similar to "C" but shows the results regarding the above prostate dataset example.
[0022] FIG. 6 shows an example of a process 600 for training a neural network based on the uniqueness of training data according to at least one embodiment. In at least one embodiment, some or all of process 600 (or any other process described herein or its variations and / or combinations) is executed under the control of one or more computer systems configured with computer-executable instructions and executed collectively by hardware, software, or a combination thereof (e.g., computer-executable instructions, one or more computer programs, or one or more applications) on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions usable to execute process 600 are not stored using only a transient signal (e.g., a propagating transient electrical or electromagnetic transmission). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include a non-transitory data storage circuit (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 600 is executed at least in part on a computer system such as those described elsewhere in the present disclosure. In at least one embodiment, process 600 is executed by one or more circuits for identifying one or more images used to train one or more neural networks based at least in part on one or more labels of one or more objects in one or more images.
[0023] In at least one embodiment, a system that executes at least a portion of process 600 includes executable code for obtaining 602 training data for training one or more neural networks. In at least one embodiment, the training data includes images, medical images, audio data, and the like. In at least one embodiment, the images are obtained from one or more cameras, imaging devices, or scanners. In at least one embodiment, the image data includes images in formats such as.tif,.jpg,.png.,.gif, and the like. In at least one embodiment, the audio data is digital audio data that includes a binary representation of an audio signal. In at least one embodiment, the audio data includes raw audio files, waveform audio files, MPEG-1 Audio Layer III (MP3) files, and / or Real Audio Metadata (RAM) files. In at least one embodiment, the training data is obtained from a server or a storage device.
[0024] In at least one embodiment, a system that executes at least a portion of process 600 includes executable code for selecting 604 a portion of training data for training one or more neural networks based on the uniqueness of the training data. In at least one embodiment, a portion of the training data is selected and used for estimating hyperparameters that can be used to train one or more neural networks. In at least one embodiment, as mentioned and described above, a hyperparameter is a parameter whose value is used to control the learning process. In at least one embodiment, hyperparameters indicate the topology of the neural network, the size of the neural network, the learning rate, and the size of the input batch. In at least one embodiment, the uniqueness of the data refers to the similarity between the training data. In at least one embodiment, the similarity between the training data is measured by identifying an object from a training image and assigning a similarity score between the object in the first image and the objects in the remaining images. In at least one embodiment, as described above, the similarity score uses a mutual information formula (shown in Equation 1 above) that calculates a score regarding how similar one object in an image is to the object in another image. In at least one embodiment, the similarity score is stored in a list. In at least one embodiment, a portion of the training data is selected based on the list. In at least one embodiment, the list includes all similarity scores for each corresponding object between the training data, and each score is stored based on a sorted and / or ranked order. In at least one embodiment, a system that executes at least a portion of process 600 includes executable code for estimating 604 hyperparameters for training one or more neural networks using a portion of the training data.
[0025] FIG. 7 shows an example of a process 700 for estimating parameters used to train a neural network based on the uniqueness of training data according to at least one embodiment. In at least one embodiment, some or all of process 700 (or any other process or its variations and / or combinations described herein) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) collectively executed by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions usable to execute process 700 are not stored using only a transient signal (e.g., a propagating transient electrical or electromagnetic transmission). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include a non-transitory data storage circuit (e.g., buffers, caches, and queues) within a transient signal transceiver. In at least one embodiment, process 700 is executed at least in part on a computer system such as those described elsewhere in the present disclosure. In at least one embodiment, process 700 is implemented by one or more circuits for identifying one or more images used to train one or more neural networks based at least in part on one or more labels of one or more objects in one or more images.
[0026] In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for identifying an object 702 from one or more training images used to train one or more neural networks. In at least one embodiment, hyperparameters for training a neural network can be estimated by using a subset of training images that more accurately reflects the entire dataset of training images. In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for first discovering a region of interest within each training image, such as a bounding box around a particular object to be recognized by the trained neural network. In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for comparing 704, for each image, the region of interest to corresponding regions of interest of all other regions of interest. In at least one embodiment, an object within a region of interest is compared to an object within a corresponding region of interest among the one or more training images.
[0027] In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for generating 706 a score (e.g., a similarity score) based on the comparison. In at least one embodiment, if the similarity of a region of interest to regions of interest in other images is lower, the less similar image may be assigned a higher similarity score because it is not redundant with respect to the other images. In at least one embodiment, this is performed for all images in the training data set, and each image is scored based on how similar the region of interest of each image is to other regions of interest of the remaining images. In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for storing 708 the similarity scores in a list, the list including a plurality of similarity scores. In at least one embodiment, the list includes scores associated with each image and is arranged in a ranked order based on the similarity scores.
[0028] In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for identifying 710 and selecting a subset of training data (e.g., proxy data) based on the similarity scores from the above list. In at least one embodiment, the subset is used to calculate hyperparameters for training a neural network. In at least one embodiment, a system that executes at least a portion of process 700 includes executable code for estimating 712 one or more hyperparameters that can be used to train the one or more neural networks based on the identified subset of training images using an alternative neural network (e.g., a proxy network representing a network smaller than the one or more neural networks to be trained). In at least one embodiment, the proxy network is a convolutional neural network such as a U-net that has a reduction in the number of channels and a reduction in residual blocks compared to the neural network to be trained. In at least one embodiment, the proxy network is created by reducing the number of levels to 5, 4, and 3 (e.g., reducing the number of encoding and decoding blocks). In at least one embodiment, by giving an importance score to each image, the data used to calculate the hyperparameters is more representative of the entire data set, and as a result, the hyperparameters can be calculated more efficiently and accurately.
[0029] In at least one embodiment, using both proxy data and a proxy network can be a powerful tool for accelerating the HPO estimation process. In at least one embodiment, a maximum training acceleration of 4.4 times can be obtained, which can reduce training over several days to a few hours. In at least one embodiment, the techniques described herein can be extended for the estimation of multiple hyperparameters and for multiple frameworks such as neural architecture search and federated learning where resources are critical for both.
[0030] Inference and Training Logic FIG. 8A shows inference and / or training logic 815 for performing inference and / or training operations with respect to one or more embodiments. Details regarding the inference and / or training logic 815 are provided below in conjunction with FIGS. 8A and / or 8B.
[0031] In at least one embodiment, the inference and / or training logic 815 may include, without limitation, code and / or data storage 801 for storing forward propagation and / or output weights, and / or input / output data, and / or other parameters for constructing neurons or layers of a neural network that are trained and / or used to infer in aspects of one or more embodiments. In at least one embodiment, the training logic 815 may include, or be coupled to, code and / or data storage 801 for storing graph code or other software for controlling timing and / or order, and code and / or data storage 801 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, the code and / or data storage 801 stores the weight parameters and / or input / output data of each layer of the neural network that is trained or used in conjunction with one or more embodiments during forward propagation of the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 801 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0032] In at least one embodiment, any portion of the code and / or data storage 801 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or the code and / or data storage 801 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or the code and / or data storage 801 is, for example, internal or external to a processor, or the selection of whether to include DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in the neural network inference and / or training, or any combination of these factors.
[0033] In at least one example, the inference and / or training logic 815 may include, without limitation, code and / or data storage 805 for storing backward propagation and / or output weights corresponding to neurons or layers of a neural network that are trained and / or used for inference in one or more example aspects, and / or input / output data. In at least one example, the code and / or data storage 805 stores the weight parameters and / or input / output data for each layer of a neural network that is trained or used in conjunction with one or more examples during backward propagation of input / output data and / or weight parameters during training and / or inference using the aspects of one or more examples. In at least one example, the training logic 815 may include or be coupled to code and / or data storage 805 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 805 has weights and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)).
[0034] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into the processor ALU based on the architecture of the neural network to which such code corresponds. In at least one embodiment, any portion of the code and / or data storage 805 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. In at least one embodiment, any portion of the code and / or data storage 805 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or data storage 805 is internal or external to, for example, the processor, or the selection of including DRAM, SRAM, flash memory, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.
[0035] In at least one embodiment, code and / or data storage 801 and code and / or data storage 805 may be separate storage structures. In at least one embodiment, code and / or data storage 801 and code and / or data storage 805 may be combined storage structures. In at least one embodiment, code and / or data storage 801 and code and / or data storage 805 may be partially combined and partially separate. In at least one embodiment, any portion of code and / or data storage 801 and code and / or data storage 805 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.
[0036] In at least one embodiment, the inference and / or training logic 815 includes, without limitation, one or more arithmetic logic units (ALUs) 810 including integer and / or floating point units to perform logical and / or arithmetic operations based at least in part on and / or indicated by training and / or inference code (e.g., graph code), the result of which may generate activations (e.g., output values from a layer or neuron within a neural network) stored in the activation storage 820, which are a function of the code and / or data storage 801 and / or the input / output and / or weight parameter data stored in the code and / or data storage 805. In at least one embodiment, the activations stored in the activation storage 820 are generated according to linear algebra calculations and / or matrix-based calculations performed by the ALU 810 in response to executing instructions or other code, where the weight values stored in the code and / or data storage 805 and / or data storage 801 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in the code and / or data storage 805, or the code and / or data storage 801, or another storage on-chip or off-chip.
[0037] In at least one embodiment, the ALU 810 is included within one or more processors, or other hardware logic devices or circuits, but in another embodiment, the ALU 810 may be external to the processor or other hardware logic device or circuit using them (e.g., a coprocessor). In at least one embodiment, the ALU 810 may be included within the execution unit of a processor, or may be distributed among execution units of processors that are either within the same processor or different processors of a different type (e.g., a central processing unit, a graphics processing unit, a fixed function unit, etc.). In at least one embodiment, the ALU 810 may be otherwise included within an ALU bank accessible by an execution unit of a processor. In at least one embodiment, the code and / or data storage 801, the code and / or data storage 805, and the activation storage 820 may share a processor or other hardware logic device or circuit, but in another embodiment, they may be in different processors or other hardware logic devices or circuits, or in any combination of the same processor or other hardware logic device or circuit and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of the activation storage 820 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory. Further, the inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuit, and may be fetched and / or processed using the fetch, decode, scheduling, execution, retirement, and / or other logic circuits of the processor.
[0038] In at least one embodiment, the activation storage 820 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 820 may be wholly or partially internal or external to one or more processors or other logic circuits. In at least one embodiment, the choice of whether the activation storage 820 is internal or external to, for example, a processor, or the choice of including DRAM, SRAM, flash memory, or some other type of storage may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being performed, the batch size of data used in neural network inference and / or training, or any combination of these factors.
[0039] In at least one embodiment, the inference and / or training logic 815 shown in FIG. 8A may be used in conjunction with application specific integrated circuits (ASICs) such as Google's TensorFlow® processing unit, Graphcore's inference processing unit (IPU), or Intel Corp's Nervana® (e.g., "Lake Crest") processor. In at least one embodiment, the inference and / or training logic 815 shown in FIG. 8A may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate array (FPGA).
[0040] FIG. 8B is a diagram showing inference and / or training logic 815 according to at least one embodiment. In at least one embodiment, the inference and / or training logic 815 may include, without limiting the hardware logic, in which computing resources are dedicated to one or more layers of neurons in a neural network for weight values or other information, or are used only in conjunction with them in other ways. In at least one embodiment, the inference and / or training logic 815 shown in FIG. 8B may be used in conjunction with application-specific integrated circuits (ASICs) such as Google's TensorFlow® processing units, Graphcore's inference processing units (IPUs), or Intel Corp's Nervana® (e.g., "Lake Crest") processors. In at least one embodiment, the inference and / or training logic 815 shown in FIG. 8B may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate array (FPGA). In at least one embodiment, the inference and / or training logic 815 includes, without limitation, code and / or data storage 801 and code and / or data storage 805, and these are used to store code (e.g., graph code), weight values, and / or bias values, gradient information, momentum values, and / or other information including other parameters or hyperparameter information. In at least one embodiment shown in FIG. 8B, each of the code and / or data storage 801 and the code and / or data storage 805 is associated with dedicated computing resources such as computing hardware 802 and computing hardware 806, respectively. In at least one embodiment, each of the computing hardware 802 and the computing hardware 806 includes one or more ALUs that execute mathematical functions such as linear algebra functions only on the information stored in the code and / or data storage 801 and the code and / or data storage 805, respectively, and the results are stored in the activation storage 820.
[0041] In at least one embodiment, each of code and / or data storage 801 and 805, and corresponding computing hardware 802 and 806, respectively corresponds to different layers of a neural network, such that the activation resulting from one “storage / computation pair 801 / 802” of code and / or data storage 801 and computing hardware 802 is provided as an input to the next “storage / computation pair 805 / 806” of code and / or data storage 805 and computing hardware 806 in order to reflect the conceptual organization of the neural network. In at least one embodiment, storage / computation pairs 801 / 802, and 805 / 806 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / computation pairs (not shown) may be included in inference and / or training logic 815 after or in parallel with storage / computation pairs 801 / 802, and 805 / 806.
[0042] Training and Introduction of Neural Network FIG. 9 shows the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 906 is trained using a training dataset 902. In at least one embodiment, the training framework 904 is the PyTorch framework, while in other embodiments, the training framework 904 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 904 trains the untrained neural network 906 and enables it to be trained using the processing resources described herein to generate a trained neural network 908. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, the training may be performed in any of a supervised, semi-supervised, or unsupervised manner.
[0043] In at least one embodiment, the untrained neural network 906 is trained using supervised learning, where the training dataset 902 includes inputs paired with the desired outputs for the inputs, or the training dataset 902 includes inputs with known outputs, and the output of the neural network 906 is scored manually. In at least one embodiment, the untrained neural network 906 is trained in a supervised manner, processes the inputs from the training dataset 902, and compares the resulting output to a set of predicted or desired outputs. In at least one embodiment, an error is then backpropagated through the untrained neural network 906. In at least one embodiment, the training framework 904 adjusts the weights that control the untrained neural network 906. In at least one embodiment, the training framework 904 includes tools that monitor how well the untrained neural network 906 converges towards a model such as a trained neural network 908 that is suitable for generating correct answers in results 914 etc. based on input data such as a new dataset 912. In at least one embodiment, the training framework 904 repeatedly trains the untrained neural network 906 while using a loss function and an adjustment algorithm such as stochastic gradient descent to adjust the weights to refine the output of the untrained neural network 906. In at least one embodiment, the training framework 904 trains the untrained neural network 906 until the untrained neural network 906 reaches the desired accuracy. In at least one embodiment, the trained neural network 908 can then be introduced to perform any number of machine learning operations.
[0044] In at least one embodiment, the untrained neural network 906 is trained using unsupervised learning, where the untrained neural network 906 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training data set 902 includes input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 906 can learn to group within the training data set 902 and determine how individual inputs relate to the untrained data set 902. In at least one embodiment, an unsupervised training can be used to generate a self-organizing map of the trained neural network 908 that can perform operations useful for reducing the dimensions of the new data set 912. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which enables identification of data points within the new data set 912 that deviate from the normal pattern of the new data set 912.
[0045] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled data and unlabeled data are mixed in the training data set 902. In at least one embodiment, the training framework 904 may be used to perform incremental learning, such as by transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 908 to adapt to the new data set 912 without forgetting the knowledge taught into the trained neural network 908 during initial training.
[0046] In at least one embodiment, the training framework 904 is a framework processed in relation to a software development toolkit such as the OpenVINO (Open Visual Inference and Neural network Optimization) toolkit. In at least one embodiment, the OpenVINO toolkit is a toolkit such as one developed by Intel Corporation of Santa Clara, California.
[0047] In at least one embodiment, OpenVINO is a toolkit for facilitating the development of applications for various tasks and operations such as human qualification emulation, speech recognition, natural language processing, recommendation systems, and / or variations thereof, particularly neural network applications. In at least one embodiment, OpenVINO supports neural networks such as convolutional neural networks (CNNs), recurrent and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries such as OpenCV, OpenCL, and / or variations thereof.
[0048] In at least one embodiment, OpenVINO supports neural network models for various tasks and operations such as classification, segmentation, object detection, face recognition, speech recognition, pose estimation (e.g., of humans and / or objects), monocular depth estimation, image inpainting, style transfer, action recognition, colorization, and / or variations thereof.
[0049] In at least one embodiment, OpenVINO includes one or more software tools and / or modules for model optimization, also referred to as the Model Optimizer. In at least one embodiment, the Model Optimizer is a command-line tool that facilitates the transition between the training and deployment of neural network models. In at least one embodiment, the Model Optimizer optimizes neural network models for execution on various devices and / or processing units such as GPUs, CPUs, PPUs, GPGPUs, and / or their variants. In at least one embodiment, the Model Optimizer generates an internal representation of the model and optimizes the model to generate an intermediate representation. In at least one embodiment, the Model Optimizer reduces the number of layers in the model. In at least one embodiment, the Model Optimizer removes layers of the model used for training. In at least one embodiment, the Model Optimizer performs various neural network operations such as modification of the input to the model (e.g., resizing the input to the model), modification of the size of the input to the model (e.g., modifying the batch size of the model), modification of the model structure (e.g., modifying the layers of the model), normalization, standardization, quantization (e.g., conversion of the weights of the model from a first representation such as floating point to a second representation such as integer), and / or their variants.
[0050] In at least one embodiment, OpenVINO includes one or more software libraries for inference, also referred to as the Inference Engine. In at least one embodiment, the Inference Engine is a C++ library or any suitable programming language library. In at least one embodiment, the Inference Engine is utilized to perform inference on input data. In at least one embodiment, the Inference Engine performs various classes for performing inference on input data and generating one or more results. In at least one embodiment, the Inference Engine implements one or more API functions to process the intermediate representation, set the input and / or output format, and / or execute the model on one or more devices.
[0051] In at least one embodiment, OpenVINO provides various functions for heterogeneous execution of one or more neural network models. In at least one embodiment, heterogeneous execution or heterogeneous computing refers to one or more computing processes and / or systems that utilize one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software functions for executing programs on one or more devices. In at least one embodiment, OpenVINO provides various software functions for executing programs and / or program portions on different devices. In at least one embodiment, OpenVINO provides various software functions for, for example, operating a first portion of code on a CPU and a second portion of code on a GPU and / or FPGA. In at least one embodiment, OpenVINO provides various software functions for executing one or more layers on one or more devices (for example, executing a first set of layers on a first device such as a GPU and a second set of layers on a second device such as a CPU).
[0052] In at least one embodiment, OpenVINO includes various functions similar to those associated with the CUDA programming model, such as various neural network model operations associated with frameworks such as TensorFlow, PyTorch, and / or variants thereof. In at least one embodiment, one or more CUDA programming model operations are executed using OpenVINO. In at least one embodiment, the various systems, methods, and / or techniques described herein are implemented using OpenVINO.
[0053] Data center FIG. 10 shows an exemplary data center 1000 in which at least one embodiment may be used. In at least one embodiment, the data center 1000 includes a data center infrastructure layer 1010, a framework layer 1020, a software layer 1030, and an application layer 1040.
[0054] As shown in FIG. 10, in at least one embodiment, the data center infrastructure layer 1010 may include a resource orchestrator 1012, grouped computing resources 1014, and node computing resources (“node C.R.”) 1016(1) to 1016(N), where “N” represents a positive integer (which may be a different integer “N” from that used in other figures). In at least one embodiment, the node C.R. 1016(1) to 1016(N) may include any number of central processing units (“CPU”) or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory storage devices 1018(1) to 1018(N) (e.g., dynamic random access memory, solid state storage, or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VM”), power modules, and cooling modules, but are not limited thereto. In at least one embodiment, one or more of the node C.R. 1016(1) to 1016(N) may be servers having one or more of the computing resources described above.
[0055] In at least one embodiment, the grouped computing resources 1014 may include separate groups of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). In at least one embodiment, the separate groups of node C.R.s within the grouped computing resources 1014 may include grouped computing resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s including a CPU or processor may be grouped within one or more racks to provide computing resources for supporting one or more workloads. In at least one embodiment, one or more racks may also include any combination of any number of power modules, cooling modules, and network switches.
[0056] In at least one embodiment, the resource orchestrator 1012 may configure or otherwise control one or more node C.R.s 1016(1)-1016(N) and / or the grouped computing resources 1014. In at least one embodiment, the resource orchestrator 1012 may include a software design infrastructure ("SDI") management entity for the data center 1000. In at least one embodiment, the resource orchestrator 812 may include hardware, software, or some combination thereof.
[0057] In at least one embodiment shown in FIG. 10, the framework layer 1020 includes a job scheduler 1022, a configuration manager 1024, a resource manager 1026, and a distributed file system 1028. In at least one embodiment, the framework layer 1020 may include a framework for supporting software 1032 of the software layer 1030 and / or one or more applications 1042 of the application layer 1040. In at least one embodiment, the software 1032 or the application 1042 may each include web-based service software or an application, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1020 may be a kind of free and open-source software web application framework, such as Apache Spark (trademark) (hereinafter "Spark"), which can use the distributed file system 1028 for large-scale data processing (e.g., "big data"), but is not limited thereto. In at least one embodiment, the job scheduler 1022 may include a Spark driver to facilitate the scheduling of workloads supported by various layers of the data center 1000. In at least one embodiment, the configuration manager 1024 may be able to configure different layers, such as the software layer 1030 and the framework layer 1020 including Spark and the distributed file system 1028 for supporting large-scale data processing. In at least one embodiment, the resource manager 1026 may be able to manage clustered or grouped computing resources mapped or allocated to support the distributed file system 1028 and the job scheduler 1022. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1014 in the data center infrastructure layer 1010.In at least one embodiment, the resource manager 1026 may manage these mappings or allocated computing resources in cooperation with the resource orchestrator 1012.
[0058] In at least one embodiment, the software 1032 included in the software layer 1030 may include software used by at least a portion of the node C.R. 1016(1)-1016(N), the grouped computing resources 1014, and / or the distributed file system 1028 of the framework layer 1020. In at least one embodiment, the one or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.
[0059] In at least one embodiment, the application 1042 included in the application layer 1040 may include one or more types of applications used by at least a portion of the node C.R. 1016(1)-1016(N), the grouped computing resources 1014, and / or the distributed file system 1028 of the framework layer 1020. In at least one embodiment, the one or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, applications including machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.) and machine learning applications, or other machine learning applications used in conjunction with one or more embodiments.
[0060] In at least one embodiment, any one of the configuration manager 1024, the resource manager 1026, and the resource orchestrator 1012 may implement any number and type of self - correction measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self - correction measures may prevent the data center operator of the data center 1000 from determining a configuration that may be defective and may eliminate parts of the data center that are not fully utilized and / or have low performance.
[0061] In at least one embodiment, the data center 1000 may include tools, services, software, or other resources for training one or more machine - learning models or for predicting or inferring information using one or more machine - learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine - learning model may be trained by calculating weight parameters according to a neural - network architecture using the software and computing resources described above with respect to the data center 1000. In at least one embodiment, a trained machine - learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1000 by using weight parameters calculated by one or more techniques described herein.
[0062] In at least one embodiment, the data center may use a CPU, an application - specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the resources described above. Further, the one or more software and / or hardware resources described above may be configured as a service to enable a user to perform training or inference of information, such as image recognition, speech recognition, or other artificial - intelligence services.
[0063] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of FIG. 10 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or use cases of the neural networks.
[0064] In at least one embodiment, one or more systems shown in FIG. 8 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 8 are utilized to perform various processes such as those described in connection with FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 8 are utilized to estimate parameters as part of one or more training processes of a neural network model, based on the uniqueness of training data, using a subset of the training data such as training images.
[0065] Autonomous vehicle FIG. 11A shows an example of an autonomous vehicle 1100 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1100 (or referred to herein as "vehicle 1100") can be a passenger vehicle such as, without limitation, a car, truck, bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 1100 may be a trailer truck of a semi - tractor for cargo transportation. In at least one embodiment, vehicle 1100 may be an aircraft, robotic vehicle, or other type of vehicle.
[0066] The autonomous vehicle may be described from the perspective of the automation level defined by the National Highway Traffic Safety Administration (NHTSA), a part of the U.S. Department of Transportation, and the Society of Automotive Engineers (SAE)'s "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (for example, Standard No. J3016-201806 issued on June 15, 2018, Standard No. J3016-201609 issued on September 30, 2016, and the old and new versions of this standard). In at least one embodiment, the vehicle 1100 may be capable of corresponding to the functionality by one or more of automation levels 1 to 5 of the autonomous driving level. For example, in at least one embodiment, the vehicle 1100 may be capable of corresponding to conditional automation (level 3), highly automated (level 4), and / or fully automated (level 5) according to the embodiment.
[0067] In at least one embodiment, the vehicle 1100 may include components such as, without limitation, a chassis, a vehicle body, wheels (2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, the vehicle 1100 may include a propulsion system 1150 such as, without limitation, an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, the propulsion system 1150 may be connected to the drive train of the vehicle 1100, and the drive train may include, without limitation, a transmission for enabling the propulsion of the vehicle 1100. In at least one embodiment, the propulsion system 1150 may be controlled in response to receiving a signal from the throttle / accelerator 1152.
[0068] In at least one embodiment, a steering system 1154, which may include a steering wheel without limitation, is used to steer a vehicle 1100 (e.g., along a desired path or route) when a propulsion system 1150 is operating (e.g., when the vehicle 1100 is moving). In at least one embodiment, the steering system 1154 may receive a signal from a steering actuator 1156. In at least one embodiment, the steering wheel may be optional with respect to fully automated (Level 5) functionality. In at least one embodiment, a brake sensor system 1146 may be used to operate vehicle brakes in response to receiving a signal from a brake actuator 1148 and / or a brake sensor.
[0069] In at least one embodiment, a controller 1136, which may include one or more system-on-chips (“SoCs”) (not shown in FIG. 11A) and / or graphics processing units (“GPUs”) without limitation, provides signals (e.g., representing commands) to one or more components and / or systems of the vehicle 1100. For example, in at least one embodiment, the controller 1136 may transmit signals for operating the vehicle brakes via a brake actuator 1148, signals for operating a steering system 1154 via a steering actuator 1156, and signals for operating a propulsion system 1150 via a throttle / accelerator 1152. In at least one embodiment, the controller 1136 may include one or more integrated (e.g., monolithic) computing devices that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in the operating vehicle 1100. In at least one embodiment, the controller 1136 may include a first controller for an autonomous driving function, a second controller for a functional safety function, a third controller for an artificial intelligence function (e.g., computer vision), a fourth controller for an infotainment function, a fifth controller for redundancy in an emergency, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the above functionalities, two or more controllers may handle a single functionality, and / or any combination thereof may be possible.
[0070] In at least one embodiment, controller 1136 provides signals for controlling one or more components and / or systems of vehicle 1100 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, the sensor data may be received from, for example, without limitation, a global navigation satellite system (GNSS) sensor 1158 (e.g., a global positioning system sensor), a RADAR sensor 1160, an ultrasonic sensor 1162, a LIDAR sensor 1164, an inertial measurement unit (IMU) sensor 1166 (e.g., an accelerometer, a gyroscope, one or more magnetic compasses, a magnetometer, etc.), a microphone 1196, a stereo camera 1168, a wide-angle camera 1170 (e.g., a fish-eye camera), an infrared camera 1172, a surround camera 1174 (e.g., a 360-degree camera), a long-range camera (not shown in FIG. 11A), a mid-range camera (not shown in FIG. 11A), a speed sensor 1144 (e.g., for measuring the speed of vehicle 1100), a vibration sensor 1142, a steering sensor 1140, a brake sensor (e.g., as part of a brake sensor system 1146), and / or other types of sensors.
[0071] In at least one embodiment, one or more of the controllers 1136 receive an input (e.g., represented by input data) from the instrument cluster 1132 of the vehicle 1100 and provide an output (e.g., represented by output data, display data, etc.) via the human-machine interface (「HMI」) display 1134, an audible annunciator, a speaker, and / or via other components of the vehicle 1100. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high-definition map (not shown in FIG. 11A), location data (e.g., the location of the vehicle 1100 on a map, etc.), direction, the location of other vehicles (e.g., an occupancy grid), information about objects sensed by the controller 1136, and the state of the objects. For example, in at least one embodiment, the HMI display 1134 may display information about the presence of one or more objects (e.g., road signs, warning signs, signal changes, etc.) and / or information about driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B in 3.22 km (2 miles), etc.).
[0072] In at least one embodiment, vehicle 1100 further includes a network interface 1124, which may use a wireless antenna 1126 and / or a modem for communicating via one or more networks. For example, in at least one embodiment, network interface 1124 may be capable of communicating via a Long-Term Evolution (LTE) network, a Wideband Code Division Multiple Access (WCDMA®) network, a Universal Mobile Telecommunications System (UMTS) network, a Global System for Mobile communications (GSM) network, an IMT-CDMA multi-carrier (CDMA2000) network, and the like. Also, in at least one embodiment, wireless antenna 1126 may use local area network protocols such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, and / or low power wide-area network (LPWAN) protocols such as LoRaWAN, SigFox to enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.).
[0073] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in the system of FIG. 11A for inference or prediction operations, at least in part based on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks.
[0074] In at least one embodiment, one or more systems shown in FIG. 11A are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 11A are utilized to perform various processes such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 11A are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0075] FIG. 11B shows an example of the camera locations and fields of view for the autonomous vehicle 1100 of FIG. 11A according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are one example of an embodiment and are not limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be positioned at different locations on the vehicle 1100.
[0076] In at least one embodiment, the camera type of the camera may include, but is not limited to, a digital camera that may be adapted to be used with the components and / or systems of the vehicle 1100. In at least one embodiment, the camera may operate at automotive safety integrity level (ASIL) B and / or another ASIL. In at least one embodiment, the camera type may be capable of corresponding to any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a color filter array of red, clear, clear, clear ("RCCC": red clear clear clear), a color filter array of red, clear, clear, blue ("RCCB: red clear clear blue"), a color filter array of red, blue, green, clear ("RBGC": red blue green clear), a color filter array of Foveon X3, a color filter array of a Bayer sensor (RGGB), a color filter array of a monochrome sensor, and / or another type of color filter array. In at least one embodiment, a clear pixel camera, such as a camera having an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.
[0077] In at least one embodiment, one or more of the cameras may be used to perform advanced driver assistance systems (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).
[0078] In at least one embodiment, one or more cameras may be attached to a mounting assembly such as a custom-designed (3D printed) assembly to eliminate stray light and reflections from inside the vehicle 1100 that may interfere with the camera's image data capture performance (e.g., reflections reflected from the dashboard to the windshield). Referring to the door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed such that the camera mounting plate fits the shape of the door mirror. In at least one embodiment, the camera may be integrated with the door mirror. In at least one embodiment, for a side view camera, the camera may also be integrated into the four pillars at each corner of the cabin in this case.
[0079] In at least one embodiment, a camera (e.g., a front camera) having a field of view that includes a portion of the environment in front of the vehicle 1100 is used for the surrounding view to facilitate identification of the front path and obstacles, and may be used with one or more of the controller 1136 and / or the control SoC to assist in providing information essential for the generation of the occupancy grid and / or the determination of the preferred vehicle path. In at least one embodiment, many of the ADAS functions similar to LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance, may be performed using the front camera. In at least one embodiment, the front camera may also be used for ADAS functions and systems including, without limitation, other functions such as lane departure warnings (“LDW”), autonomous cruise control (“ACC”), and / or traffic sign recognition.
[0080] In at least one embodiment, various cameras including a platform of a monocular camera including, for example, a CMOS:complementary metal oxide semiconductor (“complementary metal oxide semiconductor”) color imaging device may be used in a front configuration. In at least one embodiment, a wide-angle camera 1170 may be used to sense objects (e.g., pedestrians, cross-traffic, or bicycles) entering the view from the surroundings. Although only one wide-angle camera 1170 is shown in FIG. 11B, in other embodiments, the vehicle 1100 may have any number (including zero) of wide-angle cameras. In at least one embodiment, any number of long-distance cameras 1198 (e.g., a pair of stereo cameras for a long-distance view) may be used for depth-based object detection, especially for objects for which the neural network has not yet been trained. In at least one embodiment, the long-distance camera 1198 may also be used for object detection and classification, and basic object tracking.
[0081] In at least one embodiment, any number of stereo cameras 1168 may also be included in the front configuration. In at least one embodiment, one or more stereo cameras 1168 may include an integrated control unit with an expandable processing unit, which may provide a programmable logic (FPGA) and a multi-core microprocessor having an integrated controller area network (CAN) or Ethernet® interface on a single chip. In at least one embodiment, such units may be used to generate a 3D map of the environment of the vehicle 1100, including distance estimation for all points within the image. In at least one embodiment, one or more of the stereo cameras 1168 may include, without limitation, a compact stereo vision sensor, which measures the distance from the vehicle 1100 to a target object and can activate functions such as autonomous emergency braking and lane departure warning using the generated information (e.g., metadata), and may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip. In at least one embodiment, in addition to or instead of those described herein, other types of stereo cameras 1168 may be used.
[0082] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view that includes a portion of the environment to the side of the vehicle 1100 may be used for the surrounding view to provide information for creating and updating an occupancy grid and for generating a side collision warning. For example, in at least one embodiment, a surround camera 1174 (e.g., four surround cameras as shown in FIG. 11B) can be disposed on the vehicle 1100. The surround camera 1174 may include, without limitation, any number and combination of wide-angle cameras 1770, fish-eye cameras, and / or 360-degree cameras, etc. For example, in at least one embodiment, four fish-eye cameras may be disposed in front of, behind, and to the sides of the vehicle 1100. In at least one embodiment, the vehicle 1100 may use three surround cameras 1174 (e.g., left, right, and rear), and as a fourth surround camera, one or more other cameras (e.g., a front camera) may be utilized.
[0083] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view that includes a portion of the environment behind the vehicle 1100 may be used for parking assistance, the surrounding view, and for a rear collision warning, and the occupancy grid may be created and updated. In at least one embodiment, a wide variety of cameras including, but not limited to, cameras suitable as the front camera described herein (e.g., a long-range camera 1198, and / or a mid-range camera 1176, a stereo camera 1168), an infrared camera 1172, etc. may be used.
[0084] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of FIG. 11B for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0085] In at least one embodiment, one or more systems shown in FIG. 11B are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 11B are utilized to perform various processes such as those described in connection with FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 11B are utilized to estimate parameters as part of one or more training processes of a neural network model, based on the uniqueness of training data, using a subset of the training data such as training images.
[0086] FIG. 11C is a block diagram showing an exemplary system architecture of the autonomous vehicle 1100 of FIG. 11A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1100 of FIG. 11C is shown as being connected via a bus 1102. In at least one embodiment, the bus 1102 may include, without limitation, a CAN data interface (or referred to herein as the (CAN bus)). In at least one embodiment, CAN may be an internal network of the vehicle 1100 used to assist in controlling various features and functions of the vehicle 1100, such as brake actuation, acceleration, brake control, steering, windshield wipers, and the like. In at least one embodiment, the bus 1102 may be configured to have dozens or even hundreds of nodes, each having its own unique identifier (e.g., CAN ID). In at least one embodiment, the bus 1102 may be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle state indicators. In at least one embodiment, the bus 1102 may be a CAN bus compliant with ASIL B.
[0087] In at least one embodiment, in addition to or instead of CAN, FlexRay and / or Ethernet® protocol may be used. In at least one embodiment, any number of buses forming bus 1102 may exist, including, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet® buses, and / or zero or more other types of buses using different protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or to provide redundancy. For example, a first bus may be used for collision avoidance function and a second bus may be used for actuation control. In at least one embodiment, each bus of bus 1102 may communicate with any of the components of vehicle 1100, and two or more buses of bus 1102 may communicate with corresponding components. In at least one embodiment, each of any number of system-on-chips (“SoC”) 1104 (such as SoC 1104(A) and SoC 1104(B)), each of controllers 1136, and / or each computer in the vehicle may be accessible to the same input data (e.g., input from sensors of vehicle 1100) and may be connected to a common bus such as a CAN bus.
[0088] In at least one embodiment, vehicle 1100 may include one or more controllers 1136, such as those described herein with respect to FIG. 11A. In at least one embodiment, controller 1136 may be used for various functions. In at least one embodiment, controller 1136 may be coupled to any of the various other components and systems of vehicle 1100 and may be used for vehicle 1100, the artificial intelligence of vehicle 1100, the infotainment of vehicle 1100, and / or other functions.
[0089] In at least one embodiment, vehicle 1100 may include any number of SoCs 1104. In at least one embodiment, each of the SoCs 1104 may include, without limitation, a central processing unit (“CPU”) 1106, a graphics processing unit (“GPU”) 1108, a processor 1110, a cache 1112, an accelerator 1114, a data store 1116, and / or other components and features not shown. In at least one embodiment, the SoC 1104 may be used to control the vehicle 1100 in various platforms and systems. For example, in at least one embodiment, the SoC 1104 may be incorporated into a system (e.g., the system of vehicle 1100) having a high-definition (“HD”) map 1122 that can obtain map refreshes and / or updates via a network interface 1124 from one or more servers (not shown in FIG. 11C).
[0090] In at least one embodiment, the CPU 1106 may include a CPU cluster, or a CPU complex (or referred to herein as “CCPLEX”). In at least one embodiment, the CPU 1106 may include a plurality of cores and / or a level 2 (“L2”) cache. For example, in at least one embodiment, the CPU 1106 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU 1106 may include four dual-core clusters, where each cluster has a dedicated L2 cache (e.g., a 2 megabyte (MB) L2 cache). In at least one embodiment, the CPU 1106 (e.g., CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of the clusters of the CPU 1106 to be activated at any given time.
[0091] In at least one embodiment, one or more of the CPUs 1106 may implement a power management function, which may include, without limitation, one or more of the following features: individual hardware blocks can be automatically clock-gated during idle to save dynamic power; each core clock can be gated when such a core is not actively executing instructions due to the execution of a wait for interrupt ("WFI") / wait for event ("WFE") instruction; each core can be power-gated independently; when all cores are clock-gated or power-gated, each core cluster can be clock-gated independently; and / or when all cores are power-gated, each core cluster can be power-gated independently. In at least one embodiment, the CPU 1106 may further implement an extended algorithm for managing power states, where an acceptable power state and an expected wake-up time are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. In at least one embodiment, the processing cores may be supported in software by a simple sequence for entering a power state with work offloaded to microcode.
[0092] In at least one embodiment, GPU 1108 may include an integrated GPU (or, as referred to herein, an "iGPU"). In at least one embodiment, GPU 1108 may be programmable and may be efficient for parallel workloads. In at least one embodiment, GPU 1108 may use an extended tensor instruction set. In at least one embodiment, GPU 1108 may include one or more streaming microprocessors, where each streaming microprocessor may include a level 1 ("L1") cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache having a storage capacity of 512 KB). In at least one embodiment, GPU 1108 may include at least eight streaming microprocessors. In at least one embodiment, GPU 1108 may use a compute application programming interface (API). In at least one embodiment, GPU 1108 may use one or more parallel computing platforms and / or programming modules (e.g., NVIDIA's CUDA model).
[0093] In at least one embodiment, one or more of the GPUs 1108 may be power optimized for best performance in automotive and embedded use cases. For example, in one embodiment, the GPU 1108 can be fabricated on a fin field-effect transistor (FinFET) circuit. In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores partitioned into multiple blocks. For example, without limitation, 64 PF32 cores and 32 PF64 cores can be partitioned into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA tensor cores for deep learning matrix operations, a level zero (L0) instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. In at least one embodiment, the streaming microprocessor includes independent parallel data paths for integers and floating points, and realizes efficient execution of workloads by mixing computer processing and addressing calculations. In at least one embodiment, the streaming microprocessor includes an independent thread scheduling function, which may enable finer-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance and simplify programming.
[0094] In at least one embodiment, one or more of the GPU1108s include high bandwidth memory (HBM) and / or a 16GB HBM2 memory subsystem, and in some examples, may provide a peak memory bandwidth of about 900GB / second. In at least one embodiment, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as five graphics double data rate type five (GDDR5) synchronous random access memories, may be used.
[0095] In at least one embodiment, the GPU1108 may include integrated memory technology. In at least one embodiment, address translation services (ATS) support may be used to enable the GPU1108 to directly access the page table of the CPU1106. In at least one embodiment, when the GPU of the GPU1108 memory management unit (MMU) encounters a miss, an address translation request may be sent to the CPU1106. In at least one embodiment, in response, one of the CPUs of the CPU1106 may search its page table for a virtual-to-physical address mapping and send the translation back to the GPU1108. In at least one embodiment, the integrated memory technology enables a single integrated virtual address space for the memories of both the CPU1106 and the GPU1108, thereby simplifying the programming of the GPU1108 and the porting of applications to the GPU1108.
[0096] In at least one embodiment, GPU 1108 may include any number of access counters that can record the access frequency of GPU 1108 to the memory of other processors. In at least one embodiment, the access counter may assist in ensuring that memory pages are moved to the physical memory of the processor that most frequently accesses the pages, thereby improving the efficiency of the memory range shared among processors.
[0097] In at least one embodiment, one or more of SoC 1104 may include any number of caches 1112 including those described herein. For example, in at least one embodiment, cache 1112 may include a level 3 ("L3") cache that is available to both CPU 1106 and GPU 1108 (e.g., connected to both CPU 1106 and GPU 1108). In at least one embodiment, cache 1112 may include a write-back cache that can record the state of lines by using a cache coherence protocol or the like (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4MB or more of memory depending on the embodiment, although smaller cache sizes may be used.
[0098] In at least one embodiment, one or more of the SoCs 1104 may include one or more accelerators 1114 (e.g., a hardware accelerator, a software accelerator, or a combination thereof). In at least one embodiment, the SoC 1104 may include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memories. In at least one embodiment, a large on-chip memory (e.g., 4MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU 1108 and offload some of the tasks of the GPU 1108 (e.g., free up more cycles of the GPU 1108 so that other tasks can be executed). In at least one embodiment, the accelerator 1114 can be used for workloads (e.g., perception, convolutional neural network ("CNN"), recurrent neural network ("RNN"), etc.) that are stable enough to accept acceleration. In at least one embodiment, the CNN may include region-based, i.e., region convolutional neural network ("RCNN"), and (e.g., used for object detection) fast RCNN, or other types of CNNs.
[0099] In at least one embodiment, the accelerator 1114 (e.g., a hardware acceleration cluster) may include one or more deep learning accelerators (“DLAs”). In at least one embodiment, the DLA may include, without limitation, one or more Tensor processing units (“TPUs”), which may be further configured to provide up to 10 trillion operations per second for deep learning applications and inferences. In at least one embodiment, the TPU may be an accelerator configured and optimized to execute image processing functions (e.g., CNN, RCNN, etc.). In at least one embodiment, the DLA may be further optimized for a specific set of neural network types and floating point operations, as well as for inferences. In at least one embodiment, the design of the DLA can improve the performance per millimeter compared to a typical general-purpose GPU, typically far exceeding the performance of a CPU. In at least one embodiment, the TPU may execute several functions, including, for example, a single instance of a convolution function that supports INT8, INT16, and FP16 data types for both features and weights, as well as a post-processing function. In at least one embodiment, the DLA may execute neural networks, particularly CNNs, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including, without limitation, CNNs for object identification and detection using data from a camera sensor, CNNs for distance estimation using data from a camera sensor, emergency vehicle detection and identification using data from a microphone, CNNs for face recognition and vehicle owner identification using data from a camera sensor, and / or CNNs for events related to security and / or safety.
[0100] In at least one embodiment, the DLA may execute any function of the GPU 1108. For example, by using an inference accelerator, the designer may target either the DLA or the GPU 1108 for any function. For example, in at least one embodiment, the designer may concentrate the processing of CNN and floating-point operations on the DLA and leave other functions to the GPU 1108 and / or the accelerator 1114.
[0101] In at least one embodiment, the accelerator 1114 may include a programmable vision accelerator (referred to herein alternatively as a computer vision accelerator). In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS) 1138, autonomous driving, augmented reality (AR) applications, and / or virtual reality (VR) applications. In at least one embodiment, the PVA can provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include, without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.
[0102] In at least one embodiment, the RISC core may interact with an image sensor (e.g., the image sensor of any of the cameras described herein), and / or an image signal processor. In at least one embodiment, each of the RISC cores may include any amount of memory. In at least one embodiment, the RISC core may use any of a plurality of protocols depending on the embodiment. In at least one embodiment, the RISC core may execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.
[0103] In at least one embodiment, the DMA may enable the components of the PVA to access system memory independently of the CPU1106. In at least one embodiment, the DMA may support any number of features used to optimize the PVA, including but not limited to multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0104] In at least one embodiment, the vector processor can be a programmable processor that may be designed to efficiently and flexibly execute programming for computer vision algorithms, and provides signal processing capabilities. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, the vector processing subsystem may operate as the primary processing engine of the PVA, and may include a vector processing unit ("VPU"), an instruction cache, and / or vector memory (e.g., "VMEM"). In at least one embodiment, the VPU core may include a digital signal processor such as, for example, a single instruction multiple data ("SIMD"), very long instruction word ("VLIW") digital signal processor. In at least one embodiment, throughput and speed may be improved by a combination of SIMD and VLIW.
[0105] In at least one embodiment, each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to execute independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute a common computer vision algorithm on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on one image, or even execute different algorithms on consecutive images or on portions of an image. In at least one embodiment, in particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, the PVA may include additional error correction code (「ECC」) memory to enhance the overall security of the system.
[0106] In at least one embodiment, the accelerator 1114 includes an on-chip computer vision network and static random access memory ("SRAM"), and may provide high-bandwidth, low-latency SRAM for the accelerator 1114. In at least one embodiment, the on-chip memory may include at least 4MB of SRAM consisting of, for example, without limitation, eight field-configurable memory blocks, which may be accessible from either the PVA or the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus ("APB") interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and the DLA may access the memory via a backbone that provides fast access to the memory for the PVA and the DLA. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and the DLA to the memory (e.g., using the APB).
[0107] In at least one embodiment, the on-chip computer vision network may include an interface that determines that both the PVA and the DLA provide a ready signal and a valid signal before transmitting any control signal / address / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signal / address / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, the interface may comply with the standards of the International Organization for Standardization ("ISO") 26262 or the International Electrotechnical Commission ("IEC") 61508, although other standards and protocols may be used.
[0108] In at least one embodiment, one or more of the SoCs 1104 may include a hardware accelerator for real-time ray tracing. In at least one embodiment, the hardware accelerator for real-time ray tracing is used to quickly and efficiently determine the position and extent of an object (e.g., within a world model) for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of a SONAR system, for simulation of general waveform propagation, for comparison with LIDAR data for localization and / or other functions, and / or for other uses, and a real-time visualization simulation may be generated.
[0109] In at least one embodiment, the accelerator 1114 has various uses for autonomous driving. In at least one embodiment, the PVA can be used in the main processing stages of ADAS and autonomous vehicles. In at least one embodiment, the performance of the PVA is well-suited to algorithm domains that require predictable processing with low power and low latency. In other words, the PVA functions well for semi-dense or dense regular calculations that may require a predictable runtime with low latency and low power, even with a small data set. In at least one embodiment, in a vehicle 1100 or the like, the PVA can be designed to execute conventional computer vision algorithms because they can be effective for object detection and integer arithmetic operations.
[0110] For example, according to at least one embodiment of the technology, computer stereo vision may be performed using PVA. In at least one embodiment, in some examples, an algorithm based on semi-global matching may be used, but this is not limiting. In at least one embodiment, applications for level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may perform computer stereo vision functions for inputs from two monocular cameras.
[0111] In at least one embodiment, high-density optical flow may be performed using PVA. For example, in at least one embodiment, PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, PVA is used for time-of-flight depth processing, and for example, processed time-of-flight data is provided by processing raw time-of-flight data.
[0112] In at least one embodiment, for example without limitation, a DLA may be used to execute any type of network for enhancing control and driving safety, including a neural network that outputs a measure of reliability for each object detection. In at least one embodiment, the reliability may be represented or interpreted as the probability of each detection compared to other detections, or as providing its relative "weight". In at least one embodiment, the reliability measure enables the system to make further decisions regarding which detections should be considered positive detections rather than false detections. In at least one embodiment, the system may set a threshold for the reliability and consider only detections that exceed the threshold as positive detections. In embodiments where automatic emergency braking ("AEB") is used, false detections would cause the vehicle to automatically apply the emergency brakes, which is clearly undesirable. In at least one embodiment, highly reliable detections may be considered as triggers for AEB. In at least one embodiment, the DLA may execute the neural network to regress the reliability value. In at least one embodiment, the neural network may take as its input at least some subset of parameters such as, among others, the dimensions of the bounding box, the ground estimation obtained (e.g., from another subsystem), the output from the IMU sensor 1166 correlated with the orientation of the vehicle 1100, the distance, and the 3D location estimation of the object obtained from the neural network and / or other sensors (e.g., the LIDAR sensor 1164 or the RADAR sensor 1160).
[0113] In at least one embodiment, one or more of the SoCs 1104 may include a data store 1116 (e.g., a memory). In at least one embodiment, the data store 1116 may be an on-chip memory of the SoC 1104, and this memory may store neural networks executed on the GPU 1108 and / or the DLA. In at least one embodiment, the capacity of the data store 1116 may be large enough to store multiple instances of the neural network for redundancy and security. In at least one embodiment, the data store 1116 may include an L2 or L3 cache.
[0114] In at least one embodiment, one or more of the SoCs 1104 may include any number of processors 1110 (e.g., embedded processors). In at least one embodiment, the processor 1110 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and related security enforcement. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC 1104 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance with transitioning the system to a low-power state, management of the heat and temperature sensors of the SoC 1104, and / or management of the power state of the SoC 1104. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1104 may use the ring oscillator to detect the temperature of the CPU 1106, GPU 1108, and / or accelerator 1114. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine, put the SoC 1104 into a low-power state, and / or put the vehicle 1100 into a driver-safety shutdown mode (e.g., safely shut down the vehicle 1100).
[0115] In at least one embodiment, the processor 1110 may further include a set of embedded processors that can serve as an audio processing engine that enables complete hardware support for multi-channel audio via a multi-interface and a wide variety of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.
[0116] In at least one embodiment, the processor 1110 may further include an always-on processor engine that can provide the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, support peripherals (e.g., timers, and interrupt controllers), various I / O controller peripherals, and routing logic.
[0117] In at least one embodiment, the processor 1110 may further include a safety cluster engine, which may include, without limitation, a dedicated processor subsystem for handling safety management in automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, support peripherals (such as timers, and interrupt controllers, etc.), and / or routing logic. In the safe mode, in at least one embodiment, two or more cores may operate in the lockstep mode and may function as a single core having comparison logic for detecting any differences between these operations. In at least one embodiment, the processor 1110 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, the processor 1110 may further include a high-dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of the camera processing pipeline.
[0118] In at least one embodiment, the processor 1110 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final image in the window of the playback device. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1170, the surround camera 1174, and / or the in-cabin monitoring camera / sensor. In at least one embodiment, the in-cabin monitoring camera / sensor is preferably monitored by a neural network running on another instance of the SoC 1104 that is configured to identify events in the cabin and respond thereto as appropriate. In at least one embodiment, the in-cabin system may perform lip reading, without limitation, to activate cellular service, make a call, write an email, change the destination of the vehicle, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and unavailable otherwise.
[0119] In at least one embodiment, the video image synthesizer may include extended temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when there is movement in the video, the noise reduction appropriately weights the spatial information and downweights the information provided by adjacent frames. In at least one embodiment, when an image or a portion of the image does not contain movement, the temporal noise reduction performed by the video image synthesizer may use information from the previous image to reduce the noise in the current image.
[0120] In at least one embodiment, the video image synthesizer may also be configured to perform stereo parallelization on the input stereo lens frame. In at least one embodiment, the video image synthesizer may further be used to synthesize a user interface when the operating system desktop is in use, and the GPU 1108 need not continuously render a new surface. In at least one embodiment, when the GPU 1108 is powered on and performing active 3D rendering, the video image synthesizer may be used to offload the GPU 1108 to improve performance and responsiveness.
[0121] In at least one embodiment, one or more of the SoCs of the SoC 1104 may further include a camera serial interface of a mobile industry processor interface ( "MIPI") for receiving inputs from video and cameras, a high-speed interface, and / or a video input block that may be used for the input functions of cameras and related pixels. In at least one embodiment, one or more of the SoCs of the SoC 1104 may further include an input / output controller, which may be controlled by software and may be used to receive I / O signals that are not bound to a specific role.
[0122] In at least one embodiment, one or more of the SoCs of the SoC1104 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio encoders / decoders ("codecs"), power management, and / or other devices. In at least one embodiment, the SoC1104 may be used to process data from a camera (e.g., connected via a gigabit multimedia serial link and an Ethernet® channel), data from sensors (e.g., a LIDAR sensor 1164, a RADAR sensor 1160, etc. that may be connected via an Ethernet® channel), data from the bus 1102 (e.g., the speed of the vehicle 1100, the steering wheel position, etc.), data from a GNSS sensor 1158 (e.g., connected via an Ethernet® bus or a CAN bus), etc. In at least one embodiment, one or more of the SoCs of the SoC1104 may further include a dedicated high-performance large-capacity storage controller, which may include its own DMA engine and may be used to free the CPU1106 from routine data management tasks.
[0123] In at least one embodiment, the SoC1104 may be an end-to-end platform with a flexible architecture spanning automation levels 3 to 5, thereby providing a comprehensive functional safety architecture that leverages computer vision and ADAS techniques to obtain diversity and redundancy and use them efficiently, and a flexible and reliable driving software stack is provided together with deep learning tools. In at least one embodiment, the SoC1104 is faster, more reliable, and more energy- and space-efficient than conventional systems. For example, in at least one embodiment, when the accelerator 1114 is combined with the CPU1106, the GPU1108, and the data store 1116, a fast and efficient platform for level 3 to 5 autonomous vehicles can be realized.
[0124] In at least one embodiment, the computer vision algorithm may be executed on a CPU, and this algorithm may be configured using a high-level programming language such as C to execute various processing algorithms over various visual data. However, in at least one embodiment, the CPU often cannot meet the performance requirements of many computer vision applications, such as requirements regarding execution time and power consumption. In at least one embodiment, many CPUs cannot execute in real time the complex object detection algorithms used in ADAS applications within a vehicle and in realistic level 3-5 autonomous vehicles.
[0125] The embodiments described herein can enable the execution of multiple neural networks simultaneously and / or sequentially, and combine the results to enable level 3-5 autonomous driving functions. For example, in at least one embodiment, the CNN running on the DLA or an individual GPU (e.g., GPU 1120) may include text and word recognition, and enable reading and understanding traffic signs including signs that the neural network has not been specifically trained for. In at least one embodiment, the DLA may further include a neural network that can identify and interpret signs and provide a semantic understanding of the signs, and can pass the semantic understanding to a path planning module running on a CPU complex.
[0126] In at least one embodiment, for level 3, 4, or 5 operation, multiple neural networks may be executed simultaneously. For example, in at least one embodiment, a warning sign that describes "Caution: Frozen when flashing" in conjunction with the electro-optical light may be interpreted separately or collectively by several neural networks. In at least one embodiment, such a warning sign itself may be identified as a traffic sign by a first introduced neural network (e.g., a trained neural network), and the text "Frozen when flashing" may be interpreted by a second introduced neural network. When a flashing light is detected, this neural network notifies the vehicle's (preferably running on the CPU complex) route planning software that a frozen state exists. In at least one embodiment, the flashing light may be identified by operating a third introduced neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may be executed simultaneously, such as within the DLA and / or on the GPU 1108.
[0127] In at least one embodiment, a CNN for face recognition and vehicle owner identification may use data from a camera sensor to identify the presence of an approved driver and / or owner of the vehicle 1100. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and when the owner leaves such a vehicle, the vehicle may be made inoperable in a security mode. Thus, the SoC 1104 realizes security against theft and / or carjacking.
[0128] In at least one embodiment, the CNN for detecting and identifying emergency vehicles may detect and identify the sirens of emergency vehicles using data from microphone 1196. In at least one embodiment, SoC 1104 classifies environmental and urban sounds and uses a CNN to classify visual data. In at least one embodiment, the CNN executed on the DLA is trained to identify the relative speed at which an emergency vehicle is approaching (e.g., by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the area where the vehicle is operating, identified by GNSS sensor 1158. In at least one embodiment, when operating in Europe, the CNN attempts to detect European sirens, and in the case of North America, attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program for executing an emergency vehicle safety routine is used to reduce the speed of the vehicle, move it to the side of the road, stop the vehicle, and / or use ultrasonic sensor 1162 in combination to idle the vehicle until the emergency vehicle has passed.
[0129] In at least one embodiment, vehicle 1100 may include a CPU 1118 (e.g., a discrete CPU or dCPU), which may be coupled to SoC 1104 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU 1118 may include, for example, an X86 processor. CPU 1118 may be used to perform any of a variety of functions, including, for example, reconciling potentially inconsistent results between ADAS sensors and SoC 1104 and / or monitoring the state and health of controller 1136 and / or the in-vehicle infotainment system ("infotainment SoC") 1130 on the chip.
[0130] In at least one embodiment, vehicle 1100 may include a GPU 1120 (e.g., a discrete GPU or dGPU), which may be coupled to the SoC 1104 via a high-speed interconnect (e.g., an NVIDIA NVLINK channel). In at least one embodiment, the GPU 1120 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update a neural network based at least in part on inputs (e.g., sensor data) from sensors of the vehicle 1100.
[0131] In at least one embodiment, vehicle 1100 may further include a network interface 1124, which may include, without limitation, a wireless antenna 1126 (e.g., one or more wireless antennas for different communication protocols such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, the network interface 1124 may be used to enable a wireless connection to an Internet cloud service with a cloud (e.g., a server and / or other network devices), other vehicles, and / or computing devices (e.g., a passenger's client device). In at least one embodiment, a direct link may be established between vehicle 110 and another vehicle for communicating with the other vehicle, and / or an indirect link may be established (e.g., across a network and via the Internet). In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. In at least one embodiment, the vehicle-to-vehicle communication link may provide vehicle 1100 with information about vehicles in the vicinity of vehicle 1100 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1100). In at least one embodiment, such aforementioned functions may be part of a cooperative adaptive cruise control function of vehicle 1100.
[0132] In at least one embodiment, network interface 1124 may include a system-on-chip (SoC) that provides modulation and demodulation functions to enable the controller 1136 to communicate via a wireless network. In at least one embodiment, network interface 1124 may include a radio frequency (RF) front end for up-conversion from baseband to RF and down-conversion from RF to baseband. In at least one embodiment, the frequency conversion may be performed in any technically feasible manner. For example, the frequency conversion can be performed by well-known processes and / or using a superheterodyne process. In at least one embodiment, the RF front end functions may be provided by a separate chip. In at least one embodiment, the network interface may include wireless capabilities for communicating via LTE, WCDMA (registered trademark), UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0133] In at least one embodiment, vehicle 1100 may further include a data store 1128, which may include off-chip (e.g., not on SoC 1104) storage without limitation. In at least one embodiment, data store 1128 may include one or more storage elements including, without limitation, RAM, SRAM, dynamic random access memory ("DRAM"), video random-access memory ("VRAM"), flash memory, hard disk, and / or other components and / or devices capable of storing at least one bit of data.
[0134] In at least one embodiment, vehicle 1100 may further include a GNSS sensor 1158 (e.g., a GPS and / or an assisted GPS sensor) to assist with mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1158 including, for example without limitation, a GPS using a USB connector having a bridge from Ethernet (registered trademark) to serial (e.g., RS-232) may be used.
[0135] In at least one embodiment, vehicle 1100 may further include a RADAR sensor 1160. In at least one embodiment, RADAR sensor 1160 may be used by vehicle 1100 to perform long-range vehicle detection even in darkness and / or severe weather conditions. In at least one embodiment, the functional safety level of the RADAR may be ASIL B. In at least one embodiment, RADAR sensor 1160 may use the CAN bus and / or bus 1102 for control (e.g., to send data generated by RADAR sensor 1160) and to access object tracking data, and in some examples, can access an Ethernet (registered trademark) channel to access raw data. In at least one embodiment, various types of RADAR sensors may be used. For example without limitation, RADAR sensor 1160 may be suitable for use in front, rear, and side RADAR. In at least one embodiment, one or more of the RADAR sensors 1160 are pulse Doppler RADAR sensors.
[0136] In at least one embodiment, the RADAR sensor 1160 may include different configurations, such as a narrow field of view for long distances, a wide field of view for short distances, and short distances covering the sides. In at least one embodiment, the long-range RADAR may be used for an adaptive cruise control function. In at least one embodiment, the long-range RADAR system may provide a wide field of view within a range of 250 m (meters) realized by two or more independent scans. In at least one embodiment, the RADAR sensor 1160 may facilitate distinguishing between static and moving objects and may be used by the ADAS system 1138 for emergency braking assistance and forward collision warning. In at least one embodiment, the sensor 1160 included in the long-range RADAR system may include, without limitation, a plurality (e.g., six or more) of fixed RADAR antennas, and a monostatic multi-mode RADAR having high-speed CAN and FlexRay interfaces. In at least one embodiment, when there are six antennas, the four central antennas may generate a concentrated beam pattern designed to record the surroundings of the vehicle 1100 at a higher speed with minimum interference from adjacent lanes. In at least one embodiment, the other two antennas may expand the field of view and enable rapid detection of vehicles entering or exiting the lane of the vehicle 1100.
[0137] In at least one embodiment, the mid-range RADAR system may include, by way of example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, the short-range RADAR system may include any number of RADAR sensors 1160 designed to be installed at both ends of the rear bumper without limitation. When installed at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that constantly monitor the rear direction and blind spots adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used in the ADAS system 1138 for blind spot detection and / or lane change assistance.
[0138] In at least one embodiment, vehicle 1100 may further include ultrasonic sensor 1162. In at least one embodiment, ultrasonic sensor 1162, which may be disposed at a front, rear, and / or side location of vehicle 1100, may be used for parking assistance and / or for generating and updating an occupancy grid. In at least one embodiment, various ultrasonic sensors 1162 may be used, and different ultrasonic sensors 1162 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensor 1162 may operate at a functional safety level of ASIL B.
[0139] In at least one embodiment, vehicle 1100 may include LIDAR sensor 1164. In at least one embodiment, LIDAR sensor 1164 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, LIDAR sensor 1164 may operate at a functional safety level of ASIL B. In at least one embodiment, vehicle 1100 may include a plurality of LIDAR sensors 1164 (e.g., two, four, six, etc.), and these sensors may use an Ethernet channel (e.g., to provide data to a gigabit Ethernet switch).
[0140] In at least one embodiment, the LIDAR sensor 1164 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 1164 may, for example, have a claimed range of approximately 100 m, an accuracy of 2 cm to 3 cm, and support a 100 Mbps Ethernet® connection. In at least one embodiment, one or more non-protruding LIDAR sensors may be used. In such embodiments, the LIDAR sensor 1164 may include small devices that can be incorporated in front of, behind, on the sides of, and / or at corner locations of the vehicle 1100. In at least one embodiment, the LIDAR sensor 1164 of such embodiments may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, in a range of 200 m. In at least one embodiment, the LIDAR sensor 1164 mounted in the front may be configured to provide a horizontal field of view of 45 degrees to 135 degrees.
[0141] In at least one embodiment, LIDAR technologies such as 3D flash LIDAR may also be used. In at least one embodiment, the 3D flash LIDAR uses the flash of a laser as a transmission source to irradiate the area around the vehicle 1100 up to approximately 200 m at most. In at least one embodiment, the flash LIDAR unit includes, without limitation, a receptor that records the transit time of the laser pulse and the reflected light at each pixel, which corresponds to the range from the vehicle 1100 to the object. In at least one embodiment, the flash LIDAR enables a very accurate and non-distorted surrounding image to be generated for each flash of the laser. In at least one embodiment, four flash LIDARs may be introduced, one on each side of the vehicle 1100. In at least one embodiment, the 3D flash LIDAR system includes, without limitation, a LIDAR camera with a semiconductor 3D staring array (such as a non-scanning LIDAR device) without moving parts other than a fan. In at least one embodiment, the flash LIDAR device may use class I (eye-safe) laser pulses of 5 nanoseconds per frame and capture the reflected laser light as a 3D range point cloud and co-registered intensity data.
[0142] In at least one embodiment, the vehicle 1100 may further include an IMU sensor 1166. In at least one embodiment, the IMU sensor 1166 may be positioned at the center of the rear axle of the vehicle 1100. In at least one embodiment, the IMU sensor 1166 may include, without limitation, for example, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, a plurality of magnetic compasses, and / or other types of sensors. In at least one embodiment, for a six-axis application, the IMU sensor 1166 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, for a nine-axis application, the IMU sensor 1166 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0143] In at least one embodiment, the IMU sensor 1166 may be implemented as a small high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical systems (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and orientation. In at least one embodiment, the IMU sensor 1166 enables the vehicle 1100 to estimate the orientation of the vehicle 1100 without the need for input from a magnetic sensor by directly observing changes in speed and correlating them from GPS to the IMU sensor 1166. In at least one embodiment, the IMU sensor 1166 and the GNSS sensor 1158 may be combined in a single integrated unit.
[0144] In at least one embodiment, the vehicle 1100 may include a microphone 1196 installed inside and / or around the vehicle 1100. In at least one embodiment, the microphone 1196 may be used, inter alia, for the detection and identification of emergency vehicles.
[0145] In at least one embodiment, vehicle 1100 may further include any number of camera types including stereo camera 1168, wide-angle camera 1170, infrared camera 1172, surround camera 1174, long-range camera 1198, mid-range camera 1176, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of vehicle 1100. In at least one embodiment, the type of camera used may vary depending on vehicle 1100. In at least one embodiment, any combination of camera types may be used to provide the required field of view around vehicle 1100. In at least one embodiment, the number of cameras deployed may vary depending on the embodiment. For example, in at least one embodiment, vehicle 1100 may include 6 cameras, 7 cameras, 10 cameras, 12 cameras, or another number of cameras. In at least one embodiment, the cameras may support Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet® communication, by way of non-limiting example. In at least one embodiment, each camera may be as further described herein in more detail above with respect to FIGS. 11A and 11B.
[0146] In at least one embodiment, vehicle 1100 may further include vibration sensor 1142. In at least one embodiment, vibration sensor 1142 may measure the vibration of a component of vehicle 1100, such as an axle. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 1142 are used, a difference in vibration may be used to determine the amount of friction or slipperiness of the road surface (e.g., if there is a vibration difference between a power-driven axle and a freely rotating axle).
[0147] In at least one embodiment, vehicle 1100 may include an ADAS system 1138. In at least one embodiment, the ADAS system 1138 may include, without limitation, a SoC in some examples. In at least one embodiment, the ADAS system 1138 may include, without limitation, any number and any combination of autonomous / adaptive / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward crash warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functions.
[0148] In at least one embodiment, the ACC system may use a RADAR sensor 1160, a LIDAR sensor 1164, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to another vehicle directly in front of vehicle 1100 and automatically adjusts the speed of vehicle 1100 to maintain a safe distance from the vehicle in front. In at least one embodiment, the lateral ACC system performs distance maintenance and notifies vehicle 1100 to change lanes when necessary. In at least one embodiment, lateral ACC is related to other ADAS applications such as LC and CW.
[0149] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles by a wireless link or indirectly via a network connection (e.g., via the Internet) to network interface 1124 and / or wireless antenna 1126. In at least one embodiment, a direct link may be provided by a vehicle-to-vehicle (V2V) communication link, while an indirect link may be provided by an infrastructure-to-vehicle (I2V) communication link. Generally, V2V communication provides information about the vehicle immediately preceding (e.g., a vehicle in the same lane immediately in front of vehicle 1100), and I2V communication provides information about the traffic further ahead. In at least one embodiment, the CACC system may include either or both of the I2V and V2V information sources. In at least one embodiment, having information about the vehicle in front of vehicle 1100 can further enhance the reliability of the CACC system, make the traffic flow smoother, and have the potential to reduce congestion on the road.
[0150] In at least one embodiment, the FCW system is designed to advise the driver of a hazard, whereby such driver can take corrective action. In at least one embodiment, the FCW system uses a front camera and / or RADAR sensor 1160, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibration component. In at least one embodiment, the FCW system may provide a warning in the form of sound, visual warning, vibration, and / or a quick brake pulse.
[0151] In at least one embodiment, the AEB system may detect an imminent frontal collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front camera and / or RADAR sensor 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system typically first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the predicted collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-collision braking.
[0152] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as vibrations of the steering wheel or seat, to advise the driver when the vehicle 1100 crosses a lane marking. In at least one embodiment, the LDW system does not operate when the driver indicates an intentional lane departure, such as by activating the turn indicator. In at least one embodiment, the LDW system may use a front camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC that can be electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibration component. In at least one embodiment, the LKA system is a variant of the LDW system. In at least one embodiment, the LKA system provides steering input or brake control to correct the vehicle 1100 when the vehicle 1100 begins to veer out of its lane.
[0153] In at least one embodiment, the BSW system detects vehicles in the blind spot of the vehicle and warns the driver. In at least one embodiment, the BSW system may provide visual, auditory, and / or tactile alerts to indicate that a merge or lane change is not safe. In at least one embodiment, the BSW system may provide additional warnings when the driver uses the turn indicator. In at least one embodiment, the BSW system may use a rear camera and / or RADAR sensor 1160 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibration component.
[0154] In at least one embodiment, the RCTW system may provide visual, auditory, and / or tactile notifications when an object is detected outside the range of the rear camera when the vehicle 1100 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear RADAR sensors 1160, which are electrically coupled to a dedicated processor, DSP, FPGA, and / or ASIC that provides feedback to the driver, such as a display, speaker, and / or vibration component.
[0155] In at least one embodiment, conventional ADAS systems may produce false detection results, which can be annoying and distracting to the driver, but usually are not a big deal. This is because conventional ADAS systems advise the driver and enable the driver to determine whether a safety-critical condition actually exists and respond appropriately. In at least one embodiment, when the results conflict, the vehicle 1100 itself determines whether to follow the results from the primary computer (e.g., the first controller among the controllers 1136) or the results from the secondary computer (e.g., the second controller among the controllers 1136). For example, in at least one embodiment, the ADAS system 1138 may be a backup and / or secondary computer for resisting perception information to the rationality module of the backup computer. In at least one embodiment, the rationality monitor of the backup computer may execute various software for redundancy on the hardware components to detect perception errors and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1138 may be provided to the monitoring MCU. In at least one embodiment, when the output from the primary computer conflicts with the output from the secondary computer, the monitoring MCU determines how to reconcile the conflict to ensure safe operation.
[0156] In at least one embodiment, the primary computer may be configured to provide a reliability score indicating the reliability of the selected result of the primary computer to the monitoring MCU. In at least one embodiment, if the reliability score exceeds a threshold, the monitoring MCU may follow the instructions of the primary computer regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, if the reliability score does not meet the threshold and the primary computer and the secondary computer show different results (e.g., conflict), the monitoring MCU may mediate between the computers to determine an appropriate result.
[0157] In at least one embodiment, the monitoring MCU may be configured to execute a neural network trained and configured to determine, at least in part, based on the output from the primary computer and the output from the secondary computer, the conditions under which the secondary computer provides a false alarm. In at least one embodiment, the neural network of the monitoring MCU may learn when the output of the secondary computer may be trusted and when it may not be trusted. For example, in at least one embodiment, if the secondary computer is a RADAR-based FCW system, the neural network of the monitoring MCU may learn when the FCW system identifies a metallic object, such as a drain grate or manhole cover, that is not actually a hazard and triggers an alarm. In at least one embodiment, if the secondary computer is a camera-based LDW system, the neural network of the monitoring MCU may learn to disable the LDW when there are bicycles or pedestrians present and lane departure is actually the safest operation. In at least one embodiment, the monitoring MCU may include at least one of a DLA or GPU suitable for executing the neural network along with an associated memory. In at least one embodiment, the monitoring MCU may comprise and / or be included as a component of the SoC1104.
[0158] In at least one embodiment, the ADAS system 1138 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then rules), and the presence of the neural network in the monitoring MCU may improve reliability, safety, and performance. For example, in at least one embodiment, due to various implementations and intentional non-identities, the overall error tolerance of the system may be increased, particularly for errors caused by the functions of software (or the software-hardware interface). For example, in at least one embodiment, if there is a bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides a consistent overall result, the monitoring MCU may have a higher level of confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.
[0159] In at least one embodiment, the output of the ADAS system 1138 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 1138 indicates a forward collision warning due to an object immediately ahead, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have a trained, and thus error-detection-risk-reducing, unique neural network as described herein.
[0160] In at least one embodiment, vehicle 1100 may further include an infotainment SoC 1130 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as an SoC, the infotainment system SoC 1130 may, in at least one embodiment, not be an SoC and may include two or more individual components without limitation. In at least one embodiment, the infotainment SoC 1130 may include, without limitation, a combination of hardware and software, and this combination may be used to provide the vehicle 1100 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connection (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, wireless data system, vehicle-related information such as fuel level, total mileage, brake fuel level, oil level, door opening and closing, air filter information, etc.). For example, the infotainment SoC 1130 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connections, a carputer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, a heads-up display ("HUD"), an HMI display 1134, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, information from the ADAS system 1138, autonomous driving information such as vehicle operation plans, trajectories, etc., ambient environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information such as (e.g., visual and / or auditory) information may further be provided to the user of the vehicle 1100 using the infotainment SoC 1130.
[0161] In at least one embodiment, the infotainment SoC 1130 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1130 may communicate with other devices, systems, and / or components of the vehicle 1100 via the bus 1102. In at least one embodiment, the infotainment SoC 1130 may be coupled to a monitoring MCU, such that when the primary controller 1136 (e.g., the primary and / or backup computer of the vehicle 1100) fails, the GPU of the infotainment system may execute some self-driving functions. In at least one embodiment, the infotainment SoC 1130 may put the vehicle 1100 into a driver-safety stop mode, as described herein.
[0162] In at least one embodiment, the vehicle 1100 may further include an instrument cluster 1132 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, the instrument cluster 1132 may include, without limitation, a controller and / or a supercomputer (e.g., an individual controller or supercomputer). In at least one embodiment, the instrument cluster 1132 may include any number and combination of instrument sets, without limitation, a speedometer, a fuel level, a hydraulic pressure, a tachometer, an odometer, a direction indicator, a shift lever position indicator, a seat belt warning light, a parking brake warning light, an engine failure light, an auxiliary restraint system (e.g., an airbag) information, a light control, a safety system control, a navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1130 and the instrument cluster 1132. In at least one embodiment, the instrument cluster 1132 may be included as part of the infotainment SoC 1130, or vice versa.
[0163] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of FIG. 11C for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of neural networks.
[0164] In at least one embodiment, one or more systems shown in FIG. 11C are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 11C are utilized to perform various processes such as those described in connection with FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 11C are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0165] FIG. 11D is a diagram of a system for communicating between a cloud-based server and the autonomous vehicle 1100 of FIG. 11A according to at least one embodiment. In at least one embodiment, the system may include, without limitation, the server 1178, the network 1190, and any number and type of vehicles including the vehicle 1100. In at least one embodiment, the server 1178 may include, without limitation, a plurality of GPUs 1184(A)-1184(H) (collectively referred to herein as GPU 1184), PCIe switches 1182(A)-1182(D) (collectively referred to herein as PCIe switch 1182), and / or CPUs 1180(A)-1180(B) (collectively referred to herein as CPU 1180). In at least one embodiment, the GPUs 1184, CPUs 1180, and PCIe switches 1182 may be interconnected by high-speed interconnects such as, without limitation, the NVLink interface 1188 developed by NVIDIA and / or the PCIe connection 1186. In at least one embodiment, the GPUs 1184 are connected via NVLink and / or an NVS switch SoC, and the GPU 1184 and the PCIe switch 1182 are connected via a PCIe interconnect. Eight GPUs 1184, two CPUs 1180, and four PCIe switches 1182 are shown, but this is not limiting. In at least one embodiment, each of the servers 1178 may include, without limitation, any number of GPUs 1184, CPUs 1180, and / or PCIe switches 1182 in any combination. For example, in at least one embodiment, the server 1178 may include, respectively, 8, 16, 32, and / or more than 32 GPUs 1184.
[0166] In at least one embodiment, server 1178 may receive, via network 1190, image data representing an image indicating an unexpected or changed road condition, such as a recently started road construction, from a vehicle. In at least one embodiment, server 1178 may transmit, via network 1190, neural network 1192, an updated or other form of neural network 1192, and / or map information 1194 including information regarding traffic conditions and road conditions, without limitation, to a vehicle. In at least one embodiment, the update of map information 1194 may include, without limitation, updates to HD map 1122, such as information regarding construction sites, holes, detours, floods, and / or other obstacles. In at least one embodiment, neural network 1192 and / or map information 1194 may be obtained from new training and / or experience represented by data received from any number of vehicles in the environment, and / or may be obtained based at least in part on training performed at a data center (e.g., using server 1178 and / or other servers).
[0167] In at least one embodiment, using server 1178, a machine learning model (e.g., a neural network) may be trained based at least in part on training data. In at least one embodiment, the training data may be generated by a vehicle and / or may be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged and / or undergoes other preprocessing (e.g., if the associated neural network benefits from supervised learning). In at least one embodiment, any amount of training data is not tagged and / or preprocessed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by a vehicle (e.g., transmitted to the vehicle via network 1190) and / or the machine learning model may be used by server 1178 to remotely monitor the vehicle.
[0168] In at least one embodiment, server 1178 may receive data from a vehicle and apply the data to a state-of-the-art real-time neural network to enable real-time intelligent inference. In at least one embodiment, server 1178 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1184, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, server 1178 may include a deep learning infrastructure that uses a data center powered by a CPU.
[0169] In at least one embodiment, the deep learning infrastructure of server 1178 may be capable of high-speed real-time inference and may use that capability to evaluate and verify the health of the processor, software, and / or associated hardware of vehicle 1100. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1100, such as a series of images and / or objects located by vehicle 1100 in that series of images (e.g., by computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify an object and compare it to the object identified by vehicle 1100. If the results do not match and the deep learning infrastructure concludes that the AI of vehicle 1100 is malfunctioning, server 1178 may take control of the fail-safe computer of vehicle 1100, notify the occupants, and send a signal to vehicle 1100 commanding it to complete a safe parking operation.
[0170] In at least one embodiment, server 1178 may include GPU 1184 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT3 device). In at least one embodiment, by combining a server powered by a GPU with inference acceleration, real-time response can be enabled. In at least one embodiment, a server powered by a CPU, FPGA, and other processors may be used for inference, such as when performance is not as critical. In at least one embodiment, hardware structure 815 is used to execute one or more embodiments. Details regarding hardware structure 815 are provided herein in conjunction with FIGS. 8A and / or 8B.
[0171] Computer system FIG. 12 is a block diagram showing an exemplary computer system, which may be formed with a processor that may include an execution unit for executing instructions, interconnected devices and components, a system-on-chip (SoC), or some combination thereof, according to at least one embodiment. In at least one embodiment, computer system 1200 may include components such as processor 1202 for using an execution unit that includes logic for executing an algorithm for processing data in accordance with the present disclosure, such as in the embodiments described herein, without limitation. In at least one embodiment, computer system 1200 may include a processor such as a PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessor available from Intel Corporation of Santa Clara, California, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 1200 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.
[0172] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (DSP), a system-on-a-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of executing one or more instructions according to at least one embodiment.
[0173] In at least one embodiment, computer system 1200 may include, without limitation, a processor 1202, which may include, without limitation, one or more execution units 1208 for performing training and / or inference of a machine learning model by the techniques described herein. In at least one embodiment, computer system 1200 is a single-processor desktop or server system, although in another embodiment, computer system 1200 may be a multi-processor system. In at least one embodiment, processor 1202 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor, etc. In at least one embodiment, processor 1202 may be coupled to a processor bus 1210, which may transmit digital signals between processor 1202 and other components within computer system 1200.
[0174] In at least one embodiment, processor 1202 may include, without limitation, a level 1 (L1) internal cache memory (cache) 1204. In at least one embodiment, processor 1202 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to processor 1202. Other embodiments may also include a combination of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 1206 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.
[0175] In at least one embodiment, the execution unit 1208, which includes without limitation logic for performing integer and floating point operations, is also in the processor 1202. In at least one embodiment, the processor 1202 may also include a microcode (u-code) read only memory (ROM) that stores microcode for certain macro instructions. In at least one embodiment, the execution unit 1208 may include logic for handling a packed instruction set 1209. In at least one embodiment, by including the packed instruction set 1209 in the instruction set of a general purpose processor along with the associated circuitry for executing the instructions, operations used by many multimedia applications can be executed using the packed data of the processor 1202. In at least one embodiment, by performing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.
[0176] In at least one embodiment, the execution unit 1208 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, the computer system 1200 may include a memory 1220 without limitation. In at least one embodiment, the memory 1220 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or another memory device. In at least one embodiment, the memory 1220 may store instructions 1219 and / or data 1221 represented by data signals that may be executed by the processor 1202.
[0177] In at least one embodiment, a system logic chip may be coupled to a processor bus 1210 and a memory 1220. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub (“MCH”) 1216, and the processor 1202 may communicate with the MCH 1216 via the processor bus 1210. In at least one embodiment, the MCH 1216 may provide a high-bandwidth memory path 1218 to the memory 1220 for storing instructions and data and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 1216 may direct data signals between the processor 1202, the memory 1220, and other components of the computer system 1200 and may bridge data signals between the processor bus 1210, the memory 1220, and the system I / O interface 1222. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 1216 may be coupled to the memory 1220 via the high-bandwidth memory path 1218, and the graphics / video card 1212 may be coupled to the MCH 1216 via an accelerated graphics port (“AGP”) interconnect 1214.
[0178] In at least one embodiment, computer system 1200 may use a system I / O interface 1222, which is a proprietary hub interface bus for coupling MCH 1216 to an I / O controller hub (ICH) 1230. In at least one embodiment, ICH 1230 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 1220, the chipset, and processor 1202. By way of example, it may include, without limitation, an audio controller 1229, a firmware hub (“flash BIOS”) 1228, a wireless transceiver 1226, data storage 1224, a legacy I / O controller 1223 including a user input and keyboard interface 1225, a serial expansion port such as a universal serial bus (“USB”) port 1227, and a network controller 1234. In at least one embodiment, data storage 1224 may comprise a hard disk drive, a floppy (registered trademark) disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0179] In at least one embodiment, FIG. 12 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 12 may show an exemplary system-on-a-chip (SoC). In at least one embodiment, the devices shown in FIG. 12 may be interconnected by proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1200 may be interconnected using a Compute Express Link (CXL) interconnect.
[0180] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of FIG. 12 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the use cases of the neural networks.
[0181] In at least one embodiment, one or more systems shown in FIG. 12 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 12 are utilized to perform various processes such as those described in relation to FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 12 are utilized to estimate parameters as part of one or more training processes of a neural network model, using a subset of training data, such as training images, based on the uniqueness of the training data.
[0182] FIG. 13 is a block diagram showing an electronic device 1300 for utilizing a processor 1310, according to at least one embodiment. In at least one embodiment, the electronic device 1300 may be, for example and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0183] In at least one embodiment, the electronic device 1300 may include, without limitation, a processor 1310 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1310 is I 2 coupled using a bus or interface such as an I2C bus, a System Management Bus (SMBus), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (SPI), a High Definition Audio (HDA) bus, a Serial Advance Technology Attachment (SATA) bus, a Universal Serial Bus (USB) (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (UART) bus. In at least one embodiment, FIG. 13 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 13 may show an exemplary System-on-Chip (SoC). In at least one embodiment, the devices shown in FIG. 13 may be interconnected using proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 13 may be interconnected using a Compute Express Link (CXL) interconnect.
[0184] In at least one embodiment, FIG. 13 shows a display 1324, a touch screen 1325, a touch pad 1330, a Near Field Communications unit (NFC) 1345, a sensor hub 1340, a thermal sensor 1346, an Express Chipset (EC) 1335, a Trusted Platform Module (TPM) 1338, a BIOS / firmware / flash memory (BIOS, FW flash) 1322, a DSP 1360, a drive 1320 such as a Solid State Disk (SSD) or a Hard Disk Drive (HDD), a wireless local area network unit (WLAN) 1350, a Bluetooth unit 1352, a Wireless Wide Area Network unit (WWAN) 1356, a Global Positioning System (GPS) unit 1355, a camera such as a USB3.0 camera (USB3.0 camera) 1354, and / or a Low Power Double Data Rate (LPDDR) memory unit (LPDDR3) 1315 implemented, for example, in accordance with the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0185] In at least one embodiment, other components may be communicatively coupled to the processor 1310 via the components described herein. In at least one embodiment, an accelerometer 1341, an ambient light sensor (“ALS”), a compass 1343, and a gyroscope 1344 may be communicatively coupled to a sensor hub 1340. In at least one embodiment, a thermal sensor 1339, a fan 1337, a keyboard 1336, and a touch pad 1330 may be communicatively coupled to an EC 1335. In at least one embodiment, a speaker 1363, headphones 1364, and a microphone (“mic”) 1365 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 1362, and this audio unit may be communicatively coupled to a DSP 1360. In at least one embodiment, the audio unit 1362 may include, without limitation, for example, an audio coder / decoder (“codec”) and a class D amplifier. In at least one embodiment, a SIM card (“SIM”) 1357 may be communicatively coupled to a WWAN unit 1356. In at least one embodiment, components such as a WLAN unit 13650 and a Bluetooth unit 1352, as well as the WWAN unit 1356, may be implemented in a next generation form factor (“NGFF”).
[0186] In order to perform inference and / or training operations associated with one or more embodiments, an inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of FIG. 13 for inference or prediction operations, at least partially based on the training operations of the neural network, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network described herein.
[0187] In at least one embodiment, one or more systems shown in FIG. 13 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 13 are utilized to execute various processes, such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 13 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data, using a subset of the training data, such as training images.
[0188] FIG. 14 shows a computer system 1400 according to at least one embodiment. In at least one embodiment, the computer system 1400 is configured to implement the various processes and methods described throughout this disclosure.
[0189] In at least one embodiment, computer system 1400 includes, without limitation, at least one central processing unit (“CPU”) 1402, which is connected to a communication bus 1410 implemented using any suitable protocol, such as, without limitation, PCI: Peripheral Component Interconnect (“Peripheral Component Interconnect”), Peripheral Component Interconnect Express (“PCI-Express”: peripheral component interconnect express), AGP: Accelerated Graphics Port (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1400 includes, without limitation, main memory 1404 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in main memory 1404, which may take the form of random access memory (“RAM”: random access memory). In at least one embodiment, a network interface subsystem (“network interface”) 1422 provides an interface with other computing devices and networks for receiving data from other systems and transmitting data from computer system 1400 to other systems.
[0190] In at least one embodiment, computer system 1400 includes, without limitation in at least one embodiment, an input device 1408, a parallel processing system 1412, and a display device 1406, and this display device can be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light emitting diode ("LED") display, a plasma display, or other suitable display technology. In at least one embodiment, user input is received from an input device 1408 such as a keyboard, a mouse, a touch pad, a microphone, etc. In at least one embodiment, each module described herein can be placed on a single semiconductor platform to form a processing system.
[0191] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIG. 8A and / or FIG. 8B. In at least one embodiment, inference and / or training logic 815 may be used in the system of FIG. 14 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0192] In at least one embodiment, one or more systems shown in FIG. 14 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 14 are utilized to execute various processes such as those described in relation to FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 14 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0193] FIG. 15 shows a computer system 1500 according to at least one embodiment. In at least one embodiment, the computer system 1500 may include, without limitation, a computer 1510 and a USB stick 1520. In at least one embodiment, the computer 1510 may include, without limitation, any number and type of processors (not shown), as well as memory (not shown). In at least one embodiment, the computer 1510 includes, without limitation, servers, cloud instances, laptops, and desktop computers.
[0194] In at least one embodiment, the USB stick 1520 includes, without limitation, a processing unit 1530, a USB interface 1540, and USB interface logic 1550. In at least one embodiment, the processing unit 1530 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1530 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing unit 1530 comprises an application-specific integrated circuit ("ASIC") optimized to execute any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing unit 1530 is a tensor processing unit ("TPC") optimized to execute inference operations of machine learning. In at least one embodiment, the processing unit 1530 is a vision processing unit ("VPU") optimized to execute inference operations of machine vision and machine learning.
[0195] In at least one embodiment, the USB interface 1540 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1540 is a USB3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1540 is a USB3.0 Type-A connector. In at least one embodiment, the USB interface logic 1550 may include any amount and type of logic that enables the processing unit 1530 to interface with a device (e.g., computer 1510) via the USB connector 1540.
[0196] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the system of FIG. 15 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or use cases of the neural networks.
[0197] In at least one embodiment, one or more systems shown in FIG. 15 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 15 are utilized to perform various processes such as those described in connection with FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 15 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data, using a subset of the training data such as training images.
[0198] FIG. 16A shows an exemplary architecture in which a plurality of GPUs 1610(1) - 1610(N) are communicatively coupled to a plurality of multi - core processors 1605(1) - 1605(M) via high - speed links 1640(1) - 1640(N) (e.g., buses, point - to - point interconnects, etc.). In at least one embodiment, the high - speed links 1640(1) - 1640(N) support a communication throughput of 4GB / second, 30GB / second, 80GB / second, or more. In at least one embodiment, various interconnect protocols may be used, including but not limited to PCIe4.0 or 5.0, and NVLink2.0. In various drawings, "N" and "M" represent positive integers, and their values may vary from drawing to drawing.
[0199] Furthermore, in at least one embodiment, two or more of the GPUs 1610 are interconnected via high-speed links 1629(1)-1629(2), which may be implemented using a protocol / link similar to or different from that used for the high-speed links 1640(1)-1640(N). Similarly, two or more of the multi-core processors 1605 may be connected via a high-speed link 1628, which can be a symmetric multi-processor (SMP) bus operating at 20 GB / second, 30 GB / second, 120 GB / second, or higher. Alternatively, all communication between the various system components shown in FIG. 16A may be realized using a similar protocol / link (e.g., via a common interconnect fabric).
[0200] In one embodiment, each multi-core processor 1605 is communicatively coupled to processor memories 1601(1)-1601(M) via memory interconnects 1626(1)-1626(M), respectively, and each GPU 1610(1)-1610(N) is communicatively coupled to GPU memories 1620(1)-1620(N) via GPU memory interconnects 1650(1)-1650(N), respectively. In at least one embodiment, the memory interconnects 1626 and 1650 may utilize similar or different memory access technologies. By way of example, and not limitation, the processor memories 1601(1)-1601(M) and the GPU memories 1620 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or non-volatile memories such as 3D XPoint or Nano-Ram. In at least one embodiment, some portions of the processor memory 1601 may be volatile memories and other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).
[0201] As described herein, the various multi-core processors 1605 and GPUs 1610 may each be physically coupled to a specific memory 1601, 1620, and / or an integrated memory architecture may be implemented in which the virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, each of the processor memories 1601(1) - 1601(M) may comprise 64 GB of system memory address space, and each of the GPU memories 1620(1) - 1620(N) may comprise 32 GB of system memory address space, such that in this example, a total of 256 GB of addressable memory is obtained when M = 2 and N = 4. Other values of N and M are possible.
[0202] FIG. 16B shows further details of the interconnection of a multi-core processor 1607 and a graphics acceleration module 1646 according to one exemplary embodiment. In at least one embodiment, the graphics acceleration module 1646 may include one or more GPU chips integrated on a line card coupled to the processor 1607 via a high-speed link 1640 (e.g., a PCIe bus, NVLink, etc.). In at least one embodiment, the graphics acceleration module 1646 may alternatively be integrated into a package or chip with the processor 1607.
[0203] In at least one embodiment, the processor 1607 includes a plurality of cores 1660A - 1660D, each core having a translation lookaside buffer ( "TLB") 1661A - 1661D and one or more caches 1662A - 1662D. In at least one embodiment, the cores 1660A - 1660D may include various other components (not shown) for executing instructions and processing data. In at least one embodiment, the caches 1662A - 1662D may comprise level 1 (L1) and level 2 (L2) caches. Further, one or more shared caches 1656 may be included in the caches 1662A - 1662D and shared by a set of the cores 1660A - 1660D. For example, one embodiment of the processor 1607 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more of the L2 and L3 caches are shared by two adjacent cores. In at least one embodiment, the processor 1607 and the graphics acceleration module 1646 are connected to the system memory 1614, which may include the processor memories 1601(1) - 1601(M) of FIG. 16A.
[0204] In at least one embodiment, for the data and instructions stored in the various caches 1662A - 1662D, 1656, and the system memory 1614, coherence is maintained by inter - core communication via the coherence bus 1664. In at least one embodiment, for example, each cache may have associated cache coherence logic / circuitry to communicate via the coherence bus 1664 in response to detecting a read or write to a particular cache line. In at least one embodiment, a cache snooping protocol is implemented via the coherence bus 1664 to monitor cache accesses.
[0205] In at least one embodiment, the proxy circuit 1625 communicatively couples the graphics acceleration module 1646 to the coherence bus 1664 such that the graphics acceleration module 1646 can participate in the cache coherence protocol as a peer of cores 1660A - 1660D. In particular, in at least one embodiment, the interface 1635 provides a connection to the proxy circuit 1625 via the high - speed link 1640, and the interface 1637 connects the graphics acceleration module 1646 to the high - speed link 1640.
[0206] In at least one embodiment, the accelerator integration circuit 1636 provides services for cache management, memory access, content management, and interrupt management instead of the plurality of graphics processing engines 1631(1)-1631(N) of the graphics acceleration module 1646. In at least one embodiment, each of the graphics processing engines 1631(1)-1631(N) may comprise a separate graphics processing unit (GPU). In at least one embodiment, alternatively, the graphics processing engines 1631(1)-1631(N) may comprise different types of graphics processing engines, such as graphics execution units, media processing engines (e.g., video encoder / decoder), samplers, and blit engines, within a GPU. In at least one embodiment, the graphics acceleration module 1646 may be a GPU having a plurality of graphics processing engines 1631(1)-1631(N), or the graphics processing engines 1631(1)-1631(N) may be individual GPUs integrated in a common package, line card, or chip.
[0207] In at least one embodiment, the accelerator integration circuit 1636 includes a memory management unit (MMU) 1639 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 1614. In at least one embodiment, the MMU 1639 can also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In at least one embodiment, the cache 1638 can store commands and data for efficient access by the graphics processing engines 1631(1)-1631(N). In at least one embodiment, the data stored in the cache 1638 and the graphics memories 1633(1)-1633(M) are kept coherent with the core caches 1662A-1662D, 1656 and the system memory 1614, optionally using the fetch unit 1644. As described, this may be achieved via the proxy circuit 1625 instead of the cache 1638 and the memories 1633(1)-1633(M) (e.g., sending updates regarding cache line modifications / accesses in the processor caches 1662A-1662D, 1656 to the cache 1638 and receiving updates from the cache 1638).
[0208] In at least one embodiment, a set of registers 1645 stores context data for threads executed by graphics processing engines 1631(1) to 1631(N), and a context management circuit 1648 manages thread contexts. For example, the context management circuit 1648 may perform save and restore operations to save and restore the contexts of various threads during a context switch (e.g., here, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1648 may store the current register values in a specified area of memory (identified, for example, by a context pointer). Then, when returning to the context, the context management circuit 2248 may restore the register values. In at least one embodiment, an interrupt management circuit 1647 receives and processes interrupts received from system devices.
[0209] In at least one embodiment, virtual / effective addresses from the graphics processing engine 1631 are translated by the MMU 1639 into real / physical addresses of the system memory 1614. In at least one embodiment, the accelerator integration circuit 1636 supports a plurality (e.g., 4, 8, 16) of graphics accelerator modules 1646 and / or other accelerator devices. In at least one embodiment, the graphics accelerator module 1646 may be dedicated to a single application executed on the processor 1607 or may be shared among multiple applications. In at least one embodiment, there is a virtualized graphics execution environment in which the resources of the graphics processing engines 1631(1) to 1631(N) are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0210] In at least one embodiment, the accelerator integration circuit 1636 functions as a bridge to the system for the graphics acceleration module 1646 and provides address translation and cache services for system memory. Further, in at least one embodiment, the accelerator integration circuit 1636 may provide virtualization facilities for the host processor to manage the virtualization, interrupts, and memory management of the graphics processing engines 1631(1) to 1631(N).
[0211] In at least one embodiment, the hardware resources of the graphics processing engines 1631(1) to 1631(N) are explicitly mapped to the physical address space seen by the host processor 1607, so that any host processor can directly address these resources using the effective address value. In at least one embodiment, one function of the accelerator integration circuit 1636 is to physically separate the graphics processing engines 1631(1) to 1631(N) so that they appear as independent units to the system.
[0212] In at least one embodiment, each of the one or more graphics memories 1633(1) to 1633(M) is coupled to each of the graphics processing engines 1631(1) to 1631(N), and N = M. In at least one embodiment, the graphics memories 1633(1) to 1633(M) store the instructions and data being processed by the respective graphics processing engines 1631(1) to 1631(N). In at least one embodiment, the graphics memories 1633(1) to 1633(M) may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as 3D XPoint or Nano-Ram.
[0213] In at least one embodiment, a biasing technique is used to reduce data traffic through the high-speed link 1640 such that the data stored in the graphics memories 1633(1) to 1633(M) is data that will be most frequently used by the graphics processing engines 1631(1) to 1631(N), and preferably data that is not (at least frequently) used by the cores 1660A to 1660D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep data that the cores need (and thus preferably that the graphics processing engines 1631(1) to 1631(N) do not need) in the caches 1662A to 1662D, 1656, and the system memory 1614.
[0214] FIG. 16C shows another exemplary embodiment in which the accelerator integration circuit 1636 is integrated within the processor 1607. At least in this embodiment, the graphics processing engines 1631(1) to 1631(N) communicate directly with the accelerator integration circuit 1636 via the high-speed link 1640 by way of the interface 1637 and the interface 1635 (which may also be any form of bus or interface protocol in this case). In at least one embodiment, the accelerator integration circuit 1636 may perform operations similar to those described with respect to FIG. 16B, but potentially operate at a higher throughput considering its proximity to the coherence bus 1664 and the caches 1662A to 1662D, 1656. In at least one embodiment, the accelerator integration circuit supports different programming models including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1636 and a programming model controlled by the graphics acceleration module 1646.
[0215] In at least one embodiment, the graphics processing engines 1631(1) to 1631(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can concentrate other application requirements on the graphics processing engines 1631(1) to 1631(N) to implement virtualization within a VM / partition.
[0216] In at least one embodiment, the graphics processing engines 1631(1) to 1631(N) may be shared by a plurality of VM / application partitions. In at least one embodiment, the shared model may use a system hypervisor to virtualize the graphics processing engines 1631(1) to 1631(N) to enable access by each operating system. In at least one embodiment, in a single partition system without a hypervisor, the graphics processing engines 1631(1) to 1631(N) are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1631(1) to 1631(N) to provide access to each process or application.
[0217] In at least one embodiment, the graphics acceleration module 1646 or individual graphics processing engines 1631(1)-1631(N) select process elements using a process handle. In at least one embodiment, the process elements are stored in the system memory 1614 and are addressable using the translation techniques from virtual addresses to physical addresses described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the context of the host process with the graphics processing engines 1631(1)-1631(N) (i.e., calling system software to add the process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element within the process element link list.
[0218] FIG. 16D shows an exemplary accelerator integration slice 1690. In at least one embodiment, a "slice" comprises a designated portion of the processing resources of the accelerator integration circuit 1636. In at least one embodiment, an application virtual address space 1682 within the system memory 1614 stores process elements 1683. In at least one embodiment, the process elements 1683 are stored in response to GPU calls 1681 from an application 1680 executing on the processor 1607. In at least one embodiment, the process elements 1683 accommodate the process state of the corresponding application 1680. In at least one embodiment, the work descriptor (WD) 1684 contained in the process element 1683 can be a single job requested by the application or can accommodate a pointer to a queue of jobs. In at least one embodiment, the WD 1684 is a pointer to a job request queue in the application's virtual address space 1682.
[0219] In at least one embodiment, the graphics acceleration module 1646 and / or the individual graphics processing engines 1631(1)-1631(N) can be shared by all or a subset of the processes within the system. In at least one embodiment, an infrastructure for setting a process state and sending the WD1684 to the graphics acceleration module 1646 to initiate a job in a virtualized environment may be included.
[0220] In at least one embodiment, the dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns the graphics acceleration module 1646 or an individual graphics processing engine 1631. In at least one embodiment, when the graphics acceleration module 1646 is owned by a single process, when the graphics acceleration module 1646 is allocated, the hypervisor initializes the accelerator integration circuit 1636 for the owning partition, and the operating system initializes the accelerator integration circuit 1636 for the owning process.
[0221] In at least one embodiment, during operation, the WD fetch unit 1691 within the accelerator integration slice 1690 fetches the next WD 1684, including the display of work to be performed by one or more graphics processing engines of the graphics acceleration module 1646. In at least one embodiment, as shown, the data from the WD 1684 is stored in the register 1645 and may be used by the MMU 1639, the interrupt management circuit 1647, and / or the context management circuit 1648. For example, one embodiment of the MMU 1639 includes a segment / page walk circuit for accessing the segment / page table 1686 within the OS virtual address space 1685. In at least one embodiment, the interrupt management circuit 1647 may process the interrupt event 1692 received from the graphics acceleration module 1646. In at least one embodiment, when executing a graphics operation, the effective address 1693 generated by the graphics processing engines 1631(1)-1631(N) is translated to a physical address by the MMU 1639.
[0222] In at least one embodiment, the register 1645 is replicated for each of the graphics processing engines 1631(1)-1631(N) and / or the graphics acceleration module 1646 and may be initialized by the hypervisor or the operating system. In at least one embodiment, each of these replicated registers may be included in the accelerator integration slice 1690. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.
Table 1
[0223] Exemplary registers that may be initialized by the operating system are shown in Table 2.
Table 2
[0224] In one embodiment, each WD1684 is specific to a particular graphics acceleration module 1646 and / or graphics processing engines 1631(1) to 1631(N). In at least one embodiment, the WD1684 can contain all the information required for the graphics processing engines 1631(1) to 1631(N) to perform work, or can be a pointer to a memory location where the application has set up a command queue for the work to be completed.
[0225] FIG. 16E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor physical address space 1698 in which a process element list 1699 is stored. In at least one embodiment, the hypervisor physical address space 1698 is accessible via a hypervisor 1696 that virtualizes the graphics acceleration module engine of the operating system 1695.
[0226] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1646. In at least one embodiment, there are two programming models in which the graphics acceleration module 1646 is shared by multiple processes and partitions, namely, time slice sharing and graphics-directed shared.
[0227] In at least one embodiment, in this model, the system hypervisor 1696 owns the graphics acceleration module 1646 and makes its functions available to all operating systems 1695. In at least one embodiment, for the graphics acceleration module 1646 to support the virtualization by the system hypervisor 1696, the graphics acceleration module 1646 shall comply with specific requirements such as (1) the job requests of the application must be autonomous (i.e., there is no need to maintain state between jobs), or the graphics acceleration module 1646 shall provide a context save and restore mechanism; (2) the job requests of the application shall be guaranteed by the graphics acceleration module 1646 to complete within a specified amount of time including any translation errors, or the graphics acceleration module 1646 shall provide a function to preempt the processing of the job; and (3) when the graphics acceleration module 1646 is operating in a specified shared programming model, fairness shall be guaranteed between processes.
[0228] In at least one embodiment, the application 1680 needs to make a system call to the operating system 1695 with the type of the graphics acceleration module, the work descriptor (WD), the authority mask register (AMR) value, and the context save / restore area pointer (CSRP). In at least one embodiment, the type of the graphics acceleration module describes the acceleration function targeted by the system call. In at least one embodiment, the type of the graphics acceleration module may be a system-specific value. In at least one embodiment, the WD is specifically formatted for the graphics acceleration module 1646 and can be in the form of an effective address pointer pointing to the command of the graphics acceleration module 1646, a user-defined structure, an effective address pointer pointing to the command queue, or any other data structure for describing the work performed by the graphics acceleration module 1646.
[0229] In at least one embodiment, the AMR value is the AMR state for use in the current process. In at least one embodiment, the value passed to the operating system is the same as the application that sets the AMR. In at least one embodiment, if the embodiments of the accelerator integration circuit 1636 (not shown) and the graphics acceleration module 1646 do not support the user authority mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value and then pass the AMR to the hypervisor call. In at least one embodiment, the hypervisor 1696 may optionally put the AMR into the process element 1683 after applying the current authority mask override register (AMOR) value. In at least one embodiment, the CSRP is one of the registers 1645 that holds the effective address of an area within the effective address space 1682 of the application for the graphics acceleration module 1646 to save and restore the context state. In at least one embodiment, this pointer is optional if there is no need to save any state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.
[0230] Upon receiving a system call, the operating system 1695 may verify that the application 1680 is registered and has the authority to use the graphics acceleration module 1646. In at least one embodiment, the operating system 1695 then calls the hypervisor 1696 with the information shown in Table 3.
Table 3
[0231] In at least one embodiment, upon receiving a hypervisor call, hypervisor 1696 verifies that the operating system 1695 is registered and has been granted the right to use the graphics acceleration module 1646. In at least one embodiment, hypervisor 1696 then inserts process element 1683 into a process element link list of the type of the corresponding graphics acceleration module 1646. In at least one embodiment, the process element may include the information shown in Table 4.
Table 4
[0232] In at least one embodiment, the hypervisor initializes the registers 1645 of the plurality of accelerator integration slices 1690.
[0233] As shown in FIG. 16F, in at least one embodiment, an integrated memory that is addressable via a common virtual memory address space used to access physical processor memories 1601(1) to 1601(N) and GPU memories 1620(1) to 1620(N) is used. In this embodiment, operations executed by GPUs 1610(1) to 1610(N) utilize the same virtual / effective memory address space as accessing the processor memories 1601(1) to 1601(M), and vice versa, thereby simplifying programmability. In at least one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1601(1), a second portion is allocated to a second processor memory 1601(N), a third portion is allocated to GPU memory 1620(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes called the effective address space) is thereby distributed across each of the processor memory 1601 and the GPU memory 1620 such that any processor or GPU can access any physical memory with virtual addresses mapped to the physical memory.
[0234] In at least one embodiment, the bias / coherence management circuits 1694A - 1694E in one or more of MMUs 1639A - 1639E ensure cache coherence between the cache of one or more host processors (e.g., 1605) and the cache of GPU 1610, and implement a bias technique to indicate the physical memory in which a particular type of data should be stored. In at least one embodiment, multiple instances of the bias / coherence management circuits 1694A - 1694E are shown in FIG. 16F, but the bias / coherence circuit may be implemented within the MMU of one or more host processors 1605 and / or within the accelerator integration circuit 1636.
[0235] One embodiment enables the GPU memory 1620 to be mapped as part of the system memory and made accessible using shared virtual memory (SVM) technology, without incurring a performance degradation associated with full system cache coherence. In at least one embodiment, the GPU memory 1620 is accessible as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offloading. In at least one embodiment, this configuration enables the software of the host processor 1605 to set operands and access computation results without the overhead of conventional I / O DMA data copies. In at least one embodiment, such conventional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access the GPU memory 1620 without cache coherence overhead can be essential to the execution time of offloaded computations. In at least one embodiment, for example, in the presence of significant streaming write memory traffic, the cache coherence overhead can significantly reduce the effective write bandwidth seen by the GPU 1610. In at least one embodiment, the efficiency of operand setting, access to results, and GPU computation may help in determining the effectiveness of GPU offloading.
[0236] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. In at least one embodiment, for example, a bias table may be used, which may be a page granularity structure including one or two bits per memory page with a GPU (e.g., may be controlled at the granularity of the memory page). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPU memories 1620 with or without a bias cache (e.g., for caching frequently used / recently used entries of the bias table) in the GPU 1610. Alternatively, in at least one embodiment, the entire bias table may be maintained within the GPU.
[0237] In at least one embodiment, an entry of the bias table associated with each access to the GPU memory 1620 is accessed prior to the actual access to the GPU memory, resulting in the following operations. In at least one embodiment, a local request from the GPU 1610 finding its page within the GPU bias is transferred directly to the corresponding GPU memory 1620. In at least one embodiment, a local request from the GPU finding its page in the host bias is transferred to the processor 1605 (e.g., via the high-speed link described herein). In at least one embodiment, a request from the processor 1605 finding the requested page in the host processor bias completes the request in the same manner as a normal memory read. Alternatively, a request directed to a GPU-biased page may be transferred to the GPU 1610. In at least one embodiment, the GPU may then migrate the page to the host processor bias if the current page is not being used. In at least one embodiment, the bias state of the page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply a hardware-based mechanism.
[0238] In at least one embodiment, one mechanism for changing the bias state utilizes an API call (e.g., OpenCL), where this API call calls the GPU's device driver, and this device driver sends a message to the GPU (or adds a command descriptor to a queue) to change the bias state, and for some transitions, guides the GPU to perform a cache flushing operation on the host. In at least one embodiment, the cache flushing operation is used for the transition from the bias of the host processor 1605 to the GPU bias, but not for the opposite transition.
[0239] In at least one embodiment, cache coherence is maintained by the host processor 1605 temporarily rendering GPU-biased pages that cannot be cached. In at least one embodiment, to access these pages, the processor 1605 may request access from the GPU 1610, and the GPU 1610 may either immediately grant access or not grant access. In at least one embodiment, therefore, it is beneficial to make GPU-biased pages be requested by the GPU but not by the host processor 1605, or vice versa, to reduce communication between the processor 1605 and the GPU 1610.
[0240] A hardware structure 815 is used to execute one or more embodiments. Details regarding the hardware structure 815 may be provided herein in conjunction with FIGS. 8A and / or 8B.
[0241] FIG. 17 shows an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general-purpose processor cores.
[0242] FIG. 17 is a block diagram showing an exemplary system-on-chip integrated circuit 1700 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, integrated circuit 1700 includes one or more application processors 1705 (e.g., CPUs), at least one graphics processor 1710, and may further include an image processor 1715 and / or a video processor 1720, any of which may be a modular IP core. In at least one embodiment, integrated circuit 1700 includes peripheral devices or bus logic including a USB controller 1725, a UART controller 1730, an SPI / SDIO controller 1735, and an I 2 2S / I 2 2C controller 1740. In at least one embodiment, integrated circuit 1700 can include a display device 1745 coupled to one or more of a high-definition multimedia interface (HDMI™) controller 1750 and a mobile industry processor interface (MIPI) display interface 1755. In at least one embodiment, storage may be provided by a flash memory subsystem 1760 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1765 to access a SDRAM or SRAM memory device. In at least one embodiment, some integrated circuits further include an embedded security engine 1770.
[0243] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in integrated circuit 1700 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functions and / or architectures of neural networks described herein, or use cases of neural networks.
[0244] In at least one embodiment, one or more systems shown in FIG. 17 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 17 are utilized to execute various processes such as those described in relation to FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 17 are utilized to estimate parameters as part of one or more training processes of a neural network model, based on the uniqueness of training data, using a subset of training data such as training images.
[0245] FIGS. 18A-18B show exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processor / cores, peripheral device interface controllers, or general purpose processor cores.
[0246] FIG. 18A and FIG. 18B are block diagrams showing exemplary graphics processors for use within a SoC according to the embodiments described herein. FIG. 18A shows an exemplary graphics processor 1810 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. FIG. 18B shows a further exemplary graphics processor 1840 of a system-on-chip integrated circuit that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the graphics processor 1810 of FIG. 18A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1840 of FIG. 18B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1810, 1840 can be a variation of the graphics processor 1710 of FIG. 17.
[0247] In at least one embodiment, the graphics processor 1810 includes a vertex processor 1805 and one or more fragment processors 1815A - 1815N (e.g., 1815A, 1815B, 1815C, 1815D - 1815N - 1, and 1815N). In at least one embodiment, the graphics processor 1810 can execute different shader programs via separate logic, such that the vertex processor 1805 is optimized to perform operations for vertex shader programs, while the one or more fragment processors 1815A - 1815N perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1805 executes the vertex processing stage of a 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1815A - 1815N use the primitives and vertex data generated by the vertex processor 1805 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 1815A - 1815N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.
[0248] In at least one embodiment, the graphics processor 1810 further includes one or more memory management units (MMUs) 1820A-1820B, caches 1825A-1825B, and circuit interconnects 1830A-1830B. In at least one embodiment, the one or more MMUs 1820A-1820B provide a virtual-to-physical address mapping for the graphics processor 1810, including the vertex processor 1805 and / or the fragment processors 1815A-1815N, and they may reference vertex or image / text data stored in memory in addition to vertex or image / text data stored in the one or more caches 1825A-1825B. In at least one embodiment, the one or more MMUs 1820A-1820B may be synchronized with one or more other MMUs in the system, including one or more MMUs associated with the one or more application processors 1705, the image processor 1715, and / or the video processor 1720 of FIG. 17, such that each processor 1705-1720 can participate in a shared or integrated virtual memory system. In at least one embodiment, the one or more circuit interconnects 1830A-1830B enable the graphics processor 1810 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.
[0249] In at least one embodiment, the graphics processor 1840 includes one or more shader cores 1855A-1855N (e.g., 1855A, 1855B, 1855C, 1855D, 1855E, 1855F-1855N-1, and 1855N) as shown in FIG. 18B, which provide an integrated shader core architecture in which all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders, can be executed by a single core, or type, or core. In at least one embodiment, the number of shader cores can be varied. In at least one embodiment, the graphics processor 1840 includes an inter-core task manager 1845 that acts as a thread dispatcher for dispatching execution threads to one or more shader cores 1855A-1855N, and a tiling unit 1858 for accelerating tiling operations for tile-based rendering in which the rendering operation of a scene is subdivided in the image space, for example, to utilize local spatial coherence within the scene or to optimize the use of internal caches.
[0250] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in integrated circuits 18A and / or 18B for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0251] In at least one embodiment, one or more systems shown in FIGS. 18A and / or 18B are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIGS. 18A and / or 18B are utilized to execute various processes such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIGS. 18A and / or 18B are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0252] FIGS. 19A-19B show further exemplary graphics processor logic according to the embodiments described herein. FIG. 19A shows a graphics core 1900, which in at least one embodiment may be included in the graphics processor 1710 of FIG. 17, and in at least one embodiment may be integrated shader cores 1855A-1855N as in FIG. 18B. FIG. 19B shows a highly parallel general-purpose graphics processing unit ("GPGPU") 1930 suitable for introduction into a multi-chip module in at least one embodiment.
[0253] In at least one embodiment, the graphics core 1900 includes a shared instruction cache 1902, a texture unit 1918, and a cache / shared memory 1920, which are common to the execution resources within the graphics core 1900. In at least one embodiment, the graphics core 1900 can include a plurality of slices 1901A - 1901N, or per-core partitions, and the graphics processor can include a plurality of instances of the graphics core 1900. In at least one embodiment, the slices 1901A - 1901N can include support logic that includes local instruction caches 1904A - 1904N, thread schedulers 1906A - 1906N, thread dispatchers 1908A - 1908N, and sets of registers 1910A - 1910N. In at least one embodiment, the slices 1901A - 1901N can include a set of additional functional units (AFU1912A - 1912N), floating point units (FPU1914A - 1914N), integer arithmetic logic units (ALU1916A - 1916N), address calculation units (ACU1913A - 1913N), double precision floating point units (DPFPU1915A - 1915N), and matrix processing units (MPU1917A - 1917N).
[0254] In at least one embodiment, FPUs 1914A - 1914N can perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, and DPFPU 1915A - 1915N perform double - precision (64 - bit) floating - point operations. In at least one embodiment, ALUs 1916A - 1916N can perform variable - precision integer operations with 8 - bit, 16 - bit, and 32 - bit precision and can be configured to perform mixed - precision operations. In at least one embodiment, MPU 1917A - 1917N can also be configured to perform mixed - precision matrix operations including half - precision floating - point and 8 - bit integer operations. In at least one embodiment, MPU 1917A - 1917N can perform various matrix operations for accelerating machine - learning application frameworks, including enabling support for accelerating general matrix - matrix multiplication (GEMM). In at least one embodiment, AFU 1912A - 1912N can perform additional logical operations not supported by a floating - point unit or integer unit, including trigonometric operations (e.g., sine, cosine, etc.).
[0255] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in graphics core 1900 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0256] In at least one embodiment, one or more systems shown in FIG. 19A are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 19A are utilized to execute various processes such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 19A are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data such as training images.
[0257] FIG. 19B shows a general-purpose processing unit (GPGPU) 1930, which can be configured to execute high-parallel computing operations by an array of graphics processing units in at least one embodiment. In at least one embodiment, GPGPU 1930 can be directly linked to other instances of GPGPU 1930 to generate multiple GPU clusters to improve the training speed of a deep neural network. In at least one embodiment, GPGPU 1930 includes a host interface 1932 to enable connection to a host processor. In at least one embodiment, host interface 1932 is a PCI Express interface. In at least one embodiment, host interface 1932 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, GPGPU 1930 receives commands from a host processor and uses a global scheduler 1934 to distribute execution threads associated with these commands to a set of compute clusters 1936A-1936H. In at least one embodiment, compute clusters 1936A-1936H share a cache memory 1938. In at least one embodiment, cache memory 1938 can act as a high-level cache for cache memory within compute clusters 1936A-1936H.
[0258] In at least one embodiment, the GPGPU 1930 includes memories 1944A - 1944B coupled to compute clusters 1936A - 1936H via a set of memory controllers 1942A - 1942B. In at least one embodiment, memories 1944A - 1944B can include various types of memory devices, including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory, or graphics random access memory.
[0259] In at least one embodiment, each of the compute clusters 1936A - 1936H includes a set of graphics cores, such as the graphics core 1900 of FIG. 19A, and this set of graphics cores can include multiple types of integer and floating - point logic units capable of performing computational operations with various precisions, including those suitable for machine - learning computations. For example, in at least one embodiment, at least a subset of the floating - point units in each of the compute clusters 1936A - 1936H can be configured to perform 16 - bit or 32 - bit floating - point operations, while another subset of the floating - point units can be configured to perform 64 - bit floating - point operations.
[0260] In at least one embodiment, a plurality of instances of GPGPU 1930 can be configured to operate as a compute cluster. In at least one embodiment, the communication used for synchronization and data exchange by compute clusters 1936A-1936H varies across embodiments. In at least one embodiment, a plurality of instances of GPGPU 1930 communicate via host interface 1932. In at least one embodiment, GPGPU 1930 includes I / O hub 1939, which couples GPGPU 1930 to GPU link 1940, which enables direct connection to other instances of GPGPU 1930. In at least one embodiment, GPU link 1940 is coupled to a dedicated GPU-to-GPU bridge that enables communication and synchronization between multiple instances of GPGPU 1930. In at least one embodiment, GPU link 1940 is coupled to a high-speed interconnect for transmitting and receiving data to / from other GPGPUs or parallel processors. In at least one embodiment, a plurality of instances of GPGPU 1930 are located in separate data processing systems and communicate via a network device accessible via host interface 1932. In at least one embodiment, GPU link 1940 can be configured to enable connection to a host processor in addition to, or instead of, host interface 1932.
[0261] In at least one embodiment, the GPGPU 1930 can be configured to train a neural network. In at least one embodiment, the GPGPU 1930 can be used within an inference platform. In at least one embodiment where the GPGPU 1930 is used for inference, the GPGPU 1930 may include fewer compute clusters 1936A - 1936H than when the GPGPU 1930 is used for neural network training. In at least one embodiment, the memory technology associated with memories 1944A - 1944B may be different for the inference configuration and the training configuration, and a high - bandwidth memory technology is applied to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 1930 can support inference - specific instructions. For example, in at least one embodiment, the inference configuration can support one or more dot - product instructions for 8 - bit integers, which may be used during the inference operation of a pre - trained neural network.
[0262] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the GPGPU 1930 for inference or prediction operations, based at least in part on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network.
[0263] In at least one embodiment, one or more systems shown in FIG. 19B are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 19B are utilized to execute various processes such as those described in connection with FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 19B are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0264] FIG. 20 is a block diagram showing a computing system 2000 according to at least one embodiment. In at least one embodiment, the computing system 2000 includes a processing subsystem 2001 having one or more processors 2002 and a system memory 2004 that communicate via an interconnect path that may include a memory hub 2005. In at least one embodiment, the memory hub 2005 may be a separate component within a chipset component or may be integrated within one or more processors 2002. In at least one embodiment, the memory hub 2005 is coupled to an I / O subsystem 2011 via a communication link 2006. In at least one embodiment, the I / O subsystem 2011 includes an I / O hub 2007 that enables the computing system 2000 to receive input from one or more input devices 2008. In at least one embodiment, the I / O hub 2007 can enable a display controller, which may be included in one or more processors 2002, to provide output to one or more display devices 2010A. In at least one embodiment, one or more display devices 2010A coupled to the I / O hub 2007 can include local, internal, or embedded display devices.
[0265] In at least one embodiment, the processing subsystem 2001 includes one or more parallel processors 2012 coupled to a memory hub 2005 via a bus or other communication link 2013. In at least one embodiment, the communication link 2013 may use one of any number of standards-based communication link technologies or protocols, such as but not limited to PCI Express, or may be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 2012 form a parallel or vector processing system focused on computing that can include a number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, some or all of the parallel processors 2012 form a graphics processing subsystem that can output pixels to one of one or more display devices 2010A coupled via an I / O hub 2007. In at least one embodiment, the parallel processors 2012 can also include a display controller and display interface (not shown) that enable direct connection to one or more display devices 2010B.
[0266] In at least one embodiment, the system storage unit 2014 can be connected to the I / O hub 2007 to provide a storage mechanism for the computing system 2000. In at least one embodiment, an I / O switch 2016 can be used to provide an interface mechanism for enabling communication between the I / O hub 2007 and other components such as a network adapter 2018 and / or a wireless network adapter 2019 that may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 2020. In at least one embodiment, the network adapter 2018 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2019 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.
[0267] In at least one embodiment, the computing system 2000 can include other components not explicitly shown, including USB or other port connections, an optical storage drive, a video capture device, etc., which may also be connected to the I / O hub 2007. In at least one embodiment, the communication paths interconnecting the various components of FIG. 20 may be implemented using any suitable protocol such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces such as NV-Link high-speed interconnects, or other interconnect protocols.
[0268] In at least one embodiment, parallel processor 2012 incorporates circuitry optimized for graphics and video processing, including, for example, a video output circuit, and comprises a graphics processing unit (GPU). In at least one embodiment, parallel processor 2012 incorporates circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 2000 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, parallel processor 2012, memory hub 2005, processor 2002, and I / O hub 2007 may be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 2000 may be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 2000 may be integrated into a multi-chip module (MCM), and this module may be interconnected with other multi-chip modules to form a modular computing system.
[0269] Inference and / or training logic 815 is used to perform inference and / or training operations associated with one or more embodiments. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIG. 8A and / or FIG. 8B. In at least one embodiment, inference and / or training logic 815 may be used in the system of FIG. 2000 for inference or prediction operations, at least in part based on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0270] In at least one embodiment, one or more systems shown in FIG. 20 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 20 are utilized to execute various processes such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 20 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0271] Processor FIG. 21A shows a parallel processor 2100 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2100 may be implemented using one or more integrated circuit devices such as a programmable processor, an application specific integrated circuit (ASIC), or a field programmable gate array (FPGA). In at least one embodiment, the illustrated parallel processor 2100 is a variant of one or more parallel processors 2012 shown in FIG. 20 according to an exemplary embodiment.
[0272] In at least one embodiment, parallel processor 2100 includes parallel processing unit 2102. In at least one embodiment, parallel processing unit 2102 includes I / O unit 2104 that enables communication with other devices including other instances of parallel processing unit 2102. In at least one embodiment, I / O unit 2104 may be directly connected to other devices. In at least one embodiment, I / O unit 2104 is connected to other devices through the use of a hub or switch interface such as memory hub 2105. In at least one embodiment, the connection between memory hub 2105 and I / O unit 2104 forms communication link 2113. In at least one embodiment, I / O unit 2104 is connected to host interface 2106 and memory crossbar 2116, where host interface 2106 receives commands targeted for execution of processing operations and memory crossbar 2116 receives commands targeted for execution of memory operations.
[0273] In at least one embodiment, when host interface 2106 receives a command buffer via I / O unit 2104, host interface 2106 can direct a work operation for executing these commands towards front end 2108. In at least one embodiment, front end 2108 is coupled to scheduler 2110, and this scheduler is configured to distribute commands or other work items to processing cluster array 2112. In at least one embodiment, scheduler 2110 ensures that processing cluster array 2112 is properly configured and in a valid state before tasks are distributed to the clusters of processing cluster array 2112. In at least one embodiment, scheduler 2110 is implemented via firmware logic running on a microcontroller. In at least one embodiment, microcontroller-implemented scheduler 2110 can be configured to perform complex scheduling and work distribution operations at coarse and fine granularities, enabling rapid preemption of threads running in processing array 2112 and context switching. In at least one embodiment, host software can account for the scheduling workload in processing cluster array 2112 via one of a plurality of graphics processing paths. In at least one embodiment, the workload can then be automatically distributed across processing cluster array cluster 2112 by scheduler 2110 logic within the microcontroller including scheduler 2110.
[0274] In at least one embodiment, the processing cluster array 2112 can include a maximum of "N" processing clusters (e.g., cluster 2114A, cluster 2114B to cluster 2114N), where "N" represents a positive integer (which may be a different integer "N" from that used in other figures). In at least one embodiment, each of the clusters 2114A to 2114N of the processing cluster array 2112 can execute a large number of simultaneous threads. In at least one embodiment, the scheduler 2110 can use various scheduling and / or workload distribution algorithms to distribute work to the clusters 2114A to 2114N of the processing cluster array 2112, and these algorithms may vary according to the workload generated for each type of program or calculation. In at least one embodiment, the scheduling may be dynamically handled by the scheduler 2110, or may be partially assisted by the compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 2112. In at least one embodiment, the different clusters 2114A to 2114N of the processing cluster array 2112 can be allocated to process different types of programs or to execute different types of calculations.
[0275] In at least one embodiment, the processing cluster array 2112 can be configured to execute various types of parallel processing operations. In at least one embodiment, the processing cluster array 2112 is configured to execute general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 2112 can include logic for executing processing tasks including filtering of video and / or audio data, execution of modeling operations including physical operations, and execution of data conversion.
[0276] In at least one embodiment, the processing cluster array 2112 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2112 can include texture sampling logic for performing texture operations, as well as additional logic for supporting the execution of such graphics processing operations including, but not limited to, mosiac logic and other vertex processing logic. In at least one embodiment, the processing cluster array 2112 can be configured to execute graphics processing related shader programs such as, but not limited to, vertex shaders, mosiac shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2102 can transfer data from the system memory through the I / O unit 2104 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 2122) during processing and then written back to the system memory.
[0277] In at least one embodiment, when graphics processing is performed using the parallel processing unit 2102, the scheduler 2110 can be configured to divide the processing workload into tasks of approximately equal size so that the graphics processing operations can be more effectively distributed among the plurality of clusters 2114A to 2114N of the processing cluster array 2112. In at least one embodiment, a portion of the processing cluster array 2112 can be configured to perform different types of processing. For example, in at least one embodiment, for generating and displaying a rendered image, the first portion may be configured to perform vertex shading and topology generation, the second portion may be configured to perform mosaic and geometry shading, and the third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 2114A to 2114N can be stored in a buffer so that the intermediate data can be transmitted among the clusters 2114A to 2114N for further processing.
[0278] In at least one embodiment, the processing cluster array 2112 can receive processing tasks to be executed via a scheduler 2110, and the scheduler 2110 receives commands defining the processing tasks from a front end 2108. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data should be processed (e.g., which program should be executed). In at least one embodiment, the scheduler 2110 may be configured to fetch an index corresponding to a task, or may receive an index from the front end 2108. In at least one embodiment, the front end 2108 can be configured to ensure that the processing cluster array 2112 is configured in an active state before the workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is started.
[0279] In at least one embodiment, each of one or more instances of the parallel processing unit 2102 can be coupled to a parallel processor - memory 2122. In at least one embodiment, the parallel processor - memory 2122 can be accessed via a memory crossbar 2116, and the memory crossbar 2116 can receive memory requests from the processing cluster array 2112 as well as the I / O unit 2104. In at least one embodiment, the memory crossbar 2116 can access the parallel processor - memory 2122 via a memory interface 2118. In at least one embodiment, the memory interface 2118 can include a plurality of partition units (e.g., partition unit 2120A, partition units 2120B to 2120N), and each of these units can be coupled to a portion (e.g., a memory unit) of the parallel processor - memory 2122. In at least one embodiment, the number of partition units 2120A to 2120N is configured to be equal to the number of memory units, such that the first partition unit 2120A has a corresponding first memory unit 2124A, the second partition unit 2120B has a corresponding memory unit 2124B, and the Nth partition unit 2120N has a corresponding Nth memory unit 2124N. In at least one embodiment, the number of partition units 2120A to 2120N may not be equal to the number of memory units.
[0280] In at least one embodiment, the memory units 2124A - 2124N can include various types of memory devices, including dynamic random access memory (DRAM) or graphics random access memory, such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory. In at least one embodiment, the memory units 2124A - 2124N may also include, but are not limited to, 3D stacked memory including high bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 2122, a render target such as a frame buffer or texture map can be stored across the memory units 2124A - 2124N so that the partition units 2120A - 2120N can write portions of each render target in parallel. In at least one embodiment, the local instance of the parallel processor memory 2122 may be excluded to be advantageous for an integrated memory design that combines system memory and local cache memory.
[0281] In at least one embodiment, any one of clusters 2114A - 2114N of the processing cluster array 2112 can process data to be written to any one of memory units 2124A - 2124N within the parallel processor memory 2122. In at least one embodiment, the memory crossbar 2116 can be configured to transfer the output of each of clusters 2114A - 2114N to any partition unit 2120A - 2120N capable of performing further processing operations on the output, or to another cluster 2114A - 2114N. In at least one embodiment, each of clusters 2114A - 2114N can communicate with the memory interface 2118 through the memory crossbar 2116 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 2116 has a connection to the memory interface 2118 for communicating with the I / O unit 2104, as well as a connection to a local instance of the parallel processor memory 2122, enabling processing units within different processing clusters 2114A - 2114N to communicate with system memory or other memory not local to the parallel processing unit 2102. In at least one embodiment, the memory crossbar 2116 can use virtual channels to separate traffic streams between clusters 2114A - 2114N and partition units 2120A - 2120N.
[0282] In at least one embodiment, multiple instances of the parallel processing unit 2102 may be provided on a single add-in card or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2102 may be configured to interoperate even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 2102 may include a higher precision floating point unit than other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 2102 or the parallel processor 2100 can be implemented in a variety of configurations and form factors including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.
[0283] FIG. 21B is a block diagram of a partition unit 2120 according to at least one embodiment. In at least one embodiment, the partition unit 2120 is an instance of one of the partition units 2120A - 2120N of FIG. 21A. In at least one embodiment, the partition unit 2120 includes an L2 cache 2121, a frame buffer interface 2125, and a ROP: raster operations unit 2126. In at least one embodiment, the L2 cache 2121 is a read / write cache configured to perform load and store operations received from the memory crossbar 2116 and the ROP 2126. In at least one embodiment, read misses and urgent writeback requests are output by the L2 cache 2121 to the frame buffer interface 2125 to be processed. In at least one embodiment, updates are also sent to the frame via the frame buffer interface 2125 to be processed. In at least one embodiment, the frame buffer interface 2125 interfaces with one of the memory units of the parallel processor memory, such as the memory units 2124A - 2124N (e.g., within the parallel processor memory 2122 of FIG. 21).
[0284] In at least one embodiment, ROP2126 is a processing unit that performs raster operations such as stencil, z-test, blending, etc. In at least one embodiment, ROP2126 then outputs processed graphics data stored in the graphics memory. In at least one embodiment, ROP2126 includes compression logic for compressing depth or color data written to the memory and decompressing depth or color data read from the memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a plurality of compression algorithms. In at least one embodiment, the type of compression performed by ROP2126 can be changed based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a per-tile basis for depth and color data.
[0285] In at least one embodiment, ROP2126 is included within each processing cluster (e.g., clusters 2114A - 2114N in FIG. 21A) rather than within the partition unit 2120. In at least one embodiment, read and write requests for pixel data rather than pixel fragment data are transmitted via the memory crossbar 2116. In at least one embodiment, the processed graphics data may be displayed on a display device such as one of the one or more display devices 2010 of FIG. 20, may be routed to be further processed by the processor 2002, or may be routed to be further processed by one of the processing entities within the parallel processor 2100 of FIG. 21A.
[0286] FIG. 21C is a block diagram of a processing cluster 2114 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of the processing clusters 2114A - 2114N of FIG. 21A. In at least one embodiment, the processing cluster 2114 may be configured to execute a number of threads in parallel, where a "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, a single instruction multiple data (SIMD) instruction issuing technique is used to support parallel execution of a number of threads without providing a plurality of independent instruction units. In at least one embodiment, a single instruction multiple thread (SIMT) technique is used to support parallel execution of a number of threads that are overall synchronized using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.
[0287] In at least one embodiment, the operation of processing cluster 2114 can be controlled via a pipeline manager 2132 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, the pipeline manager 2132 receives instructions from the scheduler 2110 of FIG. 21A and manages the execution of these instructions via the graphics multiprocessor 2134 and / or the texture unit 2136. In at least one embodiment, the graphics multiprocessor 2134 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within the processing cluster 2114. In at least one embodiment, one or more instances of the graphics multiprocessor 2134 can be included within the processing cluster 2114. In at least one embodiment, the graphics multiprocessor 2134 can process data, and a data crossbar 2140 may be used to distribute the processed data to one of a plurality of possible destinations including other shader units. In at least one embodiment, the pipeline manager 2132 can facilitate the distribution of the processed data by specifying the destination of the processed data to be distributed via the data crossbar 2140.
[0288] In at least one embodiment, each graphics multiprocessor 2134 within the processing cluster 2114 can include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner such that new instructions can be issued before the previous instruction is completed. In at least one embodiment, the function execution logic supports various operations including integer and floating point arithmetic, comparison operations, boolean operations, bit shifts, and calculations of various algebraic functions. In at least one embodiment, different operations can be executed by leveraging the hardware of the same function unit, and any combination of function units may exist.
[0289] In at least one embodiment, the instructions sent to processing cluster 2114 configure threads. In at least one embodiment, a set of threads being executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a common program for different input data. In at least one embodiment, each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 2134. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within graphics multiprocessor 2134. In at least one embodiment, if the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during cycles in which the thread group is being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines within graphics multiprocessor 2134. In at least one embodiment, if the thread group includes more threads than the number of processing engines within graphics multiprocessor 2134, processing can be executed over consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 2134.
[0290] In at least one embodiment, the graphics multi-processor 2134 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multi-processor 2134 can forego the internal cache and use the cache memory (e.g., L1 cache 2148) within the processing cluster 2114. In at least one embodiment, each graphics multi-processor 2134 can also access the L2 cache within the partition unit (e.g., partition units 2120A - 2120N of FIG. 21A), and these caches can be shared among all processing clusters 2114 and may be used to transfer data between threads. In at least one embodiment, the graphics multi-processor 2134 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 2102 may be used as global memory. In at least one embodiment, the processing cluster 2114 includes multiple instances of the graphics multi-processor 2134 that can share common instructions and data, which may be stored in the L1 cache 2148.
[0291] In at least one embodiment, each processing cluster 2114 may include an MMU 2145 (memory management unit) configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2145 may be within the memory interface 2118 of FIG. 21A. In at least one embodiment, the MMU 2145 includes a set of page table entries (PTEs) used to map virtual addresses to the physical addresses of tiles and optionally cache line indices. In at least one embodiment, the MMU 2145 may include a translation lookaside buffer (TLB) or cache, which may be within the graphics multiprocessor 2134 or the L1 2148 cache, or within the processing cluster 2114. In at least one embodiment, the physical addresses are processed to locally distribute surface data access, enabling efficient interleaving of requests among partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.
[0292] In at least one embodiment, each graphics multi-processor 2134 is coupled to a texture unit 2136 such that the processing cluster 2114 may be configured to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, the texture data is read from an internal texture L1 cache (not shown) or from the L1 cache within the graphics multi-processor 2134 and, if necessary, fetched from the L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multi-processor 2134 outputs processed tasks to the data crossbar 2140 to provide the processed tasks to another processing cluster 2114 for further processing or stores the processed tasks in the L2 cache, local parallel processor memory, or system memory via the memory crossbar 2116. In at least one embodiment, the pre-ROP 2142 (pre-raster operation unit) is configured to receive data from the graphics multi-processor 2134 and direct the data to the ROP unit, which may be located within a partitioning unit (e.g., partitioning units 2120A - 2120N of FIG. 21A) as described herein. In at least one embodiment, the pre-ROP 2142 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.
[0293] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in the graphics processing cluster 2114 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of the neural networks.
[0294] In at least one embodiment, one or more systems shown in FIG. 21C are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 21C are utilized to perform various processes such as those described in relation to FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 21C are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data, using a subset of the training data such as training images.
[0295] FIG. 21D shows a graphics multi-processor 2134 according to at least one embodiment. In at least one embodiment, the graphics multi-processor 2134 is coupled to a pipeline manager 2132 of a processing cluster 2114. In at least one embodiment, the graphics multi-processor 2134 has an execution pipeline that includes, but is not limited to, an instruction cache 2152, an instruction unit 2154, an address mapping unit 2156, a register file 2158, one or more general purpose graphics processing unit (GPGPU) cores 2162, and one or more load / store units 2166. In at least one embodiment, the GPGPU cores 2162, and the load / store units 2166 are coupled to a cache memory 2172 and a shared memory 2170 via a memory and cache interconnect 2168.
[0296] In at least one embodiment, the instruction cache 2152 receives a stream of instructions to be executed from the pipeline manager 2132. In at least one embodiment, the instructions are cached in the instruction cache 2152 and dispatched for execution by an instruction unit 2154. In at least one embodiment, the instruction unit 2154 can dispatch instructions as a thread group (e.g., a warp), and each thread of the thread group is assigned to a different execution unit within a GPGPU core 2162. In at least one embodiment, instructions can access any of a local, shared, or global address space by specifying an address within a unified address space. In at least one embodiment, an address mapping unit 2156 can be used to translate an address in the unified address space to an individual memory address accessible by a load / store unit 2166.
[0297] In at least one embodiment, register file 2158 provides a set of registers to the functional units of graphics multiprocessor 2134. In at least one embodiment, register file 2158 provides temporary storage for operands connected to the data paths of the functional units (e.g., GPGPU cores 2162, load / store unit 2166) of graphics multiprocessor 2134. In at least one embodiment, register file 2158 is divided among each of the functional units such that each functional unit is allocated a dedicated portion of register file 2158. In one embodiment, register file 2158 is divided among different warps being executed by graphics multiprocessor 2134.
[0298] In at least one embodiment, each GPGPU core 2162 can include a floating point unit (FPU) and / or an integer arithmetic logic unit (ALU) used to execute instructions of graphics multiprocessor 2134. In at least one embodiment, the GPGPU cores 2162 may have the same architecture or different architectures. In at least one embodiment, a first portion of GPGPU core 2162 includes a single precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double precision FPU. In at least one embodiment, the FPU can perform IEEE 754-2008 standard floating point operations or enable variable precision floating point operations. In at least one embodiment, graphics multiprocessor 2134 can further include one or more fixed function units or special function units for performing specific functions such as rectangle copy or pixel blending operations. In at least one embodiment, one or more of GPGPU cores 2162 can also include fixed or special function logic.
[0299] In at least one embodiment, the GPGPU core 2162 includes SIMD logic that can execute a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 2162 can physically execute SIMD4, SIMD8, and SIMD16 instructions and can logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core may be generated at compile time by a shader compiler or may be automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads that execute the same or similar operations can be executed in parallel via a single SIMD8 logical unit.
[0300] In at least one embodiment, the memory and cache interconnect 2168 is an interconnect network that connects each functional unit of the graphics multiprocessor 2134 to the register file 2158 and the shared memory 2170. In at least one embodiment, the memory and cache interconnect 2168 is a crossbar interconnect that enables the load / store unit 2166 to perform load and store operations between the shared memory 2170 and the register file 2158. In at least one embodiment, the register file 2158 can operate at the same frequency as the GPGPU core 2162, and thus, data transfer between the GPGPU core 2162 and the register file 2158 can have a very low latency. In at least one embodiment, the shared memory 2170 can be used to enable communication between threads executed by functional units within the graphics multiprocessor 2134. In at least one embodiment, the cache memory 2172 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2136. In at least one embodiment, the shared memory 2170 can also be used as a program management cache. In at least one embodiment, threads executing on the GPGPU core 2162 can programmatically store data in the shared memory in addition to automatically cached data stored in the cache memory 2172.
[0301] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated as a core within a package or chip and communicatively coupled to the core via an internal processor bus / interconnect within the package or chip. In at least one embodiment, regardless of the method of connection of the GPU, the processor core may distribute work to such GPUs in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.
[0302] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, inference and / or training logic 815 may be used in graphics multiprocessor 2134 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks.
[0303] In at least one embodiment, one or more systems shown in FIG. 21D are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 21D are utilized to execute various processes such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 21D are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0304] FIG. 22 shows a multi-GPU computing system 2200 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 2200 can include a processor 2202 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2206A-D via a host interface switch 2204. In at least one embodiment, the host interface switch 2204 is a PCI Express switch device that couples the processor 2202 to a PCI Express bus, via which the processor 2202 can communicate with the GPGPUs 2206A-D. In at least one embodiment, the GPGPUs 2206A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 2216-. In at least one embodiment, the GPU-to-GPU link 2216 is connected to each of the GPGPUs 2206A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 2216 enables direct communication between each of the GPGPUs 2206A-D without requiring communication via the host interface bus 2204 to which the processor 2202 is connected. In at least one embodiment, when there is GPU-to-GPU traffic directed to the P2P GPU link 2216, the host interface bus 2204 is kept available to access system memory or to communicate with other instances of the multi-GPU computing system 2200, for example, via one or more network devices. In at least one embodiment, the GPGPUs 2206A-D are connected to the processor 2202 via the host interface switch 2204, and in at least one embodiment, the processor 2202 includes direct support for the P2P GPU link 2216 and can be directly connected to the GPGPUs 2206A-D.
[0305] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in a multi-GPU computing system 2200 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functionality and / or architecture of a neural network, or use cases of a neural network described herein.
[0306] In at least one embodiment, one or more systems shown in FIG. 22 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 22 are utilized to perform various processes such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 22 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0307] FIG. 23 is a block diagram of a graphics processor 2300 according to at least one embodiment. In at least one embodiment, the graphics processor 2300 includes a ring interconnect 2302, a pipeline front end 2304, a media engine 2337, and graphics cores 2380A-2380N. In at least one embodiment, the ring interconnect 2302 couples the graphics processor 2300 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2300 is one of a number of processors integrated within a multi-core processing system.
[0308] In at least one embodiment, the graphics processor 2300 receives a batch of commands via the ring interconnect 2302. In at least one embodiment, the incoming commands are interpreted by the command streamer 2303 of the pipeline front end 2304. In at least one embodiment, the graphics processor 2300 includes scalable execution logic for performing 3D geometry processing and media processing via the graphics cores 2380A - 2380N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2303 supplies the commands to the geometry pipeline 2336. In at least one embodiment, for at least some media processing commands, the command streamer 2303 supplies the commands to the video front end 2334, and the video front end 2334 is coupled to the media engine 2337. In at least one embodiment, the media engine 2337 includes a Video Quality Engine (VQE) 2330 for post - processing of video and images and a multi - format encode / decode (MFX) 2333 engine that provides hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 2336 and the media engine 2337 each generate execution threads for the thread execution resources provided by at least one graphics core 2380.
[0309] In at least one embodiment, the graphics processor 2300 includes scalable thread execution resources characterized by graphics cores 2380A - 2380N (which can be modular and may also be referred to as core slices), and each modular core 2380A - 2380N has a plurality of sub - cores 2350A - 2350N, 2360A - 2360N (which may also be referred to as core sub - slices). In at least one embodiment, the graphics processor 2300 can have any number of graphics cores 2380A. In at least one embodiment, the graphics processor 2300 includes a graphics core 2380A having at least a first sub - core 2350A and a second sub - core 2360A. In at least one embodiment, the graphics processor 2300 is a low - power processor having a single sub - core (e.g., 2350A). In at least one embodiment, the graphics processor 2300 includes a plurality of graphics cores 2380A - 2380N, each of which includes a set of first sub - cores 2350A - 2350N and a set of second sub - cores 2360A - 2360N. In at least one embodiment, each sub - core of the first sub - cores 2350A - 2350N includes at least an execution unit 2352A - 2352N and a first set of media / texture samplers 2354A - 2354N. In at least one embodiment, each sub - core of the second sub - cores 2360A - 2360N includes at least an execution unit 2362A - 2362N and a second set of samplers 2364A - 2364N. In at least one embodiment, each sub - core 2350A - 2350N, 2360A - 2360N shares a set of shared resources 2370A - 2370N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.
[0310] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding the inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the inference and / or training logic 815 may be used in the graphics processor 2300 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functionality and / or architecture of the neural network described herein, or use cases of the neural network.
[0311] In at least one embodiment, one or more systems shown in FIG. 23 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 23 are utilized to perform various processes, such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more systems shown in FIG. 23 are utilized to estimate parameters as part of one or more training processes of a neural network model, based on the uniqueness of training data, using a subset of the training data, such as training images.
[0312] FIG. 24 is a block diagram showing the micro-architecture of a processor 2400 that may include a logic circuit for executing instructions according to at least one embodiment. In at least one embodiment, the processor 2400 may execute instructions including x86 instructions, ARM instructions, special instructions for application specific integrated circuits (ASICs), and the like. In at least one embodiment, the processor 2400 may include registers for storing packed data, such as 64-bit wide MMX (trademark) registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, the MMX registers available in both integer and floating-point formats may operate on packed data elements with single instruction multiple data (“SIMD”) and streaming SIMD extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or more (collectively referred to as “SSEx”) technologies may hold operands of such packed data. In at least one embodiment, the processor 2400 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.
[0313] In at least one embodiment, the processor 2400 includes an in-order front end (“front end”) 2401 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 2401 may include several units. In at least one embodiment, an instruction prefetcher 2426 fetches instructions from memory and supplies the instructions to an instruction decoder 2428, and the instruction decoder decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2428 decodes the received instruction into one or more operations called “microinstructions” or “micro-operations” that the machine can execute (also called “micro-ops” or “uops”). In at least one embodiment, the instruction decoder 2428 parses the instruction into an opcode and corresponding data, as well as a control field, such that these are used by the microarchitecture and the operations according to at least one embodiment may be executed. In at least one embodiment, a trace cache 2430 may assemble the decoded uops into a program-order sequence or trace in a uop queue 2434 for execution. In at least one embodiment, when the trace cache 2430 encounters a complex instruction, a microcode ROM 2432 provides the uops necessary for the completion of the operation.
[0314] In at least one embodiment, there are instructions that can be translated into a single micro-op, and there are also instructions that require several micro-ops to complete the entire operation. In at least one embodiment, if more than five micro-ops are required to complete an instruction, the instruction decoder 2428 may access the microcode ROM 2432 to execute that instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops so that it can be processed in the instruction decoder 2428. In at least one embodiment, if a large number of micro-ops are required to complete such an operation, the instruction may be stored in the microcode ROM 2432. In at least one embodiment, the trace cache 2430 determines the correct micro-instruction pointer for reading the microcode sequence by referring to an entry point programmable logic array ("PLA") in order to complete one or more instructions from the microcode ROM 2432 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2432 finishes sequencing the micro-ops for an instruction, the front end 2401 of the machine may resume fetching micro-ops from the trace cache 2430.
[0315] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 2403 may prepare instructions for execution. In at least one embodiment, the out-of-order execution logic has multiple buffers to smooth the flow of instructions and change their order, optimizing performance when instructions flow down the pipeline and are scheduled for execution. In at least one embodiment, the out-of-order execution engine 2403 includes, without limitation, an allocator / register renamer 2440, a memory uop queue 2442, an integer / floating point uop queue 2444, a memory scheduler 2446, a fast scheduler 2402, a slow / general purpose floating point scheduler ("slow / general purpose FP scheduler") 2404, and a simple floating point scheduler ("simple FP scheduler") 2406. In at least one embodiment, the fast scheduler 2402, the slow / general purpose floating point scheduler 2404, and the simple floating point scheduler 2406 are also collectively referred to herein as "uop schedulers 2402, 2404, 2406". In at least one embodiment, the allocator / register renamer 2440 allocates the machine buffers and resources required by each uop for execution. In at least one embodiment, the allocator / register renamer 2440 changes the name of the logical register upon entry into the register file. In at least one embodiment, the allocator / register renamer 2440 also distributes the entry of each uop to one of two uop queues, namely the memory uop queue 2442 for memory operations and the integer / floating point uop queue 2444 for non-memory operations, in front of the memory scheduler 2446 and the uop schedulers 2402, 2404, 2406. In at least one embodiment, the uop schedulers 2402, 2404, 2406 determine when uops are ready for execution based on the availability of the sources of their dependent input register operands and the execution resources required by the uop to complete their operations.In at least one embodiment, the high-speed scheduler 2402 may schedule every half of the main clock cycle, and the low-speed / general-purpose floating-point scheduler 2404 and the simple floating-point scheduler 2406 may schedule once per clock cycle of the main processor. In at least one embodiment, the uop schedulers 2402, 2404, 2406 arbitrate dispatch ports to schedule uops for execution.
[0316] In at least one embodiment, the execution block 2411 includes, without limitation, the integer register file / bypass network 2408, the floating-point register file / bypass network (referred to herein as the "FP register file / bypass network") 2410, the address generation units (referred to herein as "AGUs": address generation unit) 2412 and 2414, the high-speed arithmetic logic units (ALUs) (referred to herein as the "high-speed ALUs") 2416 and 2418, the low-speed arithmetic logic units (referred to herein as the "low-speed ALUs") 2420, the floating-point ALUs (referred to herein as "FPs") 2422, and the floating-point move units (referred to herein as the "FP moves") 2424. In at least one embodiment, the integer register file / bypass network 2408 and the floating-point register file / bypass network 2410 are also referred to herein as the "register files 2408, 2410". In at least one embodiment, the AGUs 2412 and 2414, the high-speed ALUs 2416 and 2418, the low-speed ALUs 2420, the floating-point ALUs 2422, and the floating-point move units 2424 are also referred to herein as the "execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424". In at least one embodiment, the execution block 2411 may include any number and type of register files, bypass networks, address generation units, and execution units (including zero) in any combination, without limitation.
[0317] In at least one embodiment, register networks 2408, 2410 may be disposed between uop schedulers 2402, 2404, 2406 and execution units 2412, 2414, 2416, 2418, 2420, 2422, and 2424. In at least one embodiment, integer register file / bypass network 2408 performs integer operations. In at least one embodiment, floating point register file / bypass network 2410 performs floating point operations. In at least one embodiment, each of register networks 2408, 2410 may include, without limitation, a bypass network that may bypass or transfer just-completed results that have not yet been written to the register file to new dependent uops. In at least one embodiment, register networks 2408, 2410 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2408 may include, without limitation, two separate register files, namely one register file for lower 32-bit data and a second register file for upper 32-bit data. In at least one embodiment, since floating point instructions typically have operands in the 64-128 bit width range, floating point register file / bypass network 2410 may include, without limitation, 128-bit wide entries.
[0318] In at least one embodiment, execution units 2412, 2414, 2416, 2418, 2420, 2422, 2424 may execute instructions. In at least one embodiment, register networks 2408, 2410 store operand values of integer and floating-point data that microinstructions need to execute. In at least one embodiment, processor 2400 may include any number and combination of execution units 2412, 2414, 2416, 2418, 2420, 2422, 2424 without limitation. In at least one embodiment, floating-point ALU 2422 and floating-point shift unit 2424 may execute floating-point, MMX, SIMD, AVX, and SEE, or other operations including special machine learning instructions. In at least one embodiment, floating-point ALU 2422 includes, without limitation, a 64-bit floating-point divider and may execute division, square root, and other micro-ops. In at least one embodiment, instructions containing floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to fast ALUs 2416, 2418. In at least one embodiment, fast ALUs 2416, 2418 may execute fast operations with an effective latency of half a clock cycle. In at least one embodiment, since low-speed ALU 2420 may include, without limitation, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branch processing, most complex integer operations proceed to low-speed ALU 2420. In at least one embodiment, memory load / store operations may be executed by AGUs 2412, 2414. In at least one embodiment, fast ALU 2416, fast ALU 2418, and low-speed ALU 2420 may execute integer operations with 64-bit data operands. In at least one embodiment, fast ALU 2416, fast ALU 2418, and low-speed ALU 2420 may be implemented to support various data bit sizes including 16, 32, 128, 256, etc.In at least one embodiment, the floating point ALU 2422 and the floating point shift unit 2424 may be implemented to support a wide range of operands having various bit widths, such as 128-bit wide packed data operands, in conjunction with SIMD and multimedia instructions.
[0319] In at least one embodiment, the uop schedulers 2402, 2404, 2406 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, since uops may be scheduled and executed speculatively in the processor 2400, the processor 2400 may also include logic for handling memory misses. In at least one embodiment, when a data load misses in the data cache, there may be ongoing dependent operations in the pipeline that have passed through a scheduler with temporarily inaccurate data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use inaccurate data. In at least one embodiment, dependent operations may need to be replayed, and independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of the processor may also be designed to capture instruction sequences for text string comparison operations.
[0320] In at least one embodiment, a "register" may refer to a storage location of an on-board processor that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be accessible from outside the processor (from the perspective of a programmer). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuits within a processor using any number of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, combinations of dedicated physical registers and physically registers dynamically allocated, and the like. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.
[0321] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, some or all of inference and / or training logic 815 may be incorporated into execution block 2411 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in execution block 2411. Additionally, weight parameters may be stored in on-chip or off-chip memories and / or registers (shown or not shown) that make up the ALU of execution block 2411 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.
[0322] In at least one embodiment, one or more systems shown in FIG. 24 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more systems shown in FIG. 24 are utilized to execute various processes such as those described in connection with FIGS. 1 - 7. In at least one embodiment, one or more systems shown in FIG. 24 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0323] FIG. 25 shows a deep learning application processor 2500 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2500 uses instructions that, when executed by the deep learning application processor 2500, cause the deep learning application processor 2500 to execute some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 2500 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 2500 executes matrix multiplication operations that are "hard-wired" to hardware as a result of executing one or both of a plurality of instructions. In at least one embodiment, the deep learning application processor 2500 includes, without limitation, processing clusters 2510(1) - 2510(12), inter-chip links ("ICL") 2520(1) - 2520(12), inter-chip controllers ("ICC") 2530(1) - 2530(2), high bandwidth memory second generation ("HBM2") 2540(1) - 2540(4), memory controllers ("Mem Ctrlr") 2542(1) - 2542(4), high bandwidth memory physical layer ("HBM PHY") 2544(1) - 2544(4), management-controller central processing unit ("management-controller CPU") 2550, serial peripheral interface, between integrated circuits, and general purpose input / output blocks ("SPI, I 2C, GPIO) 2560, Peripheral Component Interconnect Express Controller and Direct Memory Access Block ("PCIe Controller and DMA") 2570, and 16-lane Peripheral Component Interconnect Express Port ("PCI Expressx16") 2580.
[0324] In at least one embodiment, the processing cluster 2510 may perform deep learning operations including inference or prediction operations based on weight parameters calculated using one or more training techniques including the techniques described herein. In at least one embodiment, each processing cluster 2510 may include any number and type of processors, without limitation. In at least one embodiment, the deep learning application processor 2500 may include any number and type of processing clusters 2500. In at least one embodiment, the inter-chip link 2520 is bidirectional. In at least one embodiment, the inter-chip link 2520 and the inter-chip controller 2530 enable multiple deep learning application processors 2500 to exchange information including activation information obtained as a result of executing one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, the deep learning application processor 2500 may include any number and type (including zero) of ICL 2520 and ICC 2530.
[0325] In at least one embodiment, HBM2 2540 provides a total of 32 gigabytes (GB) of memory. In at least one embodiment, HBM2 2540(i) is associated with both a memory controller 2542(i) and an HBM PHY 2544(i), where "i" is any integer. In at least one embodiment, any number of HBM2 2540 may provide any type and total amount of high-bandwidth memory and may be associated with any number and type (including zero) of memory controllers 2542 and HBM PHYs 2544. In at least one embodiment, SPI, I 2C, GPIO2560, the PCIe controller, and DMA2570, and / or PCIe2580, may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically feasible way.
[0326] To perform inference and / or training operations related to one or more embodiments, inference and / or training logic 815 is used. Details regarding inference and / or training logic 815 are provided herein in conjunction with FIGS. 8A and / or 8B. In at least one embodiment, the deep learning application processor is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2500. In at least one embodiment, the deep learning application processor 2500 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system, or by the deep learning application processor 2500 itself. In at least one embodiment, the processor 2500 may be used to execute one or more of the use cases of the neural networks described herein.
[0327] In at least one embodiment, one or more of the systems shown in FIG. 25 are utilized to implement a framework for training one or more neural networks based on the uniqueness of training data. In at least one embodiment, one or more of the systems shown in FIG. 25 are utilized to execute various processes, such as those described in connection with FIGS. 1-7. In at least one embodiment, one or more of the systems shown in FIG. 25 are utilized to estimate parameters as part of one or more training processes of a neural network model based on the uniqueness of training data using a subset of the training data, such as training images.
[0328] FIG. 26 is a block diagram of a neuromorphic processor 2600 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2600 receives one or more inputs from a source external to the neuromorphic processor 2600. In at least one embodiment, these inputs may be sent to one or more neurons 2602 within the neuromorphic processor 2600. In at least one embodiment, the neurons 2602 and their components may be implemented using circuitry or logic that includes one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2600 may include thousands or millions of instances of neurons 2602, without limitation, although any suitable number of neurons 2602 may be used. In at least one embodiment, each instance of a neuron 2602 may include a neuron input 2604 and a neuron output 2606. In at least one embodiment, a neuron 2602 may generate an output, and this output may be sent to the inputs of other instances of neurons 2602. For example, in at least one embodiment, the neuron inputs 2604 and the neuron outputs 2606 may be interconnected via synapses 2608.
[0329] In at least one embodiment, neuron 2602 and synapse 2608 may be interconnected such that the neuromorphic processor 2600 operates on the information received by the neuromorphic processor 2600 to process or analyze it. In at least one embodiment, neuron 2602 may transmit an output pulse (or "fire" or "spike") when the input received via neuron input 2604 exceeds a threshold. In at least one embodiment, neuron 2602 may sum or integrate the signals received at neuron input 2604. For example, in at least one embodiment, neuron 2602 may be implemented as a leaky integrate-and-fire neuron, where, when the sum (referred to as the "membrane potential") exceeds a threshold, neuron 2602 may use a transfer function such as a sigmoid function or a threshold function to generate an output (or "fire"). In at least one embodiment, the leaky integrate-and-fire neuron may sum the signals received at neuron input 2604 to form a membrane potential and may also apply a decay factor (or leak) to reduce the membrane potential. In at least one embodiment, the leaky integrate-and-fire neuron may fire if multiple input signals are received at neuron input 2604 quickly enough such that the sum exceeds the threshold (i.e., before the decay of the membrane potential is too great to prevent firing). In at least one embodiment, neuron 2602 may be implemented using circuitry or logic that receives an input, integrates the input to form a membrane potential, and decays the membrane potential. In at least one embodiment, the input may be averaged or any other suitable transfer function may be used. Further, in...
Claims
1. A processor, comprising one or more circuits for training one or more second neural networks using one or more first neural networks, based at least in part on the uniqueness of data used to train the one or more second neural networks.
2. The processor according to claim 1, wherein the one or more circuits are further for estimating parameters for training the one or more second neural networks based at least in part on the uniqueness of the data.
3. The processor according to claim 1, wherein the one or more circuits are further for performing one or more operations for indicating the uniqueness of the data by calculating a similarity between the data.
4. The one or more circuits are further for calculating the uniqueness of the data by, at least, comparing a region of interest in a training image with a corresponding region of interest in other training images of the data. The processor according to claim 1.
5. The data includes images, and the one or more circuits are further for calculating a score indicating the uniqueness of an image by comparing the image with other images of the data; and storing the score in a list including a plurality of scores. The processor according to claim 1.
6. The one or more circuits are further for selecting a subset of the data based at least in part on the list; and using a proxy neural network to estimate parameters for training the one or more second neural networks based on the selected subset of the data, wherein the proxy neural network is a smaller representation of the one or more second neural networks. The processor according to claim 5.
7. A system, comprising One or more processors for training one or more second neural networks using one or more first neural networks, based at least in part on the uniqueness of the data used to train the one or more second neural networks.
8. The one or more processors further identifying a subset of the data based on the uniqueness of the data; and estimating one or more values used to train the one or more second neural networks based at least in part on the subset of the data The system according to claim 7, for performing
9. The system according to claim 8, wherein the one or more values include at least one of a learning rate, a size of the one or more second neural networks, and a topology of the one or more second neural networks.
10. The one or more processors further identifying an object from one or more images of the data; comparing the object to a corresponding object from the one or more images; generating an indicator representing a similarity between the object and the corresponding object; and storing the indicator in a list, the list including a plurality of indicators arranged in a ranked order The system according to claim 7, for performing
11. The one or more processors further identifying a subset of the data based on the list; and using a different neural network to infer one or more values used to train the one or more second neural networks based on the subset of the data, the different neural network being a part of the one or more second neural networks The system according to claim 10, for performing
12. The one or more second neural networks include U-Net, The system of claim 11, wherein the different neural networks include a portion of the U-Net, and at least one of a residual block, a channel, and a level is reduced as compared to the U-Net.
13. A method comprising: training the one or more first neural networks to train the one or more second neural networks, at least in part based on the uniqueness of data used to train the one or more second neural networks A method comprising the steps of:
14. using the data to identify regions within an image containing an object; using the data to identify other regions within a plurality of other images containing the object; comparing the regions and the other regions to generate a similarity score; adding the similarity score to an index, the index including a ranked order based on the similarity score The method of claim 13, further comprising:
15. The method of claim 14, wherein the similarity score is ranked higher in the index based on the dissimilarity between the object in the region and the object in the other region.
16. identifying a subset of the data based on the index; using different neural networks to generate one or more values used to estimate parameters for training the one or more second neural networks based on the subset of the data The method of claim 14, further comprising:
17. The method of claim 13, further comprising performing an operation to indicate the uniqueness of the data based on mutual information between one or more medical images of the data.
18. The method of claim 13, wherein the one or more second neural networks include a convolutional neural network.
19. A machine-readable medium storing a set of instructions that, when executed by one or more processors, cause the one or more processors to perform at least Cause one or more first neural networks to be used to train one or more second neural networks based at least in part on the uniqueness of data used to train the one or more second neural networks. A machine-readable medium that causes this to be done. **Claim 20** When the additional set of instructions is executed by the one or more processors, the one or more processors are caused to Identify an object from an image of the data; Identify the object from another image of the data; Generate a score based on dissimilarity between the object from the image and the object from the other image; Add the score to a table, where the table includes a plurality of scores in sorted order. The machine-readable medium of claim 19, further comprising instructions to cause this to be done. **Claim 21** When the additional set of instructions is executed by the one or more processors, the one or more processors are caused to Use the table to select a subset of the data; Use an alternative neural network different from the one or more second neural networks to estimate one or more hyperparameters used to train the one or more second neural networks based on the subset. The machine-readable medium of claim 20, further comprising instructions to cause this to be done. **Claim 22** When the additional set of instructions is executed by the one or more processors, the machine-readable medium of claim 19, further comprising instructions to cause the one or more processors to perform an operation to indicate the uniqueness of the data by calculating dissimilarity between the data. **Claim 23** The machine-readable medium of claim 19, wherein the data includes audio data. **Claim 24** The machine-readable medium of claim 19, wherein the one or more second neural networks include U-Net. **Claim 25** A system One or more computers comprising one or more processors for training one or more second neural networks using one or more first neural networks based at least in part on the uniqueness of data used to train the one or more second neural networks.
26. The one or more processors further are for calculating hyperparameters for training the one or more second neural networks based on the uniqueness of the data, the system of claim 25.
27. The system of claim 26, wherein the hyperparameters include the structure of the one or more second neural networks and one or more variables indicating how the one or more second neural networks should be trained.
28. The one or more processors further generating a score by comparing one or more objects among a set of images obtained from the data, the score including information indicating the similarity between an object in a first image and a corresponding object in another image; and storing the score in a list containing a plurality of scores for performing, the system of claim 25.
29. The one or more processors further use the list to select a portion of the data for estimating parameters used to train the one or more second neural networks based on the plurality of scores, the system of claim 28.
30. The system of claim 25, wherein the one or more second neural networks include a recurrent neural network (RNN).
31. The system of claim 25, wherein the data includes medical images.
Citation Information
Patent Citations
Data extraction device, learning model construction device, data extraction method, learning model construction method and program
JP2022158224A