Neural network modification
By identifying and adjusting the layers that have the greatest impact on accuracy, and using feature distillation and weighted calibration loss values, a sparse quantized neural network is generated. This solves the problem of performance degradation after sparsity and quantization, and improves the accuracy and performance of the neural network.
Patent Information
- Application Number
- CN202480017318.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-10-24
AI Technical Summary
Neural network models degrade in performance after being sparse and/or quantized, resulting in reduced accuracy.
By identifying the layers that have the greatest impact on the accuracy of the neural network, weighting them with feature distillation loss and feature calibration loss, and adjusting the layers that have a great impact on accuracy during pruning and quantization, a sparse quantized neural network is generated.
This improved the accuracy of the neural network, reduced the negative impact of the quantization process on accuracy, and enhanced the overall performance of the neural network.
Smart Images

Figure CN120836034A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] At least one embodiment relates to processing resources for executing a neural network model. For example, a processor including one or more circuits is used to modify a neural network model to become sparse and / or quantized. BACKGROUND
[0002] Modifications to a neural network model can result in a degradation of neural network performance. For example, a neural network model that is made sparse and / or quantized can become inaccurate. Accordingly, performance of the neural network model can be improved. BRIEF DESCRIPTION OF DRAWINGS
[0003] Figure 1 A block diagram of a system for modifying a neural network model to become sparse and / or quantized is shown in accordance with at least one embodiment;
[0004] Figure 2 A block diagram of a system for computing values for modifying a dense floating point neural network model to become sparse and / or quantized is shown in accordance with at least one embodiment;
[0005] Figure 3 A block diagram of a system for computing one or more loss values for modifying a dense floating point neural model is shown in accordance with at least one embodiment;
[0006] Figure 4 A block diagram of a system for computing one or more loss values for modifying a sparse floating point neural model is shown in accordance with at least one embodiment;
[0007] Figure 5 A block diagram of a system for computing one or more loss values for modifying a sparse quantized neural model is shown in accordance with at least one embodiment;
[0008] Figure 6 A block diagram of a system for computing one or more loss values for modifying a sparse quantized neural model is shown in accordance with at least one embodiment;
[0009] Figure 7 A process used by one or more processors to modify a dense floating point neural model to become sparse and quantized is shown in accordance with at least one embodiment;
[0010] Figure 8A A block diagram of a driver and / or runtime used in accordance with at least one embodiment is shown;
[0011] Figure 8B A block diagram of a processor and various modules used in accordance with at least one embodiment is shown;
[0012] Figure 9ALogic is shown in accordance with at least one embodiment;
[0013] Figure 9B Logic is shown in accordance with at least one embodiment;
[0014] Figure 10 Training and deployment of a neural network model is shown in accordance with at least one embodiment;
[0015] Figure 11 An example data center system is shown in accordance with at least one embodiment;
[0016] Figure 12A An example of an autonomous vehicle is shown in accordance with at least one embodiment;
[0017] Figure 12B An example of a camera position and field of view of an autonomous vehicle is shown in accordance with at least one embodiment; Figure 12A
[0018] A block diagram of an example system architecture of an autonomous vehicle is shown in accordance with at least one embodiment; Figure 12C Figure 12A A diagram of a system for communication between one or more cloud-based servers and an autonomous vehicle is shown in accordance with at least one embodiment;
[0019] Figure 12D Figure 12A A block diagram of a computer system is shown in accordance with at least one embodiment;
[0020] Figure 13 A block diagram of a computer system is shown in accordance with at least one embodiment;
[0021] Figure 14 A block diagram of a computer system is shown in accordance with at least one embodiment;
[0022] Figure 15 A computer system is shown in accordance with at least one embodiment;
[0023] Figure 16 A computer system is shown in accordance with at least one embodiment;
[0024] Figure 17A A computer system is shown in accordance with at least one embodiment;
[0025] Figure 17B A computer system is shown in accordance with at least one embodiment;
[0026] Figure 17C A computer system is shown in accordance with at least one embodiment;
[0027] Figure 17D A computer system is shown in accordance with at least one embodiment;
[0028] Figure 17E and Figure 17F A shared programming model is shown in accordance with at least one embodiment;
[0029] Figure 18 An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment;
[0030] Figure 19A - Figure 19B An exemplary integrated circuit and associated graphics processor are shown in accordance with at least one embodiment;
[0031] Figure 20A - Figure 20B Additional exemplary graphics processor logic is shown in accordance with at least one embodiment;
[0032] Figure 21 A computer system is shown in accordance with at least one embodiment;
[0033] Figure 22A A parallel processor is shown in accordance with at least one embodiment;
[0034] Figure 22B A partition unit is shown in accordance with at least one embodiment;
[0035] Figure 22C A processing cluster is shown in accordance with at least one embodiment;
[0036] Figure 22D A graphics multiprocessor is shown in accordance with at least one embodiment;
[0037] Figure 23 A multi-GPU system is shown in accordance with at least one embodiment;
[0038] Figure 24 A graphics processor is shown in accordance with at least one embodiment;
[0039] Figure 25 is a block diagram showing a processor microarchitecture for a processor in accordance with at least one embodiment;
[0040] Figure 26 A deep learning application processor is shown in accordance with at least one embodiment;
[0041] Figure 27 is a block diagram showing an exemplary neuromorphic processor in accordance with at least one embodiment;
[0042] Figure 28 At least portions of a graphics processor are shown in accordance with one or more embodiments;
[0043] Figure 29At least portions of a graphics processor are shown in accordance with one or more embodiments;
[0044] Figure 30 At least portions of a graphics processor are shown in accordance with one or more embodiments;
[0045] Figure 31 is a block diagram of a graphics processing engine of a graphics processor in accordance with at least one embodiment;
[0046] Figure 32 is a block diagram of at least portions of a graphics processor core in accordance with at least one embodiment;
[0047] Figure 33A - Figure 33B Thread execution logic is shown in accordance with at least one embodiment, including an array of processing elements of a graphics processor core;
[0048] Figure 34 A parallel processing unit ("PPU") is shown in accordance with at least one embodiment;
[0049] Figure 35 A general processing cluster ("GPC") is shown in accordance with at least one embodiment;
[0050] Figure 36 A memory partition unit of a parallel processing unit ("PPU") is shown in accordance with at least one embodiment;
[0051] Figure 37 A streaming multiprocessor is shown in accordance with at least one embodiment;
[0052] Figure 38 is an example dataflow graph of an advanced compute pipeline in accordance with at least one embodiment;
[0053] Figure 39 is a system diagram of an example system for training, adapting, instantiating, and deploying machine learning models in an advanced compute pipeline in accordance with at least one embodiment;
[0054] Figure 40 includes an example illustration of an advanced compute pipeline for processing imaging data in accordance with at least one embodiment;
[0055] Figure 41A includes an example dataflow graph of a virtual instrument supporting an ultrasound device in accordance with at least one embodiment;
[0056] Figure 41B includes an example dataflow graph of a virtual instrument supporting a CT scanner in accordance with at least one embodiment;
[0057] Figure 42AA dataflow graph of a process for training a machine learning model is shown in accordance with at least one embodiment;
[0058] Figure 42B An example illustration of a client-server architecture to enhance annotation tools with pre-trained annotation models in accordance with at least one embodiment; and
[0059] Figure 43 Components of a system for accessing large language models are shown in accordance with at least one embodiment. DETAILED DESCRIPTION
[0060] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be apparent to one skilled in the art that the inventive concept can be practiced without one or more of these specific details. In other instances, well-known structures and components are not described in detail in order to avoid obscuring the subject matter.
[0061] In at least one embodiment, as described further herein, the one or more processors execute software that differentially weights an indication of similarity between a full precision neural network layer and a quantized version of that neural network layer during an iterative process to find an optimal combination of quantization layers. In at least one embodiment, when a layer that has the most impact on neural network accuracy is quantized too frequently and / or aggressively (e.g., using INT8 rather than INT4), the different weighting of the indication of similarity results in a greater loss value output by the overall loss function of the neural network compared to a loss value for a less impactful layer that is similarly quantized. In at least one embodiment, as the number of iterations increases, the quantization process pays more and more attention to those layers that have the most impact on neural network accuracy.
[0062] In at least one embodiment, as part of a process to modify a dense floating point neural network into a sparse quantized neural network, changes in accuracy of layers of the neural network that are pruned by deactivating weights of a certain layer are used at least in part to identify neural network layers that have the most impact on neural network accuracy. In at least one embodiment, when quantizing a neural network, information identifying neural network layers that are more impactful to neural network accuracy when modifying the neural network is used to modify precision of one or more layers when quantizing the neural network. In at least one embodiment, accuracy of a layer is based on a comparison of activations of two versions of a neural network, such as a sparse version and a dense version of the neural network.
[0063] In at least one embodiment, a user first trains a dense floating point neural network, also referred to as a dense full precision neural network. In at least one embodiment, a neural network is referred to as a model or a neural model. In at least one embodiment, after training a dense floating point neural network, a user causes a computing system or module to generate a sparse full precision neural network (NN) using an iterative process that modifies and compares the sparse neural network being generated to the trained dense floating point neural network until the sparse neural network reaches an acceptable level of accuracy. In at least one embodiment, during the iterative process, features represented as activation values and output by certain layers of a dense neural network, such as a dense floating point neural network, are compared to activation values output by corresponding layers of a sparse neural network, such as a sparse floating point neural network, by using one or more loss functions. In at least one embodiment, as differences between activated features of corresponding layers get larger, the probability of those layers having an impact on neural network accuracy decreases. In at least one embodiment, information about the impact of layers on accuracy can be used when modifying a sparse full precision neural network to a sparse quantized neural network. In at least one embodiment, information about the impact of layers on accuracy is converted to weight values (weights) that are multiplied by output values of another loss function that compares similarity between corresponding layers of a sparse floating point neural network and a sparse quantized neural network being generated. In at least one embodiment, as the impact of corresponding layers on neural network accuracy increases, the weights applied to output values of the loss function also increase. In at least one embodiment, weighted loss values are combined into a total loss value for the neural network. In at least one embodiment, different weighting of loss values conceptually maximizes accuracy of a neural network by causing the neural network to place more emphasis on layers that have the greatest impact on accuracy, which helps the neural network learn to better quantize, thereby improving its accuracy. In at least one embodiment, any one or more quantization operations described herein are applied to pruning operations, and vice versa. In at least one embodiment, different weighting of loss values enables iterative rounds of quantization to become more sensitive to impactful layers more quickly, thereby speeding up the quantization process.
[0064] In at least one embodiment, changes in accuracy caused by pruning a dense neural network are used to determine which layers have the greatest impact on neural network accuracy. In at least one embodiment, determining which layers have the greatest impact on neural network accuracy makes those most impactful layers less likely to be selected for pruning and / or quantization, or less likely to be pruned and / or quantized to an extent that causes neural network accuracy to fall below an acceptable level.
[0065] In at least one embodiment, one or more processors including one or more circuits are used to modify a neural network to be sparse and / or quantized based at least in part on feature distillation loss values for one or more dense versions (such as dense floating point versions) and one or more sparse versions (such as sparse floating point versions) of the neural network, as further described herein. In at least one embodiment, a distillation loss value is a loss value used as part of a knowledge distillation neural network modification technique. In at least one embodiment, a feature distillation loss value is referred to as a feature-based distillation loss value. In at least one embodiment, a feature distillation loss value is based at least in part on a comparison of feature maps for two versions of a neural network.
[0066] In at least one embodiment, a distillation loss value represents a difference between a prediction of an original neural model and a prediction of a version of the neural model that is modified to be sparse and / or quantized, which is further described herein. In at least one embodiment, a distillation loss value is considered a comparison of an accuracy metric for a first version of a neural network to an accuracy metric for a second version of the neural network, where accuracy metrics are further described herein along with performance metrics. Figure 2 In at least one embodiment, a loss value is referred to as a loss. In at least one embodiment, a loss value is referred to as a distillation loss. In at least one embodiment, a loss value includes a hard label distillation loss, a soft logit distillation loss, and a feature-based distillation loss.
[0067] In at least one embodiment, one or more processors including one or more circuits are used to compute soft logits and hard label predictions for one or more dense floating point versions of a neural network and one or more sparse floating point versions of the neural network, as further described herein. In at least one embodiment, one or more processors including one or more circuits are used to modify a neural network to be sparse and / or quantized using a total pruning loss based at least in part on one or more of a soft logits distillation loss, a hard label distillation loss, or a feature-based distillation loss for a dense floating point version of the neural network and a sparse floating point version of the neural network.
[0068] In at least one embodiment, one or more processors comprising one or more circuits are used to modify a neural network to become sparse and / or quantized based at least in part on a feature calibration loss value for a sparse floating point version of the neural network and a sparse quantized version of the neural network, as described further herein. In at least one embodiment, a feature calibration loss value is referred to as a feature-based calibration loss value or a calibration loss value. In at least one embodiment, a loss value is referred to as a loss. In at least one embodiment, one or more processors comprising one or more circuits are used to compute soft logits and hard label predictions for a sparse floating point version of a neural network and a sparse quantized version of the neural network, as described further herein. In at least one embodiment, one or more processors comprising one or more circuits are used to modify a neural network to become sparse and / or quantized using a total calibration loss based at least in part on one or more of a soft logits calibration loss, a hard label calibration loss, or a feature-based calibration loss for a sparse floating point version of the neural network and a sparse quantized version of the neural network. In at least one embodiment, one or more processors comprising one or more circuits are used to modify a neural network to become sparse and quantized while minimizing accuracy drop of the neural network, as described further herein.
[0069] Figure 1 A system 100, including one or more processors 104 to modify precision of one or more layers of one or more neural networks, is shown in block diagram form in accordance with at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments described herein are in combination with one or more aspects of one or more embodiments described herein, including aspects in combination with Figure 1 In at least one embodiment, one or more aspects of one or more embodiments described herein are in combination with one or more aspects of one or more embodiments described herein, including aspects in combination with Figure 2 - Figure 8B In at least one embodiment, one or more processors perform one or more operations as part of system 100. In at least one embodiment, one or more processors performing one or more operations as part of system 100 are any of the processors or combination of processors described herein, including processor 822 in combination with Figure 8B In at least one embodiment, one or more processors performing one or more operations as part of system 100 are any of the processors or combination of processors described herein, including processor 822 in combination with Figure 12C In at least one embodiment, one or more processors performing one or more operations as part of system 100 are any of the processors or combination of processors described herein, including processor 822 in combination with Figure 19A In at least one embodiment, one or more processors performing one or more operations as part of system 100 are any of the processors or combination of processors described herein, including processor 822 in combination with Figure 34A parallel processing unit ("PPU") 3400 is described. In at least one embodiment, processor 822 performs operations used by neural network quantization and sparsification module 106, such as modifying precision of one or more neural network layers of a sparse neural network. In at least one embodiment, processor 822 performs one or more operations of neural network sparsification and quantization module 106 to modify a dense floating point neural model, which is described further herein. In at least one embodiment, processor 104 performs one or more operations of system 200 of Figure 2 , such as computing total pruning loss 260. In at least one embodiment, processor 104 performs one or more operations of system 300 of Figure 3 , such as computing partial loss 320. In at least one embodiment, processor 104 performs one or more operations of system 400 of Figure 4 , such as computing weight factors 441, which are further described herein in connection with, at least Figure 4 . In at least one embodiment, weight factors are referred to as weights. In at least one embodiment, processor 104 performs one or more operations of system 500 of Figure 5 , such as computing weighted partial loss 540. In at least one embodiment, processor 104 performs one or more operations of system 600 of Figure 6 , such as computing total calibration loss 660. In at least one embodiment, processor 104 performs one or more operations of process 700, such as adjusting weight factors using operation 716. In at least one embodiment, Figure 8A , such as one or more operations for modifying a dense floating point model 110 to become sparse and quantized. In at least one embodiment, sparse and quantized neural model modification module 824 performs one or more operations of system 100, such as generating a sparse quantized model 150.
[0070] In at least one embodiment, as used in any implementation described herein, unless the context clearly indicates otherwise or clearly provides otherwise, terms such as "system" and "module" and nominalized verbs (e.g., compiler and / or other terms) refer to any combination of software logic, firmware logic, hardware logic, and / or circuits configured to provide the functionality described herein. In at least one embodiment, any combination of software logic, firmware logic, hardware logic, and / or circuits configured to provide the functionality described herein are referred to as components. In at least one embodiment, any component described herein can be combined and / or communicatively connected with at least one other component, regardless of how these components are described as being combined and / or communicatively connected in other embodiments. In at least one embodiment, software can be embodied as a software package, code, and / or an instruction set or instructions. In at least one embodiment, hardware, alone or in any combination, includes hardwired circuits, programmable circuits, state machine circuits, fixed function circuits, execution unit circuits, and / or firmware that stores instructions executed by programmable circuits. In at least one embodiment, a module can be embodied as a circuit that is part of a larger system (e.g., an integrated circuit (IC), a system on a chip (SoC), etc.) in its entirety or individually. In at least one embodiment, any one or more circuits of one or more modules and / or processors are represented as a register transfer level (RTL) representation and / or another fabless representation, which can be licensed and / or used for tape-out (tape-out is the last stage in IC design before the IC is manufactured).
[0071] In at least one embodiment, Figure 1 The depicted system 100 is a system that includes a processor 104 for modifying a dense floating point neural network model 110 to become a sparse quantized neural network model 150. In at least one embodiment, the neural network model is referred to as a neural network, a neural model, a model, an artificial intelligence (AI) model, a machine learning (ML) model, or some combination thereof. In at least one embodiment, the system includes a process, a method, an algorithm, one or more operations, one or more tasks, or some combination thereof. In at least one embodiment, executing a neural network refers to one or more processors performing one or more operations of a neural network, such as by executing an arithmetic logic unit (ALU) (such as Figure 7 The weights of one or more neural network layers are adjusted by loading and / or storing weights in the ALU 710 in the neural network.
[0072] In at least one embodiment, the system 100 includes a trained dense floating point model 110, wherein the dense floating point model is combined with at least Figure 1Further described herein. In at least one embodiment, dense float model 110 is received and / or otherwise obtained by one or more processors 104. In at least one embodiment, processors 104 include one or more neural network quantization and sparsity modules 106 that perform one or more operations to generate a sparsely quantized model 150, as described herein in at least conjunction with Figure 1 Further described herein. In at least one embodiment, neural network quantization and sparsity modules 106 generate a non-quantized, sparsely model and / or a non-sparse, quantized model, where such models are described in at least conjunction with Figure 1 Further described herein. In at least one embodiment, system 100 is installed on any hardware platform, such as an edge computing device, a desktop computer, a server, an artificial intelligence (AI) hardware platform, a workstation, a personal computer, or some combination thereof. In at least one embodiment, system 100 is installed in any computing environment, including a data center, a high performance computing center, a cloud computing environment, or some combination thereof.
[0073] In at least one embodiment, neural network quantization and sparsity modules 106 perform operations to identify changes in accuracy of layers of a neural network being pruned. In at least one embodiment, neural network quantization and sparsity modules 106 perform operations to identify neural network layers that have the greatest impact on neural network accuracy using changes in accuracy of layers identified during pruning. In at least one embodiment, neural network quantization and sparsity modules 106 use information identifying neural network layers that have a greater impact on neural network accuracy to modify precision of one or more layers when modifying the neural network during quantization of the neural network. In at least one embodiment, neural network quantization and sparsity modules 106 identify accuracy or changes in accuracy of layers based on a comparison of activations of two versions of a neural network, such as a sparsely version of a neural network and a dense version of the neural network.
[0074] In at least one embodiment, neural network quantization and sparsity module 106 trains a dense floating point neural network. In at least one embodiment, neural network quantization and sparsity module 106 performs operations for generating a sparse full-precision neural network using an iterative process of modifying and comparing a sparse neural network being generated to the trained dense floating point neural network until the sparse neural network reaches an acceptable accuracy after training the dense floating point neural network. In at least one embodiment, the iterative process used by neural network quantization and sparsity module 106 is a knowledge distillation process, which is described further herein. In at least one embodiment, neural network quantization and sparsity module 106 performs operations for comparing features output by a certain layer of a dense neural network represented as activation values to activation values output by a corresponding layer of a sparse neural network by using one or more loss functions. In at least one embodiment, neural network quantization and sparsity module 106 performs operations of generating one or more weight values (weights, weight factors) that are multiplied by output values of another loss function that compares similarity and / or accuracy of a sparse floating point neural network to a corresponding layer of a sparse quantized neural network being generated using information about how accurate a layer is. In at least one embodiment, neural network quantization and sparsity module 106 generates weights applied to output values of a layer according to an increase in how much that layer affects accuracy of a neural network. In at least one embodiment, neural network quantization and sparsity module 106 uses one or more functions to weight loss values and combine weighted loss values to generate a total loss value for a neural network.
[0075] In at least one embodiment, processor 104 is any of the processors or combinations of processors described herein, including processor 822 described in connection with Figure 8B In at least one embodiment, processor 104 is any of the processors or combinations of processors described herein, including processor 822 described in connection with Figure 12C In at least one embodiment, processor 104 is any of the processors or combinations of processors described herein, including processor 822 described in connection with Figure 19A In at least one embodiment, processor 104 is any of the processors or combinations of processors described herein, including processor 822 described in connection with Figure 34 In at least one embodiment, processor 104 is any of the processors or combinations of processors described herein, including processor 822 described in connection with
[0076] In at least one embodiment, neural model includes one or more neural models. In at least one embodiment, neural model is a deep learning network or a deep neural network (DNN). In at least one embodiment, neural model includes any neural model discussed herein, including in connection with FIG. 9 and Figure 10Cdiscussed neural model. In at least one embodiment, a portion of a neural model is a component of a neural model. In at least one embodiment, a portion of a neural model is a layer, a node, a function, a value, or some combination thereof. In at least one embodiment, a neural model includes a processor, a memory, functions, and neural model parameters for performing inference based on input data. In at least one embodiment, neural model parameters include values that affect inference performance of a neural model, such as weights, coefficients, biases, or some combination thereof. In at least one embodiment, values are referred to as data values. In at least one embodiment, neural model parameters can be modified periodically based on rewards generated during a deep learning reinforced neural model training process. In at least one embodiment, neural model parameters include parameters that indicate a learning rate for a neural model, a local iteration for a neural model, an aggregated weight for a neural model, a number of neurons for a neural model, a loss value, or some combination thereof.
[0077] In at least one embodiment, a neural model is a recurrent neural model (RNN), a convolutional neural model (CNN), a generative adversarial network (GAN), a transformer, a graph neural model (GNN), or some combination thereof. In at least one embodiment, a neural model is a CNN, such as a U-Net, with a 5-level encoder-decoder network architecture, each level including a residual block. In at least one embodiment, a residual block is a stack of neural model layers, where output of one layer is added to another layer deeper in the residual block. In at least one embodiment, a layer is a structure or network topology that includes nodes corresponding to features extracted from a dataset. In at least one embodiment, an encoder-decoder network includes two recurrent neural models, one for encoding input data and another for decoding encoded data. In at least one embodiment, a recurrent network is a neural model that uses sequential data. In at least one embodiment, a training system trains an untrained neural model using training data to generate a trained neural model, a process that uses systems and methods described herein and the like. In at least one embodiment, an untrained neural model is a neural model that has been partially trained and is to be additionally trained. In at least one embodiment, training data is a training dataset. In at least one embodiment, a trained neural model is a neural model that has undergone a process of learning from a dataset by adjusting, at least in part, weights of the neural model such that the neural model infers on previously unseen input data at a level acceptable to a user and / or an application.
[0078] In at least one embodiment, neural models such as sparse quantization model 150 are applied to one or more fields such as deep learning super sampling (DLSS), autonomous vehicle navigation, medical segmentation, object classification, natural language processing (NLP), physical modeling, data processing, image processing, healthcare, financial services, robotics, or some combination thereof.
[0079] In at least one embodiment, a dense neural model is a neural model in which every neuron in a layer is connected to every neuron in a previous layer. In at least one embodiment, a first neuron in a layer is considered to be connected to a second neuron in a previous layer when an output of the second neuron is used as an input to the first neuron. In at least one embodiment, a neuron (node) is part of a layer in a neural model. In at least one embodiment, a neuron includes a function that outputs a weighted sum of input values multiplied by corresponding weights and sends the weighted sum as an input to an activation function.
[0080] In at least one embodiment, a sparse neural model is a type of neural model in which at least a portion of that neural model, such as a set of neurons of a certain layer, is connected to only a subset of neurons of a previous layer. In at least one embodiment, a sparse neural model is a neural model in which one or more weights of one or more layers have been deactivated. In at least one embodiment, a sparse layer is a type of layer in a neural model in which at least one or more values of that layer are zero. In at least one embodiment, a sparse neural model or sparse layer is said to exhibit sparsity. In at least one embodiment, a sparse neural model is a neural model that includes at least one sparse layer, which is a layer in which at least one of its values has been changed to zero. In at least one embodiment, a sparse neural model is generated by changing one or more neural model values, such as weights, to zero. In at least one embodiment, a sparse neural model is generated by modifying an architecture of a neural model, such as by removing a certain layer of a neural model. In at least one embodiment, a sparse neural model is generated by using an activation function that increases a chance that an output of that function is zero. In at least one embodiment, a sparse neural model is generated when a number of features in a feature map of that neural model, such as a feature map in a convolutional neural model (CNN), is reduced. In at least one embodiment, changing one or more neural model weights to zero is referred to as pruning. In at least one embodiment, deactivating one or more model weights by setting those weights to zero. In at least one embodiment, deactivating a weight is referred to as deactivation. In at least one embodiment, deactivating one or more weights deactivates, removes, distills, or prunes one or more features of a neural network layer. In at least one embodiment, any modification to reduce a computational complexity and / or memory requirement of that neural model is referred to as pruning. In at least one embodiment, modifying a neural network to be sparse reduces an inference accuracy of that neural network.
[0081] In at least one embodiment, a neural model that has been modified to reduce a computational complexity and / or memory requirement of that neural model is referred to as a compressed neural model. In at least one embodiment, a neural model that has been modified to reduce a computational complexity and / or memory requirement of that neural model is referred to as an efficient neural model. In at least one embodiment, any technique used to modify a neural model to reduce a computational complexity and / or memory requirement of that neural model is referred to as a neural model compression technique, an efficient technique, or an optimization technique. In at least one embodiment, a modification to a neural model refers to any change, removal, insertion, alteration, or some combination thereof to one or more portions of a neural model.
[0082] In at least one embodiment, a quantized neural model is a neural model that includes at least one quantized value. In at least one embodiment, quantization is a conversion of a value from one data format to another data format. In at least one embodiment, quantization is a modification of precision of one or more values in a neural network. In at least one embodiment, a data format is referred to as a data type. In at least one embodiment, a quantized neural model is referred to as a low-precision neural model. In at least one embodiment, a neural model that has not been quantized is referred to as a floating-point or full-precision neural model. In at least one embodiment, quantization includes an operation that converts one or more values from a higher-precision data format, such as half-precision (binary 16), single-precision (binary 32 or FP32), or double-precision (binary 64), to a lower-precision data format, such as 8-bit integer (INT8) or 16-bit integer (INT16). In at least one embodiment, a value that has been converted from a higher-precision data format to a lower-precision data format is referred to as a quantized value. In at least one embodiment, a neural model that has at least one quantized value is referred to as exhibiting quantization. In at least one embodiment, a process of converting a quantized value to a higher-precision and / or its original data format is referred to as dequantization. In at least one embodiment, one or more portions of a quantized neural model have quantized and / or dequantized values. In at least one embodiment, a dequantized value refers to a value that has been quantized and converted back to a higher-precision data format. In at least one embodiment, a process of quantizing and / or dequantizing values can result in a decrease in neural model performance of a processor due to operations required for quantization and / or dequantization. In at least one embodiment, quantizing and / or dequantizing values can decrease accuracy of a neural model due to negative side effects such as rounding errors, underflow, overflow, poor mapping between quantized values and original values, or some combination thereof. In at least one embodiment, quantizing and / or dequantizing values can affect neural model performance due to factors such as a type of quantization technique used and a type of data format that is best suited for a particular processor. In at least one embodiment, techniques described herein minimize the impact of quantization on accuracy of a neural network by, at least in part, identifying a particular neural network layer’s impact on accuracy.
[0083] In at least one embodiment, a neural model that exhibits both sparsity and quantization is referred to as a sparse quantized model or a sparse low-precision model. In at least one embodiment, a neural model that exhibits sparsity but not quantization is referred to as a sparse floating-point model or a sparse full-precision model. In at least one embodiment, a neural model that exhibits quantization but not sparsity is referred to as a dense quantized model or a dense low-precision model. In at least one embodiment, a neural model that exhibits neither sparsity nor quantization is referred to as a dense full-precision model or a dense floating-point model.
[0084] In at least one embodiment, an optimized neural model includes two or more operations fused together. In at least one embodiment, the process of merging two or more operations into a single software kernel is referred to as operation fusion. In at least one embodiment, a neural model that includes at least two or more operations fused together is referred to as a fused neural model. In at least one embodiment, operation fusion allows activation functions to be performed as part of one or more operations in one or more previous layers, such as a convolutional layer and / or a batch normalization layer.
[0085] In at least one embodiment, an optimized neural model includes two or more fused layers. In at least one embodiment, layer fusion processes redundant and / or similar layers by combining weights of these layers with corresponding layers. In at least one embodiment, a layer fused neural model is a neural model whose at least two of its layers are fused. In at least one embodiment, layer fusion includes freezing (maintaining) values in one of two layers, freezing gradients in one of two layers; using a mean value between a pair of layers during backpropagation; interpolating data between two layers; or some combination of the above.
[0086] In at least one embodiment, layers of a neural model are convolutional layers in which at least conceptually a convolution is performed, and a convolution includes a matrix multiplication of weights with input values. In at least one embodiment, weights of a neural model include hard weights and soft weights. In at least one embodiment, soft weights are capable of changing after a round of inference by a trained neural model. In at least one embodiment, hard weights do not change after a round of inference by a trained neural model. In at least one embodiment, layer fusion results in additional computational overhead due to factors such as increased computational complexity, increased memory usage, or some combination thereof. In at least one embodiment, an optimized neural model includes some combination of sparsity, quantization, operator fusion, and layer fusion techniques.
[0087] Figure 2 System 200 is shown in block diagram form, which causes one or more processors including one or more circuits to modify a dense floating point model, such as dense floating point model (M DF ) 210, to a sparse quantized model, such as sparse quantized model (M SQ ) 250. In at least one embodiment, one or more aspects of one or more embodiments described herein in connection with Figure 2 one or more aspects of one or more embodiments described herein, including at least in connection with Figure 1 and Figure 3 - Figure 8BIn at least one embodiment, the one or more processors that perform one or more operations as part of the system 200 are any one or combination of processors described herein, including any one or combination of processors described herein. Figure 8B The processor 822 described, combined with Figure 12C The accelerator 1214 described, combined with Figure 19A The graphics processor 1910 described, and the combination Figure 34 In at least one embodiment, the processor 822 performs the following operations: modifying the sparse quantization model (M) based at least in part on performing the calculation of the feature calibration partial loss 240. SQ )250. In at least one embodiment, Figure 1 The processor 104 of the system 200 performs one or more operations, such as calculating a loss value for generating the total pruning loss 260. In at least one embodiment, the one or more processors performing one or more operations of the system 200 also perform Figure 3 One or more operations of the system 300, such as computing Figure 3 In at least one embodiment, one or more processors performing one or more operations described in system 200 also perform Figure 4 One or more operations of the system 400, such as computing Figure 4 The total pruning loss 460. In at least one embodiment, the one or more processors performing one or more operations described in system 200 also perform Figure 5 One or more operations of the system 500, such as calculating the total calibration loss 570. In at least one embodiment, the one or more processors performing one or more operations of the system 200 also perform Figure 6 In at least one embodiment, one or more processors that perform one or more operations of system 200 also perform one or more operations of process 700, such as using Figure 7 Operation 704 generates a sparse floating point neural model. In at least one embodiment, Figure 8A The API in the API 810 causes one or more processors to perform one or more operations of the system 200, such as calculating the feature-based distillation loss 222. In at least one embodiment, the sparse and quantized neural model modification module 824 performs one or more operations of the system 200, such as causing the precision of one or more values of the sparse quantized model 250 to be modified.
[0088] In at least one embodiment, system 200 includes input data for one or more neural networks. In at least one embodiment, input data is input image 202. In at least one embodiment, input image 202 is used as input data in one or more processes to modify a dense full-precision neural network, such as dense floating point model 210, to become sparse and quantized while maximizing accuracy of the sparse quantized neural network. In at least one embodiment, as part of this process, input image 202 is used as input data for dense floating point model 210, sparse floating point model 230, and sparse quantized model 250. In at least one embodiment, dense floating point model 210 is a trained neural model. In at least one embodiment, dense floating point model 210 is any type of neural network. In at least one embodiment, dense floating point model 210 is a CNN trained to perform image classification using a transformer block. In at least one embodiment, a transformer block is a block of neural network layers that is a multi-headed self-attention block including one or more self-attention mechanisms and one or more feed-forward neural network layers. In at least one embodiment, a CNN using a transformer block in image classification treats an image as a sequence of image patches, similar to the way a sequence of words is embedded in text-based tasks.
[0089] In at least one embodiment, one or more processors of system 200 perform operations to modify different versions of a neural network, such as a dense floating point version, a sparse floating point version, and a sparse quantized version, using hard label predictions, soft logits, feature maps, or some combination thereof. In at least one embodiment, one or more processors perform operations that cause modification of dense floating point model 210, such as changing one or more weights of one or more layers to zero. In at least one embodiment, one or more processors first cause modification of dense floating point model 210 so that it becomes sparse floating point model (M SF ) 230. In at least one embodiment, one or more processors first cause other types of optimization modifications to be made, such as those that cause quantization. In at least one embodiment, first making a dense floating point model sparse so that the work of quantizing the sparse model at least conceptually takes into account the sparsity of that model, which helps to maximize accuracy of that model.
[0090] In at least one embodiment, dense floating point model (M DF ) 210 is modified to become sparse floating point model (M SF)230, the technique at least conceptually considers the trained dense floating point model as a supervisory model or teacher model and its modified sparse version as a student model. In at least one embodiment, the resulting sparse floating point model (M SF )230 relies on a comparison between the outputs (such as activation values) of the corresponding layers of each model. In at least one embodiment, the modification of the dense floating point model to become a sparse floating point model is based at least in part on one or more comparisons of the activations of the sparse floating point version of the neural network with the activations of the dense floating point version of the neural network. In at least one embodiment, the modification of the dense floating point model to become a sparse floating point model is based at least in part on the deactivation of one or more weights of the sparse floating point model. In at least one embodiment, the deactivation of weights in the sparse floating point model may result in a change and / or difference in the accuracy of the layer whose weights are deactivated. In at least one embodiment, one or more processors quantize the change in accuracy of one or more layers, wherein the change in accuracy is used at least in part to cause a modification of the accuracy of one or more corresponding layers in the sparse quantized model, which will be described herein in conjunction with at least the modification of the weight factors (such as Figure 4 441). In at least one embodiment, corresponding layers refer to two layers, where a first layer exists in a first version of the neural network and a second layer is the first layer, except that it is modified and exists in the second version of the neural network. For example, in at least one embodiment, layer 2 of the first neural network corresponds to layer 2 of the second neural network, and the second neural network is a modification of the first neural network. In at least one embodiment, a greater similarity between the outputs of corresponding layers between the dense floating point model and the sparse floating point model provides that these corresponding layers are more important to the accuracy of their respective neural networks (compared to corresponding layers that show less similarity). In at least one embodiment, the similarity between corresponding layers of the dense floating point model 210 and the sparse floating point model 230 is based at least in part on the outputs of these corresponding layers, which outputs indicate which features (e.g., which features in an image) have been activated by one or more activation functions.
[0091] In at least one embodiment, one or more processors perform the following operations: compute or determine loss values using corresponding converter blocks and / or layers of dense floating point model 210 and sparse floating point model 230. In at least one embodiment, loss values based on corresponding converter blocks and / or layers of dense floating point model 210 and sparse floating point model 230 are part of a feature distillation 220. In at least one embodiment, feature distillation is at least partially a process, sometimes also referred to as knowledge distillation, that deactivates certain features corresponding to weights set to zero by a neural network while maintaining performance of that neural network above a threshold level set by a user and / or application. In at least one embodiment, performance is measured by performance indicators such as inference speed and accuracy level.
[0092] In at least one embodiment, a performance indicator is a value related to accuracy of a neural network such as classification accuracy, log loss, true positives, true negatives, false positives, false negatives, precision, area under curve, mean absolute error, mean squared error, or some combination thereof. In at least one embodiment, a performance indicator related to accuracy of a neural network is referred to as an accuracy indicator. In at least one embodiment, a performance indicator is a unit of measure such as floating point operations per second (FLOPS), tera operations per second (TOPS), throughput, or some combination thereof. In at least one embodiment, throughput is a number of operations performed in a given time period. In at least one embodiment, a performance indicator is referred to as a performance specification, performance level, or level. In at least one embodiment, a performance indicator is inference speed, i.e., a number of elements of input data, such as number of image samples, for which a neural network generates output in a given time. In at least one embodiment, inference speed is expressed in units of samples per second or inputs per second. In at least one embodiment, throughput refers to a number of operations that can be processed per unit of time. In at least one embodiment, a performance indicator is a value such as a factor, score, integer, real number, or percentage that indicates how much faster or slower a high-efficiency neural network performs on a particular hardware resource compared to a dense neural network.
[0093] In at least one embodiment, a part loss of feature distillation 220 refers to a difference or distance between features included in one or more converter blocks and / or layers of two versions of a neural network, such as a dense floating point version (teacher model) and a sparse floating point version (student model). In at least one embodiment, a transformation function is applied to features of a teacher model and a student model. In at least one embodiment, a transformation function modifies features or representations learned by a neural network. In at least one embodiment, a loss function for determining feature distillation 220 is:
[0094] Ldistill = d(T t (Ft ), T s (F s )),
[0095] where d is a distance function, T t and T s are transformation functions of a teacher model and a student model, respectively. In at least one embodiment, F t and F s represent one or more sets of features of a teacher model and a student model, respectively.
[0096] In at least one embodiment, a partial loss refers to one or more loss values, such as feature-based distillation loss 222 and feature-based calibration loss 242, that are to be combined to output a total loss value. In at least one embodiment, a partial loss refers to one or more loss values derived based on a comparison between corresponding layers and / or transformer blocks of two or more versions of a neural network. In at least one embodiment, feature distillation partial loss 220 is used by one or more processors to generate one or more weights to be applied to other partial losses, where the weights include weight factors 341. In at least one embodiment, a larger feature distillation partial loss indicates that a corresponding layer of the two versions of the neural network has less of an impact on the accuracy of its respective neural network. In at least one embodiment, a smaller feature distillation partial loss indicates that a corresponding layer of the two versions of the neural network has more of an impact on the accuracy of its respective neural network. In at least one embodiment, feature distillation partial loss 220 and weight factors 341 are further described herein in connection with process 700 of Figure 7 .
[0097] In at least one embodiment, once sparse floating point model 230 is modified to be sparse, and its level of accuracy is acceptable to a user and / or application, one or more processors perform one or more operations on quantized sparse floating point model 230 to become a sparse quantized model (M SQ ) 250. In at least one embodiment, quantized sparse floating point model 230 includes one or more comparisons between transformer blocks and / or layers of sparse floating point model 230 and versions of the model after it undergoes an iterative process of becoming a sparse quantized version of sparse floating point model 230. In at least one embodiment, the one or more comparisons between sparse floating point model 230 and its sparse quantized version is a weighted partial loss of feature calibration 240. In at least one embodiment, the weighted partial loss is generated by multiplying a weight factor with a loss function and / or loss value, as further described herein in connection with, at least, process 700 of Figure 5 , Figure 6 and Figure 7 .
[0098] In at least one embodiment, a dense floating point model (M DF ) 210 is modified to create a sparse floating point model (M SF ) and a sparse quantized model (M SQ ) 250 through an iterative knowledge distillation process that includes softmax functions 212a, 212b, 232a, 232b, 252a, and 252b that output, at least in part, a probability distribution of potential correct predictions made by these neural networks as they are modified. In at least one embodiment, softmax functions include an associated temperature setting that adjusts a distribution of output probability distributions, such as T = t and T = 1. In at least one embodiment, softmax functions are applied to hard label predictions 216, 234, and 256 of respective versions of a neural network as shown in FIG. 21. In at least one embodiment, hard label predictions are predictions of a neural network that are limited to binary choices, such as true or false. In at least one embodiment, softmax functions are applied to soft logits, where soft logits refer to outputs after a softmax function is applied to raw, un-scaled outputs of a neural network. In at least one embodiment, soft logits represent probabilities that a prediction of a neural network is accurate. Figure 2
[0099] In at least one embodiment, one or more processors combine (e.g., sum) feature distillation partial losses 220 to output a feature-based distillation loss 222 that represents a total difference in features learned and / or activated by dense floating point model 210 and sparse floating point model 230. In at least one embodiment, one or more processors combine feature calibration weighted partial losses 240 to output a feature-based calibration loss 242. In at least one embodiment, feature calibration partial losses 240 indicate a difference in prediction confidence of corresponding layers and / or transformer blocks of two versions of a neural network. In at least one embodiment, feature-based calibration loss 242 represents a total confidence of outputs of sparse floating point model 230 and sparse quantized model 250.
[0100] In at least one embodiment, one or more processors perform the following operations: apply one or more loss functions to the soft logits predictions 214, 236, and 254, and combine the outputs of these loss functions to output the soft logits predictions 214, 236, and 254. In at least one embodiment, the one or more processors use, at least in part, the soft logits predictions 214, 236, and 254 to generate a soft logits distillation loss 239 and a soft logits calibration loss 259. In at least one embodiment, the one or more processors use, at least in part, the hard label distillation predictions 216 and 234 to generate a hard label distillation loss, such as one or more of the hard label distillation losses 238. In at least one embodiment, the hard label distillation loss 238 refers to the total difference in predicted hard labels (such as binary classification) between the dense floating point model 210 and the sparse floating point model 230. In at least one embodiment, the one or more processors use the hard label predictions 234 and the hard label predictions 256 to generate a hard label calibration loss 258, which represents the difference in predicted hard labels between the sparse floating point model 230 and the quantized sparse model 250.
[0101] In at least one embodiment, the total pruning loss 260 is compared to the total calibration loss 270 as part of an iterative process to modify one or more portions of the dense floating point model 210, the sparse floating point model 230, and the sparse quantized model 250. In at least one embodiment, one or more processors combine the feature-based distillation loss 222, the hard label distillation loss 238, and the soft logits distillation loss 239 to generate the total pruning loss 260. In at least one embodiment, one or more processors combine the feature-based calibration loss 242, the hard label calibration loss 258, and the soft logits calibration loss 259 to generate the total calibration loss 270.
[0102] Figure 3 A system 300 is shown in block diagram form that causes one or more processors including one or more circuits to generate values that are used, in part, to modify a dense floating point model 310 to generate a sparse floating point model, such as Figure 4 Sparse floating point model (M SF )430. In at least one embodiment, this article combines Figure 3 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figure 1 - Figure 2 and Figure 4 - Figure 8BIn at least one embodiment, the one or more processors that perform one or more operations as part of the system 300 are any one or combination of processors described herein, including any one or combination of processors described herein. Figure 8B The processor 822 described, combined with Figure 12C The accelerator 1214 described, combined with Figure 19A The described graphics processor 1910 and the combination Figure 34 In at least one embodiment, processor 822 performs one or more operations of system 300, such as calculating weight factors 341. In at least one embodiment, Figure 1 The processor 104 of the system 300 performs one or more operations, such as calculating a loss value used in part to generate the total pruning loss 360. In at least one embodiment, the one or more processors that perform one or more operations of the system 200 also perform Figure 2 One or more operations of the system 300, such as computing Figure 3 320. In at least one embodiment, the one or more processors performing one or more operations described in system 300 also perform Figure 3 One or more operations of the system 400, such as computing Figure 5 The total pruning loss 460. In at least one embodiment, the one or more processors performing one or more operations described in system 300 also perform Figure 6 One or more operations of the system 500, such as calculating the feature-based calibration loss 542. In at least one embodiment, the one or more processors performing one or more operations of the system 300 also perform Figure 7 One or more operations of the system 600, such as calculating the weighted partial loss 640. In at least one embodiment, one or more processors that perform one or more operations of the system 300 also perform one or more operations of the process 700, such as using Figure 8A Operation 704 generates a sparse floating point neural model. In at least one embodiment, Figure 3 The API in API 810 causes one or more processors to perform one or more operations of system 300, such as calculating weight factors 341. In at least one embodiment, the sparse and quantized neural model modification module 824 performs one or more operations of system 300, such as generating weight factors 341.
[0103] In at least one embodiment, system 300 depicts a dense floating point model 310 having several layers of blocks. In at least one embodiment, dense floating point model 310 includes an image patch embedding block 312, a transformer block 314, and a final projection block 316. In at least one embodiment, image patch embedding block 312 includes multiple neural network layers. In at least one embodiment, one or more processors use image patch embedding block 312 to divide one or more input images into image patches using visual markers. In at least one embodiment, transformer block 314 is a neural network layer block that are multi-headed self-attention blocks, each block including one or more self-attention mechanisms and one or more feed-forward neural network layers. In at least one embodiment, one or more processors cause embedding block 312 to output modified input data that is subsequently received and / or otherwise obtained by a transformer block labeled as level 1 of transformer block 314. In at least one embodiment, level 1 transformer block processes modified input data and outputs that processed data to a next transformer block, such as a level 2 transformer block in transformer block 314. In at least one embodiment, one or more processors repeatedly input data into transformer blocks for processing and output. In at least one embodiment, transformer block 314 has N transformer blocks and K levels. In at least one embodiment, dense floating point model 310 includes a final projection block 316. In at least one embodiment, final projection block modifies data output by a layer into a desired format. In at least one embodiment, final projection block includes a final output of a neural model of a transformer type, such as a prediction or an inference. In at least one embodiment, final projection block is used for a task such as classifying an object identified in an image.
[0104] In at least one embodiment, one or more processors cause output of one or more layers and / or blocks of dense floating point model 310 to be used, at least in part, to compute partial loss 320, as described herein at least in connection with Figure 3 In at least one embodiment, partial loss 320 is a feature distillation partial loss 220, as described further below. In at least one embodiment, one or more processors use output of one or more layers and / or blocks of sparse floating point model 430 and output of one or more layers and / or blocks of dense floating point model 310 to compute partial loss 320. Figure 3 Figure 2 In at least one embodiment, one or more processors use output of one or more layers and / or blocks of sparse floating point model 430 and output of one or more layers and / or blocks of dense floating point model 310 to compute partial loss 320.
[0105] In at least one embodiment, one or more processors perform one or more operations based on sparse floating point model 530 and Figure 2 Figure 3 corresponding layer and / or block of the sparse quantization model 650 to be applied to one or more of the partial losses of the feature calibration 240. In at least one embodiment, the one or more processors use the weight factor 341 to weight the partial loss, as depicted by the weighted partial loss 540. Figure 3 In at least one embodiment, the one or more processors use the weight factor 341 to weight the partial loss, as depicted by the weighted partial loss 540. Figure 5
[0106] The weight factor 341 is a value that is multiplied by another value, such as a partial loss value, to increase or decrease the importance of that loss value in a total loss calculation of a neural network, such as the total calibration loss 270. Figure 6 In at least one embodiment, the weight factor value is increased to increase the importance of the partial loss value in calculating a total loss value. In at least one embodiment, the weight factor value is decreased to decrease the importance of the partial loss value in calculating a total loss value. In at least one embodiment, a partial loss value is more important because the layer that the partial loss value is from is more important to the accuracy of the one or more neural networks. In at least one embodiment, when the one or more processors modify and / or calibrate values of the sparse quantization model 250, the accuracy of the sparse quantization model is compared to the sparse floating point model 230 based at least in part on comparing outputs of corresponding layers of the sparse quantization model 250 and the sparse floating point model 230, where this comparison is quantified as Figure 2 a partial loss of the feature calibration 240, Figure 5 a weighted partial loss 540, or Figure 3 a weighted partial loss 640.
[0107] In at least one embodiment, when the one or more processors perform a minimization of a feature-based calibration loss, such as at least in combination with Figure 3 , Figure 3 and Figure 2 The weights applied to the partial loss values for feature calibration at least conceptually determine the degree to which modifications made to a particular layer of a sparse quantized model affect the accuracy of the neural model. For example, in at least one embodiment, a partial loss value to which a relatively large weight is applied (as compared to a partial loss value to which a relatively small weight is applied) causes the partial loss value to increase more quickly with modifications. In at least one embodiment, the weights quantize one or more values of the corresponding layer in the model. For example, in at least one embodiment, the weights applied to the partial loss values at least conceptually manage the degree to which hardware, firmware, software, or some combination thereof quantizes values in the neural network. In at least one embodiment, the degree to which the neural network is quantized refers to the number of values that are quantized, the precision of the data type used to quantize the values, how many layers have values that are quantized, or some combination thereof.
[0108] In at least one embodiment, the processor computes feature-based distillation loss 322 using partial losses 320 based on dense float model 310 and sparse float model 430. In at least one embodiment, in addition to hard label distillation loss 238 and soft logits distillation loss 239 of FIG. 2, feature-based distillation loss 322 is used as a factor in computing total pruning loss 360. Figure 5 In at least one embodiment, total pruning loss 360 is Figure 6 total pruning loss 260 of FIG. 2.
[0109] Figure 5 System 400 is illustrated in block diagram form, which causes one or more processors comprising one or more circuits to generate values that are used, in part, to modify sparse float model 430 (as part of an iterative process) to improve the accuracy of the model. In at least one embodiment, in conjunction with Figure 1 - Figure 4 one or more aspects of one or more embodiments described herein, including at least in conjunction with Figure 6 - Figure 8B and Figure 8B embodiments described herein. In at least one embodiment, the one or more processors perform one or more operations that are used as part of system 400. In at least one embodiment, the one or more processors that perform one or more operations that are used as part of system 400 are any of the processors or combinations of processors described herein, including processor 822 in conjunction with Figure 12C described herein, accelerator 1214 in conjunction with Figure 19A described herein, graphics processor 1910 in conjunction with Figure 34 described herein, and graphics processor 1910 in conjunction with Figure 1In at least one embodiment, processor 822 performs one or more operations of system 300, such as calculating weight factors 341. In at least one embodiment, Figure 2 The processor 104 of the system 400 performs one or more operations, such as calculating a loss value used in part to generate the total pruning loss 460. In at least one embodiment, the one or more processors performing one or more operations of the system 400 also perform Figure 3 One or more operations of the system 200, such as calculating the feature distillation partial loss 220. In at least one embodiment, the one or more processors performing one or more operations of the system 400 also perform Figure 3 One or more operations of the system 300, such as computing Figure 4 320. In at least one embodiment, one or more processors performing one or more operations described in system 400 also perform Figure 6 One or more operations of the system 500, such as calculating the feature-based calibration loss 542. In at least one embodiment, the one or more processors performing one or more operations of the system 400 also perform Figure 7 One or more operations of the system 600, such as calculating the weighted partial loss 640. In at least one embodiment, one or more processors that perform one or more operations of the system 400 also perform one or more operations of the process 700, such as using Figure 8A Operation 704 generates a sparse floating point model. In at least one embodiment, Figure 4 The API in API 810 causes one or more processors to perform one or more operations of system 400, such as calculating weight factors 441. In at least one embodiment, the sparse and quantized neural model modification module 824 performs one or more operations of system 400, such as generating weight factors 441.
[0110] In at least one embodiment, the system 400 depicts a sparse floating point model 430 having several layer blocks. In at least one embodiment, the sparse floating point model 430 includes an image block embedding block 432, a transformer block 434, and a final projection block 436. In at least one embodiment, the sparse floating point model 430 is Figure 4 In at least one embodiment, the image patch embedding block 432 is a sparse version of the dense floating point model 310 of Figure 4 In at least one embodiment, the converter block 430 is a Figure 4version of converter block 314. In at least one embodiment, final projection block 436 is a version of final projection block 316.
[0111] In at least one embodiment, the one or more processors cause the output of one or more layers and / or blocks of sparse float model 430 to be used, at least in part, in computing partial losses 420, as described herein at least in connection with Figure 2 In at least one embodiment, partial losses 420 are feature distillation partial losses 220, as described further herein. In at least one embodiment, partial losses 420 are computed using the output of one or more layers and / or blocks of dense float model 310 and the output of one or more layers and / or blocks of sparse float model 430. Figure 2 In at least one embodiment, partial losses 420 are feature distillation partial losses 220, as described further herein. In at least one embodiment, partial losses 420 are computed using the output of one or more layers and / or blocks of dense float model 310 and the output of one or more layers and / or blocks of sparse float model 430. Figure 6 In at least one embodiment, partial losses 420 are feature distillation partial losses 220, as described further herein. In at least one embodiment, partial losses 420 are computed using the output of one or more layers and / or blocks of dense float model 310 and the output of one or more layers and / or blocks of sparse float model 430.
[0112] In at least one embodiment, weight factors 441 are weight factors 341. In at least one embodiment, the one or more processors perform one or more operations to apply weight factors 441 to partial losses 420 of feature calibration 240 based on the output of one or more layers and / or blocks of dense float model 310 and the output of one or more layers and / or blocks of sparse float model 430. Figure 6 In at least one embodiment, weight factors 441 are weight factors 341. In at least one embodiment, the one or more processors perform one or more operations to apply weight factors 441 to partial losses 420 of feature calibration 240 based on the output of one or more layers and / or blocks of dense float model 310 and the output of one or more layers and / or blocks of sparse float model 430. Figure 4 In at least one embodiment, the one or more processors generate one or more weight factors 441 to be applied to partial losses 420 of feature calibration 240 based on the output of one or more layers and / or blocks of sparse float model 530 and the output of one or more layers and / or blocks of sparse quantization model 650. In at least one embodiment, the one or more processors use weight factors 441 to weight partial losses of feature calibration 240, as depicted in weighted partial losses 540. Figure 5 In at least one embodiment, the one or more processors use weight factors 441 to weight partial losses, as depicted in weighted partial losses 540. Figure 6 Figure 2 In at least one embodiment, the one or more processors use partial losses 420 of dense float model 310 and sparse float model 430 to compute feature-based distillation loss 422. In at least one embodiment, feature-based distillation loss 422 is feature-based distillation loss 322.
[0113] In at least one embodiment, the one or more processors use partial losses 420 of dense float model 310 and sparse float model 430 to compute feature-based distillation loss 422. In at least one embodiment, feature-based distillation loss 422 is feature-based distillation loss 322. Figure 6 In at least one embodiment, total pruning loss 460 is total pruning loss 360 and total pruning loss 260. Figure 6 In at least one embodiment, total pruning loss 460 is total pruning loss 360 and total pruning loss 260. Figure 5 Figure 6
[0114] Figure 1 - Figure 5 In at least one embodiment, system 500 causes one or more processors comprising one or more circuits to generate values that are used, at least in part, to modify sparse float model 530 to become sparse quantization model 650. In at least one embodiment, system 500 causes one or more processors comprising one or more circuits to generate values that are used, at least in part, to modify sparse float model 530 to become sparse quantization model 650, as described herein in connection with Figure 7 - Figure 8B In at least one embodiment, system 500 causes one or more processors comprising one or more circuits to generate values that are used, at least in part, to modify sparse float model 530 to become sparse quantization model 650. In at least one embodiment, system 500 causes one or more processors comprising one or more circuits to generate values that are used, at least in part, to modify sparse float model 530 to become sparse quantization model 650, as described herein in connection with Figure 8B One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figure 12C and Figure 19A In at least one embodiment, the one or more processors that perform one or more operations as part of the system 500 are any one or combination of processors described herein, including any one or combination of processors described herein. Figure 34 The processor 822 described, combined with Figure 1 The accelerator 1214 described, combined with Figure 2 The described graphics processor 1910 and the combination Figure 3 3400. In at least one embodiment, processor 822 performs one or more operations of system 500, such as computing weighted partial losses 540. In at least one embodiment, Figure 3 The processor 104 of the system 500 performs one or more operations, such as calculating a loss value used in part to generate the total calibration loss 570. In at least one embodiment, the one or more processors performing one or more operations of the system 500 also perform Figure 4 One or more operations of the system 200, such as calculating the feature distillation partial loss 220. In at least one embodiment, the one or more processors performing one or more operations of the system 500 also perform Figure 5 One or more operations of the system 300, such as computing Figure 7 320. In at least one embodiment, one or more processors performing one or more operations described in system 500 also perform Figure 8A One or more operations of the system 400, such as calculating the feature-based distillation loss 422. In at least one embodiment, the one or more processors performing one or more operations of the system 500 also perform Figure 5 One or more operations of the system 600, such as calculating the weighted partial loss 640. In at least one embodiment, one or more processors that perform one or more operations of the system 500 also perform one or more operations of the process 700, such as using Figure 5 Operation 704 generates a sparse floating point model. In at least one embodiment, Figure 5 The API in the API 810 causes one or more processors to perform one or more operations of the system 500, such as calculating the weighted partial loss 540. In at least one embodiment, the sparse and quantized neural model modification module 824 performs one or more operations of the system 500, such as calculating the feature-based calibration loss 542.
[0115] In at least one embodiment, system 500 depicts a sparse floating point model 530 having several layers of blocks. In at least one embodiment, sparse floating point model 530 includes an image block embedding block 532, a converter block 534, and a final projection block 536. In at least one embodiment, sparse floating point model 530 is a version of sparse floating point model 430. In at least one embodiment, image block embedding block 532 is a version of image block embedding block 432. In at least one embodiment, converter block 534 is a version of converter block 434. In at least one embodiment, final projection block 536 is a version of final projection block 416. Figure 2 Figure 5 Figure 2 Figure 5
[0116] In at least one embodiment, one or more processors cause outputs of one or more layers and / or blocks of sparse floating point model 530 to be used, at least in part, to compute weighted partial losses 540, as further described herein, e.g., in connection with Figure 5 In at least one embodiment, weighted partial losses 540 are a version of weighted partial losses 640. In at least one embodiment, weighted partial losses 540 are loss values based on calibration losses of corresponding layers and / or blocks of sparse floating point model 530 and sparse quantization model 650. In at least one embodiment, one or more processors generate weighted partial losses 540 by applying a weight factor, such as weight factor 441, to calibration losses. In at least one embodiment, processors perform one or more operations to generate one or more weighted partial losses of weighted partial losses 540 based on corresponding layers and / or blocks of sparse floating point model 530 and sparse quantization model 650. Figure 5 Figure 4 In at least one embodiment, weighted partial losses 540 are a version of weighted partial losses 640. In at least one embodiment, weighted partial losses 540 are loss values based on calibration losses of corresponding layers and / or blocks of sparse floating point model 530 and sparse quantization model 650. In at least one embodiment, one or more processors generate weighted partial losses 540 by applying a weight factor, such as weight factor 441, to calibration losses. In at least one embodiment, processors perform one or more operations to generate one or more weighted partial losses of weighted partial losses 540 based on corresponding layers and / or blocks of sparse floating point model 530 and sparse quantization model 650.
[0117] In at least one embodiment, weighted partial losses 540 are a version of weighted partial losses 640. In at least one embodiment, weighted partial losses 540 are loss values based on calibration losses of corresponding layers and / or blocks of sparse floating point model 530 and sparse quantization model 650. In at least one embodiment, one or more processors generate weighted partial losses 540 by applying a weight factor, such as weight factor 441, to calibration losses. In at least one embodiment, processors perform one or more operations to generate one or more weighted partial losses of weighted partial losses 540 based on corresponding layers and / or blocks of sparse floating point model 530 and sparse quantization model 650. Figure 5 Figure 6 In at least one embodiment, one or more processors use weighted partial losses 540 to compute feature-based calibration losses 542. In at least one embodiment, feature-based calibration losses 542 are a version of feature-based calibration losses 442. In at least one embodiment, feature-based calibration losses 542 are computed using weighted partial losses 540 and a loss function, such as loss function 442. In at least one embodiment, one or more processors use feature-based calibration losses 542 to compute a feature-based calibration loss 544, as further described herein, e.g., in connection with Figure 2 Figure 5 In at least one embodiment, one or more processors use weighted partial losses 540 to compute feature-based calibration losses 542. In at least one embodiment, feature-based calibration losses 542 are a version of feature-based calibration losses 442. In at least one embodiment, feature-based calibration losses 542 are computed using weighted partial losses 540 and a loss function, such as loss function 442. In at least one embodiment, one or more processors use feature-based calibration losses 542 to compute a feature-based calibration loss 544, as further described herein, e.g., in connection with
[0118] In at least one embodiment, one or more processors use weighted partial losses 540 to compute feature-based calibration losses 542. In at least one embodiment, feature-based calibration losses 542 are a version of feature-based calibration losses 442. In at least one embodiment, feature-based calibration losses 542 are computed using weighted partial losses 540 and a loss function, such as loss function 442. In at least one embodiment, one or more processors use feature-based calibration losses 542 to compute a feature-based calibration loss 544, as further described herein, e.g., in connection with Figure 7 In at least one embodiment, the feature-based calibration loss 542 is Figure 7 Feature-based calibration loss 642.
[0119] Figure 1 - Figure 6 A system 600 is shown in block diagram form for enabling one or more processors, including one or more circuits, to generate, in part, a program for modifying Figure 8A - Figure 8B The sparse floating point model 530 is converted into the value of the sparse quantization model 650. In at least one embodiment, the present invention combines Figure 8B One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figure 12C and Figure 19A In at least one embodiment, one or more processors perform one or more operations as part of system 600. In at least one embodiment, one or more processors that perform one or more operations as part of system 600 are any one or combination of processors described herein, including in combination with Figure 34 The processor 822 described, combined with Figure 1 The accelerator 1214 described, combined with Figure 2 The described graphics processor 1910 and the combination Figure 3 3400. In at least one embodiment, processor 822 performs one or more operations of system 600, such as computing weighted partial losses 640. In at least one embodiment, Figure 3 The processor 104 performs one or more operations of the system 600, such as calculating the weighted partial loss 640. In at least one embodiment, the one or more processors performing one or more operations of the system 600 perform Figure 4 One or more operations of the system 200, such as computing the soft logits distillation loss 239. In at least one embodiment, the one or more processors performing one or more operations of the system 600 also perform Figure 5 One or more operations of the system 300, such as computing Figure 6 In at least one embodiment, one or more processors performing one or more operations described in system 600 also perform Figure 8A One or more operations of the system 400, such as calculating the feature-based distillation loss 422. In at least one embodiment, the one or more processors performing one or more operations described in the system 600 also perform Figure 1one or more operations of system 500, such as computing feature-based calibration loss 542. In at least one embodiment, one or more processors performing one or more operations of system 600 also perform one or more operations of process 700, such as using Figure 2 sparse float model. In at least one embodiment, operation 704 of Figure 3 API in API 810 causes one or more processors to perform one or more operations of system 600, such as computing weighted partial loss 640. In at least one embodiment, sparse and quantized neural model modification module 824 performs one or more operations of system 600, such as computing feature-based calibration loss 642.
[0120] In at least one embodiment, system 600 depicts sparse quantized model 650 having several layer blocks. In at least one embodiment, sparse quantized model 650 includes image patch embedding block 632, transformer block 634, and final projection block 636. In at least one embodiment, sparse quantized model 650 is a modified version of sparse float model 430. In at least one embodiment, image patch embedding block 632 is a version of image patch embedding block 532 of Figure 2 In at least one embodiment, transformer block 634 is a version of transformer block 534 of Figure 2 In at least one embodiment, final projection block 636 is a version of final projection block 536 of Figure 2
[0121] In at least one embodiment, one or more processors cause output of one or more layers and / or blocks of sparse quantized model 650 to be used, at least in part, to compute weighted partial loss 640, as further described herein, at least in connection with Figure 3 In at least one embodiment, weighted partial loss 640 is weighted partial loss 540 of Figure 4 In at least one embodiment, weighted partial loss 640 is a weighted partial loss of feature calibration 240 of Figure 5 In at least one embodiment, one or more processors use output of one or more layers and / or blocks of sparse quantized model 650, and output of one or more layers and / or blocks of sparse float model 530 of Figure 6
[0122] In at least one embodiment, weighted partial loss 640 is weighted partial loss 540 of Figure 8A In at least one embodiment, weighted partial loss 640 is based on Figure 8B In at least one embodiment, one or more processors calculate the loss value of the calibration loss of the corresponding layer and / or block of the sparse floating point model 530 and the sparse quantization model 650. Figure 19A The weight factor 441 of is applied to the calibration loss to generate the weighted partial loss 540. In at least one embodiment, the processor performs one or more of the following operations: Figure 34 The sparse floating point model 530 and Figure 1 - Figure 7 The corresponding layers and / or blocks of the sparse quantization model 650 generate one or more weighted partial losses in the feature-calibrated weighted partial losses 640.
[0123] In at least one embodiment, the processor uses the weighted partial losses 640 to at least partially calculate the feature-based calibration loss 642. In at least one embodiment, the feature-based calibration loss 642 is Figure 1 In at least one embodiment, the feature-based calibration loss 642 is Figure 1 Feature-based calibration loss 542.
[0124] Figure 2 A process 700 is shown that causes one or more processors including one or more circuits to modify a dense floating point model to become a sparse quantized model. In at least one embodiment, the present invention is combined with Figure 3 One or more aspects of one or more embodiments described herein may be combined with one or more aspects of one or more embodiments described herein, including at least one aspect of one or more embodiments described herein. Figure 4 and Figure 4 In at least one embodiment, one or more processors perform one or more operations as part of process 700. In at least one embodiment, the one or more processors that perform one or more operations as part of process 700 are any one or combination of processors described herein, including in combination with Figure 5 The processor 822 described, combined with Figure 5 The accelerator 1214 described, combined with Figure 6 The described graphics processor 1910 and the combination Figure 6 In at least one embodiment, processor 822 performs one or more operations of process 700, such as generating weight factors for the loss function using operation 708. In at least one embodiment, Figure 7 The processor 104 performs one or more operations of process 700, such as generating a sparse floating point model using operation 704. In at least one embodiment, the one or more processors performing one or more operations of process 700 perform Figure 1one or more operations of system 200, such as computing soft logits distillation loss 239. In at least one embodiment, one or more processors performing one or more operations of process 700 also perform one or more operations of process 800, such as computing soft logits distillation loss 839. Figure 1 one or more operations of system 300, such as computing Figure 1 - Figure 8A partial loss 320. In at least one embodiment, one or more processors performing one or more operations of process 700 also perform one or more operations of process 800, such as computing partial loss 839. Figure 8B one or more operations of system 400, such as computing feature-based distillation loss 422. In at least one embodiment, one or more processors performing one or more operations of process 700 also perform one or more operations of process 800, such as computing feature-based distillation loss 842. Figure 1 - Figure 8A one or more operations of system 500, such as computing feature-based calibration loss 542. In at least one embodiment, one or more processors performing one or more operations of process 700 also perform one or more operations of process 800, such as computing feature-based calibration loss 842. Figure 8B one or more operations of system 600, such as computing feature-based calibration loss 642. In at least one embodiment, Figure 1 - Figure 8A APIs in API 810 cause one or more processors to perform one or more operations of process 700, such as generating a sparse quantized model using operation 710. In at least one embodiment, sparse and quantized neural model modification module 824 performs one or more operations of process 700, such as applying loss weights to feature calibration loss of operation 712.
[0125] In at least one embodiment, one or more processors begin process 700 with operation 702 by performing operations to train a dense floating point model. In at least one embodiment, the dense floating point model is Figure 12C dense floating point model 110 of FIG. 1, Figure 19A dense floating point model 210 of FIG. 2, Figure 34 dense floating point model 310 of FIG. 3. In at least one embodiment, training a dense floating point model includes any one or combination of neural network training techniques described herein, including techniques used in supervised learning, unsupervised learning, hybrid learning, or some combination thereof. In at least one embodiment, training a dense floating point uses any type of input data, including Figure 1 input images 202 of FIG. 2.
[0126] In at least one embodiment, one or more processors continue process 700 with operation 704 by performing operations to generate a sparse floating point model based on the dense floating point model trained with operation 702. In at least one embodiment, sparse floating point model generation module 826 performs one or more operations of process 700, such as generating a sparse floating point model based on the dense floating point model trained with operation 702. In at least one embodiment, one or more processors performing one or more operations of process 700 also perform one or more operations of process 800, such as generating a sparse floating point model based on the dense floating point model trained with operation 712. Figure 1 - Figure 8AThe sparse floating point neural network model is further described with reference to one or more processors for feature distillation partial loss 220 and partial loss 320. In at least one embodiment, generating a sparse floating point model is referred to as compressing a dense floating point model trained using operation 702.
[0127] In at least one embodiment, the one or more processors continue process 700 with operation 706 by performing an operation to initialize the sparse floating point model using the weight parameters of the dense floating point model. In at least one embodiment, initializing the weight parameters generates a starting point for the weight parameters of the neural network. In at least one embodiment, the weight parameters are initialized using any initialization technique, such as zero initialization or random initialization. In at least one embodiment, the one or more processors prune one or more weight parameters by identifying one or more weight parameters that should be set to zero. In at least one embodiment, any pruning technique can be used, including amplitude-based pruning, gradient-based pruning, derivative-based pruning, norm-based pruning, or some combination thereof.
[0128] In at least one embodiment, the one or more processors continue process 700 with operation 708 by performing operations to maintain the accuracy of the sparse floating point model generated by operation 704 above a threshold. In at least one embodiment, the one or more processors maintain or maximize the accuracy of the sparse floating point model based at least in part on hard label distillation, soft logits distillation, and feature-based distillation between the dense floating point model of operation 702 and the sparse floating point model of operation 704.
[0129] In at least one embodiment, the processor continues process 700 with operation 710 by performing an accumulation or combination of hard label distillation loss, soft logits distillation loss, and feature-based distillation loss, as described herein in at least conjunction with Figure 8A In at least one embodiment, operation 710 includes one or more processors performing operations to minimize a total sparse pruning loss with respect to weight parameters of the sparse floating point model.
[0130] In at least one embodiment, the processor continues process 700 with operation 712 by causing the sparse floating point model of operation 704 to be compressed by modifying the precision (quantization) of one or more values of the sparse floating point model, as further described herein.
[0131] In at least one embodiment, the processor continues process 700 with operation 714 by performing the following operations: calibrating a quantization scale factor by computing a loss function that quantizes the difference between the sparse floating point model and the sparse quantized model. In at least one embodiment, the scale factor is a factor used in one or more quantization functions for modifying values and / or mapping high-precision values to low-precision values. In at least one embodiment, the scale factor modifies the distribution of low-precision values. In at least one embodiment, the one or more processors perform the operation of computing a loss function to calibrate the scale factor based on the hard label predictions, soft logits, and feature maps of a plurality of corresponding layers of the sparse floating point model and the sparse quantized version of the model.
[0132] In at least one embodiment, the processor continues process 700 with operation 716 by performing self-calibration of the weighting factors multiplied by the feature-based calibration loss values based, at least conceptually, on the corresponding layers between the sparse floating point model and the sparse quantized model. In at least one embodiment, self-calibration of the weighting factors refers to using at least Figure 1 - Figure 8A The weight factor is 341. Figure 8A The weight factor is 441. Figure 8A The weighted partial loss 540 and Figure 1 - 8A The one or more distillation loss values are described by the weighted partial loss 640. In at least one embodiment, self-calibration of the weighting factors refers to adjusting the weighting factors in each iteration of the process of modifying the sparse quantization model by at least partially using the total calibration loss value.
[0133] In at least one embodiment, the weights are generated by one or more processors after the dense floating point model and the sparse floating point model exhibit similar levels of accuracy. For example, in at least one embodiment, if the feature distillation loss value for a layer and its corresponding layer in the dense floating point model is large relative to other feature distillation loss values for other layers, then the one or more processors will generate a relatively small weight factor to apply to the feature-based calibration loss value for the same layer but between the sparse floating point model and the sparse quantized model. In at least one embodiment, the relatively large feature-based calibration loss value indicates a relatively low probability that the accuracy of the particular layer in the two versions of the neural network will drop below a threshold level of accuracy for the final sparse quantized model.
[0134] In at least one embodiment, the processor continues process 700 with operation 718 by performing iterative operations that minimize one or more calibration loss functions with respect to weight parameters of the sparse quantization model, as further described herein.
[0135] Figure 9AA block diagram of a driver and / or runtime according to at least one embodiment is shown, which includes one or more libraries for providing one or more application programming interfaces (APIs). In at least one embodiment, any one processor or combination of processors executes API 810, including Figure 9A Processor 822, combined with Figure 9B The described graphics processor 1910 and the combination Figure 9A In at least one embodiment, the API 810 is further described herein. In at least one embodiment, the API 810 is invoked to enable Figure 9A In at least one embodiment, the API 810 receives Figure 9B The dense floating point model 110 or an indication thereof is taken as input and makes Figure 9B The neural network quantization and sparsification module 106 performs operations and / or causes the API function 812 to modify the dense floating point model and output or generate the sparse quantized model 150, as further described herein. In at least one embodiment, the API 810 converts the dense floating point model (such as Figure 9B One or more output values of one or more layers of a dense floating point model 210) and one or more output values of one or more layers of a sparse floating point model (such as the sparse floating point model 230) are used as inputs to cause the one or more processors to perform operations to output one or more feature-based distillation partial loss values (such as Figure 9B In at least one embodiment, the API 810 receives Figure 10 The sparse floating point model 430 or an indication thereof is taken as input and outputs a weighted partial loss 430, as at least in combination with Figure 11 Further described. In at least one embodiment, API 810 receives Figure 11 The sparse floating point model 530 or an indication thereof is taken as input and outputs a weighted partial loss 540, as at least in conjunction with Figure 11 In at least one embodiment, API 810 will Figure 9A The sparse quantization model 650 or its indication as input and output weighted partial loss 640, as at least in combination with Figure 9B In at least one embodiment, the API 810 receives a dense floating point model trained using operation 702 of process 700 and outputs a sparse floating point model based on the dense floating point model, as described herein in at least conjunction with Figure 9A - Figure 11 Further description.
[0136] In at least one embodiment, software programs 802 are software modules. In at least one embodiment, software programs 802 include one or more software modules. In at least one embodiment, one or more APIs 810 are sets of software instructions that, if executed, cause one or more processors to perform one or more computational operations. In at least one embodiment, one or more APIs 810 are distributed or otherwise provided as part of one or more libraries 806, runtimes 804, drivers 804, and / or any other grouping of software and / or executable code described further herein. In at least one embodiment, one or more APIs 810 perform one or more computational operations in response to a call by a software program 802. In at least one embodiment, a software program 802 is a collection of software code, commands, instructions, or other sequences of text to be executed that instruct a computing device to perform one or more computational operations and / or call one or more other instruction sets such as APIs 810 or API functions 812 to generate a sparse quantized neural model 812. In at least one embodiment, functionality provided by one or more APIs 810 includes API functions for generating a sparse quantized neural model 812. In at least one embodiment, a software program is a compiler.
[0137] In at least one embodiment, an API 810 is a hardware interface for one or more circuits to perform one or more computational operations. In at least one embodiment, one or more software APIs 810 described herein are implemented as one or more circuits to perform one or more techniques described herein. In at least one embodiment, one or more software programs 802 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform one or more techniques described further herein.
[0138] In at least one embodiment, a software program 802 such as a user-implemented software program utilizes one or more application programming interfaces (APIs) 810 to perform various computational operations such as memory reservations, matrix multiplications, arithmetic operations, or any computational operations performed by a parallel processing unit (PPU) such as a graphics processing unit (GPU), as described further herein. In at least one embodiment, one or more APIs 810 provide a set of callable functions 812 (referred to herein as APIs, API functions, and / or functions) that each perform one or more computational operations such as those related to parallel computing. For example, in one embodiment, one or more APIs 810 provide functions 812 for causing a scheduler to schedule instructions for execution by a processor based on latency of an interconnect coupled to the processors.
[0139] In at least one embodiment, one or more software programs 802 interact with or otherwise communicate with one or more APIs 810 to perform one or more computing operations using one or more PPUs (such as GPUs). In at least one embodiment, the one or more computing operations using the one or more PPUs include at least one or more groups of computing operations that are accelerated by being executed at least in part by the one or more PPUs. In at least one embodiment, the one or more software programs 802 interact with the one or more APIs 810 to facilitate parallel computing using remote or local interfaces.
[0140] In at least one embodiment, the interface is software instructions that, if executed, provide access to one or more functions 812 provided by the one or more APIs 810. In at least one embodiment, the software programs 802 use a native interface when a software developer compiles the one or more software programs 802 in conjunction with one or more libraries 806 that include or otherwise provide access to the one or more APIs 810. In at least one embodiment, the one or more software programs 802 are statically compiled in conjunction with precompiled libraries 806 or uncompiled source code that includes instructions for executing the one or more APIs 810. In at least one embodiment, the one or more software programs 802 are dynamically compiled and the one or more software programs are linked to the one or more precompiled libraries 806 that include the one or more APIs 810 using a linker.
[0141] In at least one embodiment, a software program 802 uses a remote interface when a software developer executes a software program that utilizes a library 806 including one or more APIs 810 or otherwise communicates with it over a network or other remote communication medium. In at least one embodiment, the one or more libraries 806 including the one or more APIs 810 are executed by a remote computing service, such as a computing resource service provider. In another embodiment, the one or more libraries 806 including the one or more APIs 810 are executed by any other computing host that provides the one or more APIs 810 to the one or more software programs 802.
[0142] In at least one embodiment, a processor executing or using one or more software programs 802 invokes, uses, executes, or otherwise implements one or more APIs 810 to allocate and otherwise manage memory to be used by the software programs 802. In at least one embodiment, one or more software programs 802 utilize one or more APIs 810 to allocate and otherwise manage memory to be used by one or more portions of the software programs 802 to accelerate using one or more PPUs, such as a GPU or any other accelerator or processor described further herein. These software programs 802 can execute based, at least in part, on latency of an interconnect coupled to one or more processors using functions 812 provided by one or more APIs 810 (in one embodiment).
[0143] In at least one embodiment, API 810 is an API that facilitates parallel computing. In at least one embodiment, API 810 is any other API described further herein. In at least one embodiment, API 810 is provided by a driver and / or runtime 804. In at least one embodiment, API 810 is provided by a CUDA user mode driver. In at least one embodiment, API 810 is provided by a CUDA runtime. In at least one embodiment, driver 804 is data values and software instructions that, if executed, perform or otherwise facilitate operations of one or more functions 812 of API 810 during loading and execution of one or more portions of software programs 802. In at least one embodiment, runtime 804 is data values and software instructions that, if executed, perform or otherwise facilitate operations of one or more API functions 812 of API 810 during execution of software programs 802. In at least one embodiment, one or more software programs 802 utilize one or more APIs 810 implemented by or otherwise provided by driver and / or runtime 804 to perform combined arithmetic operations by the one or more software programs 802 during execution by one or more PPUs, such as a GPU.
[0144] In at least one embodiment, one or more software programs 802 utilize one or more APIs 810 provided by drivers and / or runtime 804 to perform combined arithmetic operations of one or more PPUs such as a GPU. In at least one embodiment, one or more APIs 810 provide combined arithmetic operations through drivers and / or runtime 804 as described above. In at least one embodiment, one or more software programs 802 utilize one or more APIs 810 provided by drivers and / or runtime 804 to allocate or otherwise reserve one or more blocks of memory 814 of one or more PPUs such as a GPU. In at least one embodiment, one or more software programs 802 utilize one or more APIs 810 provided by drivers and / or runtime 804 to allocate or otherwise reserve blocks of memory. In at least one embodiment, one or more APIs 810 are used to perform combined mathematical functions described herein.
[0145] In at least one embodiment, to improve usability of software programs 802 and / or optimize one or more portions of said software programs 802 for acceleration by one or more PPUs such as a GPU, one or more APIs 810 provide one or more API functions 812 to perform a scheduling system usable or used by one or more computing devices as described herein. In at least one embodiment, a processor executes one or more software programs to combine two or more application programming interfaces (APIs) into a single API. In at least one embodiment, a processor uses an API to cause a scheduler to select a thread selection mechanism and / or otherwise perform operations described herein. In at least one embodiment, an API calls a scheduler to cause resource allocation. In at least one embodiment, a processor uses an example API to schedule one or more instructions to be executed by one or more processors based at least in part on latency of one or more interconnects coupled with the one or more processors.
[0146] In at least one embodiment, memory 814 is system memory 1904 of computing system 1900. In at least one embodiment, memory 814 stores data parameters, weight parameters of one or more neural models described herein. In at least one embodiment, memory 814 stores data parameters used by modules described herein such as neural network quantization and sparsity module of Figure 1 - Figure 8B .
[0147] In at least one embodiment, memory 814 is a computer-readable storage medium and / or code stored thereon in a computer program, the computer program comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of computer-readable instructions described in connection with Figure 12A the operations described are not stored only in transitory signal (e.g., a propagating transient electric or electromagnetic transmission). In at least one embodiment, a non-transitory computer- readable medium does not necessarily encompass transitory signals per se. In at least one embodiment, memory 814 is implemented as a non-transitory computer-readable storage medium that stores executable instructions that, if executed by one or more processors of a computer system, cause the computer system to modify a dense floating point model to become a sparse quantized model. Figure 12A
[0148] Figure 12A A block diagram 800 is shown, in accordance with at least one embodiment, including a processor 822 to perform any one or more operations of one or more modules and / or processes described herein, in connection with Figure 12A at least one embodiment. In at least one embodiment, processor 822 is to perform modifying a dense floating point neural model to become a sparse quantized neural model, as further described herein. In at least one embodiment, one or more aspects of one or more embodiments described in connection with Figure 12A at least one embodiment. In at least one embodiment, one or more aspects of one or more embodiments described herein, including aspects described in connection with Figure 9A at least one embodiment. In at least one embodiment, processor 822 is any of or a combination of processors described herein, including accelerator 1214 described in connection with Figure 9B at least one embodiment, graphics processor 1910 described in connection with Figure 12A at least one embodiment, and parallel processing unit (“PPU”) 3400 described in connection with Figure 1 - Figure 8B at least one embodiment. In at least one embodiment, processor 822 is processor 1902 of computing system 1900.
[0149] In at least one embodiment, one or more modules described herein, including at least the modules described in connection with Figure 12B at least one embodiment, are installed on processor 822. In at least one embodiment, an exemplary module on processor 822 is sparse and quantization module 824. In at least one embodiment, sparse and quantization module 824 performs any one or more operations of modifying a dense target neural model, as described herein in connection with Figure 12A further described. In at least one embodiment, sparse and quantized neural model 824 includes Figure 12B one or more APIs 810 for performing one or more operations related to modifying a dense target neural model to become a sparse quantized neural model, as described herein in at least Figure 12B further detail.
[0150] In at least one embodiment, an exemplary module on processor 822 is API module 826. In at least one embodiment, API module 826 performs any API described herein, including APIs described in connection with Figure 12B In at least one embodiment, API module 826 includes one or more APIs and functions described in connection with Figure 1 - Figure 8B In at least one embodiment, one or more processors perform API module 826 to perform one or more APIs that modify a dense floating point neural model to a sparse quantized model, as further described herein, including in connection with Figure 12C In at least one embodiment, API module 826 includes one or more APIs and functions described in connection with
[0151] logic
[0152] Figure 12A Logic 915 is shown, in accordance with at least one embodiment, such as described elsewhere herein, which can be used in one or more devices to perform operations such as those discussed herein. In at least one embodiment, logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, logic 915 is inferencing and / or training logic. Details regarding logic 915 are provided below in conjunction with Figure 12C and / or Figure 12A In at least one embodiment, logic refers to any combination of software logic, hardware logic, and / or firmware logic that is used to provide functionality or operations described herein, where logic can collectively or individually be embodied as circuitry that forms a portion of a larger system (e.g., an integrated circuit (IC), a system on a chip (SoC), or one or more processors (e.g., CPUs, GPUs)).
[0153] In at least one embodiment, logic 915 can include, without limitation, code and / or data storage 901 for storing forward and / or output weights and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network being trained and / or used for inferencing in aspects of one or more embodiments. In at least one embodiment, logic 915 can include or be coupled to code and / or data storage 901 for storing graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic unit(s) (ALUs)). In at least one embodiment, code, such as graph code, loads weight or other parameter information into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, code and / or data storage 901 stores weight parameters and / or input / output data for each layer of a neural network being trained or used in conjunction with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inferencing using aspects of one or more embodiments. In at least one embodiment, any portion of code and / or data storage 901 can be included with other on-chip or off-chip data storage, including a processor’s Ll, L2, or L3 cache or system memory.
[0154] In at least one embodiment, any portion of code and / or data storage 901 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 901 can be cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (e.g., Flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 901 is internal or external to a processor, e.g., or includes DRAM, SRAM, Flash or some other storage type, can depend on available storage on-chip versus off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inferencing and / or training of a neural network, or some combination of these factors.
[0155] In at least one embodiment, logic 915 can include, without limitation, code and / or data storage 905 for storing backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 905 stores weight parameters and / or input / output data for each layer of a neural network trained or used in conjunction with one or more embodiments during backward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, logic 915 can include or be coupled to code and / or data storage 905 to store graph code or other software to control timing and / or order, where weight and / or other parameter information is loaded to configure logic, including integer and / or floating point units (collectively, arithmetic logic unit(s) (ALUs)).
[0156] In at least one embodiment, code such as graph code causes weight or other parameter information to be loaded into processor ALUs based on an architecture of a neural network to which that code corresponds. In at least one embodiment, any portion of code and / or data storage 905 can be included with other on-chip or off-chip data storage, including a processor’s L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of code and / or data storage 905 can be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 905 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, a choice of whether code and / or data storage 905 is internal or external to a processor, e.g., includes DRAM, SRAM, flash memory, or some other storage type, can depend on available storage on-chip versus off-chip, latency requirements of training and / or inferencing functions being performed, batch size of data used in inference and / or training of a neural network, or some combination of these factors.
[0157] In at least one embodiment, code and / or data storage 901 and code and / or data storage 905 can be separate storage structures. In at least one embodiment, code and / or data storage 901 and code and / or data storage 905 can be the same storage structure. In at least one embodiment, code and / or data storage 901 and code and / or data storage 905 can be partially combined and partially separated. In at least one embodiment, any portion of code and / or data storage 901 and code and / or data storage 905 can be included with other on-chip or off-chip data storage, including a processor’s L1, L2, or L3 cache or system memory.
[0158] In at least one embodiment, logic 915 can include, without limitation, one or more arithmetic logic units (“ALUs”) 910 (including integer and / or floating point units) for performing logical and / or mathematical operations based, at least in part, on training and / or inference code (e.g., graph code) or instructions therefrom to produce activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 920 as a function of input / output and / or weight parameter data stored in code and / or data storage 901 and / or code and / or data storage 905. In at least one embodiment, activations stored in activation storage 920 are generated in accordance with linear algebra and / or matrix-based mathematics performed by ALUs 910 in response to execution of instructions or other code, where weight values stored in code and / or data storage 905 and / or code and / or data storage 901 are used as operands along with other values such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which can be stored in code and / or data storage 905 or code and / or data storage 901 or other on-chip or off-chip storage.
[0159] In at least one embodiment, one or more ALUs 910 are included in one or more processors or other hardware logic devices or circuits, while in another embodiment one or more ALUs 910 can be external to a processor or other hardware logic devices or circuits that use them (e.g., co-processors). In at least one embodiment, ALUs 910 can be included within execution units of a processor, or otherwise included in a bank of ALUs that are accessible by execution units of a processor, which can be within a same processor or distributed between different types of processors (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, code and / or data storage 901, code and / or data storage 905, and activation storage 920 can share a processor or other hardware logic devices or circuits, while in another embodiment they can be in different processors or other hardware logic devices or circuits or some combination of same and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 920 can be included with other on-chip or off-chip data storage, including a processor’s LI, L2, or L3 cache or system memory. Moreover, inference and / or training code can be stored with other code accessible to a processor or other hardware logic or circuits, and can be fetched and / or processed using fetch, decode, schedule, execute, exit, and / or other logic circuits of a processor.
[0160] In at least one embodiment, active storage 920 can be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash), or other storage. In at least one embodiment, active storage 920 can be entirely or partially within or outside of one or more processors or other logic circuits. In at least one embodiment, choice of whether active storage 920 is internal or external to a processor, for example, or includes DRAM, SRAM, flash, or some other storage type, can depend on available storage on-chip versus off-chip, latency requirements for performing training and / or inferencing functions, batch size of data used in inferencing and / or training neural networks, or some combination of these factors.
[0161] In at least one embodiment, Figure 12C Logic 915 shown in FIG. 9A can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as Tensorflow® Processing Unit from Google, Tensor processing unit from Graphcore® AI computer, inference processing unit (IPU) from TM Graphcore®, or a “Lake Crest”) processor from Intel Corp. In at least one embodiment, Figure 12A Logic 915 shown in FIG. 9A can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware, such as field-programmable gate array (“FPGA”).
[0162] Figure 12B Logic 915 is shown in FIG. 9A in accordance with at least one embodiment. In at least one embodiment, logic 915 is inferencing and / or training logic. In at least one embodiment, logic 915 can include, without limitation, hardware logic in which compute resources are dedicated or otherwise exclusively used in conjunction with weight values or other information corresponding to one or more layers of neurons within a neural network. In at least one embodiment, Figure 12C Logic 915 shown in FIG. 9A can be used in conjunction with an application-specific integrated circuit (ASIC), such as Tensorflow® Processing Unit from Google, Tensor processing unit from Graphcore® AI computer, inference processing unit (IPU) from TM Graphcore®, or a “Lake Crest”) processor from Intel Corp. In at least one embodiment, Figure 1 - Figure 8BThe logic 915 shown in FIG. 10 can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware such as field programmable gate arrays (FPGAs), for example. In at least one embodiment, the logic 915 includes, without limitation, code and / or data storage 901 and code and / or data storage 905, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In Figure 12D In at least one embodiment, each of the code and / or data storage 901 and the code and / or data storage 905 is associated with a dedicated computing resource, such as computing hardware 902 and computing hardware 906, respectively. In at least one embodiment, each of the computing hardware 902 and the computing hardware 906 includes one or more ALUs that perform mathematical functions (e.g., linear algebra functions) only on information stored in the code and / or data storage 901 and the code and / or data storage 905, respectively, the results of which are stored in activation storage 920.
[0163] In at least one embodiment, each of the code and / or data storage 901 and 905 and the corresponding computing hardware 902 and 906 correspond to different layers of a neural network, such that activations resulting from one storage / computing pair 901 / 902 of code and / or data storage 901 and computing hardware 902 are provided as input to the next storage / computing pair 905 / 906 of code and / or data storage 905 and computing hardware 906 in order to reflect the conceptual organization of a neural network. In at least one embodiment, each storage / computing pair 901 / 902 and 905 / 906 can correspond to more than one neural network layer. In at least one embodiment, additional storage / computing pairs (not shown) can be included in the logic 915 after or in parallel with the storage / computing pairs 901 / 902 and 905 / 906.
[0164] Neural network training and deployment
[0165] Figure 12ATraining and deployment of a deep neural network is shown in accordance with at least one embodiment. In at least one embodiment, an untrained neural network 1006 is trained using a training dataset 1002. In at least one embodiment, a training framework 1004 is a PyTorch framework, while in other embodiments, the training framework 1004 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training framework. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 and enables it to be trained using processing resources described herein to generate a trained neural network 1008. In at least one embodiment, weights can be chosen randomly or by pre-training using a deep belief network. In at least one embodiment, training can be performed in a supervised, partially supervised, or unsupervised manner.
[0166] In at least one embodiment, supervised learning is used to train the untrained neural network 1006, where the training dataset 1002 includes inputs paired with desired outputs for the inputs, or where the training dataset 1002 includes inputs with known outputs and the outputs of the neural network 1006 are manually graded. In at least one embodiment, the untrained neural network 1006 is trained in a supervised manner and processes inputs from the training dataset 1002 and compares resulting outputs to a set of desired or wanted outputs. In at least one embodiment, errors are then propagated backwards through the untrained neural network 1006. In at least one embodiment, the training framework 1004 adjusts weights that control the untrained neural network 1006. In at least one embodiment, the training framework 1004 includes tools for monitoring how well the untrained neural network 1006 is converging toward a model (such as the trained neural network 1008) that is suitable for generating correct answers (such as results 1014) based on input data (such as new dataset 1012). In at least one embodiment, the training framework 1004 repeatedly trains the untrained neural network 1006 while adjusting weights to refine outputs of the untrained neural network 1006 using a loss function and adjustment algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 until the untrained neural network 1006 reaches a desired level of accuracy. In at least one embodiment, the trained neural network 1008 can then be deployed to implement any number of machine learning operations.
[0167] In at least one embodiment, unsupervised learning is used to train untrained neural network 1006, where untrained neural network 1006 attempts to train itself using unlabeled data. In at least one embodiment, unsupervised learning training dataset 1002 will include input data without any associated output data or “ground truth” data. In at least one embodiment, untrained neural network 1006 can learn groupings within training dataset 1002 and can determine how individual inputs relate to untrained dataset 1002. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in trained neural network 1008, which can perform operations useful for reducing dimensionality of new dataset 1012. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows for identification of data points in new dataset 1012 that deviate from normal patterns of new dataset 1012.
[0168] In at least one embodiment, semi-supervised learning can be used, which is a technique where a mix of labeled and unlabeled data is included in training dataset 1002. In at least one embodiment, training framework 1004 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables trained neural network 1008 to adapt to new dataset 1012 without forgetting knowledge that was imprinted into trained neural network 1008 during initial training.
[0169] In at least one embodiment, training framework 1004 is a framework that is handled in conjunction with a software development kit such as an OpenVINO (Open Visual Inference and Neural Network Optimization) toolkit. In at least one embodiment, OpenVINO toolkit is a toolkit such as developed by Intel Corporation of Santa Clara, CA. In at least one embodiment, OpenVINO includes or uses logic 915 to perform operations described herein. In at least one embodiment, a SoC, integrated circuit, or processor uses OpenVINO to perform operations described herein.
[0170] In at least one embodiment, OpenVINO is a toolkit for facilitating development of applications, particularly neural network applications, for various tasks and operations such as human vision simulation, speech recognition, natural language processing, recommendation systems, and / or variations thereof. In at least one embodiment, OpenVINO supports neural networks such as convolutional neural networks (CNNs), recurrent neural networks, and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries, e.g., OpenCV, OpenCL, and / or variations thereof.
[0171] In at least one embodiment, OpenVINO supports neural network models for various tasks and operations such as classification, segmentation, object detection, facial recognition, speech recognition, pose estimation (e.g., of people and / or objects), monocular depth estimation, image inpainting, style transfer, action recognition, colorization, and / or variations thereof.
[0172] In at least one embodiment, OpenVINO includes one or more software tools and / or modules for model optimization, also referred to as a model optimizer. In at least one embodiment, a model optimizer is a command-line tool that facilitates conversion between training and deployment of neural network models. In at least one embodiment, a model optimizer optimizes neural network models for execution on various devices and / or processing units such as GPUs, CPUs, PPUs, GPGPUs, and / or variations thereof. In at least one embodiment, a model optimizer generates an internal representation of a model and optimizes the model to generate an intermediate representation. In at least one embodiment, a model optimizer reduces a number of layers of a model. In at least one embodiment, a model optimizer removes layers of a model used for training. In at least one embodiment, a model optimizer performs various neural network operations such as modifying an input of a model (e.g., adjusting a size of an input of a model), modifying a size of an input of a model (e.g., modifying a batch size of a model), modifying a structure of a model (e.g., modifying a layer of a model), normalizing, standardizing, quantizing (e.g., converting a weight of a model from a first representation such as a floating point to a second representation such as an integer), and / or variations thereof.
[0173] In at least one embodiment, OpenVINO includes one or more software libraries for inference, also referred to as an inference engine. In at least one embodiment, an inference engine is a C++ library or a library in any suitable programming language. In at least one embodiment, an inference engine is used to infer input data. In at least one embodiment, an inference engine implements various classes to infer input data and generate one or more results. In at least one embodiment, an inference engine implements one or more API functions to process an intermediate representation, set input and / or output formats, and / or execute a model on one or more devices.
[0174] In at least one embodiment, OpenVINO provides various capabilities for heterogeneous execution of one or more neural network models. In at least one embodiment, heterogeneous execution or heterogeneous computing refers to one or more computing processes and / or systems that utilize one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software capabilities to execute programs on one or more devices. In at least one embodiment, OpenVINO provides various software capabilities to execute programs and / or portions of programs on different devices. In at least one embodiment, OpenVINO provides various software capabilities, for example, to run a first portion of code on a CPU, a second portion of code on a GPU and / or FPGA. In at least one embodiment, OpenVINO provides various software capabilities to execute one or more layers of a neural network on one or more devices (e.g., a first set of layers on a first device (e.g., GPU) and a second set of layers on a second device (e.g., CPU)).
[0175] In at least one embodiment, OpenVINO includes various capabilities similar to those associated with CUDA programming models, such as various neural network model operations associated with frameworks such as TensorFlow, PyTorch, and / or variants thereof. In at least one embodiment, one or more CUDA programming model operations are performed using OpenVINO. In at least one embodiment, various systems, methods, and / or techniques described herein are implemented using OpenVINO.
[0176] Data Center
[0177] Figure 9A An example data center 1100 that can use at least one embodiment is shown. In at least one embodiment, data center 1100 includes a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and an application layer 1140.
[0178] In at least one embodiment, as Figure 9BAs shown, the data center infrastructure layer 1110 can include a resource orchestrator 1112, grouped computing resources 1114, and node computing resources (“node C.R.s”) 1116(1)-1116(N), where “N” represents a positive integer (which can be a different integer “N” than the integer used in other figures). In at least one embodiment, the node C.R.s 1116(1)-1116(N) can include, without limitation, any number of central processing units (“CPUs” or “processors”), including accelerators, field programmable gate arrays (FPGAs), graphics processors, and / or the like, memory storage devices 1118(1)-1118(N) (e.g., dynamic read-only memory, solid-state storage, or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, and / or the like. In at least one embodiment, one or more of the node C.R.s 1116(1)-1116(N) can be a server having one or more of the above-described computing resources.
[0179] In at least one embodiment, the grouped computing resources 1114 can include individual groups of node C.R.s housed within one or more racks (not shown) or a number of racks housed within data centers (also not shown) at various geographic locations. In at least one embodiment, individual groups of node C.R.s within the grouped computing resources 1114 can include groups of computing, network, memory, or storage resources that can be configured or allocated to support one or more workloads. In at least one embodiment, a number of node C.R.s including CPUs or processors can be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, one or more racks can also include any number of power modules, cooling modules, and network switches in any combination.
[0180] In at least one embodiment, the resource orchestrator 1112 can configure or otherwise control one or more node C.R.s 1116(1)-1116(N) and / or the grouped computing resources 1114. In at least one embodiment, the resource orchestrator 1112 can include a software design infrastructure (“SDI”) management entity for the data center 1100. In at least one embodiment, the resource orchestrator 1112 can comprise hardware, software, or some combination thereof.
[0181] In at least one embodiment, as Figure 13As shown, the framework layer 1120 includes a job scheduler 1122, a configuration manager 1124, a resource manager 1126, and a distributed file system 1128. In at least one embodiment, the framework layer 1120 may include a framework that supports software 1132 of the software layer 1130 and / or one or more applications 1142 of the application layer 1140. In at least one embodiment, the software 1132 or the application 1142 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1120 may be, but is not limited to, a type of free and open source software web application framework, such as Apache Spark, which may utilize the distributed file system 1128 for large-scale data processing (e.g., "big data"). TM (hereinafter referred to as "Spark"). In at least one embodiment, the job scheduler 1122 may include a Spark driver to facilitate scheduling workloads supported by the various layers of the data center 1100. In at least one embodiment, the configuration manager 1124 may be capable of configuring different layers, such as the software layer 1130 and the framework layer 1120 including Spark and a distributed file system 1128 for supporting large-scale data processing. In at least one embodiment, the resource manager 1126 may be capable of managing the mapping or allocation of clustered or grouped computing resources to support the distributed file system 1128 and the job scheduler 1122. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1114 at the data center infrastructure layer 1110. In at least one embodiment, the resource manager 1126 may coordinate with the resource coordinator 1112 to manage these mapped or allocated computing resources.
[0182] In at least one embodiment, the software 1132 included in the software layer 1130 may include software used by at least portions of the node CRs 1116(1)-1116(N), the grouped computing resources 1114, and / or the distributed file system 1128 of the framework layer 1120. In at least one embodiment, the one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0183] In at least one embodiment, one or more applications 1142 included in application layer 1140 can include one or more types of applications used by at least portions of node C.R.s 1116(1)-1116(N), grouped computing resources 1114, and / or distributed file system 1128 of framework layer 1120. In at least one embodiment, one or more types of applications can include, but are not limited to, any number and type of genomics applications, cognitive computing, applications, and machine learning applications including training or inferencing software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0184] In at least one embodiment, any of configuration manager 1124, resource manager 1126, and resource orchestrator 1112 can implement any number and type of self-modifying actions based on any amount and type of data obtained in any technically feasible fashion. In at least one embodiment, self-modifying actions can mitigate data center operators of data center 1100 making possibly poor configuration decisions and can avoid underutilization and / or poorly performing portions of a data center.
[0185] In at least one embodiment, data center 1100 can include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information in accordance with one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained in accordance with a neural network architecture computing weight parameters by using software and computing resources described above with respect to data center 1100. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks can be used to infer or predict information using resources described above with respect to data center 1100 by using weight parameters computed through one or more training techniques described herein.
[0186] In at least one embodiment, a data center can use CPUs, application specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using resources described above. Moreover, one or more software and / or hardware resources described above can be configured as a service for allowing users to train or perform information inference such as image recognition, speech recognition, or other artificial intelligence services.
[0187] Logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. In at least one embodiment, as described herein, logic 915 can include one or more components of a neural processing unit (NPU) or other processing logic. Figure 13 and / or Details are provided regarding logic 915. In at least one embodiment, logic 915 can be used in data center 1100 for inferencing or prediction operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0188] In at least one embodiment, regarding At least one component shown or described is used to implement a technique and / or functionality described In at least one embodiment, logic 915 performs one or more operations regarding modifying a dense floating point neural model to become a sparse quantized version of the neural model using, at least in part, weighted loss values used during a quantization operation, as described elsewhere herein.
[0189] Autonomous vehicle
[0190] An example of an autonomous vehicle 1200 is shown, in accordance with at least one embodiment. In at least one embodiment, autonomous vehicle 1200 (alternatively referred to herein as “vehicle 1200”) can be, but is not limited to, a passenger vehicle such as a car, truck, bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, vehicle 1200 can be a semi tractor-trailer truck used for hauling cargo. In at least one embodiment, vehicle 1200 can be an airplane, a robotic vehicle, or another type of vehicle.
[0191] Autonomous vehicles can be described in terms of automation levels defined by the National Highway Traffic Safety Administration (“NHTSA”), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers (“SAE”) “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles” (e.g., Standard No. J3016-201806 published on June 15, 2018, Standard No. J3016-201609 published on September 30, 2016, and previous and future versions of the standard). In at least one embodiment, vehicle 1200 can be capable of functionality according to one or more of Level 1 through Level 5 of the levels of autonomous driving. For example, in at least one embodiment, vehicle 1200 can be capable of conditional automation (Level 3), high automation (Level 4), and / or full automation (Level 5), depending on the embodiment.
[0192] In at least one embodiment, vehicle 1200 can include, without limitation, components such as a chassis, a body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 1200 can include, without limitation, a propulsion system 1250 such as a combustion engine, a hybrid electric device, a fully electric motor, and / or another propulsion system type. In at least one embodiment, propulsion system 1250 can be connected to a drivetrain of vehicle 1200, which can include, without limitation, a transmission for enabling propulsion of vehicle 1200. In at least one embodiment, propulsion system 1250 can be controlled in response to receiving signals from a throttle / accelerator 1252.
[0193] In at least one embodiment, when propulsion system 1250 is operating (e.g., when vehicle 1200 is in motion), a steering system 1254 (which can include, without limitation, a steering wheel) is used to steer vehicle 1200 (e.g., along a desired path or course). In at least one embodiment, steering system 1254 can receive signals from a steering actuator 1256. In at least one embodiment, for full automation (level 5) functionality, a steering wheel can be optional. In at least one embodiment, a brake sensor system 1246 can be used to operate vehicle brakes in response to receiving signals from brake actuators 1248 and / or brake sensors.
[0194] In at least one embodiment, one or more controllers 1236, which can include, without limitation, one or more system on a chip (“SoC”) (e.g., one or more processors) can be used to control vehicle 1200. In at least one embodiment, one or more controllers 1236 can be used to control vehicle 1200 in response to receiving signals from one or more sensors 1240 and / or one or more actuators 1242. ) and / or a graphics processing unit (“GPU”) to provide signals (e.g., representing commands) to one or more components and / or systems of vehicle 1200. For example, in at least one embodiment, one or more controllers 1236 can send signals to operate vehicle brakes via brake actuator 1248, operate steering system 1254 via one or more steering actuators 1256, and operate propulsion system 1250 via one or more throttle / accelerator 1252. In at least one embodiment, one or more controllers 1236 can include one or more on-board (e.g., integrated) computing devices that process sensor signals and output operational commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in driving vehicle 1200. In at least one embodiment, one or more controllers 1236 can include a first controller for autonomous driving functionality, a second controller for functional safety functionality, a third controller for artificial intelligence functionality (e.g., computer vision), a fourth controller for infotainment functionality, a fifth controller for redundancy in emergency situations, and / or other controllers. In at least one embodiment, a single controller may handle two or more of the above functions, two or more controllers may handle a single function, and / or any combination thereof.
[0195] In at least one embodiment, the one or more controllers 1236 provide signals for controlling one or more components and / or systems of the vehicle 1200 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, the sensor data can be received from sensors such as, but not limited to, one or more global navigation satellite system ("GNSS") sensors 1258 (e.g., one or more global positioning system sensors), one or more RADAR sensors 1260, one or more ultrasonic sensors 1262, one or more LIDAR sensors 1264, one or more inertial measurement unit (IMU) sensors 1266 (e.g., one or more accelerometers, one or more gyroscopes, one or more magnetic compasses, one or more magnetometers, etc.), one or more microphones 1296, one or more stereo cameras 1268, one or more wide-angle cameras 1270 (e.g., fisheye cameras), one or more infrared cameras 1272, one or more surround cameras 1274 (e.g., 360-degree cameras), telemetry cameras (e.g., gyroscopes), and the like. Not shown), mid-range camera ( one or more speed sensors 1244 (e.g., to measure a speed of the vehicle 1200), one or more vibration sensors 1242, one or more steering sensors 1240, one or more brake sensors (e.g., as part of a brake sensor system 1246), and / or other sensor types.
[0196] In at least one embodiment, the one or more controllers 1236 can receive input (e.g., represented by input data) from an instrument cluster 1232 of the vehicle 1200 and provide output (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 1234, audible annunciators, speakers, and / or via other components of the vehicle 1200. In at least one embodiment, output can include information such as vehicle speed, velocity, time, map data (e.g., high-definition map In at least one embodiment, the HMI display 1234 can display information about the presence of one or more objects (e.g., a street sign, a warning sign, a traffic signal change, etc.) and / or information about a driving maneuver that the vehicle has made, is making, or will make (e.g., changing lanes now, reaching exit 34B in two miles, etc.). In at least one embodiment, the one or more controllers 1236 can receive input from one or more sensors 1238 (e.g., a camera, a radar, a lidar, a sonar, etc.) and / or from one or more other vehicles (e.g., via a vehicle-to-vehicle communication system). In at least one embodiment, the one or more controllers 1236 can provide output to one or more actuators 1248 (e.g., a steering actuator, a braking actuator, a throttle actuator, etc.) to control the vehicle 1200.
[0197] In at least one embodiment, the vehicle 1200 further includes a network interface 1224 that can communicate over one or more networks using one or more wireless antennas 1226 and / or one or more modems. For example, in at least one embodiment, the network interface 1224 can be capable of communicating over a Long-Term Evolution (“LTE”), Wideband Code-Division Multiple Access (“WCDMA”), Universal Mobile Telecommunications System (“UMTS”), Global System for Mobile Communications (“GSM”), IMT-CDMA Multi-Carrier (“CDMA2000”) network, etc. In at least one embodiment, the one or more wireless antennas 1226 can also enable communication between objects (e.g., vehicles, mobile devices, etc.) in an environment using one or more local area networks (such as Bluetooth, Bluetooth Low Energy (LE), Z-Wave, ZigBee, etc.) and / or one or more Low-Power Wide-Area Networks (“LPWANs”) (such as LoRaWAN, SigFox, etc. protocols).
[0198] Logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. As described herein, operations of the logic 915 can be in response to commands received from a processor or other hardware logic. and / or Details are provided regarding logic 915. In at least one embodiment, logic 915 can be used in vehicle 1200 to perform inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases as described herein.
[0199] In at least one embodiment, At least one component shown or described is used to implement the combination In at least one embodiment, wireless antenna 1226 is used to transmit and / or receive one or more wireless signals that are used as input to a sparsely quantized neural model that assists in autonomous driving and is generated by modifying a dense floating-point neural model to a sparsely quantized version of the neural model based at least in part on a weighted loss value used during quantization, as described elsewhere herein.
[0200] According to at least one embodiment, 1200. In at least one embodiment, the cameras and respective fields of view are an example embodiment and are not intended to be limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or the cameras may be located in different locations on vehicle 1200.
[0201] In at least one embodiment, the camera type used for the camera may include, but is not limited to, a digital camera that may be suitable for use with components and / or systems of the vehicle 1200. In at least one embodiment, one or more cameras may operate at Automotive Safety Integrity Level ("ASIL") B and / or other ASILs. In at least one embodiment, the camera type may be capable of having any image capture rate, such as 60 frames per second (fps), 120fps, 240fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, other types of shutters, or combinations thereof. In at least one embodiment, the color filter array may include a red-clear-clear-clear ("RCCC") filter array, a red-clear-clear-blue ("RCCB") filter array, a red-blue-green-clear ("RBGC") filter array, a Foveon X3 filter array, a Bayer sensor ("RGGB") filter array, a monochrome sensor filter array, and / or other types of filter arrays. In at least one embodiment, a clear pixel camera, such as one having an RCCC, RCCB, and / or RBGC color filter array, may be used in an effort to increase photosensitivity.
[0202] In at least one embodiment, one or more cameras can be used to perform advanced driver assistance system (“ADAS”) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono camera can be installed to provide functionality including lane departure warning, traffic sign assist, and intelligent headlamp control. In at least one embodiment, one or more cameras (e.g., all cameras) can simultaneously record and provide image data (e.g., video).
[0203] In at least one embodiment, one or more cameras can be mounted in mounting assemblies, such as custom designed (three-dimensional (“3D”) printed) assemblies, in order to remove stray light and reflected light from within vehicle 1200 (e.g., reflected light from a dashboard reflected in a windshield mirror) that can interfere with image data capture capabilities of the cameras. With regard to a rearview mirror mounting assembly, in at least one embodiment, a rearview mirror assembly can be 3D printed custom designed such that a camera mounting plate matches a shape of the rearview mirror. In at least one embodiment, one or more cameras can be integrated into a rearview mirror. In at least one embodiment, for side view cameras, one or more cameras can also be integrated within four pillars at each corner of a cabin.
[0204] In at least one embodiment, cameras with a field of view that includes portions of an environment in front of vehicle 1200 (e.g., front-facing cameras) can be used for surround view to help identify a path and obstacles ahead, as well as to assist in providing information that is critical to generating an occupancy grid and / or determining a preferred vehicle path with the help of one or more controllers 1236 and / or a control SoC. In at least one embodiment, front-facing cameras can be used to perform many ADAS functions similar to LIDAR, including but not limited to emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, front-facing cameras can also be used for ADAS functions and systems including, but not limited to, lane departure warning (“LDW”), automatic cruise control (“ACC”), and / or other functions such as traffic sign recognition.
[0205] In at least one embodiment, a wide variety of cameras can be used in a front-facing configuration, including for example, a mono camera platform including a CMOS (“complementary metal-oxide semiconductor”) color imager. In at least one embodiment, a wide-angle camera 1270 can be used to perceive objects (e.g., pedestrians, intersection traffic, or bicycles) entering the view from the periphery. Although in Only one wide-view camera 1270 is shown, but in other embodiments there can be any number (including zero) of wide-view cameras on the vehicle 1200. In at least one embodiment, any number of long-range cameras 1298 (e.g., a pair of long-range stereo cameras) can be used for depth-based object detection, especially for objects for which a neural network has not been trained. In at least one embodiment, one or more long-range cameras 1298 can also be used for object detection and classification as well as basic object tracking.
[0206] In at least one embodiment, any number of stereo cameras 1268 can also be included in a forward-facing configuration. In at least one embodiment, one or more stereo cameras 1268 can include an integrated control unit that includes a scalable processing unit that can provide programmable logic (“FPGA”) and a multi-core microprocessor with controller area network (“CAN”) or Ethernet interfaces integrated on a single chip. In at least one embodiment, such a unit can be used to generate a 3D map of the environment of the vehicle 1200, including distance estimates for all points in the image. In at least one embodiment, one or more stereo cameras 1268 can include, without limitation, a compact stereo-vision sensor that can include, without limitation, two camera lenses (one on the left and one on the right, respectively) and an image processing chip that can measure distances from the vehicle 1200 to target objects and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane-departure warning functions. In at least one embodiment, other types of stereo cameras 1268 can be used in addition to or instead of those described herein.
[0207] In at least one embodiment, cameras with fields of view that include portions of the environment on the sides of the vehicle 1200 (e.g., side-view cameras) can be used for surround view, which provides information used to create and update an occupancy grid, as well as to generate side impact collision warnings. For example, in at least one embodiment, surround cameras 1274 (e.g., four surround cameras as shown) can be positioned on the vehicle 1200. In at least one embodiment, one or more surround cameras 1274 can include, without limitation, any number and combination of wide-view cameras, one or more fisheye cameras, one or more 360-degree cameras, and / or the like. For example, in at least one embodiment, four fisheye cameras can be located on the front, back, and sides of the vehicle 1200. In at least one embodiment, the vehicle 1200 can use three surround cameras 1274 (e.g., left, right, and rear) and can utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround view camera.
[0208] In at least one embodiment, a camera having a field of view that includes portions of the environment behind the vehicle 1200 (e.g., a rearview camera) can be used for parking assistance, surround view, rear collision warning, and creating and updating an occupancy grid. In at least one embodiment, a variety of cameras can be used, including but not limited to cameras that are also suitable as one or more forward-facing cameras (e.g., long-range camera 1298 and / or one or more mid-range cameras 1276, one or more stereo cameras 1268, one or more infrared cameras 1272, etc.), as described herein.
[0209] In at least one embodiment, At least one component shown or described is used to implement the combination In at least one embodiment, wireless antenna 1226 is used to transmit and / or receive one or more wireless signals that are used as input to a sparsely quantized neural model that assists in autonomous driving and is generated by modifying a dense floating-point neural model to a sparsely quantized version of the neural model based at least in part on a weighted loss value used during quantization, as described elsewhere herein.
[0210] is a diagram illustrating a method according to at least one embodiment A block diagram of an example system architecture for an autonomous vehicle 1200 is provided. In at least one embodiment, Each of the components, features, and systems of vehicle 1200 is shown as being connected via bus 1202. In at least one embodiment, bus 1202 may include, but is not limited to, a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, CAN can be a network internal to vehicle 1200 that assists in controlling various features and functions of vehicle 1200, such as brake actuation, acceleration, braking, steering, wipers, etc. In at least one embodiment, bus 1202 can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). In at least one embodiment, bus 1202 can be read to find steering wheel angle, ground speed, engine revolutions per minute ("RPM"), button positions, and / or other vehicle status indicators. In at least one embodiment, bus 1202 can be an ASIL B compliant CAN bus.
[0211] In at least one embodiment, FlexRay and / or Ethernet protocols can also be used in addition to or instead of CAN. In at least one embodiment, there can be any number of buses forming bus 1202, which can include, but are not limited to, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses using different protocols. In at least one embodiment, two or more buses can be used to perform different functions, and / or can be used for redundancy. For example, a first bus can be used for collision avoidance functions, and a second bus can be used for actuation control. In at least one embodiment, each of buses 1202 can communicate with any component of vehicle 1200, and two or more of buses 1202 can communicate with corresponding components. In at least one embodiment, each of any number of system on a chip (“SoC”) 1204 (e.g., SoC 1204(A) and SoC 1204(B)), each of one or more controllers 1236, and / or every computer within a vehicle can have access to same input data (e.g., input from sensors of vehicle 1200), and can be connected to a common bus, such as a CAN bus.
[0212] In at least one embodiment, vehicle 1200 can include one or more controllers 1236, such as those described herein with respect to In at least one embodiment, controllers 1236 can be used for a wide variety of functions. In at least one embodiment, controllers 1236 can be coupled to any of vehicle 1200’s various other components and systems, and can be used to control vehicle 1200, vehicle 1200’s artificial intelligence, vehicle 1200’s infotainment, and / or other functions.
[0213] In at least one embodiment, vehicle 1200 can include any number of SoCs 1204. In at least one embodiment, each of SoCs 1204 can include, without limitation, a central processing unit (“CPU(s)”) 1206, a graphics processing unit (“GPU(s)”) 1208, one or more processors 1210, one or more caches 1212, one or more accelerators 1214, one or more data stores 1216, and / or other non-illustrated components and features. In at least one embodiment, one or more SoCs 1204 can be used to control vehicle 1200 in a wide variety of platforms and systems. For example, in at least one embodiment, one or more SoCs 1204 can be combined with a high definition (“HD”) map 1222 in a system (e.g., a system of vehicle 1200) that can obtain map refreshes and / or updates from one or more servers (not shown in FIG. 12) via a network interface 1224.
[0214] In at least one embodiment, CPU(s) 1206 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). In at least one embodiment, CPU(s) 1206 can include multiple cores and / or level two (“L2”) caches. For example, in at least one embodiment, CPU(s) 1206 can include eight cores in a coherent multi-processor configuration. In at least one embodiment, CPU(s) 1206 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., a 2 megabyte (“MB”) L2 cache). In at least one embodiment, CPU(s) 1206 (e.g., a CCPLEX) can be configured to support simultaneous cluster operation, which enables any combination of clusters of CPU(s) 1206 to be active at any given time.
[0215] In at least one embodiment, one or more CPU(s) 1206 can implement power management functionality including, but not limited to, one or more of the following features: individual hardware blocks can be automatically clock-gated at idle to conserve dynamic power; each core clock can be gated when that core is not actively executing instructions due to execution of a Wait for Interrupt (“WFI”) / Wait for Event (“WFE”) instruction; each core can be independently power-gated; when all cores are clock-gated or power-gated, each core cluster can be independently clock-gated; and / or when all cores are power-gated, each core cluster can be independently power-gated. In at least one embodiment, one or more CPU(s) 1206 can further implement enhanced algorithms for managing power states where allowed power states and expected wake-up times are specified and hardware / microcode determines optimal power states to enter for cores, clusters, and CCPLEX. In at least one embodiment, processing cores can support a simplified power state entry sequence in software where work is offloaded to microcode.
[0216] In at least one embodiment, one or more GPU(s) 1208 can include an integrated GPU (alternatively referred to herein as an “iGPU”). In at least one embodiment, one or more GPU(s) 1208 can be programmable and can be efficient for parallel workloads. In at least one embodiment, one or more GPU(s) 1208 can use an enhanced tensor instruction set. In at least one embodiment, one or more GPU(s) 1208 can include one or more streaming microprocessors, where each streaming microprocessor can include a level one (“Ll”) cache (e.g., an Ll cache with at least 96 KB of storage capacity), and two or more streaming microprocessors can share an L2 cache (e.g., an L2 cache with 512 KB storage capacity). In at least one embodiment, one or more GPU(s) 1208 can include at least eight streaming microprocessors. In at least one embodiment, one or more GPU(s) 1208 can use one or more compute application programming interfaces (“APIs”). In at least one embodiment, one or more GPU(s) 1208 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA model).
[0217] In at least one embodiment, one or more GPU(s) 1208 can be power optimized for best performance in automotive and embedded use cases. For example, in at least one embodiment, one or more GPU(s) 1208 can be fabricated on Fin Field Effect Transistor (“FinFET”) circuitry. In at least one embodiment, each streaming microprocessor can include a number of mixed-precision processing cores partitioned into a number of blocks. For example, and without limitation, 64 FP32 cores and 32 FP64 cores can be partitioned into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix arithmetic, a Level-0 (“LO”) instruction cache, a scheduler (e.g., a thread-warp scheduler) or arbiter, a dispatch unit, and / or a 64KB register file. In at least one embodiment, a streaming microprocessor can include independent parallel integer and floating point datapaths for a mix of computing and addressing operations to provide efficient execution of workloads. In at least one embodiment, a streaming microprocessor can include independent thread scheduling capabilities to enable finer-grain synchronization and cooperation between parallel threads. In at least one embodiment, a streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.
[0218] In at least one embodiment, one or more GPU(s) 1208 can include a high bandwidth memory (“HBM”) and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of about 900 GB / s. In at least one embodiment, in addition to or instead of HBM memory, a synchronous graphics random access memory (“SGRAM”) such as a fifth generation graphics double data rate type synchronous random access memory (“GDDR5”) can be used.
[0219] In at least one embodiment, one or more GPUs 1208 can include unified memory technology. In at least one embodiment, address translation services (“ATS”) support can be used to allow one or more GPUs 1208 to directly access one or more CPU 1206 page tables. In at least one embodiment, when a memory management unit (“MMU”) of a GPU of one or more GPUs 1208 experiences a miss, an address translation request can be sent to one or more CPUs 1206. In response, in at least one embodiment, a CPU of one or more CPUs 1206 can look up a virtual-to-physical mapping for an address in its page tables and send the translation back to one or more GPUs 1208. In at least one embodiment, unified memory technology can allow a single unified virtual address space to be used for memory of both one or more CPUs 1206 and one or more GPUs 1208, simplifying programming for one or more GPUs 1208 and porting applications to one or more GPUs 1208.
[0220] In at least one embodiment, one or more GPUs 1208 can include any number of access counters that can track how frequently one or more GPUs 1208 are accessing memory of other processors. In at least one embodiment, one or more access counters can help ensure that memory pages are moved to physical memory of a processor that most frequently accesses the page, improving efficiency of memory ranges shared between processors.
[0221] In at least one embodiment, one or more SoCs 1204 can include any number of caches 1212, including those described herein. For example, in at least one embodiment, one or more caches 1212 can include a level three (“L3”) cache that can be used by both one or more CPUs 1206 and one or more GPUs 1208 (e.g., connected to one or more CPUs 1206 and one or more GPUs 1208). In at least one embodiment, one or more caches 1212 can include a write-back cache that can track a state of individual lines, such as by using a cache coherency protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, L3 cache can include 4 MB of memory or more, although smaller cache sizes can be used, depending on embodiment.
[0222] In at least one embodiment, one or more SoC(s) 1204 can include one or more accelerator(s) 1214 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, one or more SoC(s) 1204 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4 MB of SRAM) can enable hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, hardware acceleration cluster can be used to supplement GPU(s) 1208 and offload some tasks of GPU(s) 1208 (e.g., to free up more cycles of GPU(s) 1208 to perform other tasks). In at least one embodiment, accelerator(s) 1214 can be used for target workloads that are stable enough to be worthy of acceleration (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.). In at least one embodiment, CNNs can include region-based or region with convolutional neural networks (“RCNNs”) and fast RCNNs (e.g., as used for object detection) or other types of CNNs.
[0223] In at least one embodiment, one or more accelerators 1214 (e.g., hardware acceleration clusters) can include one or more deep learning accelerators (“DLAs”). In at least one embodiment, one or more DLAs can include, without limitation, one or more tensor processing units (“TPUs”) that can be configured to provide an additional 100 trillion operations per second for deep learning applications and inferencing. In at least one embodiment, a TPU can be an accelerator configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, one or more DLAs can be further optimized for a particular set of neural network types and floating point operations and inferencing. In at least one embodiment, design of one or more DLAs can provide higher performance per mm than a typical general purpose GPU, and often significantly outperform CPUs. In at least one embodiment, one or more TPUs can perform several functions including supporting, for example, INT8, INT16, and FP16 data types for features and weights, single instance convolution functionality, and post-processor functionality. In at least one embodiment, one or more DLAs can execute neural networks, especially CNNs, on processed or unprocessed data for any of a variety of functions quickly and efficiently, including, for example and without limitation: CNNs for object recognition and detection using data from camera sensors; CNNs for distance estimation using data from camera sensors; CNNs for emergency vehicle detection, as well as identification and detection, using data from microphones; CNNs for face recognition and vehicle owner identification using data from camera sensors; and / or CNNs for protection and / or security related events.
[0224] In at least one embodiment, one or more DLAs can perform any of functions of one or more GPU(s) 1208, and by using an inferencing accelerator, for example, a designer can target one or more DLAs or one or more GPU(s) 1208 for any function. For example, in at least one embodiment, a designer can concentrate processing and floating point operations of a CNN on one or more DLAs, and leave other functions to one or more GPU(s) 1208 and / or one or more accelerators 1214.
[0225] In at least one embodiment, one or more accelerators 1214 can include a programmable vision accelerator (“PVA”), which can alternatively be referred to herein as a computer vision accelerator. In at least one embodiment, a PVA can be designed and configured to accelerate computer vision algorithms used for advanced driver assistance systems (“ADAS”) 1238, autonomous driving, augmented reality (“AR”) applications, and / or virtual reality (“VR”) applications. In at least one embodiment, a PVA can provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA can include, without limitation, any number of reduced instruction set computer (“RISC”) cores, direct memory access (“DMA”), and / or any number of vector processors.
[0226] In at least one embodiment, RISC cores can interact with image sensors (e.g., image sensors of any camera described herein), image signal processors, etc. In at least one embodiment, each RISC core can include any number of memories. In at least one embodiment, RISC cores can use any of a number of protocols, depending on embodiment. In at least one embodiment, RISC cores can execute a real-time operating system (“RTOS”). In at least one embodiment, RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, RISC cores can include instruction caches and / or tightly coupled RAM.
[0227] In at least one embodiment, DMA can enable components of a PVA to access system memory independently of one or more CPUs 1206. In at least one embodiment, DMA can support any number of features for providing optimizations to a PVA, including, without limitation, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA can support up to six or more dimensions of addressing, which can include, without limitation, block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.
[0228] In at least one embodiment, the vector processor can be a programmable processor that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In at least one embodiment, the PVA can include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, the vector processing subsystem can operate as the main processing engine of the PVA and can include a vector processing unit ("VPU"), an instruction cache, and / or a vector memory (e.g., "VMEM"). In at least one embodiment, the VPU core can include a digital signal processor, such as, for example, a single instruction multiple data ("SIMD"), a very long instruction word ("VLIW") digital signal processor. In at least one embodiment, the combination of SIMD and VLIW can increase throughput and speed.
[0229] In at least one embodiment, each vector processor may include an instruction cache and may be coupled to dedicated memory. Thus, in at least one embodiment, each vector processor may be configured to execute independently of the other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to exploit data parallelism. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute a common computer vision algorithm, but on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may execute different computer vision algorithms simultaneously on an image, or even different algorithms on a sequence of images or portions of an image. In at least one embodiment, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each PVA. In at least one embodiment, the PVAs may include additional error correction code ("ECC") memory to enhance overall system security.
[0230] In at least one embodiment, one or more accelerators 1214 can include on-chip computer vision networks and static random access memory (“SRAM”) for providing high bandwidth, low latency SRAM for one or more accelerators 1214. In at least one embodiment, on-chip memory can include at least 4 MB of SRAM that includes, for example and without limitation, eight field-programmable memory blocks that are accessible by both PVA and DLA. In at least one embodiment, each pair of memory blocks can include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory can be used. In at least one embodiment, PVA and DLA can access memory via a backbone that provides PVA and DLA with high-speed access to memory. In at least one embodiment, a backbone can include on-chip computer vision networks that interconnect PVA and DLA to memory (e.g., using APB).
[0231] In at least one embodiment, on-chip computer vision networks can include an interface that determines that both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. In at least one embodiment, an interface can provide separate phases and separate channels for sending control signals / addresses / data, as well as burst-type communication for continuous data transfer. In at least one embodiment, although other standards and protocols can be used, an interface can comply with International Organization for Standardization (“ISO”) 26262 or International Electrotechnical Commission (“IEC”) 61508 standards.
[0232] In at least one embodiment, one or more SoC 1204 can include real-time ray tracing hardware accelerators. In at least one embodiment, real-time ray tracing hardware accelerators can be used to quickly and efficiently determine locations and ranges of objects (e.g., within a world model) to generate real-time visualizations simulations for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulations of SONAR systems, for general wave propagation simulations, for comparison with LIDAR data for positioning and / or other functions, and / or for other uses.
[0233] In at least one embodiment, one or more accelerators 1214 can have broad use for autonomous driving. In at least one embodiment, PVAs can be used for key processing stages in ADAS and autonomous vehicles. In at least one embodiment, capabilities of PVAs at low power and low latency are well matched to algorithmic domains that require predictable processing. In other words, PVAs excel at semi-dense or dense regular computations, even on small data sets that can require predictable runtimes with low latency and low power. In at least one embodiment, such as in vehicle 1200, PVAs can be designed to run classic computer vision algorithms as they can be efficient at object detection and integer math operations.
[0234] For example, in accordance with at least one embodiment of technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching based algorithm can be used in some examples, but this is not meant to be limiting. In at least one embodiment, applications for level 3-5 autonomous driving use motion estimation / stereo matching in their run (e.g., structure from motion, pedestrian recognition, lane detection, etc.). In at least one embodiment, a PVA can perform computer stereo vision functions on inputs from two monocular cameras.
[0235] In at least one embodiment, a PVA can be used to perform dense optical flow. For example, in at least one embodiment, a PVA can process raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, for example, a PVA is used for time-of-flight depth processing by processing raw time-of-flight data to provide processed time-of-flight data.
[0236] In at least one embodiment, DLA can be used to run any type of network to enhance control and driving safety, including, for example and without limitation, a neural network that outputs a measure of confidence for each object detection. In at least one embodiment, confidence can be represented or interpreted as a probability, or as providing a relative “weight” of each detection compared to other detections. In at least one embodiment, a confidence measure enables system to make further decisions as to which detections should be considered true positive detections and not false positive detections. In at least one embodiment, system can set a threshold for confidence, and only consider detections that exceed threshold as true positive detections. In embodiments using automatic emergency braking (“AEB”) systems, false positive detections would result in vehicle automatically performing emergency braking, which is obviously undesirable. In at least one embodiment, highly confident detections can be considered a trigger for AEB. In at least one embodiment, DLA can run a neural network for regression of confidence values. In at least one embodiment, neural network can take as its input at least some subset of parameters, such as bounding box dimensions, a ground plane estimate obtained (e.g., from another subsystem), output of one or more IMU sensors 1266 related to 3D position estimates of vehicle 1200 direction, distance, objects obtained from neural network and / or other sensors (e.g., one or more LIDAR sensors 1264 or one or more RADAR sensors 1260).
[0237] In at least one embodiment, one or more SoC(s) 1204 can include one or more data stores 1216 (e.g., memory). In at least one embodiment, one or more data stores 1216 can be on-chip memory of one or more SoC(s) 1204 that can store neural networks to be executed on one or more GPU(s) 1208 and / or DLA. In at least one embodiment, one or more data stores 1216 can have sufficient capacity to store multiple instances of a neural network for redundancy and safety. In at least one embodiment, one or more data stores 1216 can include one or more L2 or L3 caches.
[0238] In at least one embodiment, one or more SoC(s) 1204 can include any number of processor(s) 1210 (e.g., embedded processors). In at least one embodiment, one or more processor(s) 1210 can include a boot and power management processor that can be a dedicated processor and subsystem for handling boot power and management functions and related secure enclaves. In at least one embodiment, a boot and power management processor can be part of a boot sequence of one or more SoC(s) 1204 and can provide run-time power management services. In at least one embodiment, a boot power and management processor can provide clock and voltage programming, assist system low power state transitions, one or more SoC(s) 1204 thermal and temperature sensor management, and / or one or more SoC(s) 1204 power state management. In at least one embodiment, each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and one or more SoC(s) 1204 can use ring oscillators to detect temperature of one or more CPU(s) 1206, one or more GPU(s) 1208, and / or one or more accelerator(s) 1214. In at least one embodiment, if a temperature is determined to exceed a threshold, a boot and power management processor can enter a temperature fault routine and put one or more SoC(s) 1204 into a lower power state and / or put vehicle 1200 into a safe park mode for the driver (e.g., cause vehicle 1200 to safely park).
[0239] In at least one embodiment, one or more processor(s) 1210 can further include a group of embedded processors that can function as an audio processing engine that can be an audio subsystem that enables full hardware support for multi-channel audio through a wide and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processor core with a digital signal processor with dedicated RAM.
[0240] In at least one embodiment, one or more processor(s) 1210 can further include an always-on processor engine that can provide necessary hardware features to support low power sensor management and wake-up use cases. In at least one embodiment, an always-on processor engine can include, without limitation, a processor core, tightly coupled RAM, support peripherals (e.g., timers and interrupt controllers), various I / O controller peripherals, and routing logic.
[0241] In at least one embodiment, one or more processors 1210 can further include a safety cluster engine including, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, safety cluster engine can include, without limitation, two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores can operate in a lockstep mode and can function as a single core with comparison logic to detect any differences between their operations. In at least one embodiment, one or more processors 1210 can further include a real-time camera engine that can include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, one or more processors 1210 can further include a high dynamic range signal processor that can include, without limitation, an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0242] In at least one embodiment, one or more processors 1210 can include a video image compositor that can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions needed by a video playback application to produce a final image for a player window. In at least one embodiment, video image compositor can perform lens distortion correction on one or more wide-view cameras 1270, one or more surround cameras 1274, and / or one or more in-cabin monitoring camera sensors. In at least one embodiment, in-cabin monitoring camera sensors are preferably monitored by a neural network running on another instance of SoC 1204 that is configured to identify in-cabin events and respond accordingly. In at least one embodiment, in-cabin systems can perform, without limitation, lip reading to activate cellular service and place a phone call, dictate an email, change a destination of a vehicle, activate or change a vehicle’s infotainment system and settings, or provide voice-activated web surfing. In at least one embodiment, certain functionality is available to a driver when a vehicle is operating in an autonomous mode, otherwise it is disabled.
[0243] In at least one embodiment, video image compositor can include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, where motion occurs in a video, noise reduction appropriately weights spatial information, reducing a weight of information provided by adjacent frames. In at least one embodiment, where an image or portion of an image does not include motion, temporal noise reduction performed by video image compositor can use information from a previous image to reduce noise in a current image.
[0244] In at least one embodiment, video image compositor can also be configured to perform stereo rectification on input stereoscopic lens frames. In at least one embodiment, when operating system desktop is being used, video image compositor can also be used for user interface composition and does not require one or more GPU(s) 1208 to continuously render new surfaces. In at least one embodiment, when one or more GPU(s) 1208 are powered on and active for 3D rendering, video image compositor can be used to offload one or more GPU(s) 1208 to improve performance and responsiveness.
[0245] In at least one embodiment, one or more SoCs in SoC(s) 1204 can further include mobile industry processor interface (“MIPI”) camera serial interfaces for receiving video and input from cameras, high-speed interfaces, and / or video input blocks that can be used for camera and related pixel input functionality. In at least one embodiment, one or more SoC(s) 1204 can further include input / output controllers that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.
[0246] In at least one embodiment, one or more SoCs in SoC(s) 1204 can further include a wide range of peripheral interfaces for enabling communication with peripherals, audio encoders / decoders (“codecs”), power management, and / or other devices. In at least one embodiment, one or more SoC(s) 1204 can be used to process data from cameras (e.g., connected over Gigabit Multimedia Serial Link and Ethernet channels), sensors (e.g., one or more LIDAR sensors 1264, one or more RADAR sensors 1260, etc., which can be connected over Ethernet channels), data from bus 1202 (e.g., speed of vehicle 1200, steering wheel position, etc.), data from one or more GNSS sensors 1258 (e.g., connected over Ethernet bus or CAN bus), etc. In at least one embodiment, one or more SoCs in SoC(s) 1204 can further include dedicated high-performance mass storage controllers that can include their own DMA engines and can be used to free one or more CPU(s) 1206 from routine data management tasks.
[0247] In at least one embodiment, SoC(s) 1204 can be an end-to-end platform with a flexible architecture that spans automation levels 3-5, providing a comprehensive functional safety architecture that leverages and effectively uses computer vision and ADAS technology to implement diversity and redundancy, and provides a platform for flexible, reliable driving software stack, as well as deep learning tools. In at least one embodiment, SoC(s) 1204 can be faster, more reliable, and even more energy and spatial efficient than conventional systems. For example, in at least one embodiment, accelerator(s) 1214, when combined with CPU(s) 1206, GPU(s) 1208, and data storage(s) 1216, can provide a fast, efficient platform for level 3-5 autonomous vehicles.
[0248] In at least one embodiment, computer vision algorithms can be executed on CPUs, which can be configured using high-level programming languages (e.g., C) to perform various processing algorithms on various vision data. However, in at least one embodiment, CPUs typically cannot meet performance requirements of many computer vision applications, e.g., performance requirements related to execution time and power consumption. In at least one embodiment, many CPUs cannot execute complex object detection algorithms in real-time, which are used in on-board ADAS applications and actual level 3-5 autonomous vehicles.
[0249] Embodiments described herein allow multiple neural networks to be executed simultaneously and / or sequentially, and allow results to be combined together to implement level 3-5 autonomous driving functionality. For example, in at least one embodiment, CNNs executed on DLAs or discrete GPUs (e.g., GPU(s) 1220) can include text and word recognition, allowing for reading and understanding of traffic signs, including signs for which a neural network has not been specifically trained. In at least one embodiment, DLAs can also include neural networks capable of recognizing, interpreting, and providing semantic understanding of signs, and passing that semantic understanding to a path planning module running on a CPU complex.
[0250] In at least one embodiment, for Level 3, 4, or 5 driving, multiple neural networks can be run simultaneously. For example, in at least one embodiment, a warning sign stating "Caution: Flashing lights indicate icy conditions," along with electric lights, can be interpreted independently or collectively by several neural networks. In at least one embodiment, the warning sign itself can be recognized as a traffic sign by a first deployed neural network (e.g., an already trained neural network), and the text "Flashing lights indicate icy conditions" can be interpreted by a second deployed neural network, which notifies the vehicle's path planning software (preferably executing on a CPU complex) that icy conditions exist when flashing lights are detected. In at least one embodiment, flashing lights can be identified by operating a third deployed neural network over multiple frames, notifying the vehicle's path planning software of the presence (or absence) of flashing lights. In at least one embodiment, all three neural networks can run simultaneously, for example within the DLA and / or on one or more GPUs 1208.
[0251] In at least one embodiment, a CNN for face recognition and vehicle owner recognition can use data from the camera sensor to identify the presence of an authorized driver and / or owner of the vehicle 1200. In at least one embodiment, an always-on sensor processing engine can be used to unlock the vehicle when the owner approaches the driver's door and turns on the lights, and in security mode, can be used to disable the vehicle when the owner leaves the vehicle. In this way, one or more SoCs 1204 provide protection against theft and / or carjacking.
[0252] In at least one embodiment, a CNN for emergency vehicle detection and identification can use data from microphone 1296 to detect and identify emergency vehicle sirens. In at least one embodiment, one or more SoCs 1204 use a CNN to classify environmental and urban sounds, as well as classify visual data. In at least one embodiment, a CNN running on a DLA is trained to identify the relative approaching speed of an emergency vehicle (e.g., by using the Doppler effect). In at least one embodiment, the CNN can also be trained to identify emergency vehicles specific to the local area in which the vehicle is operating, as identified by one or more GNSS sensors 1258. In at least one embodiment, when operating in Europe, the CNN will seek to detect European sirens, while when operating in North America, the CNN will seek to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program can be used with the assistance of one or more ultrasonic sensors 1262 to execute emergency vehicle safety routines, slow the vehicle, pull over, stop the vehicle, and / or idle the vehicle until the emergency vehicle passes.
[0253] In at least one embodiment, vehicle 1200 can include one or more CPUs 1218 (e.g., one or more discrete CPUs or one or more dCPUs) that can be coupled to one or more SoCs 1204 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, one or more CPUs 1218 can include, for example, an X86 processor. One or more CPUs 1218 can be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between ADAS sensors and one or more SoCs 1204, and / or monitoring status and health of one or more controllers 1236 and / or an infotainment system on a chip (“infotainment SoC”) 1230. In at least one embodiment, SoC 1204 includes one or more interconnects, and interconnects can include a Peripheral Component Interconnect Express (PCIe).
[0254] In at least one embodiment, vehicle 1200 can include one or more GPUs 1220 (e.g., one or more discrete GPUs or one or more dGPUs) that can be coupled to one or more SoCs 1204 via a high-speed interconnect (e.g., NVIDIA’s NVLINK channel). In at least one embodiment, one or more GPUs 1220 can provide additional artificial intelligence functionality, such as by executing redundant and / or different neural networks, and can be used to train and / or update neural networks based at least in part on input (e.g., sensor data) from sensors of vehicle 1200.
[0255] In at least one embodiment, vehicle 1200 can further include a network interface 1224, which can include, without limitation, one or more wireless antennas 1226 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). In at least one embodiment, network interface 1224 can be used to enable wireless connectivity to the Internet cloud services (e.g., with servers and / or other network devices), with other vehicles, and / or with computing devices (e.g., client devices of passengers). In at least one embodiment, to communicate with other vehicles, a direct link can be established between vehicle 1200 and another vehicle and / or an indirect link can be established (e.g., through a network and the Internet). In at least one embodiment, a vehicle-to-vehicle communication link can be used to provide a direct link. In at least one embodiment, a vehicle-to-vehicle communication link can provide vehicle 1200 with information about vehicles in a vicinity of vehicle 1200 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1200). In at least one embodiment, this foregoing functionality can be part of a cooperative adaptive cruise control functionality of vehicle 1200.
[0256] In at least one embodiment, network interface 1224 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 1236 to communicate over a wireless network. In at least one embodiment, network interface 1224 can include a radio frequency front end to upconvert from baseband to radio frequency and downconvert from radio frequency to baseband. In at least one embodiment, frequency conversion can be performed in any technically feasible way. For example, frequency conversion can be performed through well-known processes and / or using a super-heterodyne process. In at least one embodiment, radio frequency front end functionality can be provided by a separate chip. In at least one embodiment, a network interface can include wireless functionality to communicate over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocol.
[0257] In at least one embodiment, vehicle 1200 can further include one or more data stores 1228, which can include, without limitation, off-chip (e.g., one or more off-chip SoCs 1204) storage. In at least one embodiment, one or more data stores 1228 can include, without limitation, one or more storage elements including RAM, SRAM, dynamic random access memory (“DRAM”), video random access memory (“VRAM”), flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.
[0258] In at least one embodiment, vehicle 1200 can further include one or more GNSS sensors 1258 (e.g., GPS and / or assisted GPS sensors) to assist in mapping, perception, occupancy grid generation, and / or path planning functionality. In at least one embodiment, any number of GNSS sensors 1258 can be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet-to-serial interface (e.g., RS-232) bridge.
[0259] In at least one embodiment, vehicle 1200 can further include one or more RADAR sensors 1260. In at least one embodiment, one or more RADAR sensors 1260 can be used by vehicle 1200 for long-range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, a RADAR functional safety level can be ASIL B. In at least one embodiment, one or more RADAR sensors 1260 can use CAN bus and / or bus 1202 (e.g., for transmission of data generated by one or more RADAR sensors 1260) for control and access to object tracking data, in certain examples can access an Ethernet channel for access to raw data. In at least one embodiment, a wide variety of RADAR sensor types can be used. For example and without limitation, one or more RADAR sensors 1260 can be suitable for front, rear, and side RADAR use. In at least one embodiment, one or more of RADAR sensors 1260 are pulse-Doppler RADAR sensors.
[0260] In at least one embodiment, one or more RADAR sensors 1260 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR can be used for adaptive cruise control functionality. In at least one embodiment, a long-range RADAR system can provide a wide field of view achieved through two or more independent scans (e.g., over a 250 m (meter) range). In at least one embodiment, one or more RADAR sensors 1260 can help distinguish between static and moving objects, and can be used by ADAS system 1238 for emergency brake assist and forward collision warning. In at least one embodiment, one or more sensors 1260 included in a long-range RADAR system can include, without limitation, a monostatic multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas, as well as high-speed CAN and FlexRay interfaces. In at least one embodiment, with six antennas, a central four antennas can create a focused beam pattern designed to record the environment around vehicle 1200 at higher speeds with minimal traffic interference from adjacent lanes. In at least one embodiment, other two antennas can expand the field of view, enabling it to quickly detect vehicles entering or leaving the lane of vehicle 1200.
[0261] In at least one embodiment, as an example, a mid-range RADAR system can include a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, a short-range RADAR system can include, without limitation, any number of RADAR sensors 1260 designed to be mounted at either end of a rear bumper. When mounted at either end of a rear bumper, in at least one embodiment, a RADAR sensor system can produce two beams that constantly monitor the vehicle’s rearward direction and blind spot nearness. In at least one embodiment, a short-range RADAR system can be used in ADAS system 1238 for blind spot detection and / or lane change assist.
[0262] In at least one embodiment, vehicle 1200 can further include one or more ultrasonic sensors 1262. In at least one embodiment, one or more ultrasonic sensors 1262, which can be positioned in front, back, and / or side positions of vehicle 1200, can be used for parking assist and / or to create and update an occupancy grid. In at least one embodiment, a wide variety of ultrasonic sensors 1262 can be used, and different ultrasonic sensors 1262 can be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, ultrasonic sensors 1262 can operate at a functional safety level of ASIL B.
[0263] In at least one embodiment, vehicle 1200 can include one or more LIDAR sensors 1264. In at least one embodiment, one or more LIDAR sensors 1264 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, one or more LIDAR sensors 1264 can operate at a functional safety level of ASIL B. In at least one embodiment, vehicle 1200 can include multiple (e.g., two, four, six, etc.) LIDAR sensors 1264 that can use Ethernet channels (e.g., to provide data to a Gigabit Ethernet switch).
[0264] In at least one embodiment, LIDAR sensor(s) 1264 can be capable of providing a list of objects and their distances for a 360 degree field of view. In at least one embodiment, commercially available LIDAR sensor(s) 1264 can have, for example, an advertised range of about 100 m, with an accuracy of 2 cm - 3 cm, and support for 100 Mbps Ethernet connectivity. In at least one embodiment, one or more non-protruding LIDAR sensors can be used. In such embodiments, LIDAR sensor(s) 1264 can include small devices that can be embedded into front, rear, side, and / or corner positions of vehicle 1200. In at least one embodiment, LIDAR sensor(s) 1264, in such embodiments, can provide up to 120 degrees of horizontal and 35 degrees of vertical field of view, with a range of 200 m, even for low reflectivity objects. In at least one embodiment, forward-mounted LIDAR sensor(s) 1264 can be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0265] In at least one embodiment, LIDAR technology such as 3D Flash LIDAR can also be used. In at least one embodiment, 3D Flash LIDAR uses a laser flash as a transmission source to illuminate up to about 200 m around vehicle 1200. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receiver that records laser pulse travel time and reflected light on each pixel, which in turn corresponds to a range from vehicle 1200 to an object. In at least one embodiment, flash LIDAR can allow for highly accurate and distortion-free images of surrounding environment to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors can be deployed, one on each side of vehicle 1200. In at least one embodiment, a 3D flash LIDAR system includes, without limitation, a solid-state 3D staring array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, a flash LIDAR device can use 5 nanosecond Class I (eye-safe) laser pulses per frame, and can capture reflected laser light as a 3D ranging point cloud and co-registered intensity data.
[0266] In at least one embodiment, vehicle 1200 can also include one or more IMU sensors 1266. In at least one embodiment, one or more IMU sensors 1266 can be located at a center of a rear axle of vehicle 1200. In at least one embodiment, one or more IMU sensors 1266 can include, without limitation, one or more accelerometers, one or more magnetometers, one or more gyroscopes, one or more magnetic compasses, and / or other sensor types. In at least one embodiment, such as a six-axis application, one or more IMU sensors 1266 can include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, such as a nine-axis application, one or more IMU sensors 1266 can include, without limitation, an accelerometer, a gyroscope, and a magnetometer.
[0267] In at least one embodiment, one or more IMU sensors 1266 can be implemented as a miniature, high-performance GPS-aided inertial navigation system (“GPS / INS”) that combines micro-electro-mechanical systems (“MEMS”) inertial sensors, high-sensitivity GPS receiver, and advanced Kalman filtering algorithms for providing estimates of position, velocity, and attitude. In at least one embodiment, one or more IMU sensors 1266 can enable vehicle 1200 to estimate its heading by observing and correlating changes in velocity from GPS to one or more IMU sensors 1266 without input from magnetic sensors. In at least one embodiment, one or more IMU sensors 1266 and one or more GNSS sensors 1258 can be combined in a single integrated unit.
[0268] In at least one embodiment, vehicle 1200 can include one or more microphones 1296 placed within and / or around vehicle 1200. In at least one embodiment, one or more microphones 1296 can be used for emergency vehicle detection and identification.
[0269] In at least one embodiment, vehicle 1200 can further include any number of camera types including one or more stereo cameras 1268, one or more wide-view cameras 1270, one or more infrared cameras 1272, one or more surround cameras 1274, one or more long-range cameras 1298, one or more mid-range cameras 1276, and / or other camera types. In at least one embodiment, cameras can be used to capture image data around an entire peripheral of vehicle 1200. In at least one embodiment, a type of camera used depends on vehicle 1200. In at least one embodiment, any combination of camera types can be used to provide necessary coverage around vehicle 1200. In at least one embodiment, a number of cameras deployed can vary from embodiment to embodiment. For example, in at least one embodiment, vehicle 1200 can include six cameras, seven cameras, ten cameras, twelve cameras, or other number of cameras. In at least one embodiment, cameras can support, by way of example but not limitation, Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet communications. In at least one embodiment, each camera is described in more detail previously herein with reference to and Each camera is described in more detail.
[0270] In at least one embodiment, vehicle 1200 can further include one or more vibration sensors 1242. In at least one embodiment, one or more vibration sensors 1242 can measure vibrations of components of vehicle 1200 (e.g., axles). For example, in at least one embodiment, changes in vibration can be indicative of changes in a road surface. In at least one embodiment, when two or more vibration sensors 1242 are used, differences between vibrations can be used to determine friction or slippage of a road surface (e.g., when there is a difference in vibration between a power driven axle and a freely rotating axle).
[0271] In at least one embodiment, vehicle 1200 can include an ADAS system 1238. In at least one embodiment, ADAS system 1238 can include, without limitation in some examples, a SoC. In at least one embodiment, ADAS system 1238 can include, without limitation, any number and combination of adaptive / automatic / autonomous cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward collision warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane-keep assist (“LKA”) systems, blind-spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functionality.
[0272] In at least one embodiment, ACC system can use one or more RADAR sensors 1260, one or more LIDAR sensors 1264, and / or any number of cameras. In at least one embodiment, ACC system can include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls a distance to another vehicle immediately in front of vehicle 1200 and automatically adjusts a speed of vehicle 1200 to maintain a safe distance from the vehicle in front. In at least one embodiment, a lateral ACC system performs distance keeping and advises vehicle 1200 to change lanes when necessary. In at least one embodiment, lateral ACC is related to other ADAS applications, such as LC and CW.
[0273] In at least one embodiment, a CACC system uses information from other vehicles, which can be received from other vehicles indirectly via a wireless link or over a network connection (e.g., over the Internet) via network interface 1224 and / or one or more wireless antennas 1226. In at least one embodiment, a direct link can be provided by a vehicle-to-vehicle (“V2V”) communication link, while an indirect link can be provided by an infrastructure-to-vehicle (“I2V”) communication link. Generally, V2V communication provides information about vehicles immediately in front (e.g., vehicles immediately in front of and in same lane as vehicle 1200), while I2V communication provides information about traffic further ahead. In at least one embodiment, a CACC system can include one or both of I2V and V2V sources of information. In at least one embodiment, a CACC system can be more reliable with information about vehicles in front of vehicle 1200, and has a potential to improve smoothness of traffic flow and reduce road congestion.
[0274] In at least one embodiment, a FCW system is designed to warn a driver of a hazard so that the driver can take corrective action. In at least one embodiment, a FCW system uses a forward-facing camera and / or one or more RADAR sensors 1260, coupled to a dedicated processor, DSP, FPGA, and / or ASIC, which are electrically coupled to provide driver feedback, such as a display, a speaker, and / or a vibrating component. In at least one embodiment, a FCW system can provide a warning, such as in the form of a sound, a visual warning, a vibration, and / or a quick brake pulse.
[0275] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and can automatically apply brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, AEB systems can use one or more forward facing cameras and / or one or more RADAR sensors 1260 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs. In at least one embodiment, when an AEB system detects a hazard, it will typically first warn a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system can automatically apply brakes in an attempt to prevent or at least mitigate the effects of a predicted collision. In at least one embodiment, AEB systems can include technologies such as dynamic brake support and / or crash imminent braking.
[0276] In at least one embodiment, a LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to warn the driver when vehicle 1200 crosses a lane marker. In at least one embodiment, a LDW system does not activate when the driver indicates an intentional lane departure, such as by activating a turn signal. In at least one embodiment, a LDW system can use a forward facing camera coupled to dedicated processors, DSPs, FPGAs, and / or ASICs that are electrically coupled to provide driver feedback such as displays, speakers, and / or vibrating components. In at least one embodiment, an LKA system is a variation of a LDW system. In at least one embodiment, if vehicle 1200 begins to drift from its lane, an LKA system provides a steering input or brake to correct vehicle 1200.
[0277] In at least one embodiment, a BSW system detects and warns the driver that the vehicle is in a blind spot of a car. In at least one embodiment, a BSW system can provide visual, audible, and / or tactile alerts to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system can provide additional warnings when a driver uses a turn signal. In at least one embodiment, a BSW system can use one or more rear facing cameras and / or one or more RADAR sensors 1260 coupled to dedicated processors, DSPs, FPGAs, and / or ASICs that are electrically coupled to driver feedback such as displays, speakers, and / or vibrating components.
[0278] In at least one embodiment, when an object is detected outside of a rear camera range while vehicle 1200 is backing up, RCTW system can provide a visual, audible, and / or tactile notification. In at least one embodiment, RCTW system includes an AEB system to ensure vehicle brakes are applied to avoid a collision. In at least one embodiment, RCTW system can use one or more rear-facing RADAR sensors 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to provide driver feedback such as a display, speaker, and / or vibrating component.
[0279] In at least one embodiment, conventional ADAS systems can be prone to false positives, which can annoy and distract drivers, but are typically not catastrophic because conventional ADAS systems alert the driver and allow that driver to decide whether the safety condition is truly present and act accordingly. In at least one embodiment, in case of a result conflict, vehicle 1200 itself decides whether to follow the result of the primary computer or the secondary computer (e.g., a first or second controller in controller 1236). For example, in at least one embodiment, ADAS system 1238 can be a backup and / or secondary computer for providing perception information to a backup computer plausibility module. In at least one embodiment, a backup computer plausibility monitor can run various software redundantly on hardware components to detect faults in perception and dynamic driving tasks. In at least one embodiment, output from ADAS system 1238 can be provided to a supervisory MCU. In at least one embodiment, if output from a primary computer and output from a secondary computer conflict, the supervisory MCU decides how to reconcile the conflict to ensure safe operation.
[0280] In at least one embodiment, a primary computer can be configured to provide a confidence score to a supervisory MCU that indicates a confidence of the primary computer in a selected result. In at least one embodiment, if the confidence score exceeds a threshold, the supervisory MCU can follow the primary computer’s indication regardless of whether the secondary computer provides a conflicting or inconsistent result. In at least one embodiment, in case the confidence score does not satisfy the threshold, and in case the primary computer and the secondary computer indicate different results (e.g., conflict), the supervisory MCU can arbitrate between the computers to determine an appropriate result.
[0281] In at least one embodiment, a supervisory MCU can be configured to run a neural network trained and configured to determine conditions under which an auxiliary computer provides false alarms based at least in part on outputs from a host computer and outputs from an auxiliary computer. In at least one embodiment, one or more neural networks in a supervisory MCU can learn when an output of an auxiliary computer can be trusted, and when it cannot. For example, in at least one embodiment, when the auxiliary computer is a RADAR-based FCW system, one or more neural networks in a supervisory MCU can learn when the FCW system is identifying metal objects that are not actually dangerous, such as drain grates or manhole covers that would trigger an alert. In at least one embodiment, when the auxiliary computer is a camera-based LDW system, a neural network in a supervisory MCU can learn to override LDW when a bicyclist or pedestrian is present and it is actually safest to lane depart. In at least one embodiment, a supervisory MCU can include at least one of a DLA or GPU suitable for running one or more neural networks with associated memory. In at least one embodiment, a supervisory MCU can include and / or be included as a component of one or more SoCs 1204.
[0282] In at least one embodiment, ADAS system 1238 can include an auxiliary computer that performs ADAS functions using traditional computer vision rules. In at least one embodiment, this auxiliary computer can use classic computer vision rules (if-then), and presence of one or more neural networks in a supervisory MCU can improve reliability, safety, and performance. For example, in at least one embodiment, diverse implementations and intentional non-identity make the overall system more fault-tolerant, especially to faults caused by software (or software-hardware interface) functionality. For example, in at least one embodiment, if there is a software bug or error in software running on a host computer, and an imperfectly identical piece of software code running on an auxiliary computer provides consistent overall results, then a supervisory MCU can have greater confidence that overall results are correct, and the bug in software or hardware on that host computer does not result in a significant error.
[0283] In at least one embodiment, outputs of ADAS system 1238 can be fed into a perception block of a host computer and / or a dynamic driving task block of a host computer. For example, in at least one embodiment, if ADAS system 1238 indicates a forward collision warning due to an object directly in front, a perception block can use this information when identifying the object. In at least one embodiment, as described herein, an auxiliary computer can have its own neural network trained in such a way that reduces a risk of false positives.
[0284] In at least one embodiment, vehicle 1200 can further include infotainment SoC 1230 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, in at least one embodiment, infotainment system SoC 1230 can not be a SoC and can include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 1230 can include, without limitation, a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., television, movies, streaming media, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation systems, rear park assist, radio data system, vehicle related information such as fuel level, total range, brake fuel level, oil level, doors open / close, air filter information, etc.) to vehicle 1200. For example, infotainment SoC 1230 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, in-car computer, in-car entertainment system, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (“HUD”), HMI display 1234, telematics equipment, control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, infotainment SoC 1230 can be further used to provide information (e.g., visual and / or audible information) to one or more users of vehicle 1200, such as information from ADAS system 1238, autonomous driving information (such as planned vehicle maneuvers), trajectory, surrounding environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0285] In at least one embodiment, infotainment SoC 1230 can include any number and type of GPU functionality. In at least one embodiment, infotainment SoC 1230 can communicate with other devices, systems, and / or components of vehicle 1200 over bus 1202. In at least one embodiment, infotainment SoC 1230 can be coupled to a supervisory MCU such that a GPU of the infotainment system can perform some autonomous driving functions in the event of a failure of one or more primary controllers 1236 (e.g., a primary computer and / or a backup computer of vehicle 1200). In at least one embodiment, infotainment SoC 1230 can place vehicle 1200 into a driver-to-safe-parking mode, as described herein.
[0286] In at least one embodiment, vehicle 1200 can further include an instrument cluster 1232 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, instrument cluster 1232 can include, without limitation, a controller and / or supercomputer (e.g., a discrete controller or supercomputer). In at least one embodiment, instrument cluster 1232 can include, without limitation, any number and combination of gauges such as speedometer, fuel level, oil pressure, tachometer, odometer, turn indicator, shift position indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, auxiliary restraint system (e.g., airbag) information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between infotainment SoC 1230 and instrument cluster 1232. In at least one embodiment, instrument cluster 1232 can be included as part of infotainment SoC 1230, and vice versa.
[0287] In at least one embodiment, regarding At least one component shown or described is used to implement a technique and / or functionality described In at least one embodiment, data store 1216 stores one or more output values of one or more layers of a dense floating point neural model and a sparse floating point neural model, as described elsewhere herein.
[0288] is a cloud-based server with FIG. 12 is a diagram of a system that enables communication between autonomous vehicles 1200. In at least one embodiment, the system can include, without limitation, one or more servers 1278, one or more networks 1290, and any number and type of vehicles, including vehicles 1200. In at least one embodiment, one or more servers 1278 can include, without limitation, a plurality of GPUs 1284(A)-1284(H) (collectively referred to herein as GPUs 1284), PCIe switches 1282(A)-1282(D) (collectively referred to herein as PCIe switches 1282), and / or CPUs 1280(A)-1280(B) (collectively referred to herein as CPUs 1280). In at least one embodiment, GPUs 1284, CPUs 1280, and PCIe switches 1282 can be interconnected with high-speed interconnects such as, for example and without limitation, NVLink interfaces 1288 and / or PCIe connections 1286 developed by NVIDIA. In at least one embodiment, GPUs 1284 are connected via NVLink and / or NVSwitch SoC connections, and GPUs 1284 and PCIe switches 1282 are connected via PCIe interconnects. Although eight GPUs 1284, two CPUs 1280, and four PCIe switches 1282 are illustrated, this is not intended to be limiting. In at least one embodiment, each of one or more servers 1278 can include, without limitation, any number of GPUs 1284, CPUs 1280, and / or PCIe switches 1282, in any combination. For example, in at least one embodiment, one or more servers 1278 can each include eight, sixteen, thirty-two, and / or more GPUs 1284.
[0289] In at least one embodiment, one or more servers 1278 can receive, from a vehicle over one or more networks 1290, image data representative of an image showing an unexpected or changing road condition, such as a road work that has recently begun. In at least one embodiment, one or more servers 1278 can transmit, to a vehicle over one or more networks 1290, an updated equalization neural network 1292 and / or map information 1294, including without limitation information about traffic and road conditions. In at least one embodiment, updates to map information 1294 can include, without limitation, updates to an HD map 1222, such as information about construction sites, potholes, detours, flooding, and / or other obstacles. In at least one embodiment, neural network 1292 and / or map information 1294 can be the result of new training and / or experience represented from data received from any number of vehicles in an environment, and / or based at least on training performed at a data center (e.g., using one or more servers 1278 and / or other servers).
[0290] In at least one embodiment, one or more servers 1278 can be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, training data can be generated by vehicles, and / or can be generated in simulations (e.g., using a game engine). In at least one embodiment, any amount of training data is labeled (e.g., in cases where an associated neural network benefits from supervised learning) and / or undergoes other pre-processing. In at least one embodiment, any amount of training data is not labeled and / or pre-processed (e.g., in cases where an associated neural network does not require supervised learning). In at least one embodiment, once a machine learning model is trained, a machine learning model can be used by vehicles (e.g., sent to vehicles over one or more networks 1290, and / or a machine learning model can be used by one or more servers 1278 to monitor vehicles remotely.
[0291] In at least one embodiment, one or more servers 1278 can receive data from vehicles and apply data to a latest real-time neural network for real-time intelligent inference. In at least one embodiment, one or more servers 1278 can include deep-learning supercomputers and / or specialized Al computers powered by one or more GPUs 1284, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, one or more servers 1278 can include deep learning infrastructure of a data center powered using CPUs.
[0292] In at least one embodiment, deep learning infrastructure of one or more servers 1278 can be capable of fast, real-time inference, and can use this capability to assess and validate health of processors, software, and / or associated hardware in vehicles 1200. For example, in at least one embodiment, deep learning infrastructure can receive periodic updates from vehicles 1200, such as sequences of images and / or objects located by vehicles 1200 in those images (e.g., via computer vision and / or other machine learning object classification techniques). In at least one embodiment, deep learning infrastructure can run its own neural network to identify objects and compare them to objects identified by vehicles 1200, and, if results do not match and deep learning infrastructure concludes that Al in vehicles 1200 is malfunctioning, one or more servers 1278 can send a signal to vehicles 1200 instructing a fail-safe computer of vehicles 1200 to take control, notify passengers, and complete a safe parking operation.
[0293] In at least one embodiment, one or more servers 1278 can include one or more GPUs 1284 and one or more programmable inference accelerators (such as NVIDIA’s TensorRT 3 devices). In at least one embodiment, a combination of GPU-powered servers and inference-accelerated servers can make real-time responses possible. In at least one embodiment, CPU-, FPGA-, and other processor-powered servers can be used for inference, such as in cases where performance is less critical. In at least one embodiment, one or more hardware structures 915 are used to perform one or more embodiments. Details are provided herein in connection with and / or Details are provided herein in connection with hardware structure 915.
[0294] Computer system
[0295] is a block diagram illustrating an exemplary computer system, which can be a system with interconnected devices and components, a system on a chip (SOC), or some combination thereof formed with a processor that can include execution units for executing instructions, according to at least one embodiment. In at least one embodiment, consistent with the present disclosure, such as in embodiments described herein, computer system 1300 can include, without limitation, components such as processor 1302 to employ execution units including logic to perform algorithms for process data. In at least one embodiment, computer system 1300 can include processors, such as Intel® Core® TM , Intel® XScale TM and / or Intel® StrongARM TM , Intel® Core® TM or Nervana® TM microprocessors, although other systems (including PCs, workstations, set-top boxes, etc. with other microprocessors) can also be used. In at least one embodiment, computer system 1300 can execute a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Wash., although other operating systems (UNIX and Linux, for example), embedded software, and / or graphical user interfaces can also be used.
[0296] Embodiments can be used in other devices such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications can include a microcontroller, a digital signal processor ("DSP"), a system on a chip, a network computer ("NetPC"), a set-top box, a network hub, a wide area
[0297] In at least one embodiment, computer system 1300 can include, but is not limited to, processor 1302, which can include, but is not limited to, one or more execution units 1308 for performing, e.g., instructions for performing machine learning model training and / or inference in accordance with techniques described herein. In at least one embodiment, computer system 1300 is a single processor desktop or server system, but in another embodiment, computer system 1300 can be a multiprocessor system. In at least one embodiment, processor 1302 can include, but is not limited to, a complex instruction set computer ("CISC") microprocessor, a reduced instruction set computing ("RISC") microprocessor, a very long instruction word ("VLIW") microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1302 can be coupled to a processor bus 1310 that can transmit data signals between processor 1302 and other components in computer system 1300.
[0298] In at least one embodiment, processor 1302 can include, but is not limited to, level 1 ("Ll") internal cache memory ("cache") 1304. In at least one embodiment, processor 1302 can have a single -level internal cache or multi-level internal cache. In at least one embodiment, cache memory can reside in the processor 1302's external. Other embodiments can also include a combination of internal and external caches depending on the specific implementation and requirements. In at least one embodiment, register file 1306 can store different types of data such as integer, floating point, status, and instruction pointer registers in various registers within processor 1302.
[0299] In at least one embodiment, execution unit 1308 also includes logic to handle a packed instruction set 1309. In at least one embodiment, by including the packed instruction set 1309 in an instruction set of a general -purpose processor, along with associated circuitry to execute the instructions, a processor can implement a packed data type that can be used for many multimedia applications in a more efficient manner than prior art processors. In at least one embodiment, operations used by many multimedia applications can be performed efficiently with packed data in processor 1302 by using the full width of a processor's data bus when performing an operation on packed data. This can eliminate the need to transfer smaller units of data across the processor's data bus to perform operations one data element at a time.
[0300] In at least one embodiment, execution unit 1308 can also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1300 can include, without limitation, a memory 1320. In at least one embodiment, memory 1320 can be a Dynamic Random Access Memory ("DRAM") device, a Static Random Access Memory ("SRAM") device, a flash memory device, or other memory device. In at least one embodiment, memory 1320 can store one or more instructions 1319 and / or data 1321 in the form of data signals representable by processor 1302.
[0301] In at least one embodiment, a system logic chip can be coupled to processor bus 1310 and memory 1320. In at least one embodiment, system logic chip can include, without limitation, a memory controller hub (“MCH”) 1316 and processor 1302 can communicate with MCH 1316 via processor bus 1310. In at least one embodiment, MCH 1316 can provide a high bandwidth memory path 1318 to memory 1320 for instruction and data storage and for storage of graphics commands, data, and textures. In at least one embodiment, MCH 1316 can direct data signals between processor 1302, memory 1320, and other components in computer system 1300 and can bridge data signals between processor bus 1310, memory 1320, and system I / O 1322. In at least one embodiment, system logic chip can provide a graphics port
[0302] In at least one embodiment, computer system 1300 can use system I / O interface 1322 as a proprietary hub interface bus to couple MCH 1316 to I / O controller hub (“ICH”) 1330. In at least one embodiment, ICH 1330 can provide a direct connection to some I / O devices and via a
[0303] In at least one embodiment, A system including interconnected hardware devices or “chips” is shown, while in other embodiments, FIG. 13 An exemplary SoC can be shown. In at least one embodiment, FIG. 13The devices illustrated in FIG. 13 can be interconnected utilizing proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1300 are interconnected using a compute express link (CXL) interconnect.
[0304] Logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. As described herein, inferencing and / or training operations FIG. 9A and / or FIG. 9B Details regarding logic 915 are provided. In at least one embodiment, logic 915 can be used in computer system 1300 to perform inferencing or prediction operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0305] In at least one embodiment, details regarding FIG. 12D-FIG. 13 At least one component shown or described is used to implement techniques and / or functionality described in connection with FIG. 1-FIG. 8B described. In at least one embodiment, processor 1302 performs one or more operations related to modifying a dense floating point neural model to a sparse quantized version of the neural model based, at least in part, on a weighted loss value used during quantization, as described elsewhere herein.
[0306] FIG. 14 is a block diagram illustrating an electronic device 1400 for utilizing processor 1410, in accordance with at least one embodiment. In at least one embodiment, electronic device 1400 can be, for example and without limitation, a laptop, a tower server, a rack server, a blade server, a laptop computer, a desktop computer, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0307] In at least one embodiment, electronic device 1400 can include, without limitation, processor 1410 communicatively coupled to any suitable number or kind of components, peripherals, modules, or devices. In at least one embodiment, processor 1410 is coupled using a bus or interface, such as an Industry Standard 2 C bus, a System Management Bus (“SMBus”), a Low Pin Count (LPC) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advanced Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, processor 1410 is coupled to one or more components of electronic device 1400 using a Compute Express Link (“CXL”) bus. FIG. 14 A system is shown that includes interconnected hardware devices or “chips,” while in other embodiments, FIG. 14 An exemplary SoC can be shown. In at least one embodiment, FIG. 14The devices shown in can be interconnected using a proprietary interconnect, a standardized interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, FIG. 14 One or more components of the system are interconnected using a Compute Express Link (CXL) interconnect.
[0308] In at least one embodiment, FIG. 14 It may include a display 1424, a touch screen 1425, a touchpad 1430, a near field communication unit (“NFC”) 1445, a sensor hub 1440, a thermal sensor 1446, an express chipset (“EC”) 1435, a trusted platform module (“TPM”) 1438, BIOS / firmware / flash memory (“BIOS, FW Flash”) 1422, a DSP 1460, a drive 1420 (such as a solid state disk (“SSD”) or a hard disk drive (“HDD”)), a wireless local area network unit (“WLAN”) 1450, a Bluetooth unit 1452, a wireless wide area network unit (“WWAN”) 1456, a global positioning system (GPS) unit 1455, a camera (“USB 3.0 camera”) 1454 (such as a USB 3.0 camera), and / or a low power double data rate (“LPDDR”) memory unit (“LPDDR3”) 1415 implemented with, for example, the LPDDR3 standard. These components may each be implemented in any suitable manner.
[0309] In at least one embodiment, other components may be communicatively coupled to processor 1410 via the components described herein. In at least one embodiment, accelerometer 1441, ambient light sensor (“ALS”) 1442, compass 1443, and gyroscope 1444 may be communicatively coupled to sensor hub 1440. In at least one embodiment, thermal sensor 1439, fan 1437, keyboard 1436, and touchpad 1430 may be communicatively coupled to EC 1435. In at least one embodiment, speaker 1463, earphone 1464, and microphone (“mic”) 1465 may be communicatively coupled to audio unit (“audio codec and class-D amplifier”) 1462, which in turn may be communicatively coupled to DSP 1460. In at least one embodiment, audio unit 1462 may include, for example, but not limited to, an audio codec / decoder (“codec”) and a class-D amplifier. In at least one embodiment, SIM card (“SIM”) 1457 may be communicatively coupled to WWAN unit 1456. In at least one embodiment, components such as the WLAN unit 1450 and the Bluetooth unit 1452 and the WWAN unit 1456 may be implemented as a next generation form factor ("NGFF").
[0310] Logic 915 is used to perform reasoning and / or training operations associated with one or more embodiments.FIG. 9A and / or FIG. 9B Details are provided regarding logic 915. In at least one embodiment, logic 915 can be used in electronic device 1400 for performing inferencing or predicting operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0311] In at least one embodiment, regarding FIG. 14 At least one component shown or described as being implemented within FIG. 1-FIG. 8B the techniques and / or functionality described. In at least one embodiment, processor 1415 performs one or more operations regarding modifying a dense floating point neural model to a sparse quantized version of the neural model based, at least in part, on a weighted loss value used during quantization, as described elsewhere herein.
[0312] FIG. 15 A computer system 1500 according to at least one embodiment is shown. In at least one embodiment, computer system 1500 is configured to implement various processes and methods described throughout this disclosure.
[0313] In at least one embodiment, computer system 1500 includes, without limitation, at least one central processing unit (“CPU”) 1502 that is connected to a communication bus 1510 implemented using any suitable protocol, such as PCI (“Peripheral Component Interconnect”), peripheral component interconnect express (“PCI-Express”), AGP (“Accelerated Graphics Port”), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1500 includes, without limitation, a main memory 1504 and control logic (e.g., implemented as hardware, software, or a combination thereof) and data are stored in the main memory 1504, which can take form of random access memory (“RAM”). In at least one embodiment, a network interface subsystem (“network interface”) 1522 provides an interface to other computing devices and networks for receiving data from other systems and for sending data to other systems using computer system 1500.
[0314] In at least one embodiment, computer system 1500 includes, without limitation, an input device 1508, parallel processing system 1512, and display device 1506, which can be implemented using a conventional cathode ray tube (“CRT”), liquid crystal display (“LCD”), light emitting diode (“LED”), plasma display, or other suitable display technologies. In at least one embodiment, user input is received from input device 1508, such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each module described herein can be located on a single semiconductor platform.
[0315] Logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. As described herein, inferencing and / or training operations FIG. 9A and / or FIG. 9B Details regarding logic 915 are provided. In at least one embodiment, logic 915 can be used in computer system 1500 to perform inferencing or prediction operations based, at least in part, on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0316] In at least one embodiment, details regarding FIG. 15 At least one component shown or described as being implemented within logic 915 is used to implement techniques and / or functionality described in connection with other FIG. 1-FIG. 8B In at least one embodiment, computer system 1500 performs one or more operations related to modifying a dense floating point neural model to a sparse quantized version of the neural model based, at least in part, on weighted loss values used during quantization, as described elsewhere herein.
[0317] FIG. 16 Computer system 1600 is shown, according to at least one embodiment. In at least one embodiment, computer system 1600 includes, without limitation, computer 1610 and USB stick 1620. In at least one embodiment, computer 1610 can include, without limitation, any number and type of processor (not shown) and memory (not shown). In at least one embodiment, computer 1610 includes, without limitation, a server, a cloud instance, a laptop computer, and a desktop computer.
[0318] In at least one embodiment, USB stick 1620 includes, without limitation, processing unit 1630, USB interface 1640, and USB interface logic 1650. In at least one embodiment, processing unit 1630 can be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, processing unit 1630 can include, without limitation, any number and type of processing core (not shown). In at least one embodiment, processing unit 1630 comprises an application-specific integrated circuit (“ASIC”) optimized to perform any number and type of operations associated with machine learning. For example, in at least one embodiment, processing unit 1630 is a tensor processing unit (“TPC”) optimized to perform machine learning inferencing operations. In at least one embodiment, processing unit 1630 is a visual processing unit (“VPU”) optimized to perform machine vision and machine learning inferencing operations.
[0319] In at least one embodiment, USB interface 1640 can be any type of USB connector or USB receptacle. For example, in at least one embodiment, USB interface 1640 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 1640 is a USB 3.0 Type-A connector. In at least one embodiment, USB interface logic 1650 can include any amount and type of logic that enables processing unit 1630 to interface with a device (e.g., computer 1610) via USB connector 1640.
[0320] Logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. As described herein, inferencing and / or training operations FIG. 9A and / or FIG. 9B Details regarding logic 915 are provided. In at least one embodiment, logic 915 can be used in computer system 1500 to perform inferencing or prediction operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0321] In at least one embodiment, details regarding FIG. 16 At least one component shown or described as integrated into computer system 1600 is used to implement techniques and / or functionality described in connection with FIG. 1-FIG. 8B In at least one embodiment, computer system 1600 performs one or more operations related to modifying a dense floating point neural model to a sparse quantized version of the neural model based, at least in part, on weighted loss values used during quantization, as described elsewhere herein.
[0322] FIG. 17A An exemplary architecture is shown in which a plurality of GPUs 1710(1)-1710(N) are communicatively coupled to a plurality of multi-core processors 1705(1)-1705(M) over high-speed links 1740(1)-1740(N) (e.g., buses, point-to-point interconnects, etc.). In at least one embodiment, high-speed links 1740(1)-1740(N) support a communication throughput of 4GB / s, 30GB / s, 80GB / s or higher. In at least one embodiment, various interconnect protocols can be used including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, “N” and “M” represent positive integers, and values of N and M can vary from figure to figure. In at least one embodiment, one or more of the plurality of GPUs 1710(1)-1710(N) include graphics processing units as described elsewhere herein. FIG. 20A and FIG. 20BOne or more graphics cores (also referred to simply as“cores”) 2000 are disclosed. In at least one embodiment, one or more graphics cores 2000 can be referred to as stream multi- processors (“SMs”), stream processors (“SPs”), stream processing units (“SPUs”), compute units (“CUs”), execution units (“EUs”), and / or slices, where in present context a slice can refer to a portion of processing resources in a processing unit (e.g., 16 cores, a ray-tracing unit, a thread director or scheduler).
[0323] Further, in at least one embodiment, two or more GPUs 1710 are interconnected over high-speed links 1729(1)-1729(2), which can be implemented using similar or different protocols / links than those used for high-speed links 1740(1)-1740(N). Similarly, two or more multi-core processors 1705 can be connected over a high-speed link 1728, which can be a symmetric multi-processor (SMP) bus that runs at 20 GB / s, 30 GB / s, 120 GB / s or higher. Alternatively, similar protocols / links (e.g., over common interconnect fabric) can be used to accomplish FIG. 17A All communication between various system components shown in FIG. 17.
[0324] In at least one embodiment, each multi-core processor 1705 is communicatively coupled to processor memories 1701(1)-1701(M) via memory interconnects 1726(1)-1726(M), respectively, and each GPU 1710(1)-1710(N) is communicatively coupled to GPU memories 1720(1)-1720(N) over GPU memory interconnects 1750(1)-1750(N), respectively. In at least one embodiment, memory interconnects 1726 and 1750 can utilize similar or different memory access technologies. By way of non-limiting example, processor memories 1701(1)-1701(M) and GPU memories 1720 can be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high-bandwidth memory (HBM), and / or can be non-volatile memories such as 3D XPoint or Nano-Ram. In at least one embodiment, certain portions of processor memories 1701 can be volatile memory while another portion can be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0325] As described herein, although individual multi-core processors 1705 and GPUs 1710 can be physically coupled to particular memories 1701, 1720, respectively, and / or a unified memory architecture can be implemented in which a virtual system address space (also referred to as an “effective address” space) is distributed among the individual physical memories. For example, processor memories 1701(1)-1701(M) can each include 64 GB of system memory address space, and GPU memories 1720(1)-1720(N) can each include 32 GB of system memory address space, resulting in a total of 256 GB of addressable memory when M = 2 and N = 4. Other values of N and M are possible.
[0326] FIG. 17B Additional details for interconnect between multi-core processor 1707 and graphics acceleration module 1746 are shown according to one exemplary embodiment. In at least one embodiment, graphics acceleration module 1746 can include one or more GPU chips integrated on a line card that is coupled via high-speed link 1740 (e.g., a PCIe bus, NVLink, etc.) to processor 1707. In at least one embodiment, graphics acceleration module 1746 can alternatively be integrated on a package with processor 1707 or on a chip with processor 1707.
[0327] In at least one embodiment, processor 1707 includes a number of cores 1760A-1760D (which can be referred to as “execution units”) each having a translation lookaside buffer (“TLB”) 1761A-1761D and one or more caches 1762A-1762D. In at least one embodiment, cores 1760A-1760D can include various other components not shown for purposes of illustration and explanation. In at least one embodiment, caches 1762A-1762D can include level one (LI) and level two (L2) caches. In addition, one or more shared caches 1756 can be included in caches 1762A-1762D and shared by groups of cores 1760A-1760D. For example, one embodiment of processor 1707 includes 24 cores each having its own LI cache, 12 shared L2 caches, and 12 shared L3 caches. In that embodiment, two adjacent cores share one or more L2 and L3 caches. In at least one embodiment, processor 1707 and graphics acceleration module 1746 are connected with system memory 1714, which can include processor memories 1701(1)-1701(M) in FIG. 17A
[0328] In at least one embodiment, coherence for data and instructions stored in respective caches 1762A-1762D, 1756, and system memory 1714 is maintained by inter-core communications over coherence bus 1764. In at least one embodiment, each cache can have cache coherence logic / circuitry associated with it to communicate over coherence bus 1764 in response to detecting a read or write to a particular cache line. In at least one embodiment, a cache snoop protocol is implemented over coherence bus 1764 to snoop cache accesses.
[0329] In at least one embodiment, agent circuit 1725 communicatively couples graphics acceleration module 1746 to coherence bus 1764, allowing graphics acceleration module 1746 to participate in cache coherence protocol as a peer to cores 1760A-1760D. In particular, in at least one embodiment, interface 1735 provides connectivity from graphics acceleration module 1746 to agent circuit 1725 over high-speed link 1740, and interface 1737 connects graphics acceleration module 1746 to high-speed link 1740.
[0330] In at least one embodiment, accelerator integration circuit 1736 provides the multiple graphics processing engines 1731(1)-1731(N) of graphics acceleration module 1746 with cache management, memory access, context management, and interrupt management services. This allows graphics processing engines 1731(1)-1731(N) to be more tightly coupled to accelerator integration circuit 1736 than to larger system agents. FIG. 20A and FIG. 20B In at least one embodiment, graphics processing engines 1731(1)-1731(N) can each comprise a separate graphics processing unit (GPU). In at least one embodiment, the multiple graphics processing engines 1731(1)-1731(N) of graphics acceleration module 1746 comprise different types of graphics processing engines, such as graphics execution units, media processing engines, samplers, and blit (block transfer) engines. In at least one embodiment, graphics acceleration module 1746 can be a GPU with the multiple graphics processing engines 1731(1)-1731(N), or the graphics processing engines 1731(1)-1731(N) can be individual GPUs integrated on a common package, line card, or chip.
[0331] In at least one embodiment, the accelerator integrated circuit 1736 includes a memory management unit (MMU) 1739 for performing various memory management functions, such as virtual to physical memory translation (also known as effective to real memory translation), and a memory access protocol for accessing system memory 1714. In at least one embodiment, the MMU 1739 may also include a translation lookaside buffer ("TLB") (not shown) for caching virtual / effective to physical / real address translations. In at least one embodiment, a cache 1738 may store commands and data for efficient access by the graphics processing engines 1731(1)-1731(N). In at least one embodiment, a fetch unit 1744 may be used to keep data stored in the cache 1738 and graphics memory 1733(1)-1733(M) consistent with the core caches 1762A-1762D, 1756, and system memory 1714. As previously described, this can be implemented on behalf of cache 1738 and memory 1733(1)-1733(M) via proxy circuit 1725 (e.g., sending updates related to modifications / accesses of cache lines on processor caches 1762A-1762D, 1756 to cache 1738 and receiving updates from cache 1738).
[0332] In at least one embodiment, a set of registers 1745 stores context data for threads executed by graphics processing engines 1731(1)-1731(N), and context management circuitry 1748 manages thread contexts. For example, context management circuitry 1748 can perform save and restore operations to save and restore the contexts of various threads during context switches (e.g., where a first thread is saved and a second thread is stored so that the second thread can be executed by the graphics processing engine). For example, context management circuitry 1748 can store current register values to a designated area in memory (e.g., identified by a context pointer) upon context switching. The register values can then be restored upon returning to the context. In at least one embodiment, interrupt management circuitry 1747 receives and processes interrupts received from system devices.
[0333] In at least one embodiment, MMU 1739 translates virtual / effective addresses from graphics processing engines 1731 into real / physical addresses in system memory 1714. In at least one embodiment, accelerator integration circuit 1736 supports a number of graphics accelerator modules 1746 and / or other accelerator devices (e.g., 4, 8, 16). In at least one embodiment, graphics accelerator modules 1746 can be dedicated to a single application executing on processor 1707 or can be shared between multiple applications. In at least one embodiment, a virtualized graphics execution environment is presented in which resources of graphics processing engines 1731(1)-1731(N) are shared between multiple applications or virtual machines (VMs). In at least one embodiment, resources can be subdivided into “slices” that are allocated to different VMs and / or applications based on processing requirements and priorities associated with the VMs and / or applications.
[0334] In at least one embodiment, accelerator integration circuit 1736 performs as a bridge to system for a system of graphics acceleration modules 1746 and provides address translation and system memory cache services. Further, in at least one embodiment, accelerator integration circuit 1736 can provide virtualization facilities for a host processor to manage virtualization of graphics processing engines 1731(1)-1731(N), interrupts, and memory management.
[0335] In at least one embodiment, because hardware resources of graphics processing engines 1731(1)-1731(N) are explicitly mapped to real address space seen by host processor 1707, any host processor can directly address these resources using effective address values. In at least one embodiment, one function of accelerator integration circuit 1736 is physical segregation of graphics processing engines 1731(1)-1731(N) so that they appear as independent units to a system.
[0336] In at least one embodiment, one or more graphics memories 1733(1)-1733(M) are respectively coupled to each graphics processing engine 1731(1)-1731(N), and N=M. In at least one embodiment, graphics memories 1733(1)-1733(M) store instructions and data being processed by each graphics processing engine 1731(1)-1731(N). In at least one embodiment, graphics memories 1733(1)-1733(M) can be volatile memory such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or can be non-volatile memory such as 3D XPoint or Nano-Ram.
[0337] In at least one embodiment, to reduce data traffic on high-speed link 1740, biasing techniques may be used to ensure that the data stored in graphics memory 1733(1)-1733(M) is the data most frequently used by graphics processing engines 1731(1)-1731(N), and preferably is not used (at least not frequently) by cores 1760A-1760D. Similarly, in at least one embodiment, the biasing mechanism attempts to keep data needed by the cores (and preferably not needed by graphics processing engines 1731(1)-1731(N)) in caches 1762A-1762D, 1756, and system memory 1714.
[0338] FIG. 17C Another exemplary embodiment is shown in which an accelerator integrated circuit 1736 is integrated within the processor 1707. In this embodiment, the graphics processing engines 1731(1)-1731(N) communicate directly with the accelerator integrated circuit 1736 via interface 1737 and interface 1735 (again, which can be any form of bus or interface protocol) over a high-speed link 1740. In at least one embodiment, the accelerator integrated circuit 1736 can perform operations related to FIG. 17B The operations described above are similar to those described above, but may have higher throughput due to its close proximity to the coherence bus 1764 and caches 1762A-1762D, 1756. In at least one embodiment, the accelerator integrated circuit supports different programming models, including a process-specific programming model (without graphics acceleration module virtualization) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integrated circuit 1736 and a programming model controlled by the graphics acceleration module 1746.
[0339] In at least one embodiment, graphics processing engines 1731(1)-1731(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can funnel other application requests to graphics processing engines 1731(1)-1731(N), thereby providing virtualization within a VM / partition.
[0340] In at least one embodiment, graphics processing engines 1731(1)-1731(N) can be shared by multiple VM / application partitions. In at least one embodiment, a shared model can use a hypervisor to virtualize graphics processing engines 1731(1)-1731(N) to allow access by each operating system. In at least one embodiment, for a single-partition system without a hypervisor, an operating system owns graphics processing engines 1731(1)-1731(N). In at least one embodiment, an operating system can virtualize graphics processing engines 1731(1)-1731(N) to provide access to each process or application.
[0341] In at least one embodiment, graphics acceleration module 1746 or individual graphics processing engines 1731(1)-1731(N) use a process handle to select a process element. In at least one embodiment, process elements are stored in system memory 1714 and can be addressed using effective to real address translation techniques described herein. In at least one embodiment, a process handle can be an implementation-specific value provided to a host process when registering its context with a graphics processing engine 1731(1)-1731(N) (i.e., calling system software to add a process element to a process element linked list). In at least one embodiment, a lower 16 bits of a process handle can be an offset into a process element linked list for a process element.
[0342] FIG. 17D An exemplary accelerator integration slice 1790 is shown. In at least one embodiment, a “slice” comprises a specified portion of processing resources of accelerator integration circuit 1736. In at least one embodiment, an application is an effective address space 1782 in system memory 1714 that stores a process element 1783. In at least one embodiment, process element 1783 is stored in response to a GPU invocation 1781 from an application 1780 executing on processor 1707. In at least one embodiment, process element 1783 contains process state for corresponding application 1780. In at least one embodiment, a work descriptor (WD) 1784 contained in process element 1783 can be a single job requested by an application or can contain a pointer to a queue of jobs. In at least one embodiment, WD 1784 is a pointer to a job request queue in an application’s effective address space 1782.
[0343] In at least one embodiment, graphics acceleration module 1746 and / or individual graphics processing engines 1731(1)-1731(N) can be shared by all or a subset of processes in a system. In at least one embodiment, a process- specific programming model can be implementation specific. In at least one embodiment, in this model a single process owns graphics acceleration module 1746 or an individual graphics processing engine 1731. In at least one embodiment, when graphics acceleration module 1746 is owned by a single process, a hypervisor initializes an accelerator integration circuit for the owned partition, and an operating system initializes an accelerator integration circuit 1736 for the owned process when graphics acceleration module 1746 is assigned.
[0344] In at least one embodiment, a process-specific programming model is implementation specific. In at least one embodiment, in this model a single process owns graphics acceleration module 1746 or an individual graphics processing engine 1731. In at least one embodiment, when graphics acceleration module 1746 is owned by a single process, a hypervisor initializes an accelerator integration circuit for the owned partition, and an operating system initializes an accelerator integration circuit 1736 for the owned process when graphics acceleration module 1746 is assigned.
[0345] In at least one embodiment, in operation, a WD fetch unit 1791 in an accelerator integration slice 1790 fetches a next WD 1784, which includes an indication of work to be completed by one or more graphics processing engines of graphics acceleration module 1746. In at least one embodiment, data from WD 1784 can be stored in registers 1745 and used by MMU 1739, interrupt management circuit 1747, and / or context management circuit 1748, as illustrated. For example, one embodiment of MMU 1739 includes segment / page walk circuitry for accessing segment / page tables 1786 within an OS virtual address space 1785. In at least one embodiment, interrupt management circuit 1747 can handle interrupt events 1792 received from graphics acceleration module 1746. In at least one embodiment, when performing graphics operations, effective addresses 1793 generated by graphics processing engines 1731(1)-1731(N) are translated to real addresses by MMU 1739.
[0346] In at least one embodiment, registers 1745 are replicated for each graphics processing engine 1731(1)-1731(N) and / or graphics acceleration module 1746, and can be initialized by a hypervisor or operating system. In at least one embodiment, each of these replicated registers can be included in an accelerator integration slice 1790. Exemplary registers that can be initialized by a hypervisor are shown in Table 1.
[0347] Table 1 - Hypervisor-Initialized Registers
[0348]
[0349]
[0350] Example registers that may be initialized by the operating system are shown in Table 2.
[0351] Table 2 - Registers initialized by the operating system
[0352]
[0353] In at least one embodiment, each WD 1784 is specific to a particular graphics acceleration module 1746 and / or graphics processing engine 1731(1)-1731(N). In at least one embodiment, it contains all the information needed by the graphics processing engine 1731(1)-1731(N) to complete its work, or it may be a pointer to a memory location where an application has set up a command queue for work to be done.
[0354] FIG. 17E 1796 virtualizes the graphics acceleration module engine for the operating system 1795.
[0355] In at least one embodiment, the shared programming model allows all processes or subsets of processes from all partitions or subsets of partitions in the system to use the graphics acceleration module 1746. In at least one embodiment, there are two programming models where the graphics acceleration module 1746 is shared by multiple processes and partitions, namely, time-sliced sharing and graphics-directed sharing.
[0356] In at least one embodiment, in this model, the hypervisor 1796 owns the graphics acceleration module 1746 and makes its functionality available to all operating systems 1795. In at least one embodiment, for the graphics acceleration module 1746 to support virtualization through the hypervisor 1796, the graphics acceleration module 1746 may adhere to certain requirements, such as (1) the application's job requests must be autonomous (i.e., no state needs to be maintained between jobs), or the graphics acceleration module 1746 must provide a context save and restore mechanism, (2) the graphics acceleration module 1746 guarantees that the application's job requests are completed within a specified amount of time, including any transition errors, or the graphics acceleration module 1746 provides the ability to preempt job processing, and (3) the graphics acceleration module 1746 must ensure fairness between processes when operating in a directed shared programming model.
[0357] In at least one embodiment, application 1780 is required to use a graphics acceleration module type, a work descriptor (WD), an authority mask register (AMR) value, and a context save / restore area pointer (CSRP) for an operating system 1795 system call. In at least one embodiment, the graphics acceleration module type describes a target acceleration function for the system call. In at least one embodiment, the graphics acceleration module type can be a system specific value. In at least one embodiment, the WD is formatted specifically for a graphics acceleration module 1746 and can take the form of a graphics acceleration module 1746 command, a valid address pointer to a user defined structure, a valid address pointer to a command queue, or the form of any other data structure describing work to be done by a graphics acceleration module 1746.
[0358] In at least one embodiment, the AMR value is the AMR state for the current process. In at least one embodiment, the value passed to the operating system is similar to how an application program sets the AMR. In at least one embodiment, if an accelerator integration circuit 1736 (not shown) and graphics acceleration module 1746 implementation does not support a user authority mask override register (UAMOR), then the operating system can apply the current UAMOR value to the AMR value before passing the AMR in a hypervisor call. In at least one embodiment, the hypervisor 1796 can selectively apply the current authority mask override register (AMOR) value before placing the AMR in the process element 1783. In at least one embodiment, the CSRP is one of registers 1745 that contains a valid address of an area in the application’s effective address space 1782 for a graphics acceleration module 1746 to save and restore context state. In at least one embodiment, this pointer is optional if there is no need to save state between jobs or when a job is preempted. In at least one embodiment, the context save / restore area can be a fixed system memory.
[0359] Upon receiving the system call, operating system 1795 can verify that application 1780 has registered and been granted authority to use graphics acceleration module 1746. Operating system 1795 then, in at least one embodiment, uses the information shown in Table 3 to call hypervisor 1796.
[0360] Table 3 - Operating System to Hypervisor Call Parameters
[0361]
[0362] In at least one embodiment, upon receiving a hypervisor call, hypervisor 1796 verifies that operating system 1795 is registered and granted permission to use graphics acceleration module 1746. Hypervisor 1796 then places process element 1783 into a process element linked list for a corresponding graphics acceleration module 1746 type, in at least one embodiment. In at least one embodiment, a process element can include information as illustrated in Table 4.
[0363] Table 4 - Process Element Information
[0364]
[0365] In at least one embodiment, hypervisor initializes a plurality of accelerator integration slice 1790 registers 1745.
[0366] As FIG. 17F illustrated, in at least one embodiment, a unified memory is used that is addressable via a common virtual memory address space for accessing physical processor memory 1701(1)-1701(N) and GPU memory 1720(1)-1720(N). In this implementation, an operation executing on a GPU 1710(1)-1710(N) accesses processor memory 1701(1)-1701(M) with the same virtual / effective memory address space and vice versa, simplifying programmability. In at least one embodiment, a first portion of a virtual / effective address space is allocated to processor memory 1701(1), a second portion is allocated to a second processor memory 1701(N), a third portion is allocated to GPU memory 1720(1), and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as an effective address space) is thereby distributed across each of processor memory 1701 and GPU memory 1720, allowing any processor or GPU to access a memory with a virtual address that maps to that memory.
[0367] In at least one embodiment, bias / coherence management circuitry 1794A-1794E within one or more MMU 1739A-1739E ensures cache coherency between one or more host processors (e.g., 1705) and caches of GPU 1710, and implements bias techniques that dictate a bias of physical memory where certain types of data should be stored. In at least one embodiment, while multiple instances of bias / coherence management circuitry 1794A-1794E are shown in FIG. 17F FIG. 16B, bias / coherence management circuitry can be implemented within an MMU of one or more host processors 1705 and / or within accelerator integration circuit 1736.
[0368] One embodiment allows GPU memory 1720 to be mapped as part of system memory and accessed using shared virtual memory (SVM) techniques, but without suffering the performance penalties associated with full system cache coherency. In at least one embodiment, the ability for GPU memory 1720 to be accessed as system memory without the heavy cache coherency overhead provides a favorable operating environment for GPU offload. In at least one embodiment, this arrangement allows software of host processor 1705 to set operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies include driver calls, interrupts, and memory mapped I / O (MMIO) accesses, which are less efficient than simple memory accesses. In at least one embodiment, the ability to access GPU memory 1720 without cache coherency overhead can be critical to the execution time of offloaded computations. In at least one embodiment, for example, with heavy streaming write memory traffic, cache coherency overhead can significantly reduce the effective write bandwidth seen by GPU 1710. In at least one embodiment, the efficiency of operand setup, the efficiency of result access, and the efficiency of GPU computation can play a role in determining the effectiveness of GPU offload.
[0369] In at least one embodiment, the selection of GPU bias and host processor bias is driven by a bias tracker data structure. In at least one embodiment, for example, a bias table can be used, which can be a page-granularity structure (e.g., controlling at the granularity of a memory page), including 1 or 2 bits per GPU-attached memory page. In at least one embodiment, with or without a bias cache in GPU 1710 (e.g., to cache frequently / recently used entries of the bias table), the bias table can be implemented in a stolen memory range of one or more GPU memories 1720. Alternatively, in at least one embodiment, the entire bias table can be maintained within the GPU.
[0370] In at least one embodiment, prior to actually accessing GPU memory, the bias table entry associated with each access to GPU-attached memory 1720 is accessed, causing the following operations. In at least one embodiment, local requests from GPU 1710 that find their pages in GPU bias are forwarded directly to corresponding GPU memory 1720. In at least one embodiment, local requests from GPU that find their pages in host bias are forwarded to processor 1705 (e.g., over a high-speed link as described herein). In at least one embodiment, requests from processor 1705 that find requested pages in host processor bias complete requests similar to normal memory reads. Alternatively, requests that point to GPU-biased pages can be forwarded to GPU 1710. In at least one embodiment, if GPU is not currently using a page, GPU can migrate the page to host processor bias. In at least one embodiment, bias state of a page can be changed by software-based mechanisms, hardware-assisted software-based mechanisms, or in the case of a limited set, purely hardware-based mechanisms.
[0371] In at least one embodiment, a mechanism for changing bias state employs an API call (e.g., OpenCL), which in turn invokes a device driver of a GPU, which in turn sends a message (or causes a command descriptor to be enqueued) to a GPU, directing the GPU to change bias state, and in certain migrations, to perform a cache flush operation in a host. In at least one embodiment, a cache flush operation is used to migrate from host processor 1705 bias to GPU bias, but not for the reverse migration.
[0372] In at least one embodiment, cache coherency is maintained by temporarily rendering GPU-biased pages that host processor 1705 cannot cache. In at least one embodiment, to access these pages, processor 1705 can request access from GPU 1710, which can or can not grant access immediately. Thus, in at least one embodiment, to reduce communication between processor 1705 and GPU 1710, it is beneficial to ensure that GPU-biased pages are pages that are needed by GPU and not by host processor 1705, and vice versa.
[0373] One or more hardware structures 915 are used to perform one or more embodiments. Details regarding one or more hardware structures 915 can be found in this document in connection with FIG. 9A and / or FIG. 9B Details regarding one or more hardware structures 915 are provided.
[0374] FIG. 18Exemplary integrated circuits and associated graphics processors in accordance with various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to the illustrated, other logic and circuitry can also be included, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.
[0375] FIG. 18 is a block diagram illustrating an exemplary system on a chip integrated circuit 1800 that can be fabricated using one or more IP cores, in accordance with at least one embodiment. In at least one embodiment, integrated circuit 1800 includes one or more application processor(s) 1805 (e.g., CPUs), at least one graphics processor 1810, and can additionally include an image processor 1815 and / or a video processor 1820, any of which can be a modular IP core. In at least one embodiment, integrated circuit 1800 includes peripheral or bus logic including a USB controller 1825, a UART controller 1830, an SPI / SDIO controller 1835, and an I2S / I2C controller 1840. In at least one embodiment, integrated circuit 1800 can include a display device 1845 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1850 and a mobile industry processor interface (MIPI) display controller 1855. In at least one embodiment, storage can be provided by a flash memory subsystem 1860 including flash memory and a flash memory controller. In at least one embodiment, memory interface can be provided via a memory controller 1865 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1870. 2 S / I 2 Ccontroller 1840. In at least one embodiment, integrated circuit 1800 can include a display device 1845 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1850 and a mobile industry processor interface (MIPI) display controller 1855. In at least one embodiment, storage can be provided by a flash memory subsystem 1860 including flash memory and a flash memory controller. In at least one embodiment, memory interface can be provided via a memory controller 1865 for access to SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits also include an embedded security engine 1870.
[0376] Logic 915 is used to perform inferencing and / or training operations associated with one or more embodiments. As described herein, inferencing can include performing one or more operations depending on FIG. 9A and / or FIG. 9B Details regarding logic 915 are provided in conjunction with the information processing system 1900, below. In at least one embodiment, logic 915 can be used in integrated circuit 1800 to infer or predict operations based, at least in part, on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0377] In at least one embodiment, details regarding at least one component shown or described are used to implement a component dependent on FIG. 17A-FIG. 18 at least one component shown or described are used to implement a component dependent on FIG. 1-FIG. 8Btechnology and / or functionality. In at least one embodiment, SOC 1800 performs one or more operations related to modifying a dense floating point neural model to become a sparse quantized version of that neural model based at least in part on weighted loss values used during quantization, as described elsewhere herein.
[0378] FIG. 19A-19B Exemplary integrated circuits and associated graphics processors according to various embodiments described herein are shown, which can be fabricated using one or more IP cores. In addition to the illustrated, other logic and circuitry can be included, including additional graphics processors / cores, peripheral interface controllers or general purpose processor cores.
[0379] FIG. 19A-19B is a block diagram illustrating an exemplary graphics processor utilized within a SoC in accordance with the embodiments described herein. FIG. 19A Exemplary graphics processor 1910 of a system on a chip integrated circuit, in accordance with at least one embodiment, is shown, which can be fabricated using one or more IP cores. FIG. 19B Additional exemplary graphics processor 1940 of a system on a chip integrated circuit, in accordance with at least one embodiment, is shown, which can be fabricated using one or more IP cores. In at least one embodiment, FIG. 19A Graphics processor 1910 is a low power graphics processor core. In at least one embodiment, FIG. 19B Graphics processor 1940 is a higher performance graphics processor core. In at least one embodiment, each graphics processor 1910, 1940 can be a FIG. 18 Variants of graphics processor 1810.
[0380] In at least one embodiment, graphics processor 1910 includes a vertex processor 1905 and one or more fragment processor(s) 1915A-1915N (e.g., 1915A, 1915B, 1915C, 1915D, through 1915N-1, and 1915N). In at least one embodiment, graphics processor 1910 can execute different shader programs via separate logic for vertex processing and / or for fragment / pixel processing. In at least one embodiment, vertex processor 1905 is optimized to execute operations for vertex shader programs, while one or more fragment processor(s) 1915A-1915N may be optimized to perform fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, vertex processor 1905 performs a vertex processing stage of a 3D graphics pipeline and generates a vertex output database of graphics primitives that are used to form graphics objects to be rendered. In at least one embodiment, one or more fragment processor(s) 1915A-1915N use graphics primitives from vertex processor 1905 and graphics data stored in graphics memory 1925A-1925B to produce a frame buffer output that is displayed on a display device. In at least one embodiment, one or more fragment processor(s) 1915A-1915N are optimized to execute a fragment shader program as provided in an OpenGL API, which can be used to perform similar operations for pixel shader programs provided in a Direct 3D API.
[0381] In at least one embodiment, graphics processor 1910 additionally includes one or more memory management units (MMUs) 1920A-1920B, one or more caches 1925A-1925B, and one or more circuit interconnects 1930A-1930B. In at least one embodiment, one or more MMUs 1920A-1920B provide translation of virtual addresses into physical addresses, as is known in the art, and can also provide other functions, such as memory protection. In at least one embodiment, one or more MMUs 1920A-1920B couple graphics processor 1910 to one or more internal memory structures that provide storage for vertex or image / texture data and program instructions. In at least one embodiment, one or more caches 1925A-1925B can accelerate data delivery from one or more memory structures to graphics processor 1910. In at least one embodiment, cache(s) 1925A-1925B may FIG. 18 In at least one embodiment, one or more circuit interconnects 1930A-1930B enable graphics processor 1910 to interface with other IP cores within a SoC, either via an internal bus, as shown, or via a direct connection.
[0382] In at least one embodiment, graphics processor 1940 includes a ring interconnect 1945, one or more GPU core complexes 1950A-1950N (e.g., 1950A, 1950B, 1950C, 1950D, through 1950N-1, and 1950N), and display engine(s) 1960A-1960B.FIG. 19B One or more shader cores 1955A-1955N (e.g., 1955A, 1955B, 1955C, 1955D, 1955E, 1955F through 1955N-1 and 1955N) are shown, which provide a unified shader core architecture in which a single core or type or core can execute all types of programmable shader code, including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can vary. In at least one embodiment, the graphics processor 1940 includes an inter-core task manager 1945 that acts as a thread dispatcher to dispatch execution threads to one or more shader cores 1955A-1955N and a tiling unit 1958 to accelerate tiling operations for tile-based rendering, in which rendering operations of a scene are subdivided in image space, for example, to exploit local spatial coherence within the scene or to optimize the use of internal caches.
[0383] Logic 915 is used to perform reasoning and / or training operations associated with one or more embodiments. FIG. 9A and / or FIG. 9B Details are provided regarding logic 915. In at least one embodiment, logic 915 may be used in graphics processors 1910 and / or 1940 to perform inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0384] In at least one embodiment, FIG. 19A-FIG. 19B At least one component shown or described is used to implement the combination FIG. 1-FIG. 8B In at least one embodiment, graphics processor 1910 performs one or more operations related to modifying a dense floating-point neural model to a sparse quantized version of the neural model based at least in part on a weighted loss value used during quantization, as described elsewhere herein.
[0385] FIG. 20A-20B Additional exemplary graphics processor logic according to embodiments described herein is shown. In at least one embodiment, FIG. 20A-20B Shown and combined FIG. 20A-20B The components described are integrated into a single system, such as a graphics processing unit (GPU), a SoC, or another type of processor. In at least one embodiment, FIG. 20A Shows that can be included in FIG. 18 Graphics core 2000 within graphics processor 1810, and in at least one embodiment, may be such as FIG. 19B Unified shader cores 1955A-1955N are shown.FIG. 20B A highly parallel general-purpose graphics processing unit ("GPGPU", which can also be referred to as a "graphics processing unit") 2030 suitable for deployment on a multi-chip module is shown in at least one embodiment. In at least one embodiment, graphics processing unit 2030 is a GPGPU that includes graphics processors. In at least one embodiment, integrated circuit 1800 includes graphics core 2000, e.g., for forming an integrated circuit and / or for forming an SoC that performs operations described herein.
[0386] In at least one embodiment, graphics core 2000 includes shared instruction cache 2002, texture units 2018, and cache / shared memory 2020 (e.g., including LI, L2, L3, last level cache, or other cache) that are common to execution resources within graphics core 2000. In at least one embodiment, graphics core 2000 can include multiple slices 2001A-2001N or partitions of each core, and graphics processor can include multiple instances of graphics core 2000. In at least one embodiment, each slice 2001A-2001N refers to graphics core 2000. In at least one embodiment, slices 2001A-2001N have sub-slices that are part of slices 2001A-2001N. In at least one embodiment, slices 2001A-2001N are independent of each other or dependent upon each other. In at least one embodiment, slices 2001A-2001N can include support logic including a local instruction cache 2004A-2004N, a thread scheduler (sequencer) 2006A-2006N, a thread dispatcher 2008A-2008N, and a set of registers 2010A-2010N. In at least one embodiment, slices 2001A-2001N can include a set of additional functional units (AFUs 2012A-2012N), floating point units (FPUs 2014A-2014N), integer arithmetic logic units (ALUs 2016A-2016N), address computation units (ACUs 2013A-2013N), double precision floating point units (DPFPUs 2015A-2015N), and matrix processing units (MPUs 2017A-2017N). In at least one embodiment, MPUs 2017A-2017N are referred to as matrix engines.
[0387] In at least one embodiment, each slice 2001A-2001N includes one or more engines for floating point and integer vector operations and one or more engines for accelerating convolution and matrix operations in Al, machine learning, or big data set workloads. In at least one embodiment, one or more slices 2001A-2001N include one or more vector engines for computing vectors (e.g., mathematical operations on vectors). In at least one embodiment, vector engines can compute vector operations in 16-bit floating point (also referred to as “FP16”), 32-bit floating point (also referred to as “FP32”), or 64-bit floating point (also referred to as “FP64”). In at least one embodiment, one or more slices 2001A-2001N include 16 vector engines that are paired with 16 matrix math units to compute matrix / tensor operations, where vector engines and math units are exposed through a matrix extension. In at least one embodiment, a slice is a designated portion of processing resources of a processing unit, e.g., 16 cores and a ray-tracing unit of a processor, or 8 cores, a thread scheduler, a thread dispatcher, and additional functional units of a processor. In at least one embodiment, graphics core 2000 includes one or more matrix engines for computing matrix operations when, for example, computing tensor operations.
[0388] In at least one embodiment, one or more slices 2001A-2001N include one or more ray-tracing units for computing ray-tracing operations (e.g., 16 ray-tracing units per slice of slices 2001A-2001N). In at least one embodiment, ray-tracing units compute ray traversal, triangle intersection, bounding volume intersection, or other ray-tracing operations.
[0389] In at least one embodiment, one or more slices 2001A-2001N include a media slice that encodes, decodes, and / or transcodes data; scales and / or formats data; and / or performs video quality operations on video data.
[0390] In at least one embodiment, one or more slices 2001 A-2001 N are linked to an L2 cache and memory structure, a link connector, a high bandwidth memory (HBM) (e.g., HBM2e, HDMI3) stack, and a media engine. In at least one embodiment, one or more slices 2001 A-2001 N include a plurality of cores (e.g., 16 cores) and a plurality of ray tracing units (e.g., 16) paired with each core. In at least one embodiment, one or more slices 2001 A-2001 N have one or more LI caches. In at least one embodiment, one or more slices 2001 A-2001 N include one or more vector engines; one or more instruction caches to store instructions; one or more LI caches to cache data; one or more shared local memories (SLMs) to store, e.g., data corresponding to instructions; one or more samplers to sample data; one or more ray tracing units to perform ray tracing operations; one or more geometry units to perform operations in a geometry pipeline and / or apply geometric transformations to vertices or polygons; one or more rasterizers to describe images in a vector graphics format (e.g., shapes) and convert them to raster images (e.g., a series of pixels, dots, ...
Claims
1. A processor comprising: one or more circuits to cause precision of one or more layers of one or more first versions of a neural network to be modified based at least in part on a change in accuracy of one or more corresponding layers of one or more second versions of the neural network caused by deactivating one or more weights of the one or more corresponding layers.
2. The processor of claim 1, wherein the one or more circuits are to quantify the change in accuracy of the one or more corresponding layers of the one or more second versions of the neural network based at least in part on one or more comparisons of activations of one or more sparse versions and one or more dense versions of the neural network.
3. The processor of claim 1, wherein the one or more circuits are to generate one or more other weights to be applied to one or more loss values based at least in part on the change in accuracy of the one or more corresponding layers of the one or more second versions of the neural network.
4. The processor of claim 1, wherein the one or more circuits are to quantify the change in accuracy of the one or more corresponding layers based at least in part on a feature alignment loss value for the one or more corresponding layers.
5. The processor of claim 1, wherein the one or more circuits are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified using a scaling factor based at least in part on one or more of hard label predictions, soft logits, or feature maps of the one or more first versions and the one or more second versions of the neural network.
6. The processor of claim 1, wherein the one or more circuits are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on a comparison of accuracy indicators of the one or more first versions of the neural network to accuracy indicators of the one or more second versions of the neural network.
7. The processor of claim 1, wherein the one or more circuits are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified using one or more of hard label distillation loss, soft logits distillation loss, or feature-based distillation loss based at least in part on the one or more first versions of the neural network and the one or more second versions of the neural network.
8. A system comprising: one or more processors to cause precision of one or more layers of one or more first versions of a neural network to be modified based at least in part on a change in accuracy of one or more corresponding layers of one or more second versions of the neural network caused by deactivating one or more weights of the one or more corresponding layers.
9. The system of claim 8, wherein the one or more processors are to quantify the change in accuracy of the one or more corresponding layers of the one or more second versions of the neural network based at least in part on one or more comparisons of activated features of one or more sparse versions and one or more dense versions of the neural network.
10. The system of claim 8, wherein the one or more processors are to generate one or more other weights to be applied to one or more loss values based at least in part on feature distillation loss values of the one or more corresponding layers of the one or more first versions and the one or more second versions of the neural network.
11. The system of claim 8, wherein the one or more processors are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on a comparison of feature calibration loss values of one or more sparse floating point versions of the neural network and corresponding layers of the one or more first versions.
12. The system of claim 8, wherein the one or more processors are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on a total calibration loss of one or more sparse floating point versions of the neural network and the one or more first versions using a scaling factor.
13. The system of claim 8, wherein the one or more processors are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on a total pruning loss of one or more dense versions of the neural network and the one or more second versions.
14. The system of claim 8, wherein the one or more processors are to replicate the one or more second versions of the neural network for use as the one or more first versions of the neural network.
15. A method comprising: causing, by one or more processors, precision of one or more layers of one or more first versions of a neural network to be modified based at least in part on a change in accuracy of one or more corresponding layers of one or more second versions of the neural network caused by deactivating one or more weights of the one or more corresponding layers.
16. The method of claim 15, wherein the one or more processors are to quantify changes in the accuracy of the one or more corresponding layers of the one or more second versions of the neural network based at least in part on one or more comparisons of activations of one or more dense full-precision versions and one or more sparse full-precision versions of the neural network.
17. The method of claim 15, wherein the one or more processors are to modify one or more other weights applied to one or more feature calibration loss values based at least in part on a total calibration loss of the one or more first versions and the one or more second versions of the neural network.
18. The method of claim 15, wherein the one or more processors are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on a comparison of one or more sparse floating-point versions of the neural network and feature distillation loss values of corresponding layers of the one or more first versions of the neural network that are one or more sparse quantized versions of the neural network.
19. The method of claim 15, wherein the one or more processors are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on one or more knowledge distillation neural network modification techniques.
20. The method of claim 15, wherein the one or more processors are to cause the precision of the one or more layers of the one or more first versions of the neural network to be modified based at least in part on a total pruning loss of one or more dense floating-point versions of the neural network and the one or more second versions of the neural network that are one or more sparse full-precision versions of the neural network.