Distributed Weight Updates for Neural Network Backpropagation

Distributed and parallel weight updates across workers in neural network training reduce memory and computational requirements, enhancing training speed and capacity for larger models.

JP7714536B2Active Publication Date: 2025-07-29NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022523930
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-11-05
Filing Date
2020-10-30
Publication Date
2025-07-29
Estimated Expiration
2040-10-30

AI Technical Summary

Technical Problem

Training neural networks is an intensive process that requires significant memory, time, and computational resources, necessitating a reduction in these resources to improve efficiency.

Method used

Implementing parallel and distributed weight updates across multiple workers, where each worker applies gradients to a subset of weights and shares the updated weights, reducing the time and memory footprint by distributing the weight update process.

Benefits of technology

This method accelerates the training process by a factor equal to the number of workers, alleviates memory capacity limitations, and enables larger models or batch sizes, making machine learning training more efficient.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007714536000005
    Figure 0007714536000005
  • Figure 0007714536000006
    Figure 0007714536000006
  • Figure 0007714536000007
    Figure 0007714536000007
Patent Text Reader

Abstract

The speed at which a neural network is trained is improved by updating the neural network weights in parallel. In at least one embodiment, after backpropagation, the gradients are distributed to multiple processors, each of which computes a portion of the updated weights of the neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to U.S. Patent Application No. 16 / 675,069, filed on November 5, 2019, entitled "Distributed Weight Updates for Backpropagation of Neural Networks", which is incorporated herein by reference in its entirety and for all purposes.

[0002] At least one embodiment relates to processing resources used to perform and facilitate the training of neural networks. For example, at least one embodiment relates to a processor or computing system used to train neural networks with various novel techniques described herein.

Background Art

[0003] Neural networks are an important part of many computer-based solutions. Training a neural network, in some instances, can be an intensive iterative process that can use significant memory, time, and computational resources. Thus, reducing the amount of memory, time, or computational resources used to train a neural network is an important problem.

Summary of the Invention

Means for Solving the Problems

[0004] Various techniques will be described with reference to the drawings.

Brief Description of the Drawings

[0005]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 10

Figure 11

Figure 12A

Figure 12B

Figure 12C

Figure 12D

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17A

Figure 17B

Figure 17C

Figure 17D

Figure 17E

Figure 17F

Figure 18

Figure 19A

Figure 19B

Figure 20A

Figure 20B

Figure 21

Figure 22A

Figure 22B

Figure 22C

Figure 22D

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29

Figure 30

Figure 31

Figure 32

Figure 33A

Figure 33B

Figure 34

Figure 35

Figure 36

Figure 37

[0006] This specification describes systems and methods for improving the training of machine learning models by enabling parallel updates of node weights. In at least one embodiment, multiple workers perform forward and backward propagation to create a set of gradients in parallel. In at least one embodiment, a worker may be a thread, process, processor, processor core, or parallel processing circuit that executes instructions in parallel with other workers. In at least one embodiment, the gradients are distributed across the workers, and each worker is assigned a subset of the weights to which the gradients are applied. In at least one embodiment, each worker applies the gradients to the assigned subset of weights in parallel with other workers. In at least one embodiment, after the gradients are applied, the updated weights are distributed among the workers, such that each worker has a complete set of updated weights. In at least one embodiment, updating the weights in parallel with multiple workers improves the speed at which the network can be trained. In at least one embodiment, the backward propagation process is repeatedly iterated until training is complete.

[0007] In at least one embodiment, training a deep learning model with a backpropagation algorithm is an iterative process that includes three stages: a forward propagation pass, a backpropagation pass, and a weight update, repeated for each iteration. In at least one embodiment, the training is distributed across multiple workers in a parallel manner such that the work for the forward and backward propagation passes is divided across the workers. In at least one embodiment, the weight update is performed at least partially in parallel by the workers. In at least one embodiment, some or all of the weight updates can be performed redundantly by multiple workers. In at least one embodiment, the weight update portion of training a deep learning model is accelerated by a factor approximately equal to the number of workers. In at least one embodiment, a worker may be a thread, process, processor, processor core, or a processor with a multi-processor graphical processing unit (“GPU”).

[0008] In at least one embodiment, in a synchronous data parallel distributed deep learning ("DL") training regime, each worker of a plurality of workers can access a local copy of the weights of a machine learning model. In at least one embodiment, the weight gradients are computed per worker in the backward pass, and are all-reduced across all workers such that each worker has its own copy of the sum of the weight gradients across the workers. In at least one embodiment, the workers then use this to repeatedly update their own copies of the weights, ensuring that the workers maintain the same copy of the updated weights as they proceed to the next iteration.

[0009] In at least one embodiment, the all-reduce operation is interleaved with a reduce-scatter operation before the weights are updated and a call to an all-gather after the weights are updated. In at least one embodiment, reduce-scatter leads to each worker computing the sum of its own 1 / k slice of the weight gradients across the workers (where k is the number of workers). In at least one embodiment, each of the k workers updates only the weights corresponding to the slice for which the worker has the sum across the workers. In at least one embodiment, this reduces the weight update time by a factor of k (the number of workers). In at least one embodiment, the updated weights are then all-gathered across the workers, such that each worker receives an updated copy of the new weights.

[0010] In at least one embodiment, the all-reduce is implemented as a combination of reduce-scatter and all-gather, moving the same number of bytes as the algorithm proposed in this example, so the net memory traffic remains unchanged. In at least one embodiment, one difference is that instead of all-gathering the summed weight gradients, the all-gather is applied to the updated weights, enabling a distributed form of weight update.

[0011] In at least one embodiment, the method described herein has the advantage of reducing the total memory footprint per worker. In at least one embodiment, the optimization algorithm needs to maintain a persistent state over training iterations that exceed only the weights and weight gradients. In at least one embodiment, this state is approximately proportional to the size of the weights and includes tensors such as higher-precision copies of the weights, or first and second moment inferences of the gradients. In at least one embodiment, with the distributed optimization algorithm, each of the "k" workers only needs to maintain a 1 / k slice of these tensors corresponding to the weights for which it is responsible for updating.

[0012] In at least one embodiment, the synchronous data parallel DL training scheme focuses on optimizing the optimization code per worker while maintaining the principle of duplicate updates across all workers. In at least one embodiment, the model parallel distributed training scheme benefits from smaller optimization steps per worker due to the fact that each worker only owns a fraction of the model's weights. In at least one embodiment, this comes with substantially increased complexity in other parts of the execution flow.

[0013] In at least one embodiment, as DL training scales to more GPUs (usually across nodes), the cost of weight updates becomes a prominent performance bottleneck for many models. In at least one embodiment, this is addressed by improving the performance of the weight update kernel on a per-worker basis. However, in at least one embodiment, unlike other parts of the computation that become faster as more workers are added, this is a fixed cost from an end-to-end runtime perspective. In at least one embodiment, this method enables the cost of optimization to decrease as the scale increases and makes algorithmic changes to reduce that cost as a fraction of the total runtime. In at least one embodiment, this optimization is important to make several machine learning training models competitive.

[0014] In at least one embodiment, the techniques described herein have the effect of alleviating some of the memory capacity limitations currently observed when training large models on a GPU. In at least one embodiment, by distributing the FP32 master copy of the weights in mixed precision training and some of the other persistent states required for optimizations such as SGD with momentum and Adam, a non-trivial amount of memory is freed up, potentially enabling larger models or larger batch sizes for existing models.

[0015] FIG. 1 shows an example of a machine learning process in which weight updates are performed redundantly and in parallel, according to one embodiment. In at least one embodiment, a computer system trains a machine learning model comprising a set of nodes or neurons, each having an associated weight or coefficient. In at least one embodiment, the weights 102 are updated during the training process in response to training the data. In at least one embodiment, training the data includes inputs and outputs, and the training process creates a trained machine learning model that infers a function that fits the training data. In at least one embodiment, forward and backward propagation processes are used to generate a set of gradients 104 for the weights 102. In at least one embodiment, in forward propagation, input data is provided to the machine learning model to produce output values, and the output values are backpropagated to infer a set of gradients 104.

[0016] In at least one embodiment, an all-reduce process is performed in which the gradients 104 are distributed to a set of workers. In at least one embodiment, the worker applies the gradient 104 to the previous weights to create updated weights 108. In at least one embodiment, each worker applies the gradient 106 redundantly to all the weights so as to process and generate information that each worker fits. In at least one embodiment, the updated weights 108 are then used during the next round of forward and backward propagation.

[0017] In at least one embodiment, a worker is implemented as an individual thread that executes on a computer system such as the computer system described below. In at least one embodiment, a worker is implemented as an individual process on a graphical processing unit such as the graphical processing unit described below. In at least one embodiment, an individual worker is implemented as a process on a multiprocessor computer system. In at least one embodiment, an individual worker executes on a separate core of a multi-core processor.

[0018] In at least one embodiment, process 100 enables parallel processing of forward and backward propagation to create weight gradients. In at least one embodiment, the rate at which gradients are created is increased by adding additional workers to the training task. In at least one embodiment, weight updates are performed redundantly by each worker, thereby taking a consistent period regardless of the number of workers applied.

[0019] In at least one embodiment, the number of weights in a machine-learned model is at least partially based on the number of nodes and layers. In at least one embodiment, the model has many layers and nodes, and weight updates consume a significant amount of processing time.

[0020] In at least one embodiment, FIG. 1 shows a synchronous data parallel distributed deep learning ("DL") training form in which each worker of a plurality of workers can access a local copy of the weights of a machine learning model. In at least one embodiment, weight gradients are calculated for each worker in the reverse pass and are all-reduced across all workers such that each worker has its own copy of the sum of the weight gradients across the workers. In at least one embodiment, the workers can use a copy of the weight gradients to redundantly update each respective copy of the weights, ensuring that each worker maintains a matching copy of the updated weights for future iterations.

[0021] Figure 2 shows an example of a machine learning process 200 in which weight updates are distributed across a set of workers and performed in parallel. In at least one embodiment, a machine-learned model that includes a set of nodes has an associated set of weights 202. In at least one embodiment, the set of weights is a set of numerical coefficients that, when applied to the inputs of a neural network and propagated through the neural network, approximate a function defined by a set of test data used to train the model. In at least one embodiment, test data is propagated forward through the model to produce results, and an error that is determined based at least in part on the predicted results within the test data is propagated backward to produce a gradient estimate. In at least one embodiment, the gradient estimate is used to adjust the weights of the model.

[0022] In at least one embodiment, the set of workers operates in each way to produce a set of gradients 204 in parallel. In at least one embodiment, the set of gradients 204 is produced by applying a gradient descent algorithm. In at least one embodiment, the gradient is a slope that can be expressed as a ratio between the produced error and the input parameters of the neural network. In at least one embodiment, the gradient represents the relationship between the change in the input and the produced error. In at least one embodiment, the set of gradients 204 includes a set of partial derivative functions of the contribution of each input parameter to the change in the error.

[0023] In at least one embodiment, the reducing scatter operation is performed before the weights are updated 206, and the call to the all-gather is performed after the weights are updated. In at least one embodiment, the reducing scatter leads to each worker computing the sum across its own 1 / k slice of the weight gradients (where k is the number of workers). In at least one embodiment, the various assignments of work can be distributed to the workers by available processing power, bandwidth, memory, or other computing resources. In at least one embodiment, the number of workers (k) each updates these weights corresponding to the assigned slices. In at least one embodiment, this operation is performed in parallel by the workers, thereby reducing the period required to complete the weight update. In at least one embodiment, the period required to update the weights is reduced by a factor of K, where K is the number of workers.

[0024] In at least one embodiment, the worker is implemented as an individual thread running on a computer system such as the computer system described below. In at least one embodiment, the worker is implemented as an individual process on a graphical processing unit such as the graphical processing unit described below. In at least one embodiment, the individual worker is implemented as a process on a multiprocessor computer system. In at least one embodiment, the individual worker runs on a separate core of a multi-core processor.

[0025] In at least one embodiment, after each worker updates its portion of the weights, the updated portions are distributed among the workers such that each worker has a set of fully updated weights 208. In at least one embodiment, the updated weights are then all-gathered across the workers such that each worker obtains a fully updated copy of the new weights for the neural network.

[0026] Figure 3 shows an example of a machine learning process 300 in which weight updates are scattered and performed in parallel. In at least one embodiment, a set of initial weights is distributed across four workers, and a corresponding set of gradients 302 is computed by the workers. In at least one embodiment, a worker can be implemented as an individual thread running on a computer system, an individual process on a graphical processing unit, or a process on a multiprocessor computer system. In at least one embodiment, the reduction scattering operation divides the weight update operation into approximately the same number of parts that are distributed to the workers. In at least one embodiment, at an intermediate state 304, partial gradients are distributed across the workers. In at least one embodiment, a first worker is assigned a first part of a weight update 308, a second worker is assigned a second part of a weight update 310, a third worker is assigned a third part of a weight update 312, and a fourth worker is assigned a fourth part of a weight update 314. In at least one embodiment, the parts of the weight updates are distributed evenly across the workers. In at least one embodiment, the parts of the weight updates are non-overlapping.

[0027] In at least one embodiment, each worker processes its assigned part of the weight update by applying the gradient to the existing weights. In at least one embodiment, each worker distributes updated weight parts to the other workers so that each worker has a conforming copy of the fully updated weights. In at least one embodiment, each worker updates the weights within a region of shared memory accessible to all workers. In at least one embodiment, the gradients and updated weights are distributed among the workers using interprocess communication. In at least one embodiment, the completed parts are collected in a finished state 306 so that each worker has a conforming copy of the updated weights.

[0028] Figure 4 shows an example of a machine learning process 400 in which weight updates are scattered and performed in parallel and duplicate. In at least one embodiment, a set of initial weights is distributed across four workers, and a corresponding set of gradients 402 is calculated by the workers. In at least one embodiment, a worker can be implemented as an individual thread running on a computer system, an individual process on a graphical processing unit, or a process on a multiprocessor computer system. In at least one embodiment, the reduction scattering operation divides the weight update operation into approximately the same number of parts distributed to the workers. In at least one embodiment, at an intermediate state 404, partial gradients are distributed across the workers. In at least one embodiment, the first worker is assigned a first part of the weight update 408, the second worker is assigned a second part of the weight update 410, the third worker is assigned a third part of the weight update 412, and the fourth worker is assigned a fourth part of the weight update 414. In at least one embodiment, the parts of the weight update are evenly distributed across the set of workers such that each weight is calculated by two or more workers. In at least one embodiment, the workers are split into pairs, and the weight updates are evenly distributed across the pairs. In at least one embodiment, each worker is assigned a part of the weight update that partially overlaps at least a part of the weight updates of two other workers.

[0029] In at least one embodiment, each worker processes its assigned part of the weight update by applying the gradient to the existing weights. In at least one embodiment, each worker distributes the updated weight parts to the other workers such that each worker has a conforming copy of the fully updated weights. In at least one embodiment, each worker updates the weights within a region of shared memory accessible to all workers. In at least one embodiment, the gradients and the updated weights are distributed among the workers using interprocess communication. In at least one embodiment, the completed parts are collected in a finishing state 406 such that each worker has a conforming copy of the updated weights.

[0030] FIG. 5 shows an example of a machine learning process 500 that is scattered across a set of workers and performed in parallel based on a processing bandwidth where weight updates are available, according to one embodiment. In at least one embodiment, a set of initial weights is distributed across four workers, and a corresponding set of gradients 502 is calculated by the workers. In at least one embodiment, a worker can be implemented as an individual thread running on a computer system, an individual process on a graphical processing unit, or a process on a multiprocessor computer system. In at least one embodiment, the reduction scatter operation divides the weight update operation into approximately the same number of parts that are distributed to the workers. In at least one embodiment, at an intermediate state 504, partial gradients are distributed across the workers. In at least one embodiment, a first worker is assigned a first part of the weight update 508, a second worker is assigned a second part of the weight update 510, a third worker is assigned a third part of the weight update 512, and a fourth worker is assigned a fourth part of the weight update 514. In at least one embodiment, the parts of the weight update are distributed across the workers such that all weights are determined and the weights are not calculated redundantly. In at least one embodiment, the amount of weight update assigned to each worker is determined based on the characteristics of each worker. In at least one embodiment, the weight update is determined based on the amount of processing bandwidth available to each worker. In at least one embodiment, work is distributed to the workers on an incremental basis, and when each worker completes its assigned part, additional work is assigned by an executive process adjustment allocation of the weight update work.

[0031] In at least one embodiment, each worker processes its assigned portion of the weight update by applying a gradient to the existing weights. In at least one embodiment, each worker distributes updated weight portions to other workers such that each worker has a conforming copy of the fully updated weights. In at least one embodiment, each worker updates the weights within a region of shared memory accessible to all workers. In at least one embodiment, the gradients and updated weights are distributed among the workers using interprocess communication. In at least one embodiment, the completed portions are collected in a finished state 506 such that each worker has a conforming copy of the updated weights.

[0032] FIG. 6 illustrates an example of a process 600 for training a machine learning model in which weight updates are performed in parallel by a computer system, according to one embodiment. In at least one embodiment, at block 602, the computer system begins training the machine learned model by propagating a set of input values from a set of training data through the machine learned model to produce a set of outputs. In at least one embodiment, a set of gradients is produced by backpropagating the difference between the produced outputs and the training data 604. In at least one embodiment, the gradients describe the relationship between the change in the input values and the error terms produced by the machine learned model. In at least one embodiment, the gradients are produced in parallel by a plurality of workers. In at least one embodiment, the workers can execute on separate threads, processors, processor cores, or processes on a computer system. In at least one embodiment, the gradients are produced at least partially in parallel by the workers.

[0033] In at least one embodiment, at block 606, the gradients are distributed across a set of workers. In at least one embodiment, the gradients are distributed through an interprocess communication mechanism. In at least one embodiment, the gradients are distributed across a set of workers using shared memory. In at least one embodiment, the gradients are aggregated by an executor and the redistributed gradients are sent to each of the workers.

[0034] In at least one embodiment, at block 608, the computer system performing the training analyzes a set of weights associated with the machine-learned model and divides the task of updating the weights into a set of weight update portions to be performed by the workers. In at least one embodiment, at block 610, each worker calculates a portion of the weight update in parallel using the gradients. In at least one embodiment, the weight update operations are evenly distributed across the workers and the weight update operations are performed in parallel. In at least one embodiment, the weight update operations can be distributed across the workers as illustrated and described with reference to FIGS. 3-5.

[0035] In at least one embodiment, at block 612, each worker produces the result of the weight update operation distributed to the other workers such that each worker owns a conforming copy of the complete set of updated weights. In at least one embodiment, at block 614, each worker merges the updated portion of the weights obtained from the other workers with its portion of the updated weights to produce the complete set of updated weights. In at least one embodiment, the weights are assembled within a shared memory region accessible to all workers. In at least one embodiment, the shared memory region is a semiconductor memory having an addressable region mapped within an address space accessible to all workers. In at least one embodiment, the shared memory is a file on a storage volume such as a hard disk. In at least one embodiment, the storage volume is a hard disk volume. In at least one embodiment, the storage volume is a solid state memory device such as an SD drive.

[0036] In at least one embodiment, at block 614, process 600 returns to block 602 and proceeds to the next iteration of the training process.

[0037] FIG. 7 shows an example of a process 700 for training a machine learning model that is distributed to individual workers based on the processing power available for weight updates as a result of being performed by a computer system, according to one embodiment. In at least one embodiment, at block 702, the computer system begins training the machine learned model by propagating a set of input values from a set of training data through the machine learned model in order to produce a set of outputs. In at least one embodiment, a set of gradients is produced by backpropagating the difference between the produced outputs and the training data 704. In at least one embodiment, the gradients describe the relationship between the change in the input values and the error terms produced by the machine learned model. In at least one embodiment, the gradients are produced in parallel by a plurality of workers. In at least one embodiment, the workers can execute on separate threads, processors, processor cores, or processes on the computer system. In at least one embodiment, the gradients are produced at least partially in parallel by the workers.

[0038] In at least one embodiment, at block 706, the gradients are distributed across a set of workers. In at least one embodiment, the gradients are distributed through an interprocess communication mechanism. In at least one embodiment, the gradients are distributed across a set of workers using shared memory. In at least one embodiment, the gradients are aggregated by an executive process that redistributes the gradients to each worker. In at least one embodiment, the executive process is a process that coordinates the overall process of training the neural network.

[0039] In at least one embodiment, at block 708, the executive process determines the amount of processing bandwidth available to each worker. In at least one embodiment, the amount of processing bandwidth available to each worker is determined by executing a test task on each worker. In at least one embodiment, the amount of available processing vendor is determined based on the clock speed, processor type, or process priority associated with each worker.

[0040] In at least one embodiment, at block 710, the computer system performing the training analyzes a set of weights associated with the machine-learned model and divides the task of updating the weights into a set of weight update portions to be performed by the workers. In at least one embodiment, the weight update operations are distributed across the workers and the weight update operations are performed in parallel. In at least one embodiment, the weight update operations can be distributed across the workers as illustrated and described with reference to FIGS. 3-5. In at least one embodiment, the work is distributed to the workers in proportion to the amount of processing bandwidth as determined above. In at least one embodiment, at block 712, each worker uses the gradients to compute portions of the weight updates in parallel.

[0041] In at least one embodiment, at block 714, the workers produce the results of the weight update operations distributed to other workers such that each worker owns a conforming copy of the complete set of updated weights. In at least one embodiment, at block 716, each worker merges the updated portion of the weights obtained from other workers with its portion of the updated weights to produce a complete set of updated weights. In at least one embodiment, the weights are assembled within a shared memory region accessible to all workers. In at least one embodiment, the shared memory region is a semiconductor memory having an addressable region mapped within an address space accessible to all workers. In at least one embodiment, the shared memory is a file on a storage such as a hard disk. In at least one embodiment, the storage is a hard disk storage. In at least one embodiment, the storage is a solid state memory device such as an SD drive.

[0042] In at least one embodiment, at block 716, process 700 returns to block 702 and proceeds to the next iteration of the training process.

[0043] FIG. 8 shows an example of a process 800 for training a machine learning model in which result implementation by a computer system causes weight updates to be incrementally distributed to individual workers. In at least one embodiment, at block 802, an executive process divides the task of updating the weights of a machine learned model into a set of work units. In at least one embodiment, the number of work units is substantially larger than the number of available workers on which the task of updating the weights is performed. In at least one embodiment, at block 802, individual work units are assigned by the executive to each available worker. In at least one embodiment, the executive is a process that controls the actions of multiple workers running on one or more processors. In at least one embodiment, the executive can start, stop, and assign work to one or more workers.

[0044] In at least one embodiment, at decision block 804, a worker executes each work unit until at least one worker has completed its respective work unit and becomes idle. In at least one embodiment, at block 806, an idle worker is identified and additional work units are assigned to the idle worker. In at least one embodiment, at decision block 808, if there are additional work units to be assigned, execution returns to decision block 804 and the executive waits until the additional worker becomes idle. In at least one embodiment, if there are no additional work units, execution proceeds to block 810. In at least one embodiment, at block 810, execution waits until all workers have completed their assigned tasks. In at least one embodiment, when all workers have completed their assigned tasks, the process of updating the weights of the machine learned model is complete. In at least one embodiment, worker utilization can be improved when a heterogeneous set of workers is utilized in parallel to perform the weight update task by assigning work units incrementally.

[0045] In at least one embodiment, a worker may be a parallel processing circuit that executes instructions in parallel with a thread, process, processor, processor core, or other worker. The worker operates on a computer system to enable tasks to be performed simultaneously as described below and illustrated in FIGS. 9-37. In at least one embodiment, for example, the worker can be implemented on a core of a graphics processing unit such as the GPU illustrated and described in FIG. 22D. In at least one embodiment, the techniques described herein can be implemented on a computer system such as the computer system shown in FIGS. 13-17D and described in the associated description. In at least one embodiment, the techniques described herein are implemented using executable instructions stored on a physical computer-readable memory, and as a result of the executable instructions being executed on one or more processors of the computer system, the system implements the techniques described herein.

[0046] Inference and training logics FIG. 9A shows inference and / or training logic 915 used to perform inference and / or training operations with respect to one or more embodiments. Details regarding the inference and / or training logic 915 are provided below in conjunction with FIGS. 9A and / or 9B.

[0047] In at least one example, the inference and / or training logic 915 may include, without limitation, code and / or data storage 901 for storing forward propagation and / or output weights, and / or input / output data, and / or other parameters for constructing neurons or layers of a neural network that are trained and / or used to infer in aspects of one or more examples. In at least one example, the training logic 915 may include, or be coupled to, code and / or data storage 901 for storing graph code or other software for controlling timing and / or order, and code and / or data storage 901 may have weight and / or other parameter information loaded to configure logic including integer and / or floating point units (collectively arithmetic logic units (ALUs)). In at least one example, code such as graph code loads weight or other parameter information to the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one example, the code and / or data storage 901 stores weight parameters and / or input / output data for each layer of a neural network that is trained or used in conjunction with one or more examples while propagating the input / output data and / or weight parameters forward during training and / or inference using aspects of one or more examples. In at least one example, any portion of the code and / or data storage 901 may be included with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.

[0048] In at least one embodiment, any portion of the code and / or data storage 901 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or the code and / or data storage 901 may be cache memory, dynamic randomly addressable memory (“DRAM”), static randomly addressable memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the selection of whether the code and / or the code and / or data storage 901 is internal or external to, for example, a processor, or the selection of whether it is composed of DRAM, SRAM, flash, or some other type of storage may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.

[0049] In at least one embodiment, the inference and / or training logic 915 may include, without limitation, code and / or data storage 905 for storing backpropagation and / or output weights, and / or input / output data corresponding to neurons or layers of a neural network that are trained and / or used for inference in one or more embodiments. In at least one embodiment, the code and / or data storage 905 stores the weight parameters and / or input / output data of each layer of a neural network that is trained or used in conjunction with one or more embodiments while backpropagating the input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, the training logic 915 may include, or be coupled to, code and / or data storage 905 for storing graph code or other software for controlling timing and / or order, and the code and / or data storage 905 has weight and / or other parameter information loaded therein to configure logic including integer and / or floating point units (collectively referred to as arithmetic logic units (ALUs)). In at least one embodiment, code such as graph code loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, any portion of the code and / or data storage 905 may be included with other on-chip or off-chip data storage including the L1, L2, or L3 cache of the processor, or system memory. In at least one embodiment, any portion of the code and / or data storage 905 may be internal or external to one or more processors, or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 905 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage.In at least one embodiment, the choice of whether the code and / or data storage 905 is, for example, internal or external to the processor, or is composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to the on-chip versus off-chip available storage, the latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or any combination of these factors.

[0050] In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be separate storage structures. In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be the same storage structure. In at least one embodiment, the code and / or data storage 901 and the code and / or data storage 905 may be partially the same storage structure and partially separate storage structures. In at least one embodiment, any portion of the code and / or data storage 901 and the code and / or data storage 905 may be included together with the L1, L2, or L3 cache of the processor, or other on-chip or off-chip data storage including system memory.

[0051] In at least one embodiment, the inference and / or training logic 915 may include one or more arithmetic logic units (“ALUs”) 910, including, without limitation, integer and / or floating point units, for performing logical and / or arithmetic operations based at least in part on or indicated by the training and / or inference code (e.g., graph code), the results of which may generate activations (e.g., output values from layers or neurons in a neural network) stored in activation storage 920, which are functions of input / output and / or weight parameter data stored in code and / or data storage 901 and / or code and / or data storage 905. In at least one embodiment, the activations stored in activation storage 920 are generated according to linear algebra and / or matrix-based calculations performed by ALU 910 in response to executing instructions or other code, where weight values stored in code and / or data storage 905 and / or data 901 are used as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data storage 905, or code and / or data storage 901, or in another storage, on-chip or off-chip.

[0052] In at least one embodiment, ALU 910 is included within one or more processors or other hardware logic devices or circuits, while in other embodiments, ALU 910 may be external to the processors or other hardware logic devices or circuits that use them (e.g., a coprocessor). In at least one embodiment, ALU 910 may be included within an execution unit of a processor or may otherwise be included within an ALU bank accessible by execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.). In at least one embodiment, data storage 901, code and / or data storage 905, and activation storage 920 may be in the same processor or other hardware logic devices or circuits, while in other embodiments, they may be in different processors or other hardware logic devices or circuits, or some combination of the same processor or other hardware logic devices or circuits and different processors or other hardware logic devices or circuits. In at least one embodiment, any portion of activation storage 920 may be included with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Additionally, inference and / or training code may be stored with other code accessible to the processor or other hardware logic or circuitry, and may be fetched and / or processed using the processor's fetch, decode, schedule, execute, retire, and / or other logic.

[0053] In at least one embodiment, the activation storage 920 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the activation storage 920 may be fully or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the selection of whether the activation storage 920 is, for example, inside or outside a processor, or the selection of being composed of DRAM, SRAM, flash, or some other type of storage, may be determined according to on-chip versus off-chip available storage, latency requirements of the training and / or inference functions being executed, the batch size of the data used in neural network inference and / or training, or a combination from any of these factors. In at least one embodiment, the inference and / or training logic 915 shown in FIG. 9A may be used in conjunction with an application-specific integrated circuit (ASIC) such as a TensorFlow® processing unit from Google, an inference processing unit (IPU) from Graphcore®, or a Nervana® (e.g., "Lake Crest") processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 915 shown in FIG. 9A may be used in conjunction with other hardware such as central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or field programmable gate array (FPGA).

[0054] FIG. 9B illustrates inference and / or training logic 915 according to at least one various embodiment. In at least one embodiment, the inference and / or training logic 915 may include, without limitation, hardware logic in which computational resources are dedicated to, or otherwise used only in conjunction with, weight values or other information corresponding to one or more layers of neurons in a neural network. In at least one embodiment, the inference and / or training logic 915 illustrated in FIG. 9B may be used in conjunction with an application-specific integrated circuit (ASIC), such as a Tensorflow® processing unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corporation. In at least one embodiment, the inference and / or training logic 915 illustrated in FIG. 9B may be used in conjunction with other hardware, such as central processing unit (CPU) hardware, graphics processing unit (“GPU”) hardware, or a field-programmable gate array (FPGA). In at least one embodiment, inference and / or training logic 915 may include, without limitation, code and / or data storage 901 and code and / or data storage 905, which may be used to store code (e.g., graph code), weight and / or bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. In at least one embodiment shown in FIG. 9B , each of code and / or data storage 901 and code and / or data storage 905 is associated with dedicated computational resources, such as compute hardware 902 and compute hardware 906, respectively. In at least one embodiment, compute hardware 902 and compute hardware 906 each include one or more ALUs that perform mathematical functions, such as linear algebraic functions, solely on the information stored in code and / or data storage 901 and code and / or data storage 905, respectively, with the results stored in activation storage 920.

[0055] In at least one embodiment, each of code and / or data storage 901 and 905, and corresponding compute hardware 902 and 906, respectively corresponds to different layers of a neural network, such that the activation resulting from one "storage / compute pair 901 / 902" of code and / or data storage 901 and compute hardware 902 is provided as an input to the next "storage / compute pair 905 / 906" of code and / or data storage 905 and compute hardware 906 in order to reflect the conceptual organization of the neural network. In at least one embodiment, storage / compute pairs 901 / 902, and 905 / 906 may correspond to two or more layers of a neural network. In at least one embodiment, additional storage / compute pairs (not shown) may be included in inference and / or training logic 915 after, or in parallel with, storage / compute pairs 901 / 902, and 905 / 906.

[0056] Training and Introduction of Neural Networks FIG. 10 shows the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, an untrained neural network 91006 is trained using a training data set 1002. In at least one embodiment, the training framework 1004 is the PyTorch framework, while in other embodiments, the training framework 1004 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 and enables it to be trained using the processing resources described herein to generate a trained neural network 1008. In at least one embodiment, the weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, the training may be performed in any of a supervised, semi-supervised, or unsupervised manner.

[0057] In at least one embodiment, the untrained neural network 1006 is trained using supervised learning, where the training data set 1002 includes inputs paired with the desired outputs for the inputs, or the training data set 1002 includes inputs having known outputs, and the output of the neural network 1006 is manually scored. In at least one embodiment, the untrained neural network 1006 is trained in a supervised manner, processes the inputs from the training data set 1002, and compares the resulting output to a set of predicted or desired outputs. In at least one embodiment, an error is then backpropagated through the untrained neural network 1006. In at least one embodiment, the training framework 1004 adjusts the weights that control the untrained neural network 1006. In at least one embodiment, the training framework 1004 includes a tool that monitors how well the untrained neural network 1006 converges towards a model such as a trained neural network 1008 that is suitable for generating correct answers in results 1014 etc. based on known input data such as a new data set 1012. In at least one embodiment, the training framework 1004 repeatedly trains the untrained neural network 1006 while adjusting the weights to refine the output of the untrained neural network 1006 using a loss function and an adjustment algorithm such as stochastic gradient descent. In at least one embodiment, the training framework 1004 trains the untrained neural network 1006 until the untrained neural network 1006 reaches the desired accuracy. In at least one embodiment, the trained neural network 1008 can then be introduced to implement any number of machine learning operations.

[0058] In at least one embodiment, the untrained neural network 1006 is trained using unsupervised learning, where the untrained neural network 1006 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training data set 1002 includes input data that has no associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 1006 can learn to group within the training data set 1002 and can determine how individual inputs relate to the untrained data set 1002. In at least one embodiment, unsupervised training can be used to generate a self-organizing map, which is a trained neural network 1008 of a type that can perform operations useful for reducing the dimensions of a new data set 1012. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which enables identification of data points within the new data set 1012 that deviate from the normal pattern of the new data set 1012.

[0059] In at least one embodiment, semi-supervised learning may be used, which is a technique in which labeled data and unlabeled data are mixed in the training data set 1002. In at least one embodiment, the training framework 1004 may be used to perform incremental learning, such as by transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 1008 to adapt to a new data set 1012 without forgetting the knowledge taught into the network during initial training.

[0060] Data Center FIG. 11 shows an exemplary data center 1100 in which at least one embodiment may be used. In at least one embodiment, the data center 1100 includes a data center infrastructure layer 1110, a framework layer 1120, a software layer 1130, and an application layer 1140.

[0061] In at least one embodiment, as shown in FIG. 11, the data center infrastructure layer 1110 may include a resource orchestrator 1112, grouped computing resources 1114, and node computing resources ("node C.R.") 1116(1) to 1116(N), where "N" represents any positive integer. In at least one embodiment, the node C.R. 1116(1) to 1116(N) may include any number of central processing units ("CPU") or other processors (including accelerators, field programmable gate arrays (FPGA), graphics processors, etc.), memory devices (e.g., dynamic read-only memory), storage devices (e.g., semiconductor drives or disk drives), network input / output ("NW I / O") devices, network switches, virtual machines ("VM"), power modules, and cooling modules, but are not limited thereto. In at least one embodiment, one or more of the node C.R. 1116(1) to 1116(N) may be a server having one or more of the computing resources described above.

[0062] In at least one embodiment, the grouped computing resources 1114 may include separate groups of node C.R.s housed within one or more racks (not shown), or multiple racks housed in a data center at various graphical locations (also not shown). Separate groups of node C.R.s within the grouped computing resources 1114 may include grouped compute resources, network resources, memory resources, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, some node C.R.s that include a CPU or processor may be grouped within one or more racks to provide compute resources for supporting one or more workloads. In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.

[0063] In at least one embodiment, the resource orchestrator 1112 may configure or otherwise control one or more node C.R.s 1116(1) - 1116(N) and / or the grouped computing resources 1114. In at least one embodiment, the resource orchestrator 1112 may include a software design infrastructure ("SDI") management entity for the data center 1100. In at least one embodiment, the resource orchestrator may include hardware, software, or some combination thereof.

[0064] In at least one embodiment shown in FIG. 11 , framework layer 1120 includes job scheduler 1132, configuration manager 1134, resource manager 1136, and distributed file system 1138. In at least one embodiment, framework layer 1120 may include frameworks to support software 1131 in software layer 1130 and / or one or more applications 1142 in application layer 1140. In at least one embodiment, software 1131 or applications 1142 may each include web-based service software or applications, such as those offered by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 1120 may be a type of free and open-source software web application framework, such as, but not limited to, Apache Spark™ (hereinafter “Spark”), which can use distributed file system 1138 for large-scale data processing (e.g., “big data”). In at least one embodiment, job scheduler 1132 may include a Spark driver to facilitate scheduling of workloads supported by various tiers of data center 1100. In at least one embodiment, configuration manager 1134 may be capable of configuring different tiers, such as software tier 1130 and framework tier 1120, which includes Spark and distributed file system 1138 to support large-scale data processing. In at least one embodiment, resource manager 1136 may be capable of managing clustered or grouped computing resources that are mapped or allocated to support distributed file system 1138 and job scheduler 1132. In at least one embodiment, the clustered or grouped computing resources may include grouped computing resources 1114 in data center infrastructure tier 1110.In at least one embodiment, the resource manager 1136 may manage these mappings or allocated computing resources in cooperation with the resource orchestrator 1112.

[0065] In at least one embodiment, the software 1132 included in the software layer 1130 may include software used by at least a portion of the node C.R. 1116(1)-1116(N), the grouped computing resources 1114, and / or the distributed file system 1138 of the framework layer 1120. One or more types of software may include, but are not limited to, Internet web page search software, email virus scan software, database software, and streaming video content software.

[0066] In at least one embodiment, the application 1142 included in the application layer 1140 may include one or more types of applications used by at least a portion of the node C.R. 1116(1)-1116(N), the grouped computing resources 1114, and / or the distributed file system 1138 of the framework layer 1120. One or more types of applications may include, but are not limited to, any number of genomics applications, recognition computing, and software for training or inference, machine learning application including machine learning framework software (such as PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.

[0067] In at least one embodiment, any one of the configuration manager 1134, the resource manager 1136, and the resource orchestrator 1112 may implement any number and type of self-corrective measures based on any amount and type of data obtained in any technically feasible manner. In at least one embodiment, the self-corrective measures may prevent the data center operator of the data center 1100 from determining configurations that may be defective and may eliminate parts of the data center that are not being fully utilized and / or have low performance.

[0068] In at least one embodiment, the data center 1100 may include tools, services, software, or other resources for training one or more machine learning models or for predicting or inferring information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weight parameters according to a neural network architecture using the software and computing resources described above with respect to the data center 1100. In at least one embodiment, a trained machine learning model corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1100 by using weight parameters calculated by one or more techniques described herein.

[0069] In at least one embodiment, the data center may use a CPU, an application specific integrated circuit (ASIC), a GPU, an FPGA, or other hardware to perform training and / or inference using the resources described above. Further, the one or more software and / or hardware resources described above may be configured as a service to enable a user to perform training or inference of information, such as image recognition, speech recognition, or other artificial intelligence services.

[0070] Using the inference and / or training logic 915, inference and / or training operations associated with one or more embodiments are performed. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of FIG. 11 for inference or prediction operations, at least in part based on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network. In at least one embodiment, the training logic 915 includes two or more processing cores for separately training portions of the neural network in parallel as described herein.

[0071] Autonomous vehicle FIG. 12A shows an example of an autonomous vehicle 1200 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1200 (or referred to herein as "vehicle 1200") can be a passenger vehicle such as, without limitation, a car, truck, bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, the vehicle 1200 may be a trailer truck of a semi-tractor for cargo transportation. In at least one embodiment, the vehicle 1200 may be an aircraft, a robotic vehicle, or another type of vehicle.

[0072] The autonomous vehicle may be described from the perspective of the automation level defined by the National Highway Traffic Safety Administration (NHTSA), a part of the US Department of Transportation, and the Society of Automotive Engineers (SAE)'s "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806 issued on June 15, 2018, Standard No. J3016-201609 issued on September 30, 2016, and the old and new versions of this standard). In one or more embodiments, vehicle 1200 may be capable of corresponding to the functionality according to one or more of automation levels 1 to 5 of the autonomous driving level. For example, in at least one embodiment, vehicle 1200 may be capable of corresponding to conditional automation (level 3), highly automated (level 4), and / or fully automated (level 5) according to the embodiment.

[0073] In at least one embodiment, vehicle 1200 may include components such as, without limitation, a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of the vehicle. In at least one embodiment, vehicle 1200 may include a propulsion system 1250 such as, without limitation, an internal combustion engine, a hybrid power plant, a fully electric engine, and / or another type of propulsion system. In at least one embodiment, propulsion system 1250 may be coupled to the drive train of vehicle 1200, and the drive train may include, without limitation, a transmission for enabling the propulsion of vehicle 1200. In at least one embodiment, propulsion system 1250 may be controlled in response to receiving a signal from throttle / accelerator 1252.

[0074] In at least one embodiment, the steering system 1254, which may include a steering wheel without limitation, is used to steer the vehicle 1200 (e.g., along a desired path or route) when the propulsion system 1250 is operating (e.g., when the vehicle is in motion). In at least one embodiment, the steering system 1254 may receive a signal from a steering actuator 1256. The steering wheel may be optional with respect to fully automated (level 5) functionality. In at least one embodiment, a brake sensor system 1246 may be used to operate the vehicle brakes in response to receiving a signal from a brake actuator 1248 and / or a brake sensor.

[0075] In at least one embodiment, controller 1236, which may include one or more system-on-chips (“SoCs”) (not shown in FIG. 12A) and / or graphics processing units (“GPUs”) without limitation, provides signals (e.g., representing commands) to one or more components and / or systems of vehicle 1200. For example, in at least one embodiment, controller 1236 may transmit signals for operating the vehicle brakes via brake actuator 1248, signals for operating steering system 1254 via steering actuator 1256, and signals for operating propulsion system 1250 via throttle / accelerator 1252. Controller 1236 may include one or more integrated (e.g., monolithic) computing devices (e.g., supercomputers) that process sensor signals and output operation commands (e.g., signals representing commands) to enable autonomous driving and / or assist a human driver in operating vehicle 1200. In at least one embodiment, controller 1236 may include a first controller 1236 for autonomous driving functions, a second controller 1236 for functional safety functions, a third controller 1236 for artificial intelligence functions (e.g., computer vision), a fourth controller 1236 for infotainment functions, a fifth controller 1236 for redundancy in emergencies, and / or other controllers. In at least one embodiment, a single controller 1236 may handle two or more of the above functionalities, two or more controllers 1236 may handle a single functionality, and / or any combination thereof may be possible.

[0076] In at least one embodiment, controller 1236 provides signals for controlling one or more components and / or systems of vehicle 1200 in response to sensor data (e.g., sensor inputs) received from one or more sensors. In at least one embodiment, the sensor data may be received from, for example and without limitation, a global navigation satellite system (“GNSS”) sensor 1258 (e.g., a global positioning system sensor), a RADAR sensor 1260, an ultrasonic sensor 1262, a LIDAR sensor 1264, an inertial measurement unit (“IMU”) sensor 1266 (e.g., an accelerometer, a gyroscope, a magnetic compass, a magnetometer, etc.), a microphone 1296, a stereo camera 1268, a wide-angle camera 1270 (e.g., a fish-eye camera), an infrared camera 1272, a surround camera 1274 (e.g., a 360-degree camera), a long-range camera (not shown in FIG. 12A), a mid-range camera (not shown in FIG. 12A), a speed sensor 1244 (e.g., for measuring the speed of vehicle 1200), a vibration sensor 1242, a steering sensor 1240, a brake sensor (e.g., as part of a brake sensor system 1246), and / or other types of sensors.

[0077] In at least one embodiment, one or more of the controllers 1236 receive an input (e.g., represented by input data) from the instrument cluster 1232 of the vehicle 1200 and provide an output (e.g., represented by output data, display data, etc.) via the human-machine interface (「HMI」) display 1234, an audible annunciator, a speaker, and / or via other components of the vehicle 1200. In at least one embodiment, the output may include information such as vehicle speed, speed, time, map data (e.g., a high-definition map (not shown in FIG. 12A), location data (e.g., the location of the vehicle 1200 on a map, etc.), direction, the location of other vehicles (e.g., an occupancy grid), information about objects sensed by the controller 1236, and the state of the objects. For example, in at least one embodiment, the HMI display 1234 may display information about the presence of one or more objects (e.g., road signs, warning signs, signal changes, etc.) and / or information about driving operations that the vehicle has performed, is performing, or will perform (e.g., currently changing lanes, exiting at Exit 34B in 3.22 km (2 miles), etc.).

[0078] In at least one embodiment, vehicle 1200 further includes network interface 1224, which may use a wireless antenna 1226 and / or a modem for communicating over one or more networks. For example, in at least one embodiment, network interface 1224 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile communications ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000"), etc. Additionally, in at least one embodiment, wireless antenna 1226 may enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low power wide-area networks ("LPWAN") such as LoRaWAN, SigFox, etc.

[0079] Inference and / or training logic 915 is used to perform inference and / or training operations associated with one or more embodiments. More details regarding inference and / or training logic 915 are provided herein in conjunction with FIG. 9A and / or FIG. 9B . In at least one embodiment, inference and / or training logic 915 may be used in the system of FIG. 12A for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functionality and / or architecture, or neural network use cases described herein. In at least one embodiment, the neural network used by autonomous vehicle 1200 may operate using inference with one or more neural networks each trained by two or more processing cores for separately training portions of the neural network in parallel, as described herein.

[0080] 12B illustrates example camera locations and fields of view for autonomous vehicle 1200 of FIG. 12A, according to at least one embodiment. In at least one embodiment, the cameras and their respective fields of view are an example example and are not limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or cameras may be positioned at different locations on vehicle 1200.

[0081] In at least one embodiment, the camera type of the camera may include, but is not limited to, a digital camera that may be adapted to be used with the components and / or systems of the vehicle 1200. The camera may operate at automotive safety integrity level (ASIL) B and / or another ASIL. In at least one embodiment, the camera type may be capable of corresponding to any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the camera may be capable of using a rolling shutter, a global shutter, another type of shutter, or a combination thereof. In at least one embodiment, the color filter array may include a color filter array of red, clear, clear, clear (RCCC: red clear clear clear), a color filter array of red, clear, clear, blue (RCCB: red clear clear blue), a color filter array of red, blue, green, clear (RBGC: red blue green clear), a color filter array of Foveon X3, a color filter array of a Bayer sensor (RGGB), a color filter array of a monochrome sensor, and / or another type of color filter array. In at least one embodiment, a clear pixel camera, such as a camera having an RCCC, RCCB, and / or RBGC color filter array, may be used to increase light sensitivity.

[0082] In at least one embodiment, one or more of the cameras may be used to perform advanced driver assistance systems (ADAS) functions (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-functional mono-camera may be installed to provide functions including lane departure warning, traffic sign assistance, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all of the cameras) may simultaneously record and provide image data (e.g., video).

[0083] In at least one embodiment, one or more of the cameras may be attached to a mounting assembly such as a custom-designed (three-dimensional (3D) printed) assembly to eliminate stray light and reflections from inside the vehicle (e.g., reflections reflected from the dashboard to the windshield) that may interfere with the camera's image data capture performance. Referring to the door mirror mounting assembly, in at least one embodiment, the door mirror assembly may be custom 3D printed such that the camera mounting plate conforms to the shape of the door mirror. In at least one embodiment, the camera may be integrated with the door mirror. For a side view camera, in at least one embodiment, the camera may also be integrated into the four pillars at each corner of the cabin in this case.

[0084] In at least one embodiment, a camera (e.g., a front-facing camera) having a field of view that includes a portion of the environment ahead of vehicle 1200 may be used for a surroundings view to facilitate identification of the forward path and obstacles, and may be used in conjunction with controller 1236 and / or one or more of the control SoCs to assist in providing information essential for generating an occupancy grid and / or determining a preferred vehicle path. In at least one embodiment, the front-facing camera may be used to perform many of the same ADAS functions as LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance. In at least one embodiment, the front-facing camera may also be used for ADAS features and systems, including, without limitation, other features such as lane departure warnings ("LDW"), autonomous cruise control ("ACC"), and / or traffic sign recognition.

[0085] In at least one embodiment, various cameras may be used in a front-facing configuration, including, for example, a monocular camera platform including a CMOS (complementary metal oxide semiconductor) color imager. In at least one embodiment, a wide-angle camera 1270 may be used to sense objects (e.g., pedestrians, cross traffic, or bicyclists) coming into view from the periphery. While FIG. 12B shows only one wide-angle camera 1270, in other embodiments, there may be any number (including zero) of wide-angle cameras 1270 on the vehicle 1200. In at least one embodiment, any number of long-range cameras 1298 (e.g., a pair of long-view stereo cameras) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the long-range camera 1298 may also be used for object detection and classification, as well as basic object tracking.

[0086] In at least one embodiment, any number of stereo cameras 1268 may also be included in the front configuration. In at least one embodiment, one or more stereo cameras 1268 may include an integrated control unit with an expandable processing unit, which control unit may provide a programmable logic (FPGA) and a multi-core microprocessor having an integrated controller area network (CAN) or Ethernet® interface on a single chip. In at least one embodiment, such units may be used to generate a 3D map of the environment of the vehicle 1200, including distance estimation for all points within the image. In at least one embodiment, one or more of the stereo cameras 1268 may include, without limitation, a compact stereo vision sensor, which sensor may measure the distance from the vehicle 1200 to a target object and may include, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip that can use the generated information (e.g., metadata) to activate functions such as autonomous emergency braking and lane departure warning. In at least one embodiment, in addition to or instead of those described herein, other types of stereo cameras 1268 may be used.

[0087] In at least one embodiment, a camera (e.g., a side-view camera) having a field of view that includes a portion of the environment to the side of vehicle 1200 may be used for the surrounding view to provide information for occupancy grid creation and updating and for generating side collision warnings. For example, in at least one embodiment, surround cameras 1274 (e.g., four surround cameras 1274 as shown in FIG. 12B) may be disposed on vehicle 1200. The surround cameras 1274 may include, without limitation, any number and combination of wide-angle cameras 1270, fisheye cameras, and / or 360-degree cameras, etc. For example, in at least one embodiment, four fisheye cameras may be disposed in front of, behind, and to the sides of vehicle 1200. In at least one embodiment, vehicle 1200 may use three surround cameras 1274 (e.g., left, right, and rear), and as a fourth surround camera, one or more other cameras (e.g., a front camera) may be utilized.

[0088] In at least one embodiment, a camera (e.g., a rear-view camera) having a field of view that includes a portion of the environment behind vehicle 1200 may be used for parking assistance, the surrounding view, and rear collision warnings, and occupancy grid creation and updating may be performed. In at least one embodiment, a wide variety of cameras including, but not limited to, cameras suitable as the front cameras described herein (e.g., long-range camera 1298, and / or mid-range camera 1276, stereo camera 1268), infrared camera 1272, etc. may be used.

[0089] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of FIG. 12B for inference operations, at least in part based on weight parameters calculated using the training operations of the neural network, the functionality and / or architecture of the neural network, or the use cases of the neural network as described herein. In at least one embodiment, the neural network used by the autonomous vehicle 1200 may operate using inferences by one or more neural networks, each trained by two or more processing cores for separately training portions of the neural network in parallel as described herein.

[0090] FIG. 12C is a block diagram showing an exemplary system architecture of the autonomous vehicle 1200 of FIG. 12A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1200 of FIG. 12C is shown as being connected via a bus 1202. In at least one embodiment, the bus 1202 may include, without limitation, a CAN data interface (or referred to herein as the (CAN bus)). In at least one embodiment, CAN may be an in-vehicle network used to assist in controlling various features and functions of the vehicle 1200, such as brake actuation, acceleration, brake control, steering, windshield wipers, and the like. In at least one embodiment, the bus 1202 may be configured to have dozens or even hundreds of nodes, each having its own unique identifier (e.g., CAN ID). In at least one embodiment, the bus 1202 may be read to find the steering wheel angle, ground speed, engine revolutions per minute (RPM), button position, and / or other vehicle state indicators. In at least one embodiment, the bus 1202 may be a CAN bus compliant with ASIL B.

[0091] In at least one embodiment, in addition to or instead of CAN, FlexRay and / or Ethernet® may be used. In at least one embodiment, any number of buses 1202 may be present, which may include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet® buses, and / or zero or more other types of buses using other protocols. In at least one embodiment, two or more buses 1202 may be used to perform different functions and / or to provide redundancy. For example, a first bus 1202 may be used for a collision avoidance function and a second bus 1202 may be used for actuation control. In at least one embodiment, each bus 1202 may communicate with any of the components of vehicle 1200, and two or more buses 1202 may communicate with the same component. In at least one embodiment, each of any number of system-on-chips (“SoC”) 1204, each of controllers 1236, and / or each computer in the vehicle may be accessible to the same input data (e.g., input from sensors of vehicle 1200) and may be connected to a common bus such as a CAN bus.

[0092] In at least one embodiment, vehicle 1200 may include one or more controllers 1236, such as those described herein with respect to FIG. 12A. Controllers 1236 may be used for various functions. In at least one embodiment, controller 1236 may be coupled to any of the various other components and systems of vehicle 1200 and may be used for control of vehicle 1200, the artificial intelligence of vehicle 1200, and / or the infotainment of vehicle 1200.

[0093] In at least one embodiment, vehicle 1200 may include any number of SoCs 1204. Each of the SoCs 1204 may include, without limitation, a central processing unit (“CPU”) 1206, a graphics processing unit (“GPU”) 1208, a processor 1210, a cache 1212, an accelerator 1214, a data store 1216, and / or other components and features not shown. In at least one embodiment, the SoCs 1204 may be used to control the vehicle 1200 in various platforms and systems. For example, in at least one embodiment, the SoC 1204 may be incorporated into a system (such as the system of the vehicle 1200) having a high definition (“HD”) map 1222 that can obtain map refreshes and / or updates via a network interface 1224 from one or more servers (not shown in FIG. 12C).

[0094] In at least one embodiment, the CPU 1206 may include a CPU cluster, or a CPU complex (or referred to herein as “CCPLEX”). In at least one embodiment, the CPU 1206 may include a plurality of cores and / or a level 2 (“L2”) cache. For example, in at least one embodiment, the CPU 1206 may include eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU 1206 may include four dual-core clusters, where each cluster has a dedicated L2 cache (such as a 2MB L2 cache). In at least one embodiment, the CPU 1206 (such as CCPLEX) may be configured to support simultaneous cluster operation that allows any combination of the clusters of the CPU 1206 to be activated at any given time.

[0095] In at least one embodiment, one or more of the CPUs 1206 may implement a power management function, which may include, without limitation, one or more of the following features: individual hardware blocks can be automatically clock-gated during idle to save dynamic power; each core clock can be gated when the core is not actively executing instructions due to the execution of a wait for interrupt ("WFI") / wait for event ("WFE") instruction; each core can be independently power-gated; when all cores are clock-gated or power-gated, each core cluster can be independently clock-gated; and / or when all cores are power-gated, each core cluster can be independently power-gated. In at least one embodiment, the CPU 1206 may further implement an extended algorithm for managing power states, where the allowed power states and the expected wake-up times are specified, and the hardware / microcode determines the best power state for the cores, clusters, and CCPLEX to enter. In at least one embodiment, the processing core may support, in software, a simple sequence for entering a power state with the work offloaded to microcode.

[0096] In at least one embodiment, GPU 1208 may include an integrated GPU (or referred to herein as "iGPU"). In at least one embodiment, GPU 1208 may be programmable and may be efficient for parallel workloads. In at least one embodiment, GPU 1208 may use an extended tensor instruction set. In one embodiment, GPU 1208 may include one or more streaming microprocessors, where each streaming microprocessor may include a level 1 ("L1") cache (e.g., an L1 cache having a storage capacity of at least 96 KB), and two or more of the streaming microprocessors may share an L2 cache (e.g., an L2 cache having a storage capacity of 512 KB). In at least one embodiment, GPU 1208 may include at least eight streaming microprocessors. In at least one embodiment, GPU 1208 may use a compute application programming interface (API). In at least one embodiment, GPU 1208 may use one or more parallel computing platforms and / or programming modules (e.g., NVIDIA's CUDA).

[0097] In at least one embodiment, one or more of the GPUs 1208 may be power optimized for best performance in automotive and embedded use cases. For example, in one embodiment, the GPU 1208 can be fabricated on a fin field-effect transistor (“FinFET”). In at least one embodiment, each streaming microprocessor may incorporate a number of mixed-precision processing cores partitioned into multiple blocks. For example, without limitation, 64 PF32 cores and 32 PF64 cores can be partitioned into four processing blocks. In at least one embodiment, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA TENSOR cores for deep learning matrix operations, a level zero (“L0”) instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. In at least one embodiment, the streaming microprocessor includes independent parallel data paths for integers and floating points, and realizes efficient execution of the workload by mixing computer processing and addressing calculations. In at least one embodiment, the streaming microprocessor includes an independent thread scheduling function, which may enable finer-grained synchronization and cooperation between parallel threads. In at least one embodiment, the streaming microprocessor may include a combination of an L1 data cache and a shared memory unit to improve performance and simplify programming.

[0098] In at least one embodiment, one or more of the GPUs 1208 include high bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem, and in some examples, may provide a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, in addition to, or instead of, HBM memory, synchronous graphics random-access memory (SGRAM), such as five graphics double data rate type five (GDDR5) synchronous random access memories, may be used.

[0099] In at least one embodiment, the GPU 1208 may include integrated memory technology. In at least one embodiment, address translation services (ATS) support may be used to enable the GPU 1208 to directly access the page table of the CPU 1206. In at least one embodiment, when the GPU 1208 memory management unit (MMU) encounters a miss, an address translation request may be sent to the CPU 1206. In at least one embodiment, in response, the CPU 1206 may search its page table for a virtual-to-physical address mapping and send the translation back to the GPU 1208. In at least one embodiment, the integrated memory technology enables a single integrated virtual address space for the memories of both the CPU 1206 and the GPU 1208, thereby simplifying the programming of the GPU 1208 and the porting of applications to the GPU 1208.

[0100] In at least one embodiment, GPU 1208 may include any number of access counters that can record the frequency of access of GPU 1208 to the memory of other processors. In at least one embodiment, the access counter may assist in ensuring that memory pages are moved to the physical memory of the processor that most frequently accesses the pages, thereby improving the efficiency of the memory range shared among processors.

[0101] In at least one embodiment, one or more of SoC 1204 may include any number of caches 1212, including those described herein. For example, in at least one embodiment, cache 1212 may include a level 3 (“L3”) cache that is available to both CPU 1206 and GPU 1208 (e.g., connected to both CPU 1206 and GPU 1208). In at least one embodiment, cache 1212 may include a write-back cache that can record the state of lines, such as by using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, the L3 cache may include 4MB or more, depending on the embodiment, although smaller cache sizes may be used.

[0102] In at least one embodiment, one or more of the SoCs 1204 may include one or more accelerators 1214 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoCs 1204 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, the large on-chip memory (e.g., 4 MB of SRAM) may enable the hardware acceleration cluster to accelerate neural networks and other calculations. In at least one embodiment, the hardware acceleration cluster may be used to complement the GPU 1208 and offload some of the GPU 1208's tasks (e.g., freeing up more cycles for the GPU 1208 to perform other tasks). In at least one embodiment, accelerator 1214 may be used for targeted workloads that are stable enough to accommodate acceleration (e.g., perception, convolutional neural networks (“CNNs”), recurrent neural networks (“RNNs”), etc.). In at least one embodiment, CNNs may include region-based, i.e., regional convolutional neural networks (“RCNNs”), and Fast RCNNs (e.g., used for object detection), or other types of CNNs.

[0103] In at least one embodiment, the accelerator 1214 (e.g., a hardware acceleration cluster) may include a deep learning accelerator (DLA). The DLA may include, without limitation, one or more Tensor processing units (TPUs), which may be further configured to provide up to 10 trillion operations per second for deep learning applications and inferences. In at least one embodiment, the TPU may be an accelerator configured and optimized to execute image processing functions (e.g., CNN, RCNN, etc.). The DLA may be further optimized for a specific set of neural network types and floating point operations, as well as for inferences. In at least one embodiment, the design of the DLA can improve the performance per millimeter over a typical general-purpose GPU, typically far exceeding the performance of a CPU. In at least one embodiment, the TPU may execute several functions, including, for example, a single instance of a convolution function and a post-processing function that support INT8, INT16, and FP16 data types for both features and weights. In at least one embodiment, the DLA may execute a neural network, particularly a CNN, quickly and efficiently on processed or unprocessed data for any of a variety of functions, including, without limitation, a CNN for object identification and detection using data from a camera sensor, a CNN for distance estimation using data from a camera sensor, a CNN for emergency vehicle detection, identification, and detection using data from a microphone 1296, a CNN for face recognition and vehicle owner identification using data from a camera sensor, and / or a CNN for security and / or safety-related events.

[0104] In at least one embodiment, the DLA may execute any function of the GPU 1208. For example, by using an inference accelerator, the designer may target either the DLA or the GPU 1208 for any function. For example, in at least one embodiment, the designer may concentrate the processing of CNN and floating-point operations on the DLA and leave other functions to the GPU 1208 and / or other accelerators 1214.

[0105] In at least one embodiment, the accelerator 1214 (e.g., a hardware acceleration cluster) may include a programmable vision accelerator (referred to herein alternatively as a computer vision accelerator). In at least one embodiment, the PVA may be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS) 1238, autonomous driving, augmented reality (AR) applications, and / or virtual reality (VR) applications. The PVA may maintain a balance between performance and flexibility. For example, in at least one embodiment, each PVA may include any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors, without limitation.

[0106] In at least one embodiment, the RISC core may interact with, for example, an image sensor (e.g., the image sensor of any of the cameras described herein) and / or an image signal processor. In at least one embodiment, each of the RISC cores may include any amount of memory. In at least one embodiment, the RISC core may use any of a plurality of protocols depending on the embodiment. In at least one embodiment, the RISC core may execute a real-time operating system (“RTOS”). In at least one embodiment, the RISC core may be implemented using one or more integrated circuit devices, application specific integrated circuits (“ASICs”), and / or memory devices. For example, in at least one embodiment, the RISC core may include an instruction cache and / or tightly coupled RAM.

[0107] In at least one embodiment, the DMA may enable the components of the PVA to access the system memory independently of the CPU1206. In at least one embodiment, the DMA may support any number of features used to optimize the PVA, including but not limited to multidimensional addressing and / or circular addressing. In at least one embodiment, the DMA may support up to six or more addressing dimensions, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.

[0108] In at least one embodiment, the vector processor may be a programmable processor that may be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, the PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, the PVA core may include a processor subsystem, a DMA engine (e.g., two DMA engines), and / or other peripheral devices. In at least one embodiment, the vector processing subsystem may operate as the primary processing engine of the PVA and may include a vector processing unit ("VPU"), an instruction cache, and / or a vector memory (e.g., "VMEM"). In at least one embodiment, the VPU may include a digital signal processor, such as a single instruction, multiple data ("SIMD"), very long instruction word ("VLIW") digital signal processor. In at least one embodiment, the combination of SIMD and VLIW may improve throughput and speed.

[0109] In at least one embodiment, each of the vector processors may include an instruction cache and may be coupled to dedicated memory. As a result, in at least one embodiment, each of the vector processors may be configured to execute independently of other vector processors. In at least one embodiment, the vector processors included in a particular PVA may be configured to use data parallel processing. For example, in at least one embodiment, multiple vector processors included in a single PVA may execute the same computer vision algorithm on different regions of an image. In at least one embodiment, the vector processors included in a particular PVA may simultaneously execute different computer vision algorithms on the same image, or even execute different algorithms on consecutive images or on portions of an image. In at least one embodiment, in particular, any number of PVAs may be included in a hardware acceleration cluster, and any number of vector processors may be included in each of the PVAs. In at least one embodiment, the PVA may include additional error correction code ("ECC") memory to enhance the overall security of the system.

[0110] In at least one embodiment, the accelerator 1214 (e.g., a hardware acceleration cluster) may include an on-chip computer vision network and static random access memory (“SRAM”) to provide high-bandwidth, low-latency SRAM for the accelerator 1214. In at least one embodiment, the on-chip memory may include, for example, without limitation, at least 4 MB of SRAM consisting of eight field-configurable memory blocks, which may be accessible from both the PVA and the DLA. In at least one embodiment, each pair of memory blocks may include an advanced peripheral bus (“APB”) interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used. In at least one embodiment, the PVA and DLA may access the memory through a backbone that provides the PVA and DLA with high-speed access to the memory. In at least one embodiment, the backbone may include an on-chip computer vision network that interconnects the PVA and DLA to the memory (e.g., using the APB).

[0111] In at least one embodiment, the on-chip computer vision network may include an interface that determines whether both the PVA and DLA provide ready and enable signals before transmitting any control signals / addresses / data. In at least one embodiment, the interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-based communication for continuous data transfer. In at least one embodiment, the interface may conform to International Organization for Standardization ("ISO") 26262 or International Electrotechnical Commission ("IEC") 61508 standards, although other standards and protocols may be used.

[0112] In at least one embodiment, one or more of the SoCs 1204 may include a hardware accelerator for real-time ray tracing. In at least one embodiment, the hardware accelerator for real-time ray tracing is used to quickly and efficiently determine the position and extent of an object (e.g., within a world model) for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of a SONAR system, for simulation of general waveform propagation, for comparison with LIDAR data for localization and / or other functions, and / or for real-time visualization simulation for other uses.

[0113] In at least one embodiment, the accelerator 1214 (e.g., a hardware accelerator cluster) has various uses for autonomous driving. In at least one embodiment, the PVA may be a programmable vision accelerator that can be used in the main processing stages of ADAS and autonomous vehicles. In at least one embodiment, the performance of the PVA is well-suited to algorithm domains that require predictable processing with low power and low latency. In other words, the PVA functions well for semi-dense or dense regular calculations that require a predictable runtime with low latency and low power, even with a small data set. In at least one embodiment, in an autonomous vehicle such as the vehicle 1200, the PVA is designed to execute conventional computer vision algorithms because they are effective for object detection and integer arithmetic.

[0114] For example, according to at least one embodiment of the technology, computer stereo vision may be performed using PVA. In at least one embodiment, in some examples, an algorithm based on semi-global matching may be used, but this is not limiting. In at least one embodiment, applications for level 3-5 autonomous driving use motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.) on the fly. In at least one embodiment, PVA may perform computer stereo vision functions for inputs from two monocular cameras.

[0115] In at least one embodiment, high-density optical flow may be performed using PVA. For example, in at least one embodiment, PVA can process raw RADAR data (e.g., using 4D fast Fourier transform) to provide processed RADAR data. In at least one embodiment, PVA is used for time-of-flight depth processing, and for example, processed time-of-flight data is provided by processing raw time-of-flight data.

[0116] In at least one embodiment, without limitation, for example, the DLA may be used to execute any type of network for enhancing control and driving safety, including a neural network that outputs a measure of reliability for each object detection. In at least one embodiment, the reliability may be represented or interpreted as the probability of each detection compared to other detections, or as providing its relative "weight". In at least one embodiment, the reliability enables the system to make further decisions regarding which detections should be considered positive detections rather than false detections. For example, in at least one embodiment, the system may set a threshold for the reliability and consider only detections that exceed the threshold as positive detections. In embodiments where automatic emergency braking ("AEB") is used, false detections would cause the vehicle to automatically apply the emergency brake, which is clearly undesirable. In at least one embodiment, highly reliable detections may be considered as triggers for AEB. In at least one embodiment, the DLA may execute a neural network to regress the confidence value. In at least one embodiment, the neural network may take as its input at least some subset of parameters such as, among other things, the dimensions of the bounding box, the ground estimation obtained (e.g., from another subsystem), the output from the IMU sensor 1266 correlated with the orientation of the vehicle 1200, the distance, and the 3D location estimation of the object obtained from the neural network and / or other sensors (e.g., the LIDAR sensor 1264 or the RADAR sensor 1260).

[0117] In at least one embodiment, one or more of the SoCs 1204 may include a data store 1216 (e.g., memory). In at least one embodiment, the data store 1216 may be on-chip memory of the SoC 1204, and this memory may store neural networks executed on the GPU 1208 and / or DLA. In at least one embodiment, the capacity of the data store 1216 may be large enough to store multiple instances of neural networks for redundancy and security. In at least one embodiment, the data store 1212 may include an L2 or L3 cache.

[0118] In at least one embodiment, one or more of the SoCs 1204 may include any number of processors 1210 (e.g., embedded processors). The processors 1210 may include a boot and power management processor, which may be a dedicated processor and subsystem for handling boot power and management functions and related security execution. In at least one embodiment, the boot and power management processor may be part of the boot sequence of the SoC 1204 and may provide runtime power management services. In at least one embodiment, the boot power and management processor may provide clock and voltage programming, assistance in transitioning the system to a low-power state, management of the thermal and temperature sensors of the SoC 1204, and / or management of the power state of the SoC 1204. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 1204 may use the ring oscillator to detect the temperature of the CPU 1206, GPU 1208, and / or accelerator 1214. In at least one embodiment, if it is determined that the temperature exceeds a threshold, the boot and power management processor may enter a temperature fault routine, put the SoC 1204 into a low-power state, and / or put the vehicle 1200 into a driver-safety stop mode (e.g., safely stop the vehicle 1200).

[0119] In at least one embodiment, the processor 1210 may further include a set of embedded processors that can serve as an audio processing engine. In at least one embodiment, the audio processing engine may be an audio subsystem that enables complete hardware support for multi-channel audio via multiple interfaces and a wide variety of flexible audio I / O interfaces. In at least one embodiment, the audio processing engine is a dedicated processor core having a digital signal processor with dedicated RAM.

[0120] In at least one embodiment, the processor 1210 may further include an always-on processor engine that can provide the hardware features necessary to support low-power sensor management and startup use cases. In at least one embodiment, the always-on processor engine may include, without limitation, a processor core, tightly coupled RAM, support peripherals (e.g., timers, and interrupt controllers), various I / O controller peripherals, and routing logic.

[0121] In at least one embodiment, the processor 1210 may further include a safety cluster engine, which may include, without limitation, a dedicated processor subsystem for handling safety management in automotive applications. In at least one embodiment, the safety cluster engine may include, without limitation, two or more processor cores, tightly coupled RAM, support peripherals (such as timers, and interrupt controllers, etc.), and / or routing logic. In the safe mode, in at least one embodiment, two or more cores may operate in a lockstep mode and function as a single core having comparison logic for detecting any differences between these operations. In at least one embodiment, the processor 1210 may further include a real-time camera engine, which may include, without limitation, a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, the processor 1210 may further include a high dynamic range signal processor, which may include, without limitation, an image signal processor that is a hardware engine that is part of the camera processing pipeline.

[0122] In at least one embodiment, the processor 1210 may include a video image synthesizer, which may be a processing block (e.g., implemented on a microprocessor) that implements the video post-processing functions required by the video playback application to generate the final image in the window of the playback device. In at least one embodiment, the video image synthesizer may perform lens distortion correction on the wide-angle camera 1270, the surround camera 1274, and / or the in-cabin monitoring camera / sensor. In at least one embodiment, the in-cabin monitoring camera / sensor is preferably monitored by a neural network running on another instance of the SoC 1204, which is configured to identify in-cabin events and respond thereto as appropriate. In at least one embodiment, the in-cabin system may perform lip reading, without limitation, to activate cellular service, make a call, write an email, change the destination of the vehicle, activate or change the vehicle's infotainment system and settings, or provide voice-activated web surfing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in autonomous mode and unavailable otherwise.

[0123] In at least one embodiment, the video image synthesizer may include extended temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment, when there is motion in the video, the noise reduction appropriately weights the spatial information to attenuate the weight of the information provided by adjacent frames. In at least one embodiment, when an image or a portion of an image does not contain motion, the temporal noise reduction performed by the video image synthesizer may use information from the previous image to reduce the noise in the current image.

[0124] In at least one embodiment, the video image combiner may also be configured to perform stereo rectification on the input stereo lens frames. In at least one embodiment, the video image combiner may also be used to combine the user interface when the operating system desktop is in use, eliminating the need for the GPU 1208 to continually render new surfaces. In at least one embodiment, the video image combiner may be used to offload the GPU 1208 when it is powered on and actively performing 3D rendering, improving performance and responsiveness.

[0125] In at least one embodiment, one or more of the SoCs 1204 may further include a mobile industry processor interface ("MIPI") camera serial interface for receiving input from video and cameras, a high-speed interface, and / or a video input block that may be used for camera and associated pixel input functions. In at least one embodiment, one or more of the SoCs 1204 may further include an input / output controller, which may be controlled by software and may be used to receive I / O signals that are not tied to a specific role.

[0126] In at least one embodiment, one or more of the SoCs 1204 may further include a wide range of peripheral device interfaces to enable communication with peripheral devices, audio encoders / decoders (“codecs”), power management, and / or other devices. The SoC 1204 may be used to process data from cameras (connected, for example, via Gigabit Multimedia Serial Link and Ethernet®), sensors (such as LIDAR sensor 1264, RADAR sensor 1260, etc., which may be connected via Ethernet®), data from bus 1202 (such as the speed, steering wheel position, etc., of vehicle 1200), data from GNSS sensor 1258 (connected, for example, via Ethernet® or CAN bus), and the like. In at least one embodiment, one or more of the SoCs 1204 may further include a dedicated high-performance large-capacity storage controller, which may include its own DMA engine and may be used to free the CPU 1206 from routine data management tasks.

[0127] In at least one embodiment, the SoC 1204 may be an end-to-end platform with a flexible architecture spanning automation levels 3 - 5, thereby providing a comprehensive functional safety architecture that leverages computer vision and ADAS techniques for diversity and redundancy and uses them efficiently, and providing a flexible and reliable driving software stack along with deep learning tools. In at least one embodiment, the SoC 1204 is faster, more reliable, and more energy- and space-efficient than conventional systems. For example, in at least one embodiment, when accelerator 1214 is combined with CPU 1206, GPU 1208, and data store 1216, a fast and efficient platform for level 3 - 5 autonomous vehicles can be realized.

[0128] In at least one embodiment, the computer vision algorithm may be executed on a CPU, and this algorithm may be configured using a high-level programming language such as the C programming language to execute various processing algorithms over various visual data. However, in at least one embodiment, the CPU often cannot meet the performance requirements of many computer vision applications, such as requirements regarding execution time and power consumption. In at least one embodiment, many CPUs cannot execute in real time the complex object detection algorithms used in ADAS applications within a vehicle and in realistic level 3-5 autonomous vehicles.

[0129] The embodiments described herein can enable multiple neural networks to be executed simultaneously and / or sequentially, and combine the results to enable level 3-5 autonomous driving functions. For example, in at least one embodiment, the CNN executed on a DLA or an individual GPU (e.g., GPU 1220) may include text and word recognition, and enable a supercomputer to read and understand traffic signs including signs that the neural network has not been specifically trained for. In at least one embodiment, the DLA may further include a neural network that can identify and interpret signs and provide a semantic understanding of the signs, and can pass the semantic understanding to a path planning module executed on a CPU complex.

[0130] In at least one embodiment, for level 3, 4, or 5 operation, multiple neural networks may be executed simultaneously. For example, in at least one embodiment, a warning sign that reads "Caution: Frozen when flashing" in conjunction with the electro-optical may be interpreted separately or collectively by several neural networks. In at least one embodiment, the sign itself may be identified as a traffic sign by a first pre-introduced neural network (e.g., a trained neural network), the text "Frozen when flashing" may be interpreted by a second pre-introduced neural network, and when a flashing light is detected, this neural network notifies the vehicle's (preferably running on the CPU complex) route planning software that a frozen state exists. In at least one embodiment, the flashing light may be identified by operating a third pre-introduced neural network over multiple frames, and the presence (or absence) of the flashing light is notified to the vehicle's route planning software. In at least one embodiment, all three neural networks may be executed simultaneously, such as within the DLA and / or on the GPU 1208.

[0131] In at least one embodiment, a CNN for face recognition and vehicle owner identification may use data from a camera sensor to identify the presence of an approved driver and / or owner of the vehicle 1200. In at least one embodiment, an always-on sensor processing engine may be used to unlock the vehicle and turn on the lights when the owner approaches the driver's door, and make the vehicle inoperable in a security mode when the owner leaves the vehicle. Thus, the SoC 1204 realizes security against theft and / or carjacking.

[0132] In at least one embodiment, the CNN for detecting and identifying emergency vehicles may detect and identify the sirens of emergency vehicles using data from microphone 1296. In at least one embodiment, SoC 1204 classifies environmental and urban sounds and uses a CNN to classify visual data. In at least one embodiment, the CNN executed on the DLA is trained to identify the relative speed at which an emergency vehicle is approaching (e.g., by using the Doppler effect). In at least one embodiment, the CNN may also be trained to identify emergency vehicles specific to the area where the vehicle is operating, as identified by GNSS sensor 1258. In at least one embodiment, when operating in Europe, the CNN attempts to detect European sirens, and when in the United States, attempts to identify only North American sirens. In at least one embodiment, when an emergency vehicle is detected, a control program for executing an emergency vehicle safety routine is used to reduce the speed of the vehicle, move it to the side of the road, stop the vehicle, and / or use ultrasonic sensor 1262 to idle the vehicle until the emergency vehicle has passed.

[0133] In at least one embodiment, vehicle 1200 may include a CPU 1218 (e.g., a discrete CPU or dCPU), which may be coupled to SoC 1204 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU 1218 may include, for example, an X86 processor. CPU 1218 may be used to perform any of a variety of functions, including, for example, arbitrating potentially inconsistent results between the ADAS sensors and SoC 1204 and / or monitoring the state and health of controller 1236 and / or the in-vehicle infotainment system (the "IVI SoC") 1230 on the chip.

[0134] In at least one embodiment, vehicle 1200 may include a GPU 1220 (e.g., a discrete GPU or dGPU), which may be coupled to the SoC 1204 via a high-speed interconnect (e.g., NVIDIA's NVLINK). In at least one embodiment, the GPU 1220 may provide additional artificial intelligence capabilities, such as by executing redundant and / or different neural networks, and may be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from the sensors of the vehicle 1200.

[0135] In at least one embodiment, vehicle 1200 may further include a network interface 1224, which may include, without limitation, a wireless antenna 1226 (e.g., one or more wireless antennas 1226 for different communication protocols such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, the network interface 1224 may be used to enable a wireless connection via the Internet with a cloud (e.g., a server and / or other network devices), other vehicles, and / or computing devices (e.g., a passenger's client device). In at least one embodiment, a direct link may be established between vehicle 120 and other vehicles for communicating with the other vehicles, and / or an indirect link may be established (e.g., across a network and via the Internet). In at least one embodiment, the direct link may be provided using a vehicle-to-vehicle communication link. The vehicle-to-vehicle communication link may provide vehicle 1200 with information about vehicles in the vicinity of vehicle 1200 (e.g., vehicles in front of, to the side of, and / or behind vehicle 1200). In at least one embodiment, the functions described above may be part of a cooperative adaptive cruise control function of vehicle 1200.

[0136] In at least one embodiment, network interface 1224 may include a system-on-chip (SoC) that provides modulation and demodulation functions to enable the controller 1236 to communicate via a wireless network. In at least one embodiment, network interface 1224 may include a radio frequency (RF) front end for up-conversion from baseband to RF and down-conversion from RF to baseband. In at least one embodiment, the frequency conversion may be performed in any technically feasible manner. For example, the frequency conversion can be performed by well-known processes and / or using a superheterodyne process. In at least one embodiment, the RF front end functions may be provided by a separate chip. In at least one embodiment, the network interface may include wireless capabilities for communicating via LTE, WCDMA®, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0137] In at least one embodiment, vehicle 1200 may further include a data store 1228, which may include off-chip (e.g., not on SoC 1204) storage without limitation. In at least one embodiment, data store 1228 may include one or more storage elements including, without limitation, RAM, SRAM, dynamic random access memory (“DRAM”), video random-access memory (“VRAM”), flash, hard disk, and / or other components and / or devices capable of storing at least one bit of data.

[0138] In at least one embodiment, vehicle 1200 may further include a GNSS sensor 1258 (e.g., a GPS and / or assisted GPS sensor) to assist with mapping, perception, occupancy grid generation, and / or path planning functions. In at least one embodiment, any number of GNSS sensors 1258 may be used, including for example without limitation a GPS that uses a USB connector having a bridge from Ethernet (registered trademark) to serial (e.g., RS-232).

[0139] In at least one embodiment, vehicle 1200 may further include a RADAR sensor 1260. The RADAR sensor 1260 may be used by the vehicle 1200 to perform long-range vehicle detection, even in darkness and / or harsh weather conditions. In at least one embodiment, the functional safety level of the RADAR may be ASIL B. The RADAR sensor 1260 may use the CAN and / or bus 1202 for control (e.g., to send data generated by the RADAR sensor 1260) and to access object tracking data, and in some examples, may access Ethernet (registered trademark) to access raw data. In at least one embodiment, various types of RADAR sensors may be used. For example without limitation, the RADAR sensor 1260 may be suitable for use in front, rear, and side RADAR. In at least one embodiment, one or more of the RADAR sensors 1260 are pulse Doppler RADAR sensors.

[0140] In at least one embodiment, the RADAR sensor 1260 may include different configurations, such as a narrow field of view for long distances, a wide field of view for short distances, and short distances covering the sides. In at least one embodiment, the long-range RADAR may be used for an adaptive cruise control function. In at least one embodiment, the long-range RADAR system may provide a wide field of view, such as within a range of 250 m realized by two or more independent scans. In at least one embodiment, the RADAR sensor 1260 may be made easier to distinguish between static and moving objects and may be used by the ADAS system 1238 for emergency braking assistance and forward collision warning. The sensor 1260 included in the long-range RADAR system may include, without limitation, a plurality of (e.g., six or more) fixed RADAR antennas and a monostatic multimode RADAR having high-speed CAN and FlexRay interfaces. In at least one embodiment, when there are six antennas, the four central antennas may generate a concentrated beam pattern designed to record the surroundings of the vehicle 1200 at a higher speed with minimal interference from adjacent lanes. In at least one embodiment, the other two antennas may expand the field of view and enable rapid detection of vehicles entering or leaving the lane of the vehicle 1200.

[0141] In at least one embodiment, the mid-range RADAR system may include, by way of example, a range of up to 160 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, the short-range RADAR system may include any number of RADAR sensors 1260 designed to be installed at both ends of the rear bumper, without limitation. When installed at both ends of the rear bumper, in at least one embodiment, the RADAR sensor system may generate two beams that constantly monitor the rear and the blind spots adjacent to the vehicle. In at least one embodiment, the short-range RADAR system may be used in the ADAS system 1238 for blind spot detection and / or lane change assistance.

[0142] In at least one embodiment, vehicle 1200 may further include an ultrasonic sensor 1262. The ultrasonic sensor 1262 may be disposed in front of, behind, and / or on the side of the vehicle 1200 and may be used for parking assistance and / or for generating and updating an occupancy grid. In at least one embodiment, various ultrasonic sensors 1262 may be used, and different ultrasonic sensors 1262 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensor 1262 may operate at a functional safety level of ASIL B.

[0143] In at least one embodiment, vehicle 1200 may include a LIDAR sensor 1264. The LIDAR sensor 1264 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensor 1264 may be at a functional safety level of ASIL B. In at least one embodiment, vehicle 1200 may include a plurality of LIDAR sensors 1264 (e.g., two, four, six, etc.), and these sensors may use Ethernet (registered trademark) (e.g., to provide data to a gigabit Ethernet (registered trademark) switch).

[0144] In at least one embodiment, the LIDAR sensor 1264 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, a commercially available LIDAR sensor 1264 may, for example, have a claimed range of approximately 100 m, an accuracy of 2 cm to 3 cm, and support a 100 Mbps Ethernet® connection. In at least one embodiment, one or more non-protruding LIDAR sensors 1264 may be used. In such embodiments, the LIDAR sensor 1264 may be implemented as a small device that can be incorporated in front of, behind, on the sides, and / or at the corners of the vehicle 1200. In at least one embodiment, the LIDAR sensor 1264 of such embodiments may provide a horizontal field of view of up to 120 degrees and a vertical field of view of 35 degrees, even for low-reflectivity objects, in a range of 200 m. In at least one embodiment, the LIDAR sensor 1264 mounted in the front may be configured to provide a horizontal field of view of 45 degrees to 135 degrees.

[0145] In at least one embodiment, LIDAR technologies such as 3D flash LIDAR may also be used. The 3D flash LIDAR uses a laser flash as a transmission source to irradiate the surroundings of the vehicle 1200 up to approximately 200 m at most. In at least one embodiment, the flash LIDAR unit includes, without limitation, a receptor that records the time-of-flight of the laser pulse and the reflected light at each pixel, which corresponds to the range from the vehicle 1200 to the object. In at least one embodiment, the flash LIDAR enables a very accurate and distortion-free surrounding image to be generated for each laser flash. In at least one embodiment, four flash LIDARs may be introduced, one on each side of the vehicle 1200. In at least one embodiment, the 3D flash LIDAR system includes, without limitation, a LIDAR camera of a semiconductor 3D staring array (for example, a non-scanning LIDAR device) without moving parts other than a fan. In at least one embodiment, the flash LIDAR device may use a Class I (eye-safe) laser pulse of 5 nanoseconds per frame and capture the reflected laser light in the form of a 3D range point cloud and co-registered intensity data.

[0146] In at least one embodiment, the vehicle may further include an IMU sensor 1266. In at least one embodiment, the IMU sensor 1266 may be positioned at the center of the rear axle of the vehicle 1200 in at least one embodiment. In at least one embodiment, the IMU sensor 1266 may include, for example without limitation, an accelerometer, a magnetometer, a gyroscope, a magnetic compass, and / or other types of sensors. In at least one embodiment, for a 6-axis application, the IMU sensor 1266 may include, without limitation, an accelerometer and a gyroscope. In at least one embodiment, for a 9-axis application, the IMU sensor 1266 may include, without limitation, an accelerometer, a gyroscope, and a magnetometer.

[0147] In at least one embodiment, the IMU sensor 1266 may be implemented as a small, high-performance GPS-aided inertial navigation system (GPS / INS) that combines a micro-electro-mechanical systems (MEMS) inertial sensor, a high-sensitivity GPS receiver, and an advanced Kalman filtering algorithm to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor 1266 enables the vehicle 1200 to estimate its orientation without the need for input from a magnetic sensor by directly observing speed changes and correlating them from GPS to the IMU sensor 1266. In at least one embodiment, the IMU sensor 1266 and the GNSS sensor 1258 may be combined in a single integrated unit.

[0148] In at least one embodiment, the vehicle 1200 may include a microphone 1296 installed inside and / or around the vehicle 1200. In at least one embodiment, the microphone 1296 may be used, among other things, for the detection and identification of emergency vehicles.

[0149] In at least one embodiment, vehicle 1200 may further include any number of camera types, such as stereo camera 1268, wide-angle camera 1270, infrared camera 1272, surround camera 1274, long-range camera 1298, medium-range camera 1276, and / or other camera types. In at least one embodiment, the cameras may be used to capture image data around the entire perimeter of vehicle 1200. In at least one embodiment, the type of cameras used may vary depending on vehicle 1200. In at least one embodiment, any combination of camera types may be used to provide the necessary field of view around vehicle 1200. In at least one embodiment, the number of cameras may vary depending on the embodiment. For example, in at least one embodiment, vehicle 1200 may include six cameras, seven cameras, ten cameras, twelve cameras, or another number of cameras. The cameras may support, by way of non-limiting example, Gigabit Multimedia Serial Link (“GMSL”) and / or Gigabit Ethernet®. In at least one embodiment, each camera is described in more detail above herein with respect to FIGS. 12A and 12B.

[0150] In at least one embodiment, vehicle 1200 may further include vibration sensor 1242. Vibration sensor 1242 may measure the vibration of components of vehicle 1200, such as an axle. For example, in at least one embodiment, a change in vibration may indicate a change in the road surface. In at least one embodiment, if two or more vibration sensors 1242 are used, the difference in vibration may be used to determine the amount of friction or slip on the road surface (e.g., if there is a vibration difference between a powered axle and a freely rotating axle).

[0151] In at least one embodiment, vehicle 1200 may include an ADAS system 1238. The ADAS system 1238 may include, without limitation, a SoC in some examples. In at least one embodiment, the ADAS system 1238 may include, without limitation, any number and any combination of autonomous / adaptive / automatic cruise control (“ACC”) systems, cooperative adaptive cruise control (“CACC”) systems, forward crash warning (“FCW”) systems, automatic emergency braking (“AEB”) systems, lane departure warning (“LDW”) systems, lane keep assist (“LKA”) systems, blind spot warning (“BSW”) systems, rear cross-traffic warning (“RCTW”) systems, collision warning (“CW”) systems, lane centering (“LC”) systems, and / or other systems, features, and / or functions.

[0152] In at least one embodiment, the ACC system may use a RADAR sensor 1260, a LIDAR sensor 1264, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, the longitudinal ACC system monitors and controls the distance to the vehicle immediately in front of vehicle 1200 and automatically adjusts the speed of vehicle 1200 to maintain a safe distance from the vehicle ahead. In at least one embodiment, the lateral ACC system performs distance maintenance and notifies vehicle 1200 to change lanes when necessary. In at least one embodiment, the lateral ACC is related to other ADAS applications such as LC and CW.

[0153] In at least one embodiment, the CACC system uses information from other vehicles, which may be received from other vehicles by a wireless link or indirectly via a network connection (e.g., via the Internet) to the network interface 1224 and / or the wireless antenna 1226. In at least one embodiment, a direct link may be provided by a vehicle-to-vehicle ("V2V") communication link, while an indirect link may be provided by an infrastructure-to-vehicle ("I2V") communication link. Generally, the concept of V2V communication provides information about the immediately preceding vehicle (e.g., a vehicle in the same lane immediately in front of vehicle 1200), and the concept of I2V communication provides information about the traffic even further ahead. In at least one embodiment, the CACC system may include either or both of the I2V and V2V information sources. In at least one embodiment, having information about the vehicle in front of vehicle 1200 can further enhance the reliability of the CACC system, make the traffic flow smoother, and have the potential to reduce congestion on the road.

[0154] In at least one embodiment, the FCW system is designed to advise the driver of a hazard, whereby the driver can take corrective measures. In at least one embodiment, the FCW system uses a front camera and / or the RADAR sensor 1260, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that are electrically coupled to driver feedback such as a display, speaker, and / or a vibrating component. In at least one embodiment, the FCW system may provide warnings in the form of sound, visual warnings, vibrations, and / or quick brake pulses.

[0155] In at least one embodiment, the AEB system may detect an impending frontal collision with another vehicle or other object and automatically apply the brakes if the driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may use a front camera and / or RADAR sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when the AEB system detects a hazard, the AEB system typically first advises the driver to take corrective action to avoid the collision, and if the driver does not take corrective action, the AEB system may automatically apply the brakes to prevent the predicted collision or at least mitigate its impact. In at least one embodiment, the AEB system may include techniques such as dynamic brake support and / or pre-crash braking.

[0156] In at least one embodiment, the LDW system provides visual, auditory, and / or tactile warnings, such as steering wheel or seat vibrations, to advise the driver when vehicle 1200 crosses a lane marker. In at least one embodiment, the LDW system does not activate if the driver indicates an intentional lane departure by activating the turn indicator. In at least one embodiment, the LDW system may use a front camera, which is coupled to a dedicated processor, DSP, FPGA, and / or ASIC that can be electrically coupled to driver feedback such as a display, speaker, and / or vibration component. In at least one embodiment, the LKA system is a variant of the LDW system. The LKA system provides steering input or brake control to correct vehicle 1200 when vehicle 1200 begins to drift out of a lane.

[0157] In at least one embodiment, the BSW system detects a vehicle in a blind spot of a motor vehicle and warns the driver. In at least one embodiment, the BSW system may provide visual, audible, and / or tactile alerts to indicate that a merge or lane change is not safe. In at least one embodiment, the BSW system may provide additional warnings when the driver uses a turn indicator. In at least one embodiment, the BSW system may use a rear camera and / or a RADAR sensor 1260 coupled to a dedicated processor, DSP, FPGA, and / or ASIC, and these dedicated processors, DSPs, FPGAs, and / or ASICs are electrically coupled to feedback to the driver such as a display, a speaker, and / or a vibration component.

[0158] In at least one embodiment, the RCTW system may provide visual, audible, and / or tactile notifications when an object is detected outside the range of a rear camera when the vehicle 1200 is reversing. In at least one embodiment, the RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a collision. In at least one embodiment, the RCTW system may use one or more rear RADAR sensors 1260, which are coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to feedback to the driver such as a display, a speaker, and / or a vibration component.

[0159] In at least one embodiment, conventional ADAS systems may sometimes produce false detection results, which can be annoying and distracting to the driver, but usually are not a big deal. This is because conventional ADAS systems are designed to allow the driver to determine whether there is truly a safety-critical situation and how to respond appropriately. In at least one embodiment, when the results conflict, the vehicle 1200 itself determines whether to follow the results from the primary computer (e.g., the first controller 1236) or the results from the secondary computer (e.g., the second controller 1236). For example, in at least one embodiment, the ADAS system 1238 may be a backup and / or secondary computer for resisting perception information to the rationality module of the backup computer. In at least one embodiment, the rationality monitor of the backup computer may execute various software for redundancy on the hardware components to detect perception errors and dynamic driving tasks. In at least one embodiment, the output from the ADAS system 1238 may be provided to the monitoring MCU. In at least one embodiment, when the output from the primary computer conflicts with the output from the secondary computer, the monitoring MCU determines how to reconcile the conflict to ensure safe operation.

[0160] In at least one embodiment, the primary computer may be configured to provide a reliability score indicating the reliability of the selected result of the primary computer to the monitoring MCU. In at least one embodiment, when the reliability score exceeds a threshold, the monitoring MCU may follow the instructions of the primary computer regardless of whether the secondary computer provides conflicting or inconsistent results. In at least one embodiment, when the reliability score does not meet the threshold and the primary computer and the secondary computer show different results (e.g., conflict), the monitoring MCU may mediate between the computers to determine an appropriate result.

[0161] In at least one embodiment, a neural network trained and configured to determine, at least in part based on outputs from a primary computer and a secondary computer, conditions under which the secondary computer provides a false alarm may be configured to be executed by a monitoring MCU. In at least one embodiment, the neural network of the monitoring MCU may learn when the output of the secondary computer may be trusted and when it may not be trusted. For example, in at least one embodiment, when the secondary computer is a RADAR-based FCW system, the neural network of the monitoring MCU may learn when the FCW system identifies a metallic object, such as a drain grate or manhole cover, that is not actually a hazard, which triggers an alarm. In at least one embodiment, when the secondary computer is a camera-based LDW system, the neural network of the monitoring MCU may learn to disable the LDW when bicycles or pedestrians are present and lane departure is actually the safest operation. In at least one embodiment, the monitoring MCU may include at least one of a DLA or a GPU suitable for executing the neural network along with an associated memory. In at least one embodiment, the monitoring MCU may comprise and / or be included as a component of the SoC1204.

[0162] In at least one embodiment, the ADAS system 1238 may include a secondary computer that executes ADAS functions using conventional rules of computer vision. In at least one embodiment, the secondary computer may use conventional computer vision rules (if-then rules), and the presence of the neural network in the monitoring MCU may improve reliability, safety, and performance. For example, in at least one embodiment, due to various implementations and intentional non-identities, the overall error tolerance of the system is increased, particularly with respect to errors caused by the functions of software (or the software-hardware interface). For example, in at least one embodiment, if there is a bug or error in the software running on the primary computer and the non-identical software code running on the secondary computer provides the same overall result, the monitoring MCU may have a higher level of confidence that the overall result is correct and that the bug in the software or hardware on the primary computer is not causing a critical error.

[0163] In at least one embodiment, the output of the ADAS system 1238 may be supplied to the perception block of the primary computer and / or the dynamic driving task block of the primary computer. For example, in at least one embodiment, if the ADAS system 1238 indicates a forward collision warning due to an object immediately ahead, the perception block may use this information when identifying the object. In at least one embodiment, the secondary computer may have a trained, and thus unique, neural network that reduces the risk of false detection, as described herein.

[0164] In at least one embodiment, vehicle 1200 may further include an infotainment SoC 1230 (e.g., an in-vehicle infotainment system (IVI)). Although the infotainment system 1230 is illustrated and described as an SoC, in at least one embodiment, it may not be an SoC and may include two or more individual components without limitation. In at least one embodiment, the infotainment SoC 1230 may include, without limitation, a combination of hardware and software, and this combination may be used to provide the vehicle 1200 with audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephone (e.g., hands-free calling), network connection (e.g., LTE, Wi-Fi, etc.), and / or information services (e.g., navigation system, rear parking assistance, wireless data system, vehicle-related information such as fuel level, total mileage, brake fuel level, oil level, door opening and closing, air filter information, etc.). For example, the infotainment SoC 1230 may include a radio, a disc player, a navigation system, a video player, USB and Bluetooth connections, a carputer, in-vehicle entertainment, Wi-Fi, steering wheel audio control, hands-free voice control, a heads-up display ("HUD"), an HMI display 1234, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, features, and / or systems), and / or other components. In at least one embodiment, information from the ADAS system 1238, autonomous driving information such as vehicle operation plans and trajectories, ambient environment information (e.g., intersection information, vehicle information, road information, etc.), and / or other information (e.g., visual and / or auditory) may further be provided to the user of the vehicle using the infotainment SoC 1230.

[0165] In at least one embodiment, the infotainment SoC 1230 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1230 may communicate with other devices, systems, and / or components of the vehicle 1200 via a bus 1202 (e.g., a CAN bus, Ethernet®, etc.). In at least one embodiment, the infotainment SoC 1230 may be coupled to a monitoring MCU, such that when the primary controller 1236 (e.g., the primary and / or backup computer of the vehicle 1200) fails, the GPU of the infotainment system may execute some self-driving functions. In at least one embodiment, the infotainment SoC 1230 may put the vehicle 1200 into a driver-safe stop mode as described herein.

[0166] In at least one embodiment, the vehicle 1200 may further include an instrument cluster 1232 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). The instrument cluster 1232 may include, without limitation, a controller and / or a supercomputer (e.g., an individual controller or supercomputer). In at least one embodiment, the instrument cluster 1232 may include any number and combination of instrument sets, without limitation, a speedometer, a fuel level, a hydraulic pressure, a tachometer, an odometer, a direction indicator, a shift lever position indicator, a seat belt warning light, a parking brake warning light, an engine failure light, an auxiliary restraint system (e.g., an airbag) information, a light control, a safety system control, a navigation information, etc. In some examples, information may be displayed and / or shared between the infotainment SoC 1230 and the instrument cluster 1232. In at least one embodiment, the instrument cluster 1232 may be included as part of the infotainment SoC 1230, or vice versa.

[0167] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of FIG. 12C for inference operations, at least in part based on the training operations of the neural network, the functionality and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network as described herein. In at least one embodiment, the neural network used by the autonomous vehicle 1200 may operate using inferences by one or more neural networks each trained by two or more processing cores for separately training portions of the neural network in parallel as described herein.

[0168] FIG. 12D is a diagram of a system 1276 for communicating between a cloud - based server and the autonomous vehicle 1200 of FIG. 12A according to at least one embodiment. In at least one embodiment, system 1276 may include, without limitation, server 1278, network 1290, and any number and type of vehicles including vehicle 1200. Server 1278 may include, without limitation, a plurality of GPUs 1284(A) - 1284(H) (collectively referred to herein as GPU 1284), PCIe switches 1282(A) - 1282(H) (collectively referred to herein as PCIe switch 1282), and / or CPUs 1280(A) - 1280(B) (collectively referred to herein as CPU 1280). The GPUs 1284, CPUs 1280, and PCIe switches 1282 may be interconnected by high - speed interconnects such as, without limitation, an NVLink interface 1288 developed by NVIDIA and / or a PCIe connection 1286. In at least one embodiment, the GPUs 1284 are connected to each other via an NVLink and / or an NVSwitch SoC, and the GPU 1284 and the PCIe switch 1282 are connected via a PCIe interconnect. In at least one embodiment, eight GPUs 1284, two CPUs 1280, and four PCIe switches 1282 are illustrated, but this is not limiting. In at least one embodiment, each of the servers 1278 may include, without limitation, any number of GPUs 1284, CPUs 1280, and / or PCIe switches 1282 in any combination. For example, in at least one embodiment, server 1278 may include eight, sixteen, thirty - two, and / or more GPUs 1284 each.

[0169] In at least one embodiment, server 1278 may receive, via network 1290, image data representing an image indicating an unexpected or changed road condition, such as recently started road construction, from a vehicle. In at least one embodiment, server 1278 may transmit, via network 1290, neural network 1292, updated neural network 1292, and / or map information 1294 including information regarding traffic conditions and road conditions, without limitation, to a vehicle. In at least one embodiment, the update of map information 1294 may include, without limitation, updates to HD map 1222, such as information regarding construction sites, holes, detours, floods, and / or other obstacles. In at least one embodiment, neural network 1292, updated neural network 1292, and / or map information 1294 may be obtained from new training and / or experience represented by data received from any number of vehicles in the environment and / or may be obtained based at least in part on training performed at a data center (e.g., using server 1278 and / or other servers).

[0170] In at least one embodiment, using server 1278, a machine learning model (e.g., a neural network) may be trained based at least in part on training data. The training data may be generated by a vehicle and / or may be generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged and / or otherwise pre-processed (e.g., if the associated neural network benefits from supervised learning). In at least one embodiment, any amount of training data is not tagged and / or pre-processed (e.g., if the associated neural network does not require supervised learning). In at least one embodiment, once the machine learning model is trained, the machine learning model may be used by a vehicle (e.g., transmitted to the vehicle via network 1290) and / or the machine learning model may be used by server 1278 to remotely monitor the vehicle.

[0171] In at least one embodiment, server 1278 may receive data from a vehicle and apply the data to a state-of-the-art real-time neural network to enable real-time intelligent inference. In at least one embodiment, server 1278 may include a deep learning supercomputer and / or a dedicated AI computer powered by GPU 1284, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, server 1278 may include a deep learning infrastructure that uses a data center powered by a CPU.

[0172] In at least one embodiment, the deep learning infrastructure of server 1278 may be capable of high-speed real-time inference and may use that capability to evaluate and verify the health of the processor, software, and / or associated hardware of vehicle 1200. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1200, such as a series of images and / or objects located by vehicle 1200 in that series of images (e.g., by computer vision and / or other machine learning object classification techniques). In at least one embodiment, the deep learning infrastructure may run its own neural network to identify an object and compare it to the object identified by vehicle 1200. If the results do not match and the deep learning infrastructure concludes that the AI of vehicle 1200 is malfunctioning, server 1278 may take control from the fail-safe computer of vehicle 1200, notify the occupants, and send a signal to vehicle 1200 instructing it to complete a safe parking operation.

[0173] In at least one embodiment, server 1278 may include GPU 1284 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT3). In at least one embodiment, by combining a server powered by a GPU with inference acceleration, real-time response can be enabled. In at least one embodiment, such as when performance is not as critical, a server powered by a CPU, FPGA, and other processors may be used for inference. In at least one embodiment, hardware structure 915 is used to execute one or more embodiments. Details regarding hardware structure 915 are provided herein in conjunction with FIGS. 9A and / or 9B.

[0174] Computer system FIG. 13 is a block diagram showing an exemplary computer system, which may be a system having interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof 1300 formed with a processor that may include an execution unit for executing instructions, according to at least one embodiment. In at least one embodiment, computer system 1300 may include, without limitation, components such as processor 1302 for using an execution unit that includes logic for executing an algorithm for processing data in accordance with the present disclosure, such as in the embodiments described herein. In at least one embodiment, computer system 1300 may include a processor such as a PENTIUM® processor family, Xeon™, Itanium®, XScale™, and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessor available from Intel Corporation of Santa Clara, California, although other systems (including PCs having other microprocessors, engineering workstations, set-top boxes, etc.) may be used. In at least one embodiment, computer system 1300 may execute a version of the WINDOWS® operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX® and Linux®), embedded software, and / or graphical user interfaces may be used.

[0175] Embodiments may be used in other devices such as portable devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants (PDAs), and portable PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (DSP), a system-on-chip, network computers (NetPCs), set-top boxes, network hubs, wide area network (WAN) switches, or any other system capable of executing one or more instructions according to at least one embodiment.

[0176] In at least one embodiment, computer system 1300 may include, without limitation, processor 1302, which may include, without limitation, one or more execution units 1308 for performing training and / or inference of a machine learning model according to the techniques described herein. In at least one embodiment, system 1300 is a single-processor desktop or server system, but in another embodiment, system 1300 may be a multi-processor system. In at least one embodiment, processor 1302 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other processor device, such as a digital signal processor. In at least one embodiment, processor 1302 may be coupled to processor bus 1310, which may transmit digital signals between processor 1302 and other components within computer system 1300.

[0177] In at least one embodiment, processor 1302 may include, without limitation, a level 1 (L1) internal cache memory (cache) 1304. In at least one embodiment, processor 1302 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may be external to processor 1302. Other embodiments may also include a combination of both internal and external caches, depending on the particular implementation and requirements. In at least one embodiment, register file 1306 may store different types of data in various registers including, without limitation, integer registers, floating point registers, status registers, and instruction pointer registers.

[0178] In at least one embodiment, execution unit 1308, which includes logic for performing integer and floating point operations without limitation, is also in processor 1302. Processor 1302 may also include a microcode (u-code) read only memory (ROM) that stores microcode for certain macro instructions. In at least one embodiment, execution unit 1308 may include logic for handling a packed instruction set 1309. In at least one embodiment, by including the packed instruction set 1309 in the instruction set of general purpose processor 1302 along with associated circuitry for executing instructions, operations used by many multimedia applications can be executed using the packed data of general purpose processor 1302. In one or more embodiments, by performing operations on packed data using the full width of the processor's data bus, many multimedia applications can be accelerated and executed more efficiently, thereby eliminating the need to transfer smaller units of data between the processor's data buses to perform one or more operations on one data element at a time.

[0179] In at least one embodiment, execution unit 1308 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1300 may include memory 1320 without limitation. In at least one embodiment, memory 1320 may be implemented as a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or other memory device. Memory 1320 may store instructions 1319 and / or data 1321 represented by data signals that may be executed by processor 1302.

[0180] In at least one embodiment, the system logic chip may be coupled to the processor bus 1310 and the memory 1320. In at least one embodiment, the system logic chip may include, without limitation, a memory controller hub ("MCH") 1316, and the processor 1302 may communicate with the MCH 1316 via the processor bus 1310. In at least one embodiment, the MCH 1316 may provide a high-bandwidth memory path 1318 to the memory 1320 for storing instructions and data and for storing graphics commands, data, and textures. In at least one embodiment, the MCH 1316 may direct data signals between the processor 1302, the memory 1320, and other components of the computer system 1300 and may bridge data signals between the processor bus 1310, the memory 1320, and the system I / O 1322. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, the MCH 1316 may be coupled to the memory 1320 via the high-bandwidth memory path 1318, and the graphics / video card 1312 may be coupled to the MCH 1316 via an Accelerated Graphics Port ("AGP") interconnect 1314.

[0181] In at least one embodiment, computer system 1300 may use a system I / O 1322, which is a proprietary hub interface bus for coupling MCH 1316 to an I / O controller hub ("ICH") 1330. In at least one embodiment, ICH 1330 may provide direct connections to several I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripheral devices to memory 1320, a chipset, and processor 1302. By way of example, an audio controller 1329, a firmware hub ("flash BIOS") 1328, a wireless transceiver 1326, a data storage 1324, a legacy I / O controller 1323 including a user input and keyboard interface, a serial expansion port such as a universal serial bus ("USB"), and a network controller 1334 may be included, without limitation. Data storage 1324 may comprise a hard disk drive, a floppy (registered trademark) disk drive, a CD-ROM device, a flash memory device, or other mass storage device.

[0182] In at least one embodiment, FIG. 13 shows a system including interconnected hardware devices or "chips", while in other embodiments, FIG. 13 may show an exemplary system-on-chip ("SoC"). In at least one embodiment, the devices shown in FIG. 13 may be interconnected by a proprietary interconnect, a standard interconnect (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of computer system 1300 may be interconnected using a compute express link (CXL) interconnect.

[0183] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of FIG. 13 for inference or prediction operations, at least in part based on weight parameters calculated using the training operations of the neural network, the functionality and / or architecture of the neural network, or the use cases of the neural network as described herein. In at least one embodiment, the neural network used by the electronic device 1200 may operate using inferences by one or more neural networks each trained by two or more processing cores for separately training portions of the neural network in parallel as described herein.

[0184] FIG. 14 is a block diagram showing an electronic device 1400 that utilizes a processor 1410 according to at least one embodiment. In at least one embodiment, the electronic device 1400 may be, for example, without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.

[0185] In at least one embodiment, system 1400 may include, without limitation, a processor 1410 communicatively coupled to any suitable number or type of components, peripherals, modules, or devices. In at least one embodiment, processor 1410 is coupled using a bus or interface such as an I°C bus, a System Management Bus (“SMBus”), a Low Pin Count (“LPC”) bus, a Serial Peripheral Interface (“SPI”), a High Definition Audio (“HDA”) bus, a Serial Advance Technology Attachment (“SATA”) bus, a Universal Serial Bus (“USB”) (versions 1, 2, 3), or a Universal Asynchronous Receiver / Transmitter (“UART”) bus. In at least one embodiment, FIG. 14 shows a system including interconnected hardware devices or “chips,” while in other embodiments, FIG. 14 may show an exemplary System on Chip (“SoC”). In at least one embodiment, the devices shown in FIG. 14 may be interconnected using proprietary interconnects, standard interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of FIG. 14 may be interconnected using a Compute Express Link (CXL) interconnect.

[0186] In at least one embodiment, FIG. 14 shows a display 1424, a touch screen 1425, a touch pad 1430, a Near Field Communications unit (NFC) 1445, a sensor hub 1440, a thermal sensor 1446, an Express Chipset (EC) 1435, a Trusted Platform Module (TPM) 1438, a BIOS / firmware / flash memory (BIOS, FW flash) 1422, a DSP 1460, a drive such as a Solid State Disk (SSD) or a Hard Disk Drive (HDD) (SSD or HDD) 1420, a wireless local area network unit (WLAN) 1450, a Bluetooth unit 1452, a Wireless Wide Area Network unit (WWAN) 1456, a Global Positioning System (GPS) 1455, a camera such as a USB3.0 camera (USB3.0 camera) 1454, or a Low Power Double Data Rate (LPDDR) memory unit (LPDDR3) 1415 implemented, for example, in accordance with the LPDDR3 standard. These components may be implemented in any suitable manner, respectively.

[0187] In at least one embodiment, other components may be communicatively coupled to the processor 1410 via the components described above. In at least one embodiment, an accelerometer 1441, an ambient light sensor ("ALS") 1442, a compass 1443, and a gyroscope 1444 may be communicatively coupled to a sensor hub 1440. In at least one embodiment, a thermal sensor 1439, a fan 1437, a keyboard 1446, and a touch pad 1430 may be communicatively coupled to an EC 1435. In at least one embodiment, a speaker 1463, headphones 1464, and a microphone ("mic") 1465 may be communicatively coupled to an audio unit (audio codec and class D amplifier) 1464, and this audio unit may be communicatively coupled to a DSP 1460. In at least one embodiment, the audio unit 1464 may include, for example and without limitation, an audio coder / decoder ("codec") and a class D amplifier. In at least one embodiment, a SIM card ("SIM") 1457 may be communicatively coupled to a WWAN unit 1456. In at least one embodiment, components such as a WLAN unit 1450 and a Bluetooth unit 1452, and a WWAN 1456 may be implemented in a next generation form factor ("NGFF": Next Generation Form Factor).

[0188] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, inference and / or training logic 915 may be used in the system of FIG. 14 for inference or prediction operations, at least partially based on weight parameters calculated using the training operations of the neural network, the functionality and / or architecture of the neural network, or the use cases of the neural network as described herein. In at least one embodiment, the neural network used by computer system 1500 may operate using inferences by one or more neural networks, each trained by two or more processing cores for separately training portions of the neural network in parallel as described herein.

[0189] FIG. 15 shows a computer system 1500 according to at least one embodiment. In at least one embodiment, computer system 1500 is configured to implement the various processes and methods described throughout this disclosure.

[0190] In at least one embodiment, computer system 1500 includes, without limitation, at least one central processing unit ("CPU") 1502, which is connected to a communication bus 1510 implemented using any suitable protocol, such as, without limitation, PCI: Peripheral Component Interconnect ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express": peripheral component interconnect express), AGP: Accelerated Graphics Port ("Accelerated Graphics Port"), HyperTransport, or any other bus or point-to-point communication protocol. In at least one embodiment, computer system 1500 includes, without limitation, main memory 1504 and control logic (implemented, for example, as hardware, software, or a combination thereof), and data is stored in main memory 1504, which may take the form of random access memory ("RAM": random access memory). In at least one embodiment, a network interface subsystem ("network interface") 1522 provides an interface with other computing devices and networks for receiving data from and transmitting data from computer system 1500 to other systems.

[0191] In at least one embodiment, computer system 1500 includes, without limitation in at least one embodiment, an input device 1508, a parallel processing system 1512, and a display device 1506, which can be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is received from an input device 1508 such as a keyboard, a mouse, a touch pad, a microphone, etc. In at least one embodiment, each of the above modules can be placed on a single semiconductor platform to form a processing system.

[0192] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIG. 9A and / or FIG. 9B. In at least one embodiment, training logic 915 may be used in the system of FIG. 15 for inference or prediction operations, at least partially based on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network. In at least one embodiment, the neural network used by computer system 1600 can operate using inference by one or more neural networks, each trained by two or more processing cores for separately training portions of the neural network in parallel as described herein.

[0193] FIG. 16 shows a computer system 1600 according to at least one embodiment. In at least one embodiment, the computer system 1600 may include, without limitation, a computer 1610 and a USB stick 1620. In at least one embodiment, the computer system 1610 may include, without limitation, any number and type of processors (not shown), as well as memory. In at least one embodiment, the computer 1610 may include, without limitation, servers, cloud instances, laptops, and desktop computers.

[0194] In at least one embodiment, the USB stick 1620 may include, without limitation, a processing unit 1630, a USB interface 1640, and USB interface logic 1650. In at least one embodiment, the processing unit 1630 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 1630 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing core 1630 comprises an application specific integrated circuit ( "ASIC") optimized to perform any amount and type of operations related to machine learning. For example, in at least one embodiment, the processing core 1630 is a tensor processing unit ( "TPC") optimized to perform inference operations of machine learning. In at least one embodiment, the processing core 1630 is a vision processing unit ( "VPU") optimized to perform inference operations of machine vision and machine learning.

[0195] In at least one embodiment, the USB interface 1640 may be any type of USB connector or USB socket. For example, in at least one embodiment, the USB interface 1640 is a USB 3.0 Type-C socket for data and power. In at least one embodiment, the USB interface 1640 is a USB 3.0 Type-A connector. In at least one embodiment, the USB interface logic 1650 may include any amount and type of logic that enables the processing unit 1630 to interface with a device (such as computer 1610) via the USB connector 1640.

[0196] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the system of FIG. 16 for inference or prediction operations, based at least in part on the training operations of the neural network, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network described herein. In at least one embodiment, workers operating in parallel can be implemented to train a neural network as described above, using an architecture as shown in FIG. 17A.

[0197] FIG. 17A shows an exemplary architecture in which a plurality of GPUs 1710-1713 are communicatively coupled to a plurality of multi-core processors 1705-1706 via high-speed links 1740-1743 (e.g., buses, point-to-point interconnects, etc.). In one embodiment, the high-speed links 1740-1743 support a communication throughput of 4 GB / second, 30 GB / second, 80 GB / second, or more. Various interconnect protocols may be used, including but not limited to PCIe 4.0 or 5.0, and NVLink 2.0.

[0198] Furthermore, in one embodiment, two or more of GPUs 1710-1713 are interconnected via high-speed links 1729-1730, which may be implemented using the same or different protocols / links as those used for high-speed links 1740-1743. Similarly, two or more of multi-core processors 1705-1706 may be connected via high-speed link 1728, which can be a symmetric multi-processor (SMP) bus operating at 20 GB / second, 30 GB / second, 120 GB / second, or more. Alternatively, all communications between the various system components shown in FIG. 17A may be realized using the same protocol / link (e.g., via a common interconnect fabric).

[0199] In one embodiment, each multi-core processor 1705-1706 is communicatively coupled to processor memories 1701-1702 via memory interconnects 1726-1727, respectively, and each GPU 1710-1713 is communicatively coupled to GPU memories 1720-1723 via GPU memory interconnects 1750-1753, respectively. Memory interconnects 1726-1727 and 1750-1753 may utilize the same or different memory access technologies. By way of example and not limitation, processor memories 1701-1702 and GPU memories 1720-1723 may be volatile memories such as dynamic random access memory (DRAM) (including stacked DRAM), graphics double data rate SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or non-volatile memories such as 3D XPoint or Nano-Ram. In one embodiment, some portions of processor memories 1701-1702 may be volatile memories and other portions may be non-volatile memories (e.g., using a two-level memory (2LM) hierarchy).

[0200] As described herein, the various processors 1705 - 1706 and GPUs 1710 - 1713 may each be physically coupled to specific memories 1701 - 1702, 1720 - 1723, or an integrated memory architecture may be implemented where the same virtual system address space (also referred to as the "effective address" space) is distributed among the various physical memories. For example, each of the processor memories 1701 - 1702 may include 64GB of system memory address space, and each of the GPU memories 1720 - 1723 may include 32GB of system memory address space (in this example, resulting in a total of 256GB of addressable memory).

[0201] FIG. 17B shows further details of the interconnection of a multi - core processor 1707 and a graphics acceleration module 1746 according to one exemplary embodiment. The graphics acceleration module 1746 may include one or more GPU chips integrated on a line card coupled to the processor 1707 via a high - speed link 1740. Alternatively, the graphics acceleration module 1746 may be integrated on the same package or chip as the processor 1707.

[0202] In at least one embodiment, the illustrated processor 1707 includes a plurality of cores 1760A - 1760D, each core having a translation lookaside buffer 1761A - 1761D and one or more caches 1762A - 1762D. In at least one embodiment, cores 1760A - 1760D may include various other components (not shown) for executing instructions and processing data. Caches 1762A - 1762D may comprise level 1 (L1) and level 2 (L2) caches. Further, one or more shared caches 1756 may be included in caches 1762A - 1762D and shared by a set of cores 1760A - 1760D. For example, one embodiment of processor 1707 includes 24 cores, each core having its own L1 cache, 12 shared L2 caches, and 12 shared L3 caches. In this embodiment, one or more of the L2 and L3 caches are shared by two adjacent cores. Processor 1707 and graphics acceleration module 1746 are connected to system memory 1714, which may include the processor memories 1701 - 1702 of FIG. 17A.

[0203] For the various caches 1762A - 1762D, 1756, and the data and instructions stored in system memory 1714, coherence is maintained by inter - core communication via coherence bus 1764. For example, each cache may have associated cache coherence logic / circuitry to communicate via coherence bus 1764 in response to detecting a read or write to a particular cache line. In one implementation, a cache snooping protocol is implemented via coherence bus 1764 to monitor cache accesses.

[0204] In one embodiment, the proxy circuit 1725 communicatively couples the graphics acceleration module 1746 to the coherence bus 1764 so that the graphics acceleration module 1746 can participate in the cache coherence protocol as a peer of cores 1760A - 1760D. In particular, the interface 1735 provides a connection to the proxy circuit 1725 via a high-speed link 1740 (e.g., a PCIe bus, NVLink, etc.), and the interface 1737 connects the graphics acceleration module 1746 to the link 1740.

[0205] In one implementation, the accelerator integration circuit 1736 provides services for cache management, memory access, content management, and interrupt management instead of the plurality of graphics processing engines 1731, 1732, N of the graphics acceleration module 1746. Each of the graphics processing engines 1731, 1732, N may comprise a separate graphics processing unit (GPU). Alternatively, the graphics processing engines 1731, 1732, N may comprise different types of graphics processing engines, such as a graphics execution unit, a media processing engine (e.g., a video encoder / decoder), a sampler, and a blit engine, etc., within a GPU. In at least one embodiment, the graphics acceleration module 1746 may be a GPU having a plurality of graphics processing engines 1731 - 1732, N, or the graphics processing engines 1731 - 1732, N may be individual GPUs integrated in a common package, line card, or chip.

[0206] In one embodiment, the accelerator integration circuit 1736 includes a memory management unit (MMU) 1739 for performing various memory management functions, such as virtual-to-physical memory translation (also referred to as effective-to-real memory translation), and a memory access protocol for accessing the system memory 1714. The MMU 1739 can also include a translation lookaside buffer (TLB) (not shown) for caching virtual / effective to physical / real address translations. In one implementation, the cache 1738 stores commands and data so that the graphics processing engines 1731-1732 can access them efficiently. In one embodiment, the data stored in the cache 1738 and the graphics memories 1733-1734, M are kept coherent with the core caches 1762A-1762D, 1756, and the system memory 1714. As described above, this may be achieved via the proxy circuit 1725 instead of the cache 1738 and the memories 1733-1734, M (e.g., sending updates regarding cache line modifications / accesses in the processor caches 1762A-1762D, 1756 to the cache 1738 and receiving updates from the cache 1738).

[0207] The set of registers 1745 stores context data for the threads executed by the graphics processing engines 1731-1732, N, and the context management circuit 1748 manages the thread contexts. For example, the context management circuit 1748 may perform save and restore operations to save and restore the contexts of various threads during a context switch (e.g., here, the first thread is saved and the second thread is stored so that the second thread can be executed by the graphics processing engine). For example, during a context switch, the context management circuit 1748 may store the current register values in a specified area of memory (identified, for example, by a context pointer). Then, when returning to the context, the context management circuit 1748 may restore the register values. In one embodiment, the interrupt management circuit 1747 receives and processes interrupts received from system devices.

[0208] In one implementation, the virtual / effective addresses from the graphics processing engine 1731 are translated by the MMU 1739 into real / physical addresses of the system memory 1714. One embodiment of the accelerator integration circuit 1736 supports a plurality (e.g., 4, 8, 16) of graphics accelerator modules 1746, and / or other accelerator devices. The graphics accelerator module 1746 may be dedicated to a single application executed on the processor 1707, or may be shared among multiple applications. In one embodiment, there is a virtualized graphics execution environment in which the resources of the graphics processing engines 1731-1732, N are shared among multiple applications or virtual machines (VMs). In at least one embodiment, the resources may be subdivided into "slices" that are allocated to different VMs and / or applications based on processing requirements and the priorities associated with the VMs and / or applications.

[0209] In at least one embodiment, the accelerator integration circuit 1736 functions as a bridge to the system for the graphics acceleration module 1746 and provides address translation and cache services for system memory. Further, the accelerator integration circuit 1736 may provide a virtualization facility for the host processor to manage virtualization, interrupts, and memory management of the graphics processing engines 1731-1732.

[0210] Since the hardware resources of the graphics processing engines 1731-1732, N are explicitly mapped to the physical address space seen by the host processor 1707, any host processor can directly address these resources using the effective address value. In one embodiment, one function of the accelerator integration circuit 1736 is to physically separate the graphics processing engines 1731-1732, N so that they appear as independent units to the system.

[0211] In at least one embodiment, one or more graphics memories 1733-1734, M are each coupled to a respective one of the graphics processing engines 1731-1732, N. The graphics memories 1733-1734, M store instructions and data processed by their respective graphics processing engines 1731-1732, N. The graphics memories 1733-1734, M may be volatile memories such as DRAM (including stacked DRAM), GDDR memory (e.g., GDDR5, GDDR6), or HBM, and / or non-volatile memories such as 3D XPoint or Nano-Ram.

[0212] In one embodiment, to reduce data traffic through link 1740, data stored in graphics memories 1733-1734, M is made to be the data most frequently used by graphics processing engines 1731-1732, N, and preferably data not used (or at least not frequently used) by cores 1760A-1760D. Similarly, the bias mechanism attempts to keep data required by the cores (and thus preferably not required by graphics processing engines 1731-1732, N) in caches 1762A-1762D, 1756 of the cores, and system memory 1714.

[0213] FIG. 17C shows another exemplary embodiment in which the accelerator integration circuit 1736 is integrated within the processor 1707. At least in this embodiment, graphics processing engines 1731-1732, N communicate directly with the accelerator integration circuit 1736 via the high-speed link 1740 through interfaces 1737 and 1735 (again, any form of bus or interface protocol can be utilized). The accelerator integration circuit 1736 may perform the same operations as described with respect to FIG. 17B, but potentially operate at a higher throughput considering its proximity to the coherence bus 1764 and caches 1762A-1762D, 1756. At least one embodiment supports different programming models including a dedicated process programming model (without virtualization of the graphics acceleration module) and a shared programming model (with virtualization), which may include a programming model controlled by the accelerator integration circuit 1736 and a programming model controlled by the graphics acceleration module 1746.

[0214] In at least one embodiment, the graphics processing engines 1731-1732, N are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can concentrate other application requirements on the graphics processing engines 1731-1732, N to implement virtualization within the VM / partition.

[0215] In at least one embodiment, the graphics processing engines 1731-1732, N may be shared by multiple VM / application partitions. In at least one embodiment, the shared model may use a system hypervisor to virtualize the graphics processing engines 1731-1732, N to enable access by each operating system. In a single partition system without a hypervisor, the graphics processing engines 1731-1732, N are owned by the operating system. In at least one embodiment, the operating system can virtualize the graphics processing engines 1731-1732, N to provide access to each process or application.

[0216] In at least one embodiment, the graphics acceleration module 1746 or individual graphics processing engines 1731-1732, N use a process handle to select process elements. In at least one embodiment, the process elements are stored in the system memory 1714 and can be addressed using the translation technique from the effective address to the physical address described herein. In at least one embodiment, the process handle may be an implementation-specific value provided to the host process when registering the context of the host process with the graphics processing engines 1731-1732, N (i.e., calling system software to add a process element to the process element link list). In at least one embodiment, the lower 16 bits of the process handle may be the offset of the process element within the process element link list.

[0217] FIG. 17D shows an exemplary accelerator integration slice 1790. As used herein, a "slice" comprises a designated portion of the processing resources of the accelerator integration circuit 1736. The application effective address space 1782 within the system memory 1714 stores a process element 1783. In one embodiment, the process element 1783 is stored in response to a GPU call 1781 from an application 1780 executing on the processor 1707. The process element 1783 accommodates the process state of the corresponding application 1780. A work descriptor (WD) 1784 accommodated by the process element 1783 can be a single job requested by the application, or can accommodate a pointer to a queue of jobs. In at least one embodiment, the WD 1784 is a pointer to a job request queue in the address space 1782 of the application.

[0218] The graphics acceleration module 1746 and / or the individual graphics processing engines 1731 - 1732, N can be shared by all or a subset of the processes within the system. In at least one embodiment, infrastructure for setting the process state and sending the WD 1784 to the graphics acceleration module 1746 to initiate a job in a virtualized environment may be included.

[0219] In at least one embodiment, a dedicated process programming model is implementation - specific. In this model, a single process owns the graphics acceleration module 1746 or an individual graphics processing engine 1731. Since the graphics acceleration module 1746 is owned by a single process, when the graphics acceleration module 1746 is allocated, the hypervisor initializes the accelerator integration circuit 1736 for the owning partition, and the operating system initializes the accelerator integration circuit 1736 for the owning process.

[0220] During operation, the WD fetch unit 1791 within the accelerator integration slice 1790 fetches the next WD 1784, which includes the display of work to be performed by one or more graphics processing engines of the graphics acceleration module 1746. As shown, the data from the WD 1784 is stored in the register 1745 and may be used by the MMU 1739, the interrupt management circuit 1747, and / or the context management circuit 1748. For example, one embodiment of the MMU 1739 includes a segment / page walk circuit for accessing the segment / page table 1786 within the OS virtual address space 1785. The interrupt management circuit 1747 may process the interrupt event 1792 received from the graphics acceleration module 1746. When executing a graphics operation, the effective address 1793 generated by the graphics processing engines 1731 - 1732, N is translated to a physical address by the MMU 1739.

[0221] In one embodiment, the same set of registers 1745 is replicated for each of the graphics processing engines 1731 - 1732, N, and / or the graphics acceleration module 1746 and may be initialized by the hypervisor or the operating system. Each of these replicated registers may be included in the accelerator integration slice 1790. Exemplary registers that may be initialized by the hypervisor are shown in Table 1.

Table 1

[0222] Exemplary registers that may be initialized by the operating system are shown in Table 2.

Table 2

[0223] In one embodiment, each WD1784 is specific to a particular graphics acceleration module 1746 and / or graphics processing engines 1731 - 1732, N. The WD1784 can contain all the information required for the graphics processing engines 1731 - 1732, N to perform work, or can be a pointer to a memory location where the application has set up a command queue for the work to be completed.

[0224] Figure 17E shows further details of an exemplary embodiment of the shared model. This embodiment includes a hypervisor physical address space 1798 in which a process element list 1799 is stored. The hypervisor physical address space 1798 is accessible via a hypervisor 1796 that virtualizes the graphics acceleration module engine of the operating system 1795.

[0225] In at least one embodiment, the shared programming model enables all or a subset of processes from all or a subset of partitions within the system to use the graphics acceleration module 1746. There are two programming models in which the graphics acceleration module 1746 is shared by multiple processes and partitions: time slice sharing and graphics-directed shared.

[0226] In this model, the system hypervisor 1796 owns the graphics acceleration module 1746 and makes its functions available to all operating systems 1795. To support the virtualization by the system hypervisor 1796, the graphics acceleration module 1746 may comply with the following: 1) The job requests of the application must be autonomous (i.e., there is no need to maintain the state between jobs), or the graphics acceleration module 1746 must provide a mechanism for saving and restoring the context. 2) The job requests of the application are guaranteed by the graphics acceleration module 1746 to be completed within a specified amount of time, including any translation errors, or the graphics acceleration module 1746 provides a function to preempt the processing of the job. 3) When the graphics acceleration module 1746 operates in a specified shared programming model, fairness must be guaranteed among processes.

[0227] In at least one embodiment, application 1780 needs to make a system call to operating system 1795, along with the type of graphics acceleration module 1746, a work descriptor (WD), a privilege mask register (AMR) value, and a context save / restore area pointer (CSRP). In at least one embodiment, the type of graphics acceleration module 1746 describes the acceleration function targeted by the system call. In at least one embodiment, the type of graphics acceleration module 1746 may be a system-specific value. In at least one embodiment, the WD is specifically formatted for graphics acceleration module 1746 and can be in the form of a command for graphics acceleration module 1746, a virtual address pointer pointing to a user-defined structure, a virtual address pointer pointing to a command queue, or any other data structure for describing the work performed by graphics acceleration module 1746. In one embodiment, the AMR value is the AMR state for use by the current process. In at least one embodiment, the value passed to the operating system is the same as the application that sets the AMR. If the implementations of accelerator integration circuit 1736 and graphics acceleration module 1746 do not support a user privilege mask override register (UAMOR), the operating system may apply the current UAMOR value to the AMR value and then pass the AMR to a hypervisor call. Hypervisor 1796 may optionally apply the current privilege mask override register (AMOR) value and then put the AMR into process element 1783. In at least one embodiment, the CSRP is one of registers 1745 that holds the virtual address of an area within application address space 1782 for graphics acceleration module 1746 to save and restore the context state. This pointer is optional if no state needs to be saved between jobs or when a job is preempted. In at least one embodiment, the context save / restore area may be pinned system memory.

[0228] Upon receiving a system call, the operating system 1795 may verify that the application 1780 is registered and has been granted the right to use the graphics acceleration module 1746. The operating system 1795 then makes a call to the hypervisor 1796 with the information shown in Table 3.

Table 3

[0229] Upon receiving a hypervisor call, the hypervisor 1796 verifies that the operating system 1795 is registered and has been granted the right to use the graphics acceleration module 1746. The hypervisor 1796 then inserts the process element 1783 into the process element link list of the corresponding type of graphics acceleration module 1746. The process element may include the information shown in Table 4.

Table 4

[0230] In at least one embodiment, the hypervisor initializes the registers 1745 of the plurality of accelerator integration slices 1790.

[0231] As shown in FIG. 17F, in at least one embodiment, an integrated memory is used that is addressable via a common virtual memory address space used to access physical processor memories 1701-1702 and GPU memories 1720-1723. In this implementation, operations executed on GPUs 1710-1713 utilize the same virtual / effective memory address space as accessing the processor memories 1701-1702, and vice versa, thereby simplifying programmability. In one embodiment, a first portion of the virtual / effective address space is allocated to processor memory 1701, a second portion is allocated to a second processor memory 1702, a third portion is allocated to GPU memory 1720, and so on. In at least one embodiment, the entire virtual / effective memory space (sometimes referred to as the effective address space) is thereby distributed across each of the processor memories 1701-1702 and GPU memories 1720-1723 such that any processor or GPU can access any physical memory with virtual addresses mapped to the physical memory.

[0232] In one embodiment, bias / coherence management circuits 1794A-1794E in one or more of MMUs 1739A-1739E ensure cache coherence between the caches of one or more host processors (e.g., 1705) and the caches of GPUs 1710-1713 and implement a bias technique to indicate the physical memory in which a particular type of data should be stored. Although multiple instances of bias / coherence management circuits 1794A-1794E are shown in FIG. 17F, the bias / coherence circuit may be implemented within the MMU of one or more host processors 1705 and / or within the accelerator integration circuit 1736.

[0233] One embodiment enables the memory with GPU 1720-1723 to be mapped as part of the system memory and made accessible using the shared virtual memory (SVM) technique, without incurring performance degradation associated with full system cache coherence. In at least one embodiment, the memory with GPU 1720-1723 is accessible as system memory without cumbersome cache coherence overhead, providing a beneficial operating environment for GPU offloading. With this configuration, the host processor 1705 software can set operands and access computation results without the overhead of conventional I / O DMA data copying. Such conventional copies require driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are less efficient than simple memory accesses. In at least one embodiment, the ability to access the memory with GPU 1720-1723 without cache coherence overhead can be essential to the execution time of offloaded computations. For example, in the presence of significant streaming write memory traffic, cache coherence overhead can significantly reduce the effective write bandwidth seen by GPUs 1710-1713. In at least one embodiment, the efficiency of operand setting, access to results, and GPU computation can help in determining the effectiveness of GPU offloading.

[0234] In at least one embodiment, the selection of the GPU bias and the host processor bias is determined by a bias tracker data structure. For example, a bias table may be used, and this table may be a page granularity structure that includes one or two bits per memory page with a GPU (i.e., it may be controlled at the granularity of the memory page). In at least one embodiment, the bias table may be implemented in a stolen memory range of one or more GPUs with memory 1720-1723 in a state where a bias cache (e.g., for caching frequently used / recently used entries of the bias table) is present or not present in GPUs 1710-1713. Alternatively, the entire bias table may be maintained within the GPU.

[0235] In at least one embodiment, an entry of the bias table associated with each access to GPUs with memory 1720-1723 is accessed prior to the actual access to the GPU memory, resulting in the following operations. First, local requests from GPUs 1710-1713 that find their pages within the GPU bias are transferred directly to the corresponding GPUs with memory 1720-1723. Local requests from GPUs that find their pages within the host bias are transferred to the processor 1705 (e.g., via the high-speed link described above). In one embodiment, a request from the processor 1705 that finds the requested page within the host processor bias completes the request in the same manner as a normal memory read. Alternatively, requests directed to GPU-biased pages may be transferred to GPUs 1710-1713. In at least one embodiment, the GPU may then migrate the page to the host processor bias if the current page is not being used. In at least one embodiment, the bias state of a page can be changed by either a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited set of cases, simply a hardware-based mechanism.

[0236] One mechanism for changing the bias state utilizes an API call (e.g., OpenCL), where this API call calls the device driver of the GPU, and this device driver sends a message to the GPU (or adds a command descriptor to the queue) to change the bias state, and for some transitions, guides the GPU to perform a cache flushing operation on the host. In at least one embodiment, the cache flushing operation is used for the transition from the bias of the host processor 1705 to the GPU bias, but not for the reverse transition.

[0237] In one embodiment, cache coherence is maintained by the host processor 1705 temporarily rendering GPU-biased pages that cannot be cached. To access these pages, the processor 1705 may request access from the GPU 1710, and the GPU 1710 may either immediately grant access or not grant access. Thus, to reduce communication between the processor 1705 and the GPU 1710, it is beneficial to make the GPU-biased pages be requested by the GPU but not by the host processor 1705, or vice versa.

[0238] A hardware structure 915 is used to execute one or more embodiments. Details regarding the hardware structure (x)915 are provided herein in conjunction with FIGS. 9A and / or 9B.

[0239] FIG. 18 shows an exemplary integrated circuit and associated graphics processor that can be fabricated using one or more IP cores according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general-purpose processor cores.

[0240] FIG. 18 is a block diagram showing an exemplary system-on-chip integrated circuit 1800 that can be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 1800 includes one or more application processors 1805 (e.g., CPU), at least one graphics processor 1810, and may further include an image processor 1815 and / or a video processor 1820, any of which may be a modular IP core. In at least one embodiment, the integrated circuit 1800 includes a peripheral device or bus logic including a USB controller 1825, a UART controller 1830, an SPI / SDIO controller 1835, and an I²S / I²C controller 1840. In at least one embodiment, the integrated circuit 1800 can include a display device 1845 coupled to one or more of a high-definition multimedia interface (HDMI) controller 1850 and a mobile industry processor interface (MIPI) display interface 1855. In at least one embodiment, storage may be provided by a flash memory subsystem 1860 including a flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 1865 to access a SDRAM or SRAM memory device. In at least one embodiment, some integrated circuits further include an embedded security engine 1870.

[0241] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the integrated circuit 1800 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functionality and / or architecture of the neural network, or use cases of the neural network described herein. In at least one embodiment, workers operating in parallel can be implemented to train a neural network as described above using an architecture as shown in FIGS. 19A and 19B.

[0242] FIGS. 19A and 19B show exemplary integrated circuits and associated graphics processors that can be fabricated using one or more IP cores, according to various embodiments described herein. In addition to what is shown, in at least one embodiment, other logic and circuitry may be included, including additional graphics processors / cores, peripheral device interface controllers, or general purpose processor cores.

[0243] FIG. 19A and FIG. 19B are block diagrams showing exemplary graphics processors for use within a SoC according to the embodiments described herein. FIG. 19A shows an exemplary graphics processor 1910 of a system on a chip integrated circuit that can be manufactured using one or more IP cores according to at least one embodiment. FIG. 19B shows an additional exemplary graphics processor 1940 of a system on a chip integrated circuit that can be manufactured using one or more IP cores according to at least one embodiment. In at least one embodiment, the graphics processor 1910 of FIG. 19A is a low-power graphics processor core. In at least one embodiment, the graphics processor 1940 of FIG. 19B is a high-performance graphics processor core. In at least one embodiment, each of the graphics processors 1910, 1940 can be a variant of the graphics processor 1810 of FIG. 18.

[0244] In at least one embodiment, the graphics processor 1910 includes a vertex processor 1905 and one or more fragment processors 1915A - 1915N (e.g., 1915A, 1915B, 1915C, 1915D - 1915N - 1, and 1915N). In at least one embodiment, the graphics processor 1910 can execute different shader programs via separate logic, whereby the vertex processor 1905 is optimized to execute operations for vertex shader programs, while the one or more fragment processors 1915A - 1915N execute fragment (e.g., pixel) shading operations for fragment or pixel shader programs. In at least one embodiment, the vertex processor 1905 executes the vertex processing stage of the 3D graphics pipeline and generates primitives and vertex data. In at least one embodiment, the fragment processors 1915A - 1915N use the primitives and vertex data generated by the vertex processor 1905 to generate a frame buffer to be displayed on a display device. In at least one embodiment, the fragment processors 1915A - 1915N are optimized to execute fragment shader programs provided in the OpenGL API, and the OpenGL API may be used to perform operations similar to pixel shader programs provided in the Direct 3D API.

[0245] In at least one embodiment, the graphics processor 1910 further includes one or more memory management units (MMUs) 1920A - 1920B, caches 1925A - 1925B, and circuit interconnects 1930A - 1930B. In at least one embodiment, one or more MMUs 1920A - 1920B provide virtual - to - physical address mapping for the graphics processor 1910, including the vertex processor 1905 and / or the fragment processors 1915A - 1915N, and they may reference vertex or image / text data stored in memory in addition to vertex or image / text data stored in one or more caches 1925A - 1925B. In at least one embodiment, one or more MMUs 1920A - 1920B may be synchronized with one or more other MMUs in the system, including one or more MMUs associated with one or more of the application processors 1805, image processors 1815, and / or video processors 1820 of FIG. 18, such that each processor 1805 - 1820 can participate in a shared or integrated virtual memory system. In at least one embodiment, one or more circuit interconnects 1930A - 1930B enable the graphics processor 1910 to interface with other IP cores within the SoC via the internal bus of the SoC or via a direct connection.

[0246] In at least one embodiment, the graphics processor 1940 includes one or more of the MMUs 1920A-1920B, caches 1925A-1925B, and circuit interconnects 1930A-1930B of the graphics processor 1910 of FIG. 19A. In at least one embodiment, the graphics processor 1940 includes one or more shader cores 1955A-1955N (e.g., 1955A, 1955B, 1955C, 1955D, 1955E, 1955F-1955N-1, and 1955N), which provide an integrated shader core architecture where a single core, or type, or core can execute all types of programmable shader code including shader program code for implementing vertex shaders, fragment shaders, and / or compute shaders. In at least one embodiment, the number of shader cores can be varied. In at least one embodiment, the graphics processor 1940 includes a core-to-core task manager 1945 that acts as a thread dispatcher for dispatching execution threads to one or more of the shader cores 1955A-1955N, and a tiling unit 1958 for accelerating tiling operations for tile-based rendering where the rendering operation of a scene is subdivided in the image space, for example, to utilize local spatial coherence within the scene or to optimize the use of internal caches.

[0247] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in integrated circuits 19A and / or 19B for inference or prediction operations, at least in part based on weight parameters calculated using the training operations, functions and / or architectures of the neural networks described herein, or the use cases of neural networks. In at least one embodiment, worker(s) operating in parallel to train a neural network as described above can be implemented using graphics processor logic as shown in FIGS. 20A-20B.

[0248] FIGS. 20A-20B show further exemplary graphics processor logic according to embodiments described herein. FIG. 20A shows a graphics core 2000, which in at least one embodiment may be included in the graphics processor 1810 of FIG. 18, and in at least one embodiment may be integrated shader cores 1955A-1955N as in FIG. 19B. FIG. 20B shows a highly parallel general-purpose graphics processing unit 2030 suitable for introduction into a multi-chip module in at least one embodiment.

[0249] In at least one embodiment, the graphics core 2000 includes a shared instruction cache 2002, a texture unit 2018, and a cache / shared memory 2020, which are common to the execution resources within the graphics core 2000. In at least one embodiment, the graphics core 2000 can include a plurality of slices 2001A - 2001N, or per-core partitions, and the graphics processor can include a plurality of instances of the graphics core 2000. The slices 2001A - 2001N can include support logic that includes local instruction caches 2004A - 2004N, thread schedulers 2006A - 2006N, thread dispatchers 2008A - 2008N, and sets of registers 2010A - 2010N. In at least one embodiment, the slices 2001A - 2001N can include a set of additional function units (AFU2012A - 2012N), floating point units (FPU2014A - 2014N), integer arithmetic logic units (ALU2016 - 2016N), address calculation units (ACU2013A - 2013N), double precision floating point units (DPFPU2015A - 2015N), and matrix processing units (MPU2017A - 2017N).

[0250] In at least one embodiment, FPUs 2014A - 2014N can perform single - precision (32 - bit) and half - precision (16 - bit) floating - point operations, and DPFPU 2015A - 2015N perform double - precision (64 - bit) floating - point operations. In at least one embodiment, ALUs 2016A - 2016N can perform variable - precision integer operations with 8 - bit, 16 - bit, and 32 - bit precision and can be configured to perform mixed - precision operations. In at least one embodiment, MPU 2017A - 2017N can also be configured to perform mixed - precision matrix operations including half - precision floating - point and 8 - bit integer operations. In at least one embodiment, MPU 2017A - 2017N can perform various matrix operations for accelerating machine - learning application frameworks, including enabling support for accelerating general matrix - matrix multiplication (GEMM). In at least one embodiment, AFUs 2012A - 2012N can perform additional logical operations not supported by floating - point units or integer units, including trigonometric operations (e.g., sine, cosine, etc.).

[0251] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, inference and / or training logic 915 may be used in the graphics core 2000 for inference or prediction operations, at least partially based on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks. In at least one embodiment, graphics - processor logic as shown in FIGS. 20A - 20B can be used to implement workers that operate in parallel to train neural networks as described above.

[0252] FIG. 20B shows a general-purpose processing unit (GPGPU) 2030, which can be configured to perform high-parallel computing operations by an array of graphics processing units in at least one embodiment. In at least one embodiment, GPGPU 2030 can be directly linked to other instances of GPGPU 2030 to generate multiple GPU clusters to improve the training speed of deep neural networks. In at least one embodiment, GPGPU 2030 includes a host interface 2032 to enable connection with a host processor. In at least one embodiment, host interface 2032 is a PCI Express interface. In at least one embodiment, host interface 2032 can be a vendor-specific communication interface or communication fabric. In at least one embodiment, GPGPU 2030 receives commands from the host processor and uses a global scheduler 2034 to distribute execution threads associated with these commands to a set of compute clusters 2036A-2036H. In at least one embodiment, compute clusters 2036A-2036H share a cache memory 2038. In at least one embodiment, cache memory 2038 can act as a high-level cache for cache memory within compute clusters 2036A-2036H.

[0253] In at least one embodiment, GPGPU 2030 includes memories 2044A-2044B coupled to compute clusters 2036A-2036H via a set of memory controllers 2042A-2042B. In at least one embodiment, memories 2044A-2044B can include various types of memory devices, including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory, or graphics random access memory.

[0254] In at least one embodiment, each of compute clusters 2036A - 2036H includes a set of graphics cores, such as graphics core 2000 of FIG. 20A, and this set of graphics cores can include multiple types of integer and floating - point logic units that can perform computational operations with various precisions, including those suitable for machine - learning computations. For example, in at least one embodiment, at least a subset of the floating - point units in each of compute clusters 2036A - 2036H can be configured to perform 16 - bit or 32 - bit floating - point operations, while another subset of the floating - point units can be configured to perform 64 - bit floating - point operations.

[0255] In at least one embodiment, a plurality of instances of GPGPU2030 can be configured to operate as a compute cluster. In at least one embodiment, the communication used by compute clusters 2036A - 2036H for synchronization and data exchange varies across embodiments. In at least one embodiment, a plurality of instances of GPGPU2030 communicate via host interface 2032. In at least one embodiment, GPGPU2030 includes an I / O hub 2039 that couples GPGPU2030 to a GPU link 2040 that enables direct connection to other instances of GPGPU2030. In at least one embodiment, GPU link 2040 is coupled to a dedicated GPU - to - GPU bridge that enables communication and synchronization between multiple instances of GPGPU2030. In at least one embodiment, GPU link 2040 is coupled to a high - speed interconnect for sending and receiving data to / from other GPGPUs or parallel processors. In at least one embodiment, a plurality of instances of GPGPU2030 are located in separate data processing systems and communicate via a network device accessible via host interface 2032. In at least one embodiment, GPU link 2040 can be configured to enable connection to a host processor in addition to, or instead of, host interface 2032.

[0256] In at least one embodiment, the GPGPU 2030 can be configured to train a neural network. In at least one embodiment, the GPGPU 2030 can be used within an inference platform. In at least one embodiment where the GPGPU 2030 is used for inference, the GPGPU may include fewer compute clusters 2036A - 2036H than when the GPGPU is used for neural network training. In at least one embodiment, the memory technology associated with memories 2044A - 2044B may be different for the inference configuration and the training configuration, and high - bandwidth memory technology is applied to the training configuration. In at least one embodiment, the inference configuration of the GPGPU 2030 can support inference - specific instructions. For example, in at least one embodiment, the inference configuration can support one or more dot - product instructions for 8 - bit integers, which may be used during the inference operation of a pre - trained neural network.

[0257] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the GPGPU 2030 for inference or prediction operations, based at least in part on the training operations of the neural network described herein, the functions and / or architecture of the neural network, or the weight parameters calculated using the use cases of the neural network. In at least one embodiment, a computing system 2100 as shown in FIG. 21 can be used to implement workers operating in parallel to train a neural network as described above.

[0258] FIG. 21 is a block diagram showing a computing system 2100 according to at least one embodiment. In at least one embodiment, the computing system 2100 includes a processing subsystem 2101 having one or more processors 2102 and a system memory 2104 that communicate via an interconnect path that may include a memory hub 2105. In at least one embodiment, the memory hub 2105 may be a separate component within a chipset component or may be integrated within one or more processors 2102. In at least one embodiment, the memory hub 2105 is coupled to an I / O subsystem 2111 via a communication link 2106. In at least one embodiment, the I / O subsystem 2111 includes an I / O hub 2107 that enables the computing system 2100 to receive input from one or more input devices 2108. In at least one embodiment, the I / O hub 2107 can enable a display controller that may be included in one or more processors 2102 to provide output to one or more display devices 2110A. In at least one embodiment, one or more display devices 2110A coupled to the I / O hub 2107 can include local, internal, or embedded display devices.

[0259] In at least one embodiment, the processing subsystem 2101 includes one or more parallel processors 2112 coupled to a memory hub 2105 via a bus or other communication link 2113. In at least one embodiment, the communication link 2113 may be one of a number of communication link technologies or protocols based on any number of standards, such as but not limited to PCI Express, or may be a vendor-specific communication interface or communication fabric. In at least one embodiment, the one or more parallel processors 2112 include a computationally focused parallel or vector processing system that can include a number of processing cores and / or processing clusters, such as a many integrated core (MIC) processor. In at least one embodiment, the one or more parallel processors 2112 form a graphics processing subsystem that can output pixels to one of one or more display devices 2110A coupled via an I / O hub 2107. In at least one embodiment, the one or more parallel processors 2112 can also include a display controller and display interface (not shown) that enables direct connection to one or more display devices 2110B.

[0260] In at least one embodiment, the system storage unit 2114 can be connected to the I / O hub 2107 to provide a storage mechanism for the computing system 2100. In at least one embodiment, an I / O switch 2116 can be used to provide an interface mechanism for enabling communication between the I / O hub 2107 and other components such as a network adapter 2118 and / or a wireless network adapter 2119 that may be integrated into the platform, as well as various other devices that can be added via one or more add-in devices 2120. In at least one embodiment, the network adapter 2118 can be an Ethernet adapter or another wired network adapter. In at least one embodiment, the wireless network adapter 2119 can include one or more of Wi-Fi, Bluetooth, Near Field Communication (NFC), or other network devices including one or more wireless radios.

[0261] In at least one embodiment, the computing system 2100 can include other components not explicitly shown, including USB or other port connections, optical storage drives, video capture devices, etc., which may also be connected to the I / O hub 2107. In at least one embodiment, the communication paths interconnecting the various components of FIG. 21 may be implemented using any suitable protocol such as a PCI (Peripheral Component Interconnect) based protocol (e.g., PCI-Express), or other buses or point-to-point communication interfaces such as NV-Link high-speed interconnects, or other interconnect protocols.

[0262] In at least one embodiment, one or more parallel processors 2112 incorporate circuitry optimized for graphics and video processing, including, for example, a video output circuit, and constitute a graphics processing unit (GPU). In at least one embodiment, one or more parallel processors 2112 incorporate circuitry optimized for general-purpose processing. In at least one embodiment, the components of computing system 2100 may be integrated with one or more other system elements on a single integrated circuit. For example, in at least one embodiment, one or more parallel processors 2112, memory hub 2105, processor 2102, and I / O hub 2107 can be integrated into a system-on-chip (SoC) integrated circuit. In at least one embodiment, the components of computing system 2100 can be integrated into a single package to form a system-in-package (SIP) configuration. In at least one embodiment, at least a portion of the components of computing system 2100 can be integrated into a multi-chip module (MCM), and this module can be interconnected with other multi-chip modules to form a modular computing system.

[0263] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, inference and / or training logic 915 may be used in the system of FIG. 2100 for inference or prediction operations, at least partially based on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks. In at least one embodiment, as shown in FIG. 22A, parallel processors may be used to implement workers that operate in parallel to train neural networks as described above.

[0264] Processor FIG. 22A shows a parallel processor 2200 according to at least one embodiment. In at least one embodiment, various components of the parallel processor 2200 may be implemented using one or more integrated circuit devices such as programmable processors, application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In at least one embodiment, the illustrated parallel processor 2200 is a variant of one or more of the parallel processors 2112 shown in FIG. 21 according to an exemplary embodiment.

[0265] In at least one embodiment, parallel processor 2200 includes parallel processing units 2202. In at least one embodiment, parallel processing unit 2202 includes an I / O unit 2204 that enables communication with other devices including other instances of parallel processing unit 2202. In at least one embodiment, I / O unit 2204 may be directly connected to other devices. In at least one embodiment, I / O unit 2204 is connected to other devices through the use of a hub or switch interface such as memory hub 2105. In at least one embodiment, the connection between memory hub 2105 and I / O unit 2204 forms communication link 2113. In at least one embodiment, I / O unit 2204 is connected to host interface 2206 and memory crossbar 2216, where host interface 2206 receives commands targeted for execution of processing operations and memory crossbar 2216 receives commands targeted for execution of memory operations.

[0266] In at least one embodiment, when host interface 2206 receives a command buffer via I / O unit 2204, host interface 2206 can direct a work operation for executing these commands towards front end 2208. In at least one embodiment, front end 2208 is coupled to scheduler 2210, and this scheduler is configured to distribute commands or other work items to processing cluster array 2212. In at least one embodiment, scheduler 2210 ensures that processing cluster array 2212 is properly configured and in an effective state before tasks are distributed to processing cluster array 2212 of processing cluster array 2212. In at least one embodiment, scheduler 2210 is implemented via firmware logic running on a microcontroller. In at least one embodiment, microcontroller-implemented scheduler 2210 can be configured to execute complex scheduling and work distribution operations at a coarse granularity and a fine granularity, enabling rapid preemption of threads running in processing array 2212 and context switching. In at least one embodiment, host software can prove the scheduling workload in processing array 2212 via one of a plurality of graphics processing doorbells. In at least one embodiment, then, the workload can be automatically distributed across processing array 2212 by scheduler 2210 logic within the microcontroller including scheduler 2210.

[0267] In at least one embodiment, the processing cluster array 2212 can include up to "N" processing clusters (e.g., cluster 2214A, cluster 2214B - cluster 2214N). In at least one embodiment, each of the clusters 2214A - 2214N of the processing cluster array 2212 can execute a large number of simultaneous threads. In at least one embodiment, the scheduler 2210 can use various scheduling and / or work distribution algorithms to distribute work to the clusters 2214A - 2214N of the processing cluster array 2212, and these algorithms may vary according to the workload generated for each type of program or calculation. In at least one embodiment, the scheduling may be dynamically handled by the scheduler 2210, or may be partially assisted by the compiler logic during the compilation of the program logic configured to be executed by the processing cluster array 2212. In at least one embodiment, the different clusters 2214A - 2214N of the processing cluster array 2212 can be distributed to process different types of programs or execute different types of calculations.

[0268] In at least one embodiment, the processing cluster array 2212 can be configured to perform various types of parallel processing operations. In at least one embodiment, the processing cluster array 2212 is configured to perform general-purpose parallel computing operations. For example, in at least one embodiment, the processing cluster array 2212 can include logic for performing processing tasks including filtering of video and / or audio data, execution of modeling operations including physical operations, and execution of data conversion.

[0269] In at least one embodiment, the processing cluster array 2212 is configured to perform parallel graphics processing operations. In at least one embodiment, the processing cluster array 2212 can include texture sampling logic for performing texture operations, as well as mosaic logic and other vertex processing logic, and additional logic for supporting the performance of such graphics processing operations, although not limited thereto. In at least one embodiment, the processing cluster array 2212 can be configured to execute graphics processing related shader programs such as, but not limited to, vertex shaders, mosaic shaders, geometry shaders, and pixel shaders. In at least one embodiment, the parallel processing unit 2202 can transfer data from the system memory through the I / O unit 2204 for processing. In at least one embodiment, during processing, the transferred data can be stored in on-chip memory (e.g., parallel processor memory 2222) during processing and then written back to the system memory.

[0270] In at least one embodiment, when graphics processing is performed using the parallel processing unit 2202, the scheduler 2210 can be configured to divide the processing workload into tasks of approximately equal size so as to more effectively distribute the graphics processing operations among the plurality of clusters 2214A - 2214N of the processing cluster array 2212. In at least one embodiment, a portion of the processing cluster array 2212 can be configured to perform different types of processing. For example, in at least one embodiment, for generating and displaying a rendered image, the first portion may be configured to perform vertex shading and topology generation, the second portion may be configured to perform mosaic and geometry shading, and the third portion may be configured to perform pixel shading or other screen space operations. In at least one embodiment, intermediate data generated by one or more of the clusters 2214A - 2214N can be stored in a buffer so that the intermediate data can be transmitted among the clusters 2214A - 2214N for further processing.

[0271] In at least one embodiment, the processing cluster array 2212 can receive processing tasks to be executed via a scheduler 2210, and the scheduler 2210 receives commands defining the processing tasks from a front end 2208. In at least one embodiment, the processing tasks can include an index of data to be processed, such as surface (patch) data, primitive data, vertex data, and / or pixel data, as well as state parameters and commands defining how the data is to be processed (e.g., which program to execute). In at least one embodiment, the scheduler 2210 may be configured to fetch an index corresponding to a task or may receive an index from the front end 2208. In at least one embodiment, the front end 2208 can be configured to ensure that the processing cluster array 2212 is configured in an active state before a workload specified by an incoming command buffer (e.g., a batch buffer, a push buffer, etc.) is started.

[0272] In at least one embodiment, each of one or more instances of the parallel processing unit 2202 can be coupled to a parallel processor - memory 2222. In at least one embodiment, the parallel processor - memory 2222 can be accessed via a memory crossbar 2216, and the memory crossbar 2216 can receive memory requests from the processing cluster array 2212 as well as the I / O unit 2204. In at least one embodiment, the memory crossbar 2216 can access the parallel processor - memory 2222 via a memory interface 2218. In at least one embodiment, the memory interface 2218 can include a plurality of partition units (e.g., partition unit 2220A, partition unit 2220B - partition unit 2220N), and each of these units can be coupled to a portion (e.g., a memory unit) of the parallel processor - memory 2222. In at least one embodiment, the number of partition units 2220A - 2220N is configured to be equal to the number of memory units, such that the first partition unit 2220A has a corresponding first memory unit 2224A, the second partition unit 2220B has a corresponding memory unit 2224B, and the Nth partition unit 2220N has a corresponding Nth memory unit 2224N. In at least one embodiment, the number of partition units 2220A - 2220N may not be equal to the number of memory devices.

[0273] In at least one embodiment, the memory units 2224A-2224N can include various types of memory devices, including dynamic random access memory (DRAM), such as synchronous graphics random access memory (SGRAM) including graphics double data rate (GDDR) memory, or graphics random access memory. In at least one embodiment, the memory units 2224A-2224N may also include 3D stacked memory including, but not limited to, high bandwidth memory (HBM). In at least one embodiment, to efficiently use the available bandwidth of the parallel processor memory 2222, a render target, such as a frame buffer or texture map, can be stored across the memory units 2224A-2224N so that the partition units 2220A-2220N can write portions of each render target in parallel. In at least one embodiment, the local instance of the parallel processor memory 2222 may be excluded to be advantageous for an integrated memory design that combines system memory and local cache memory.

[0274] In at least one embodiment, any one of clusters 2214A to 2214N of the processing cluster array 2212 can process data to be written to any one of memory units 2224A to 2224N in the parallel processor memory 2222. In at least one embodiment, the memory crossbar 2216 can be configured to transfer the output of each of clusters 2214A to 2214N to any partition unit 2220A to 2220N capable of performing further processing operations on the output, or to another cluster 2214A to 2214N. In at least one embodiment, each of clusters 2214A to 2214N can communicate with the memory interface 2218 through the memory crossbar 2216 to read from or write to various external memory devices. In at least one embodiment, the memory crossbar 2216 has a connection to the memory interface 2218 for communicating with the I / O unit 2204, as well as a connection to a local instance of the parallel processor memory 2222, enabling processing units in different processing clusters 2214A to 2214N to communicate with the system memory or other memories not local to the parallel processing unit 2202. In at least one embodiment, the memory crossbar 2216 can use virtual channels to separate traffic streams between clusters 2214A to 2214N and partition units 2220A to 2220N.

[0275] In at least one embodiment, multiple instances of the parallel processing unit 2202 may be provided on a single add-in card, or multiple add-in cards may be interconnected. In at least one embodiment, different instances of the parallel processing unit 2202 may be configured to interoperate even if they have different numbers of processing cores, different amounts of local parallel processor memory, and / or other different configurations. For example, in at least one embodiment, some instances of the parallel processing unit 2202 can include a higher-precision floating-point unit compared to other instances. In at least one embodiment, a system incorporating one or more instances of the parallel processing unit 2202 or the parallel processor 2200 can be implemented in various configurations and form factors including, but not limited to, desktop, laptop, or portable personal computers, servers, workstations, game consoles, and / or embedded systems.

[0276] FIG. 22B is a block diagram of a partition unit 2220 according to at least one embodiment. In at least one embodiment, the partition unit 2220 is an instance of one of the partition units 2220A - 2220N of the partition unit of FIG. 22A. In at least one embodiment, the partition unit 2220 includes an L2 cache 2221, a frame buffer interface 2225, and a ROP: raster operations unit 2226. The L2 cache 2221 is a read / write cache configured to perform load and store operations received from a memory crossbar 2216 and a ROP 2226. In at least one embodiment, read misses and urgent write-back requests are output by the L2 cache 2221 to the frame buffer interface 2225 to be processed. In at least one embodiment, updates are also sent to the frame via the frame buffer interface 2225 to be processed. In at least one embodiment, the frame buffer interface 2225 interfaces with one of the memory units of a parallel processor memory, such as the memory units 2224A - 2224N (e.g., within the parallel processor memory 2222 of FIG. 22).

[0277] In at least one embodiment, the ROP 2226 is a processing unit that performs raster operations such as stencil, z - test, and blending. In at least one embodiment, the ROP 2226 then outputs processed graphics data stored in the graphics memory. In at least one embodiment, the ROP 2226 includes compression logic for compressing depth or color data written to the memory and decompressing depth or color data read from the memory. In at least one embodiment, the compression logic can be lossless compression logic that utilizes one or more of a plurality of compression algorithms. The type of compression performed by the ROP 2226 can be changed based on the statistical characteristics of the data being compressed. For example, in at least one embodiment, delta color compression is performed on a tile - by - tile basis for depth and color data.

[0278] In at least one embodiment, ROP2226 is included not within partition unit 2220 but within each processing cluster (e.g., clusters 2214A - 2214N of FIG. 22). In at least one embodiment, read and write requests for pixel data, rather than pixel fragment data, are transmitted via memory crossbar 2216. In at least one embodiment, processed graphics data may be displayed on a display device such as one of the one or more display devices 2110 of FIG. 21, may be routed so as to be further processed by processor 2102, or may be routed so as to be further processed by one of the processing entities within parallel processor 2200 of FIG. 22A.

[0279] FIG. 22C is a block diagram of processing cluster 2214 within a parallel processing unit according to at least one embodiment. In at least one embodiment, the processing cluster is an instance of one of processing clusters 2214A - 2214N of FIG. 22. In at least one embodiment, processing cluster 2214 may be configured to execute a number of threads in parallel, where the term "thread" refers to an instance of a particular program executing on a particular set of input data. In at least one embodiment, a single instruction multiple data (SIMD) instruction issue technique is used to support parallel execution of a number of threads without providing a plurality of independent instruction units. In at least one embodiment, a single instruction multiple thread (SIMT) technique is used to support parallel execution of a number of threads that are overall synchronized using a common instruction unit configured to issue instructions to a set of processing engines within each processing cluster.

[0280] In at least one embodiment, the operation of processing cluster 2214 can be controlled via a pipeline manager 2232 that distributes processing tasks to SIMT parallel processors. In at least one embodiment, pipeline manager 2232 receives instructions from scheduler 2210 of FIG. 22 and manages the execution of these instructions via graphics multiprocessor 2234 and / or texture unit 2236. In at least one embodiment, graphics multiprocessor 2234 is an exemplary instance of a SIMT parallel processor. However, in at least one embodiment, various types of SIMT parallel processors with different architectures may be included within processing cluster 2214. In at least one embodiment, one or more instances of graphics multiprocessor 2234 can be included within processing cluster 2214. In at least one embodiment, graphics multiprocessor 2234 can process data, and data crossbar 2240 may be used to distribute the processed data to one of a plurality of possible destinations including other shader units. In at least one embodiment, pipeline manager 2232 can facilitate the distribution of processed data by specifying the destination of the processed data that will be distributed through data crossbar 2240.

[0281] In at least one embodiment, each graphics multiprocessor 2234 within processing cluster 2214 can include the same set of function execution logic (e.g., arithmetic logic units, load store units, etc.). In at least one embodiment, the function execution logic can be configured in a pipelined manner such that new instructions can be issued before the previous instruction has completed. In at least one embodiment, the function execution logic supports various operations including integer and floating point arithmetic, comparison operations, boolean operations, bit shifts, and the calculation of various algebraic functions. In at least one embodiment, different operations can be executed by leveraging the hardware of the same function units, and any combination of function units may exist.

[0282] In at least one embodiment, the instructions sent to processing cluster 2214 configure threads. In at least one embodiment, a set of threads being executed across a set of parallel processing engines is a thread group. In at least one embodiment, the thread group executes a program on different input data. In at least one embodiment, each thread within the thread group can be assigned to a different processing engine within graphics multiprocessor 2234. In at least one embodiment, the thread group may include fewer threads than the number of processing engines within graphics multiprocessor 2234. In at least one embodiment, if the thread group includes fewer threads than the number of processing engines, one or more of the processing engines may be idle during cycles in which the thread group is being processed. In at least one embodiment, the thread group may also include more threads than the number of processing engines within graphics multiprocessor 2234. In at least one embodiment, if the thread group includes more threads than the number of processing engines within graphics multiprocessor 2234, processing can be executed over consecutive clock cycles. In at least one embodiment, multiple thread groups can be executed simultaneously on graphics multiprocessor 2234.

[0283] In at least one embodiment, the graphics multi-processor 2234 includes an internal cache memory for performing load and store operations. In at least one embodiment, the graphics multi-processor 2234 can forego the internal cache and use the cache memory (e.g., L1 cache 2248) within the processing cluster 2214. In at least one embodiment, each graphics multi-processor 2234 can also access the L2 cache within a partition unit (e.g., partition units 2220A-2220N of FIG. 22), and these caches can be shared among all processing clusters 2214 and may be used to transfer data between threads. In at least one embodiment, the graphics multi-processor 2234 can also access off-chip global memory, which can include one or more of local parallel processor memory and / or system memory. In at least one embodiment, any memory external to the parallel processing unit 2202 may be used as global memory. In at least one embodiment, the processing cluster 2214 includes multiple instances of the graphics multi-processor 2234 that can share common instructions and data, and these may be stored in the L1 cache 2248.

[0284] In at least one embodiment, each processing cluster 2214 may include an MMU 2245 (memory management unit) configured to map virtual addresses to physical addresses. In at least one embodiment, one or more instances of the MMU 2245 may be within the memory interface 2218 of FIG. 22. In at least one embodiment, the MMU 2245 includes a set of page table entries (PTEs) used to map virtual addresses to physical addresses of tiles (tiling is described in detail) and optionally cache line indices. In at least one embodiment, the MMU 2245 may include a translation lookaside buffer (TLB) or cache, which may be within the graphics multiprocessor 2234 or L1 cache, or within the processing cluster 2214. In at least one embodiment, the physical address is processed to locally distribute surface data access, enabling efficient interleaving of requests among partition units. In at least one embodiment, a cache line index may be used to determine whether a cache line request is a hit or a miss.

[0285] In at least one embodiment, each graphics multi-processor 2234 is coupled to a texture unit 2236 such that the processing cluster 2214 may be configured to perform texture mapping operations, such as determining texture sample positions, reading texture data, and filtering texture data. In at least one embodiment, texture data is read from an internal texture L1 cache (not shown) or from an L1 cache within the graphics multi-processor 2234 and, if necessary, fetched from an L2 cache, local parallel processor memory, or system memory. In at least one embodiment, each graphics multi-processor 2234 outputs processed tasks to a data crossbar 2240 to provide the processed tasks to another processing cluster 2214 for further processing or stores the processed tasks in an L2 cache, local parallel processor memory, or system memory via a memory crossbar 2216. In at least one embodiment, a pre-ROP 2242 (pre-raster operation unit) is configured to receive data from the graphics multi-processor 2234 and direct the data to a ROP unit, which may be located within a partitioning unit (e.g., partitioning units 2220A - 2220N of FIG. 22) as described herein. In at least one embodiment, the pre-ROP 2242 unit may perform optimizations for color blending, organize pixel color data, and perform address translation.

[0286] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, inference and / or training logic 915 may be used in the graphics processing cluster 2214 for inference or prediction operations, based at least in part on weight parameters calculated using the training operations, functions and / or architectures of neural networks described herein, or use cases of neural networks. In at least one embodiment, a graphics multiprocessor as shown in FIG. 22D may be used to implement workers operating in parallel to train a neural network as described above.

[0287] FIG. 22D shows a graphics multiprocessor 2234 according to at least one embodiment. In at least one embodiment, graphics multiprocessor 2234 is coupled to a pipeline manager 2232 of processing cluster 2214. In at least one embodiment, graphics multiprocessor 2234 has an execution pipeline that includes, but is not limited to, an instruction cache 2252, an instruction unit 2254, an address mapping unit 2256, a register file 2258, one or more general purpose graphics processing unit (GPGPU) cores 2262, and one or more load / store units 2266. GPGPU cores 2262 and load / store units 2266 are coupled to cache memory 2272 and shared memory 2270 via a memory and cache interconnect 2268.

[0288] In at least one embodiment, the instruction cache 2252 receives a stream of instructions to be executed from the pipeline manager 2232. In at least one embodiment, the instructions are cached in the instruction cache 2252 and dispatched to be executed by the instruction unit 2254. In at least one embodiment, the instruction unit 2254 can dispatch instructions as a thread group (e.g., a warp), and each thread of the thread group is assigned to a different execution unit within the GPGPU core 2262. In at least one embodiment, instructions can access any of the local, shared, or global address spaces by specifying an address within the unified address space. In at least one embodiment, the address mapping unit 2256 can be used to translate an address in the unified address space to an individual memory address that the load / store unit 2266 can access.

[0289] In at least one embodiment, the register file 2258 provides a set of registers to the functional units of the graphics multiprocessor 2234. In at least one embodiment, the register file 2258 provides temporary storage for operands connected to the data paths of the functional units (e.g., GPGPU core 2262, load / store unit 2266) of the graphics multiprocessor 2234. In at least one embodiment, the register file 2258 is divided among the respective functional units such that each functional unit is allocated a dedicated portion of the register file 2258. In one embodiment, the register file 2258 is divided among different warps being executed by the graphics multiprocessor 2234.

[0290] In at least one embodiment, each of the GPGPU cores 2262 can include a floating-point unit (FPU) and / or an integer arithmetic logic unit (ALU) that are used to execute the instructions of the graphics multiprocessor 2234. The GPGPU cores 2262 can have the same architecture or different architectures. In at least one embodiment, a first portion of the GPGPU core 2262 includes a single-precision FPU and an integer ALU, and a second portion of the GPGPU core includes a double-precision FPU. In at least one embodiment, the FPU can implement the IEEE 754-2008 standard for floating-point operations or enable variable-precision floating-point operations. In at least one embodiment, the graphics multiprocessor 2234 can further include one or more fixed-function units or special-function units for performing specific functions such as rectangle copy or pixel blending operations. In at least one embodiment, one or more of the GPGPU cores can also include fixed or special-function logic.

[0291] In at least one embodiment, the GPGPU core 2262 includes SIMD logic that can execute a single instruction on multiple data sets. In at least one embodiment, the GPGPU core 2262 can physically execute SIMD4, SIMD8, and SIMD16 instructions and can logically execute SIMD1, SIMD2, and SIMD32 instructions. In at least one embodiment, the SIMD instructions for the GPGPU core may be generated during compilation by a shader compiler or may be automatically generated when executing a program written and compiled for a single program multiple data (SPMD) or SIMT architecture. In at least one embodiment, multiple threads of a program configured for the SIMT execution model can be executed via a single SIMD instruction. For example, in at least one embodiment, eight SIMT threads performing the same or similar operations can be executed in parallel via a single SIMD8 logical unit.

[0292] In at least one embodiment, the memory and cache interconnect 2268 is an interconnect network that connects each functional unit of the graphics multiprocessor 2234 to the register file 2258 and the shared memory 2270. In at least one embodiment, the memory and cache interconnect 2268 is a crossbar interconnect that enables the load / store unit 2266 to implement load and store operations between the shared memory 2270 and the register file 2258. In at least one embodiment, the register file 2258 can operate at the same frequency as the GPGPU core 2262, and thus, the data transfer between the GPGPU core 2262 and the register file 2258 is very low latency. In at least one embodiment, the shared memory 2270 can be used to enable communication between threads executed by functional units within the graphics multiprocessor 2234. In at least one embodiment, the cache memory 2272 can be used, for example, as a data cache to cache texture data communicated between the functional units and the texture unit 2236. In at least one embodiment, the shared memory 2270 can also be used as a program management cache. In at least one embodiment, threads executing on the GPGPU core 2262 can store data programmatically in the shared memory in addition to automatically cached data stored in the cache memory 2272.

[0293] In at least one embodiment, the parallel processor or GPGPU described herein is communicatively coupled to a host / processor core to accelerate graphics operations, machine learning operations, pattern analysis operations, and various general-purpose GPU (GPGPU) functions. In at least one embodiment, the GPU may be communicatively coupled to the host processor / core via a bus or other interconnect (e.g., a high-speed interconnect such as PCIe or NVLink). In at least one embodiment, the GPU may be integrated into the same package or die as the core and communicatively coupled to the core via an internal (i.e., internal to the package or die) processor bus / interconnect. In at least one embodiment, regardless of the method of connecting the GPU, the processor core may distribute work to the GPU in the form of a sequence of commands / instructions included in a work descriptor. In at least one embodiment, the GPU then uses dedicated circuitry / logic to efficiently process these commands / instructions.

[0294] In order to perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIG. 9A and / or FIG. 9B. In at least one embodiment, inference and / or training logic 915 may be used in graphics multiprocessor 2234 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks. In at least one embodiment, a multi-GPU computing system as shown in FIG. 23 may be used to implement workers operating in parallel to train a neural network as described above.

[0295] FIG. 23 shows a multi-GPU computing system 2300 according to at least one embodiment. In at least one embodiment, the multi-GPU computing system 2300 can include a processor 2302 coupled to a plurality of general-purpose graphics processing units (GPGPUs) 2306A-D via a host interface switch 2304. In at least one embodiment, the host interface switch 2304 is a PCI Express switch device that couples the processor 2302 to a PCI Express bus, via which the processor 2302 can communicate with the GPGPUs 2306A-D. The GPGPUs 2306A-D can be interconnected via a set of high-speed point-to-point GPU-to-GPU links 2316. In at least one embodiment, the GPU-to-GPU link 2316 is connected to each of the GPGPUs 2306A-D via a dedicated GPU link. In at least one embodiment, the P2P GPU link 2316 enables direct communication between each of the GPGPUs 2306A-D without requiring communication via the host interface bus 2304 to which the processor 2302 is connected. In at least one embodiment, when there is GPU-to-GPU traffic destined for the P2P GPU link 2316, the host interface bus 2304 is kept available to access system memory or to communicate with other instances of the multi-GPU computing system 2300, for example, via one or more network devices. In at least one embodiment, the GPGPUs 2306A-D are connected to the processor 2302 via the host interface switch 2304, and in at least one embodiment, the processor 2302 includes direct support for the P2P GPU link 2316 and can be directly connected to the GPGPUs 2306A-D.

[0296] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the multi-GPU computing system 2300 for inference or prediction operations, based at least in part on the training operations of the neural networks described herein, the functions and / or architectures of the neural networks, or the weight parameters calculated using the use cases of the neural networks. In at least one embodiment, a worker operating in parallel to train a neural network as described above can be implemented using a graphics processor as shown in FIG. 24.

[0297] FIG. 24 is a block diagram of a graphics processor 2400 according to at least one embodiment. In at least one embodiment, the graphics processor 2400 includes a ring interconnect 2402, a pipeline front end 2404, a media engine 2437, and graphics cores 2480A - 2480N. In at least one embodiment, the ring interconnect 2402 couples the graphics processor 2400 to other graphics processors or other processing units including one or more general-purpose processor cores. In at least one embodiment, the graphics processor 2400 is one of a number of processors integrated within a multi-core processing system.

[0298] In at least one embodiment, the graphics processor 2400 receives a batch of commands via the ring interconnect 2402. In at least one embodiment, incoming commands are interpreted by the command streamer 2403 of the pipeline front end 2404. In at least one embodiment, the graphics processor 2400 includes scalable execution logic for performing 3D geometry processing and media processing via the graphics cores 2480A - 2480N. In at least one embodiment, for 3D geometry processing commands, the command streamer 2403 supplies the commands to the geometry pipeline 2436. In at least one embodiment, for at least some media processing commands, the command streamer 2403 supplies the commands to the video front end 2434, and the video front end 2434 is coupled to the media engine 2437. In at least one embodiment, the media engine 2437 includes a Video Quality Engine (VQE) 2430 for post - processing of video and images, and a multi - format encode / decode (MFX) 2433 engine that provides hardware - accelerated encoding and decoding of media data. In at least one embodiment, the geometry pipeline 2436 and the media engine 2437 each generate execution threads for the thread execution resources provided by at least one graphics core 2480A.

[0299] In at least one embodiment, the graphics processor 2400 includes a scalable thread execution resource characterized by modular cores 2480A - 2480N (which may also be referred to as core slices), and each modular core has a plurality of sub - cores 2450A - 2450N, 2460A - 2460N (which may also be referred to as core sub - slices). In at least one embodiment, the graphics processor 2400 can have any number of graphics cores 2480A - 2480N. In at least one embodiment, the graphics processor 2400 includes a graphics core 2480A having at least a first sub - core 2450A and a second sub - core 2460A. In at least one embodiment, the graphics processor 2400 is a low - power processor having a single sub - core (e.g., 2450A). In at least one embodiment, the graphics processor 2400 includes a plurality of graphics cores 2480A - 2480N, each of which includes a first set of sub - cores 2450A - 2450N and a second set of sub - cores 2460A - 2460N. In at least one embodiment, each sub - core of the first set of sub - cores 2450A - 2450N includes at least an execution unit 2452A - 2452N and a first set of media / texture samplers 2454A - 2454N. In at least one embodiment, each sub - core of the second set of sub - cores 2460A - 2460N includes at least an execution unit 2462A - 2462N and a second set of samplers 2464A - 2464N. In at least one embodiment, each of the sub - cores 2450A - 2450N, 2460A - 2460N shares a set of shared resources 2470A - 2470N. In at least one embodiment, the shared resources include a shared cache memory and pixel operation logic.

[0300] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, the inference and / or training logic 915 may be used in the graphics processor 2400 for inference or prediction operations, at least partially based on weight parameters calculated using the training operations, functions and / or architectures of neural networks described herein, or use cases of neural networks. In at least one embodiment, a worker operating in parallel to train a neural network as described above can be implemented using a microarchitecture as shown in FIG. 25.

[0301] FIG. 25 is a block diagram showing a microarchitecture of a processor 2500 that may include logic circuitry for executing instructions, according to at least one embodiment. In at least one embodiment, the processor 2500 may execute instructions including x86 instructions, AMR instructions, special instructions for application specific integrated circuits (ASICs), etc. In at least one embodiment, the processor 2510 may include registers for storing packed data, such as 64-bit wide MMX (trademark) registers in a microprocessor enabled with MMX technology by Intel Corporation of Santa Clara, California. In at least one embodiment, MMX registers available in both integer and floating point formats may operate on packed data elements with single instruction multiple data (“SIMD”) and streaming SIMD extensions (“SSE”) instructions. In at least one embodiment, 128-bit wide XMM registers related to SSE2, SSE3, SSE4, AVX, or more (collectively referred to as “SSEx”) technologies may hold operands of such packed data. In at least one embodiment, the processor 2510 may execute instructions to accelerate machine learning or deep learning algorithms, training, or inference.

[0302] In at least one embodiment, the processor 2500 includes an in-order front end (“front end”) 2501 that fetches instructions to be executed and prepares instructions for later use in the processor pipeline. In at least one embodiment, the front end 2501 may include several units. In at least one embodiment, an instruction prefetcher 2526 fetches instructions from memory and supplies the instructions to an instruction decoder 2528, and the instruction decoder decodes or interprets the instructions. For example, in at least one embodiment, the instruction decoder 2528 decodes the received instruction into one or more operations called “microinstructions” or “micro-operations” that the machine can execute (also called “micro-ops” or “uops”). In at least one embodiment, the instruction decoder 2528 parses the instruction into an opcode and corresponding data, as well as a control field, such that these are used by the microarchitecture and the operations according to at least one embodiment may be executed. In at least one embodiment, a trace cache 2530 may assemble the decoded uops into a program-order sequence or trace in a uop queue 2534 for execution. In at least one embodiment, when the trace cache 2530 encounters a complex instruction, a microcode ROM 2532 provides the uops necessary for the completion of the operation.

[0303] In at least one embodiment, there are instructions that can be translated into a single micro-op, and there are also instructions that require several micro-ops to complete all operations. In at least one embodiment, if five or more micro-ops are required to complete an instruction, the instruction decoder 2528 may access the microcode ROM 2532 to execute the instruction. In at least one embodiment, an instruction may be decoded into a small number of micro-ops so that it can be processed in the instruction decoder 2528. In at least one embodiment, if a large number of micro-ops are required to complete an operation, the instruction may be stored in the microcode ROM 2532. In at least one embodiment, the trace cache 2530 determines the correct micro-instruction pointer for reading the microcode sequence by referring to an entry point programmable logic array ("PLA") to complete one or more instructions from the microcode ROM 2532 according to at least one embodiment. In at least one embodiment, after the microcode ROM 2532 finishes sequencing the micro-ops for an instruction, the front end 2501 of the machine may resume fetching micro-ops from the trace cache 2530.

[0304] In at least one embodiment, an out-of-order execution engine ("out-of-order engine") 2503 may prepare instructions for execution. In at least one embodiment, out-of-order execution logic has multiple buffers to smooth the flow of instructions and change their order, optimizing performance when instructions are scheduled to flow down a pipeline and be executed. The out-of-order execution engine 2503 includes, without limitation, an allocator / register renamer 2540, a memory uop queue 2542, an integer / floating point uop queue 2544, a memory scheduler 2546, a fast scheduler 2502, a slow / general purpose floating point scheduler ("slow / general purpose FP scheduler") 2504, and a simple floating point scheduler ("simple FP scheduler") 2506. In at least one embodiment, the fast scheduler 2502, the slow / general purpose floating point scheduler 2504, and the simple floating point scheduler 2506 are also collectively referred to herein as "uop schedulers 2502, 2504, 2506". The allocator / register renamer 2540 allocates the machine buffers and resources required by each uop for execution. In at least one embodiment, the allocator / register renamer 2540 changes the name of the logical register upon entry into the register file. In at least one embodiment, the allocator / register renamer 2540 also distributes the entry of each uop to one of two uop queues, namely the memory uop queue 2542 for memory operations and the integer / floating point uop queue 2544 for non-memory operations, ahead of the memory scheduler 2546 and the uop schedulers 2502, 2504, 2506. In at least one embodiment, the uop schedulers 2502, 2504, 2506 determine when uops are ready for execution based on the availability of the sources of their dependent input register operands and the execution resources required by the uop to complete their operations.In at least one embodiment, the high-speed scheduler 2502 of at least one embodiment may schedule every half of the main clock cycle, and the low-speed / general-purpose floating-point scheduler 2504 and the simple floating-point scheduler 2506 may schedule once per clock cycle of the main processor. In at least one embodiment, the uop schedulers 2502, 2504, 2506 arbitrate dispatch ports to schedule uops for execution.

[0305] In at least one embodiment, the execution block b11 includes, without limitation, an integer register file / bypass network 2508, a floating-point register file / bypass network (referred to herein as the "FP register file / bypass network") 2510, address generation units (AGUs) 2512 and 2514, high-speed arithmetic logic units (ALUs) (referred to herein as "high-speed ALUs") 2516 and 2518, low-speed arithmetic logic units (referred to herein as "low-speed ALUs") 2520, floating-point ALUs (FPs) 2522, and floating-point move units (referred to herein as "FP moves") 2524. In at least one embodiment, the integer register file / bypass network 2508 and the floating-point register file / bypass network 2510 are also referred to herein as the "register files 2508, 2510". In at least one embodiment, the AGUs 2512 and 2514, the high-speed ALUs 2516 and 2518, the low-speed ALUs 2520, the floating-point ALUs 2522, and the floating-point move units 2524 are also referred to herein as the "execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524". In at least one embodiment, the execution block b11 may include any number and type of register files, bypass networks, address generation units, and execution units (including zero) in any combination without limitation.

[0306] In at least one embodiment, register files 2508, 2510 may be disposed between uop schedulers 2502, 2504, 2506 and execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524. In at least one embodiment, integer register file / bypass network 2508 performs integer operations. In at least one embodiment, floating-point register file / bypass network 2510 performs floating-point operations. In at least one embodiment, each of register files 2508, 2510 may include, without limitation, a bypass network that may bypass or transfer a just-completed result not yet written to the register file to new dependent uops. In at least one embodiment, register files 2508, 2510 may communicate data with each other. In at least one embodiment, integer register file / bypass network 2508 may include, without limitation, two separate register files, namely, one register file for lower 32-bit data and a second register file for upper 32-bit data. In at least one embodiment, since floating-point instructions typically have operands with a width of 64 to 128 bits, floating-point register file / bypass network 2510 may include, without limitation, entries with a width of 128 bits.

[0307] In at least one embodiment, execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524 may execute instructions. In at least one embodiment, register files 2508 and 2510 store integer and floating-point data operand values required by microinstructions to execute. In at least one embodiment, processor 2500 may include any number and combination of execution units 2512, 2514, 2516, 2518, 2520, 2522, and 2524, without limitation. In at least one embodiment, floating-point ALU 2522 and floating-point move unit 2524 may execute floating-point, MMX, SIMD, AVX, and other operations, including SEE, or special machine learning instructions. In at least one embodiment, the floating-point ALU 2522 may include, without limitation, a 64-bit floating-point divider to perform division, square root, and remaining micro-ops. In at least one embodiment, instructions involving floating-point values may be handled by floating-point hardware. In at least one embodiment, ALU operations may be passed to the high-speed ALUs 2516, 2518. In at least one embodiment, the high-speed ALUs 2516, 2518 may perform high-speed operations with an effective latency of half a clock cycle. In at least one embodiment, the low-speed ALU 2520 may include, without limitation, integer execution hardware for long-latency type operations such as multipliers, shifts, flag logic, and branching, with most complex integer operations proceeding to the low-speed ALU 2520. In at least one embodiment, memory load / store operations may be performed by the ALUs 2512, 2514. In at least one embodiment, fast ALU 2516, fast ALU 2518, and slow ALU 2520 may perform integer operations on 64-bit data operands. In at least one embodiment, fast ALU 2516, fast ALU 2518, and slow ALU 2520 may be implemented to support various data bit sizes, including 16, 32, 128, 256, etc. In at least one embodiment, floating-point ALU 2522 and floating-point move unit 2524 may be implemented to support wide operands having various bit widths.In at least one embodiment, floating-point ALU 2522 and floating-point move unit 2524 may operate on 128-bit wide packed data operands in conjunction with SIMD and multimedia instructions.

[0308] In at least one embodiment, the uop schedulers 2502, 2504, 2506 dispatch dependent operations before the parent load finishes execution. In at least one embodiment, because uops may be speculatively scheduled and executed in the processor 2500, the processor 2500 may also include logic to handle memory misses. In at least one embodiment, if a data load misses in the data cache, there may be dependent operations in progress in the pipeline past the scheduler that have temporarily incorrect data. In at least one embodiment, a replay mechanism tracks and re-executes instructions that use the incorrect data. In at least one embodiment, the dependent operations may need to be replayed, and the independent operations may be allowed to complete. In at least one embodiment, the scheduler and replay mechanism of at least one embodiment of a processor may also be designed to capture instruction sequences for text string comparison operations.

[0309] In at least one embodiment, the term "register" may refer to a storage location of an on-board processor that can be used as part of an instruction to identify an operand. In at least one embodiment, a register may be one that can be used from outside the processor (from the perspective of a programmer). In at least one embodiment, a register may not be limited to a particular type of circuit. Rather, in at least one embodiment, a register may store data, provide data, and perform the functions described herein. In at least one embodiment, the registers described herein may be implemented by circuits within a processor using any number of different techniques, such as dedicated physical registers, physical registers dynamically allocated using register renaming, a combination of dedicated physical registers and physically registers dynamically allocated, and the like. In at least one embodiment, an integer register stores 32-bit integer data. The register file of at least one embodiment also includes eight multimedia SIMD registers for packed data.

[0310] In order to perform inference and / or training operations related to one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, some or all of inference and / or training logic 915 may be incorporated into EXE block 2511 and other memories or registers shown or not shown. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs shown in EXE block 2511. Additionally, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that make up the ALU of EXE block 2511 for performing one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0311] FIG. 26 shows a deep learning application processor 2600 according to at least one embodiment. In at least one embodiment, the deep learning application processor 2600 uses instructions that, when executed by the deep learning application processor 2600, cause the deep learning application processor 2600 to execute some or all of the processes and techniques described throughout this disclosure. In at least one embodiment, the deep learning application processor 2600 is an application specific integrated circuit (ASIC). In at least one embodiment, the application processor 2600 executes matrix multiplication operations that are "hard-wired" to hardware as a result of executing one or more instructions or both. In at least one embodiment, the deep learning application processor 2600 includes, without limitation, processing clusters 2610(1)-2610(12), inter-chip links ("ICLs") 2620(1)-2620(12), inter-chip controllers ("ICCs") 2630(1)-2630(2), high bandwidth memory second generation ("HBM2") 2640(1)-2640(4), memory controllers ("Mem Ctrlrs") 2642(1)-2642(4), high bandwidth memory physical layer ("HBM PHY") 2644(1)-2644(4), management-controller central processing unit ("management-controller CPU") 2650, serial peripheral interface, inter-integrated circuit, and general purpose input / output blocks ("SPI, I2C, GPIO") 2660, peripheral component interconnect express controller and direct memory access block ("PCIe controller and DMA") 2670, and 16-lane peripheral component interconnect express port ("PCI Expressx16") 2680.

[0312] In at least one embodiment, processing cluster 2610 may perform deep learning operations including inference or prediction operations based on weight parameters calculated using one or more training techniques including the techniques described herein. In at least one embodiment, each processing cluster 2610 may include any number and type of processors, without limitation. In at least one embodiment, deep learning application processor 2600 may include any number and type of processing clusters 2600. In at least one embodiment, inter-chip link 2620 is bi-directional. In at least one embodiment, inter-chip link 2620 and inter-chip controller 2630 enable multiple deep learning application processors 2600 to exchange information including activation information resulting from execution of one or more machine learning algorithms embodied in one or more neural networks. In at least one embodiment, deep learning application processor 2600 may include any number and type (including zero) of ICLs 2620 and ICCs 2630.

[0313] In at least one embodiment, HBM2 2640 provides a total of 32 gigabytes (GB) of memory. HBM2 2640(i) is associated with both memory controller 2642(i) and HBM PHY 2644(i). In at least one embodiment, any number of HBM2s 2640 may provide any type and total amount of high bandwidth memory and may be associated with any number and type (including zero) of memory controllers 2642 and HBM PHYs 2644. In at least one embodiment, SPI, I2C, GPIO 2660, PCIe controller, and DMA 2670, and / or PCIe 2680 may be replaced with any number and type of blocks enabling any number and type of communication standards in any technically feasible manner.

[0314] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIG. 9A and / or FIG. 9B. In at least one embodiment, the deep learning application processor 2600 is used to train a machine learning model, such as a neural network, to predict or infer information provided to the deep learning application processor 2600. In at least one embodiment, the deep learning application processor 2600 is used to infer or predict information based on a trained machine learning model (e.g., a neural network) that has been trained by another processor or system or by the deep learning application processor 2600 itself. In at least one embodiment, the processor 2600 may be used to execute one or more of the use cases of the neural networks described herein. In at least one embodiment, a neuromorphic processor as shown in FIG. 27 can be used to implement workers that operate in parallel to train a neural network as described above.

[0315] FIG. 27 is a block diagram of a neuromorphic processor 2700 according to at least one embodiment. In at least one embodiment, the neuromorphic processor 2700 receives one or more inputs from a source external to the neuromorphic processor 2700. In at least one embodiment, these inputs may be sent to one or more neurons 2702 within the neuromorphic processor 2700. In at least one embodiment, the neurons 2702 and their components may be implemented using circuitry or logic that includes one or more arithmetic logic units (ALUs). In at least one embodiment, the neuromorphic processor 2700 may include thousands or millions of instances of neurons 2702, without limitation, although any suitable number of neurons 2702 may be used. In at least one embodiment, each instance of a neuron 2702 may include a neuron input 2704 and a neuron output 2706. In at least one embodiment, a neuron 2702 may generate an output, and this output may be sent to the inputs of other instances of neurons 2702. For example, in at least one embodiment, the neuron inputs 2704 and the neuron outputs 2706 may be interconnected via synapses 2708.

[0316] In at least one embodiment, neurons 2702 and synapses 2708 may be interconnected such that neuromorphic processor 2700 operates to process or analyze information received by neuromorphic processor 2700. In at least one embodiment, neuron 2702 may send an output pulse (or "fire" or "spike") when input received via neuron input 2704 exceeds a threshold. In at least one embodiment, neuron 2702 may sum or integrate signals received at neuron input 2704. For example, in at least one embodiment, neuron 2702 may be implemented as a leaky integrate-and-fire neuron, where if the sum (referred to as the "membrane potential") exceeds a threshold, neuron 2702 may generate an output (or "fire") using a transfer function such as a sigmoid function or a threshold function. In at least one embodiment, a leaky integrate-and-fire neuron may sum signals received at neuron input 2704 into a membrane potential and may apply a decay factor (or leakage) to reduce the membrane potential. In at least one embodiment, a leaky integrate-and-fire neuron may fire if multiple input signals are received at neuron input 2704 quickly enough to exceed a threshold (i.e., before the membrane potential decays too little to cause firing). In at least one embodiment, neuron 2702 may be implemented using circuitry or logic that receives inputs, integrates the inputs into a membrane potential, and decays the membrane potential. In at least one embodiment, the inputs may be averaged, or any other suitable transfer function may be used. Further, in at least one embodiment, neuron 2702 may include, without limitation, comparator circuitry or logic that generates an output spike at neuron 2706 when the result of applying the transfer function to neuron 2704 exceeds a threshold. In at least one embodiment, neuron 2702 may ignore previously received input information when firing, for example, by resetting the membrane potential to 0 or another suitable default value.In at least one embodiment, when the membrane potential is reset to 0, neuron 2702 may resume normal operation after a suitable period (or refractory period).

[0317] In at least one embodiment, neurons 2702 may be interconnected through synapses 2708. In at least one embodiment, synapses 2708 may be operative to transmit signals from the output of a first neuron 2702 to the input of a second neuron 2702. In at least one embodiment, neurons 2702 may transmit information via more than one instance of synapses 2708. In at least one embodiment, one or more instances of neuron outputs 2706 may be connected via an instance of synapses 2708 to an instance of neuron inputs 2704 of the same neuron 2702. In at least one embodiment, an instance of neuron 2702 that generates an output that will be transmitted via an instance of synapses 2708 may be referred to as a “presynaptic neuron” with respect to that instance of synapses 2708. In at least one embodiment, an instance of neuron 2702 that receives an input that will be transmitted via an instance of synapses 2708 may be referred to as a “postsynaptic neuron” with respect to that instance of synapses 2708. In at least one embodiment, an instance of neuron 2702 may receive inputs from one or more instances of synapses 2708 and may also transmit outputs via one or more instances of synapses 2708, so a single instance of neuron 2702 may thus be both a “presynaptic neuron” and a “postsynaptic neuron” with respect to various instances of synapses 2708.

[0318] In at least one embodiment, neurons 2702 may be organized into one or more layers. Each instance of neuron 2702 may have one neuron output 2706 that can fan out to one or more neuron inputs 2704 through one or more synapses 2708. In at least one embodiment, the neuron output 2706 of neurons 2702 in the first layer 2710 may be connected to the neuron inputs 2704 of neurons 2702 in the second layer 2712. In at least one embodiment, layer 2710 may be referred to as a “feed-forward” layer. In at least one embodiment, each instance of neuron 2702 in an instance of the first layer 2710 may fan out to each instance of neuron 2702 in the second layer 2712. In at least one embodiment, the first layer 2710 may be referred to as a “fully-connected feed-forward layer”. In at least one embodiment, each instance of neuron 2702 in an instance of the second layer 2712 may fan out to fewer instances of neuron 2702 in the third layer 2714 than all instances of neuron 2702 in the third layer 2714. In at least one embodiment, the second layer 2712 may be referred to as a “sparsely-connected feed-forward layer”. In at least one embodiment, neurons 2702 in the second layer 2712 may fan out to neurons 2702 in a plurality of other layers, including neurons 2702 in the (same) second layer 2712. In at least one embodiment, the second layer 2712 may be referred to as a “recurrent layer”. The neuromorphic processor 2700 may include any suitable combination of recurrent layers and feed-forward layers, including but not limited to both sparsely-connected feed-forward layers and fully-connected feed-forward layers.

[0319] In at least one embodiment, the neuromorphic processor 2700 may include, without limitation, a reconfigurable interconnect architecture for connecting synapses 2708 to neurons 2702, or dedicated hard-wired interconnects. In at least one embodiment, the neuromorphic processor 2700 may include, without limitation, circuitry or logic that can distribute synapses to different neurons 2702 as needed, based on a neural network topology and the fan-in / fan-out of the neurons. For example, in at least one embodiment, the synapses 2708 may be connected to the neurons 2702 using an interconnect fabric such as a network-on-chip or using dedicated connections. In at least one embodiment, the synapse interconnects and their components may be implemented using circuitry or logic.

[0320] FIG. 28 is a block diagram of a processing system according to at least one embodiment. In at least one embodiment, the system 2800 includes one or more processors 2802 and one or more graphics processors 2808 and may be a single-processor desktop system, a multi-processor workstation system, or a server system having a number of processors 2802 or processor cores 2807. In at least one embodiment, the system 2800 is a processing platform integrated within a system-on-chip (SoC) integrated circuit for use in mobile devices, portable devices, or embedded devices.

[0321] In at least one embodiment, system 2800 may include, or be incorporated in, a server-based gaming platform, a game console including a game and media console, a mobile gaming console, a portable gaming console, or an online gaming console. In at least one embodiment, system 2800 is a mobile phone, a smartphone, a tablet computing device, or a mobile Internet device. In at least one embodiment, processing system 2800 may also include, be coupled to, or be integrated within wearable devices such as smartwatch wearable devices, smart eyewear devices, augmented reality devices, or virtual reality devices. In at least one embodiment, processing system 2800 is a television or set-top box device having one or more processors 2802 and a graphical interface generated by one or more graphics processors 2808.

[0322] In at least one embodiment, each of the one or more processors 2802 includes one or more processor cores 2807 for processing instructions that, when executed, perform operations for system and user software. In at least one embodiment, each of the one or more processor cores 2807 is configured to process a particular instruction set 2809. In at least one embodiment, instruction set 2809 may facilitate computing via a complex instruction set computing (CISC), reduced instruction set computing (RISC), or very long instruction word (VLIW). In at least one embodiment, the processor cores 2807 may each process different instruction sets 2809, which may include instructions that facilitate emulation of other instruction sets. In at least one embodiment, processor core 2807 may also include other processing devices such as a digital signal processor (DSP).

[0323] In at least one embodiment, processor 2802 includes cache memory 2804. In at least one embodiment, processor 2802 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory is shared among various components of processor 2802. In at least one embodiment, processor 2802 also uses an external cache (e.g., a level 3 (L3) cache or a last level cache (LLC)) (not shown), and this cache may be shared among processor cores 2807 using known cache coherence techniques. In at least one embodiment, a register file 2806 is further included in processor 2802, and this register file may include different types of registers (e.g., integer registers, floating point registers, status registers, and instruction pointer registers) for storing different types of data. In at least one embodiment, register file 2806 may include general-purpose registers or other registers.

[0324] In at least one embodiment, one or more processors 2802 are coupled to one or more interface buses 2810 to transmit communication signals such as address, data, or control signals between the processor 2802 and other components within the system 2800. In at least one embodiment, the interface bus 2810 can be a processor bus such as a version of a Direct Media Interface (DMI) bus in one embodiment. In at least one embodiment, the interface 2810 is not limited to the DMI bus and may include one or more peripheral component interconnect buses (e.g., PCI, PCI Express), a memory bus, or other types of interface buses. In at least one embodiment, the processor 2802 includes an integrated memory controller 2816 and a platform controller hub 2830. In at least one embodiment, the memory controller 2816 facilitates communication between the memory device and other components of the system 2800, while the platform controller hub (PCH) 2830 provides connections to I / O devices via a local I / O bus.

[0325] In at least one embodiment, the memory device 2820 can be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, a phase change memory device, or any other memory device having suitable performance to serve as a process memory. In at least one embodiment, the memory device 2820 operates as a system memory for the system 2800 and can store data 2822 and instructions 2821 for use when one or more processors 2802 execute an application or process. In at least one embodiment, the memory controller 2816 is also coupled to an optional external graphics processor 2812, and this graphics processor may communicate with one or more graphics processors 2808 within the processor 2802 to perform graphics and media operations. In at least one embodiment, the display device 2811 can be connected to the processor 2802. In at least one embodiment, the display device 2811 can include one or more of an internal display device such as a mobile electronic device or a laptop device, or an external display device attached via a display interface (e.g., a display port, etc.). In at least one embodiment, the display device 2811 can include a head-mounted display (HMD) such as a stereoscopic display device for use in a virtual reality (VR) application or an augmented reality (AR) application.

[0326] In at least one embodiment, the platform controller hub 2830 enables peripheral devices to be connected to the memory device 2820 and the processor 2802 via a high-speed I / O bus. In at least one embodiment, the I / O peripheral devices include, but are not limited to, an audio controller 2846, a network controller 2834, a firmware interface 2828, a wireless transceiver 2826, a touch sensor 2825, and a data storage device 2824 (e.g., a hard disk drive, flash memory, etc.). In at least one embodiment, the data storage device 2824 can be connected via a storage interface (e.g., SATA) or a peripheral bus such as a Peripheral Component Interconnect bus (e.g., PCI, PCI Express). In at least one embodiment, the touch sensor 2825 can include a touch screen sensor, a pressure sensor, or a fingerprint sensor. In at least one embodiment, the wireless transceiver 2826 can be a WiFi transceiver, a Bluetooth transceiver, or a mobile network transceiver such as a 3G, 4G, or Long Term Evolution (LTE) transceiver. In at least one embodiment, the firmware interface 2828 enables communication with system firmware and can be, for example, a Unified Extensible Firmware Interface (UEFI). In at least one embodiment, the network controller 2834 can enable a network connection to a wired network. In at least one embodiment, a high-performance network controller (not shown) is coupled to the interface bus 2810. In at least one embodiment, the audio controller 2846 is a multi-channel high-definition audio controller. In at least one embodiment, the system 2800 includes an optional legacy I / O controller 2840 for coupling legacy (e.g., Personal System 2 (PS / 2)) devices to the system.In at least one embodiment, the platform controller hub 2830 can also be connected to connection input devices of one or more universal serial bus (USB) controllers 2842, such as a combination of a keyboard and a mouse 2843, a camera 2844, or other USB input devices.

[0327] In at least one embodiment, instances of the memory controller 2816 and the platform controller hub 2830 may be integrated into an individual external graphics processor such as the external graphics processor 2812. In at least one embodiment, the platform controller hub 2830 and / or the memory controller 2816 may be external to one or more processors 2802. For example, in at least one embodiment, the system 2800 can include an external memory controller 2816 and a platform controller hub 2830, which may be configured as a memory controller hub and a peripheral device controller hub in a system chipset that communicates with the processor 2802.

[0328] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding inference and / or training logic 915 are provided herein in conjunction with FIG. 9A and / or FIG. 9B. In at least one embodiment, some or all of inference and / or training logic 915 may be incorporated within graphics processor 2800. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in pipeline 2812. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIG. 9A or FIG. 9B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of graphics processor 2800 to perform one or more of the machine learning algorithms, neural network architectures, use cases, or training techniques described herein.

[0329] FIG. 29 is a block diagram of a processor 2900 having one or more processor cores 2902A - 2902N, an integrated memory controller 2914, and an integrated graphics processor 2908, according to at least one embodiment. In at least one embodiment, processor 2900 can include a lesser number of additional cores, including additional core 2902N represented by the dashed square. In at least one embodiment, each of processor cores 2902A - 2902N includes one or more internal cache units 2904A - 2904N. In at least one embodiment, each processor core can also access one or more shared cache units 2906.

[0330] In at least one embodiment, internal cache units 2904A - 2904N, and shared cache unit 2906 represent the cache memory hierarchy within processor 2900. In at least one embodiment, cache memory units 2904A - 2904N may include at least one level of cache for instructions and data within each processor core, as well as one or more levels of shared intermediate-level cache such as level 2 (L2), level 3 (L3), level 4 (L4), or other levels of cache, where the highest level of cache before external memory is classified as the LLC. In at least one embodiment, cache coherence logic maintains coherence among the various cache units 2906 and 2904A - 2904N.

[0331] In at least one embodiment, processor 2900 may also include a set of one or more bus controller units 2916 and system agent core 2910. In at least one embodiment, one or more bus controller units 2916 manage a set of peripheral buses such as one or more PCI or PCI Express buses. In at least one embodiment, system agent core 2910 provides management functions for the various processor components. In at least one embodiment, system agent core 2910 includes one or more integrated memory controllers 2914 for managing access to various external memory devices (not shown).

[0332] In at least one embodiment, one or more of the processor cores 2902A-2902N include support for simultaneous multithreading. In at least one embodiment, the system agent core 2910 includes components for coordinating and operating cores 2902A-2902N during multithreaded processing. In at least one embodiment, the system agent core 2910 may further include a power control unit (PCU), which includes logic and components for adjusting the power state of one or more of the processor cores 2902A-2902N and the graphics processor 2908.

[0333] In at least one embodiment, the processor 2900 further includes a graphics processor 2908 for performing graphics processing operations. In at least one embodiment, the graphics processor 2908 is coupled to a shared cache unit 2906 and a system agent core 2910 that includes one or more integrated memory controllers 2914. In at least one embodiment, the system agent core 2910 also includes a display controller 2911 for causing the output of the graphics processor to be provided to one or more attached displays. In at least one embodiment, the display controller 2911 may also be a separate module coupled to the graphics processor 2908 via at least one interconnect, or may be integrated within the graphics processor 2908.

[0334] In at least one embodiment, a ring-based interconnect unit 2912 is used to couple the internal components of the processor 2900. In at least one embodiment, alternative interconnect units such as point-to-point interconnects, switch interconnects, or other techniques may be used. In at least one embodiment, the graphics processor 2908 is coupled to the ring interconnect 2912 via an I / O link 2913.

[0335] In at least one embodiment, the I / O link 2913 represents at least one of a variety of I / O interconnects that facilitate communication between various processor components and high-performance embedded memory modules 2918, such as eDRAM modules, including an on-package I / O interconnect. In at least one embodiment, each of the processor cores 2902A-2902N and the graphics processor 2908 use the embedded memory module 2918 as a shared last-level cache.

[0336] In at least one embodiment, the processor cores 2902A-2902N are homogeneous cores that execute a common instruction set architecture. In at least one embodiment, the processor cores 2902A-2902N are heterogeneous from the perspective of the instruction set architecture (ISA), where one or more of the processor cores 2902A-2902N execute a common instruction set, but one or more other cores of the processor cores 2902A-2902N execute a subset of the common instruction set, or a different instruction set. In at least one embodiment, the processor cores 2902A-2902N are heterogeneous from the perspective of the microarchitecture, where one or more cores with relatively high power consumption are coupled with one or more cores with lower power consumption. In at least one embodiment, the processor 2900 can be implemented on one or more chips or as a SoC integrated circuit.

[0337] To perform inference and / or training operations associated with one or more embodiments, inference and / or training logic 915 is used. Details regarding the inference and / or training logic 915 are provided herein in conjunction with FIGS. 9A and / or 9B. In at least one embodiment, some or all of the inference and / or training logic 915 may be incorporated into the graphics processor 2910. For example, in at least one embodiment, the training and / or inference techniques described herein may use one or more of the ALUs embodied in the 3D pipeline 2812, the graphics core 2915A, the shared function logic 2916, the graphics core 2915B, the shared function logic 2920, or other logic of FIG. 29. Further, in at least one embodiment, the inference and / or training operations described herein may be performed using logic other than the logic shown in FIGS. 9A or 9B. In at least one embodiment, the weight parameters may be stored in on-chip or off-chip memory and / or registers (shown or not shown) that configure the ALU of the graphics processor 2910 to perform one or more of the machine learning algorithm...

Claims

1. A processor comprising a first processing core that executes software for training different parts of a neural network in parallel on two or more second processing cores, wherein the neural network is at least partially trained by generating a weight update by combining a plurality of partial weight updates created in parallel by the two or more second processing cores, wherein the plurality of partial weight updates are at least, identifying a plurality of subsets of network nodes of the neural network, creating a partial weight update for each subset in the plurality of subsets created by, wherein an individual subset of the plurality of subsets includes a quantity of node weights, wherein the quantity of the node weights is determined based at least in part on an amount of processing power available to a second processing core assigned to process the individual subset relative to another second processing core assigned to process another subset, a processor.

2. The processor according to claim 1, wherein the two or more second processing cores apply one or more gradients to different sets of nodes of the neural network.

3. The processor according to claim 1, wherein the processor further divides a weight update operation into a plurality of partial weight update operations and distributes individual partial weight update operations to the two or more second processing cores.

4. The processor according to claim 3, wherein the partial weight update operation is created by dividing an initial weight and a gradient update into a plurality of distinct parts.

5. The processor according to claim 3, wherein each partial weight update is executed using a different second processing core.

6. The processor according to claim 3, wherein the processor further collects the partial weight updates to create the weight update.

7. A system comprising a first processor that executes software for separately training different parts of a neural network in parallel on one or more second processors, and one or more memories for storing the neural network wherein the neural network, generates a plurality of partial weight updates in parallel using a plurality of workers of the one or more second processors, and combines the plurality of partial weight updates to create a weight gradient update ​ trained at least in part by The plurality of partial weight updates includes at least identifying a plurality of subsets of network nodes of the neural network; generating a partial weight update for each subset in the plurality of subsets; It is created by a respective subset of the plurality of subsets comprising a quantity of node weights; wherein the node weight quantity is determined based at least in part on the amount of processing power available to a worker assigned to process the respective subset relative to other workers assigned to process other subsets.

8. The system of claim 7 , wherein the partial weight updates are produced in parallel using the workers.

9. The neural network comprises: forward propagating inputs through said neural network to produce outputs; determining an error based at least in part on a difference between the output and a predicted value; backpropagating the error to determine a gradient on which the plurality of partial weight updates are at least partially based. The system of claim 7 is trained at least in part by:

10. 8. The system of claim 7, wherein the plurality of subsets are non-overlapping subsets of weights of the neural network.

11. the system determining a set of gradients for each input of a set of input values; The system of claim 7 , wherein the set of gradients is distributed to each of the one or more second processors.

12. 1. A method comprising: training a neural network by causing a first processor to execute software to at least partially train different portions of the neural network separately in parallel using a plurality of second processors, the method comprising: The neural network comprises: computing different portions of the weight updates in parallel using the plurality of second processors; aggregating the different portions of the weight updates to produce new weight values for the neural network. trained at least in part by The different portions of the weight updates include at least identifying a plurality of subsets of network nodes of the neural network; Creating each of the different portions of the weight update for each subset in the plurality of subsets Created by An individual subset of the plurality of subsets includes a quantity of node weights, The quantity of node weights is determined based at least in part on the amount of processing power available to a second processor assigned to process the individual subset relative to another second processor assigned to process another subset, a method. **Claim 13** The method of claim 12, wherein the neural network is trained at least in part by distributing gradient information to a plurality of workers of the plurality of second processors. **Claim 14** The neural network, Forward propagating an input through the neural network to produce an output, Determining an error based at least in part on the output, Determining a gradient based at least in part on the error by at least backpropagating the error, wherein the different portions of the weight update are based at least in part on the gradient The method of claim 13, wherein the neural network is trained at least in part by **Claim 15** The gradient is distributed to each of the plurality of workers, The method of claim 14, wherein the plurality of workers calculate the different portions of the weight update. **Claim 16** The method of claim 13, wherein each worker of the plurality of workers executes on a different second processor. **Claim 17** The method of claim 13, wherein each worker of the plurality of workers executes in parallel on a graphical processing unit. **Claim 18** The method of claim 12, wherein the number of the different portions of the weight update matches the number of one or more second processors available to a computer system. **Claim 19** The method of claim 12, wherein the weight update is divided into non-overlapping groups of node weights to create the different portions. **Claim 20** A speech processing system comprising a neural network that takes a digital representation of sound as input and identifies elements of human speech, wherein different portions of the neural network are trained separately and in parallel by a plurality of second processors by causing software to execute on a first processor to recognize human speech, a speech processing system, The neural network, generating a plurality of weight updates in parallel using the plurality of second processors, combining the plurality of weight updates into a single weight update and thereby at least partially training, wherein the plurality of weight updates at least identify a plurality of subsets of network nodes of the neural network, create weight updates for each subset in the plurality of subsets and are created by wherein an individual subset of the plurality of subsets includes a quantity of node weights, wherein the quantity of node weights is determined at least in part based on an amount of processing power available to a second processor assigned to process the individual subset relative to other second processors assigned to process other subsets, a speech processing system.

21. As a result of executable instructions stored in a memory in the speech processing system being executed by the plurality of second processors, the speech processing system is caused to, at least obtain data indicative of speech from a microphone, process the data using the neural network to identify spoken words represented by the data, perform an action at least in part based on the identification of the spoken words, the speech processing system according to claim 20.

22. The speech processing system according to claim 21, wherein the action is a navigation request processed by a navigation system of a vehicle.

Citation Information

Patent Citations

  • Methods and apparatus for training an artificial neural network for use in speech recognition

    US20150371132A1

  • Communication optimizations for distributed machine learning

    US20190205745A1

  • Transitioning between prior dialog contexts with automated assistants

    WO2019172878A1