Distributed derivation
By partitioning and distributing training data and activations across processors based on their positions, the method addresses inefficiencies in neural network training, achieving balanced workload and improved computational efficiency in distributed systems.
Patent Information
- Application Number
- DE102025101042
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-16
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-17
AI Technical Summary
Existing computer systems struggle to efficiently distribute and manage memory and computational resources when training neural networks, particularly in large-scale distributed systems, leading to inefficiencies and bottlenecks due to varying workload demands among processors.
The method involves partitioning training data and neural network activations across multiple processors, distributing contiguous portions of information based on their positions within the data sequence, and balancing workload by distributing sections from opposite ends of the dataset to ensure equal computational demands across processors, facilitating simultaneous communication and parallel processing.
This approach reduces memory bottlenecks and workload imbalances, enabling more efficient training of neural networks by ensuring each processor has a balanced computational load, thereby enhancing the training efficiency and reducing the need for frequent recalculations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
TECHNICAL FIELD
[0001] At least one embodiment relates to computing resources used to train and operate neural networks. For example, at least one embodiment relates to the distribution of input data to processors used to train and operate neural networks. BACKGROUND
[0002] Performing computational operations may require significant memory, time, or other resources. Computer programs can be organized so that different components can be executed in different ways, in different orders, and using a variety of computer systems. Despite advances in computer hardware that speed up or otherwise assist the execution of the various components of a computer program, these advances generally fail to account for the various ways computer programs can be structured and the various ways elements of computer programs can be distributed among computer systems. The memory, time, and / or computing resources used to perform computational operations can be improved. BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 is a block diagram illustrating a computer system in which partitioned training data and activations are used to train a neural network, according to at least one embodiment; Fig. 2 is a block diagram illustrating a computer system in which data and activations are partitioned according to partitioning criteria, in accordance with at least one embodiment; Fig. 3 is a block diagram illustrating a process for partitioning training data and activations of a neural network according to at least one embodiment; Fig. 4 is a block diagram illustrating forward and backward propagation in a neural network, according to at least one embodiment; Fig. 5 is a block diagram showing a first portion of a computer system in which partitioned training data and neural network activations are used to train a neural network, according to at least one embodiment; Fig. 6 is a block diagram showing a second portion of a computer system in which partitioned training data and neural network activations are used to train a neural network, according to at least one embodiment; Fig. 7 is a block diagram illustrating a system for splitting input data among multiple processors to distribute the workload among the processors, according to at least one embodiment; Fig. 8 is a block diagram illustrating how workload balancing is achieved in connection with self-attention computations according to at least one embodiment; Fig. 9 is a block diagram illustrating a process for calculating self-attention for an input sequence according to at least one embodiment; Fig. 10 is a visual representation of the workload of different processors as a result of distributing equal portions from opposite ends of the input sequence, according to at least one embodiment; Fig. 11 is a block diagram illustrating the use of partitioned training data and neural network activations, according to at least one embodiment; Fig. 12 is a block diagram showing a processor and modules according to at least one embodiment; Fig. 13A shows logic according to at least one embodiment; Fig. 13B shows the logic according to at least one embodiment; Fig. 14 shows the training and deployment of a neural network according to at least one embodiment; Fig. 15 shows an example of a data center system according to at least one embodiment; Fig. 16A shows an example of an autonomous vehicle according to at least one embodiment; Fig. Figure 16B shows an example of camera positions and fields of view for the autonomous vehicle of Fig. 16A according to at least one embodiment; Fig. 16C is a block diagram illustrating an example system architecture for the autonomous vehicle of Fig. 16A according to at least one embodiment; Fig. 16D is a diagram illustrating a system for communication between one or more cloud-based servers and the autonomous vehicle of Fig. 16A according to at least one embodiment; Fig. 17 is a block diagram illustrating a computer system according to at least one embodiment; Fig. 18 is a block diagram showing a computer system according to at least one embodiment; Fig. 19 shows a computer system according to at least one embodiment; Fig. 20 shows a computer system according to at least one embodiment; Fig. 21A shows a computer system according to at least one embodiment; Fig. 21B shows a computer system according to at least one embodiment; Fig. 21C shows a computer system according to at least one embodiment; Fig. 21D shows a computer system according to at least one embodiment; Fig. 21E and Fig. 21F show a common programming model according to at least one embodiment; Fig. 22 shows exemplary integrated circuits and associated graphics processors according to at least one embodiment; Fig. 23A and Fig. 23B illustrate exemplary integrated circuits and associated graphics processors according to at least one embodiment; Fig. 24A and Fig. 24B illustrate additional exemplary graphics processor logic according to at least one embodiment; Fig. 25 shows a computer system according to at least one embodiment; Fig. 26A shows a parallel processor according to at least one embodiment; Fig. 26B shows a partition unit according to at least one embodiment; Fig. 26C shows a processing cluster according to at least one embodiment; Fig. 26D shows a graphics multiprocessor according to at least one embodiment; Fig. 27 shows a system with multiple graphics processing units (GPU) according to at least one embodiment; Fig. 28 shows a graphics processor according to at least one embodiment; Fig. 29 is a block diagram illustrating a processor microarchitecture for a processor according to at least one embodiment; Fig. 30 shows a deep learning application processor according to at least one embodiment; Fig. 31 is a block diagram illustrating an exemplary neuromorphic processor according to at least one embodiment; Fig. 32 shows at least portions of a graphics processor according to one or more embodiments; Fig. 33 shows at least portions of a graphics processor according to one or more embodiments; Fig. 34 shows at least portions of a graphics processor according to one or more embodiments; Fig. 35 is a block diagram of a graphics processing engine of a graphics processor in accordance with at least one embodiment; Fig. 36 is a block diagram of at least portions of a graphics processor core according to at least one embodiment; Fig. 37A and Fig. 37B illustrates thread execution logic including an array of processing elements of a graphics processor core, according to at least one embodiment; Fig. 38 shows a parallel processing unit ("PPU") according to at least one embodiment; Fig. 39 shows a general processing cluster (“GPC”) according to at least one embodiment; Fig. 40 shows a memory partition unit of a parallel processing unit ("PPU") according to at least one embodiment; Fig. 41 shows a streaming multiprocessor according to at least one embodiment; Fig. 42 is an exemplary data flow diagram for an advanced computing pipeline, according to at least one embodiment; Fig. 43 is a system diagram for an example system for training, adapting, instantiating, and deploying machine learning models in an advanced computing pipeline, according to at least one embodiment; Fig. 44 includes an example illustration of an advanced computer pipeline 4310A for processing image data in accordance with at least one embodiment; Fig. 45A includes an example data flow diagram of a virtual instrument supporting an ultrasound device, according to at least one embodiment; Fig. 45B includes an example data flow diagram of a virtual instrument supporting a CT scanner, in accordance with at least one embodiment; Fig. 46A shows a data flow diagram for a process for training a machine learning model, in accordance with at least one embodiment; Fig. 46B is an example illustration of a client-server architecture for enhancing annotation tools with pre-trained annotation models, according to at least one embodiment; and Fig. 47 shows components of a system for accessing a large language model according to at least one embodiment. DETAILED DESCRIPTION
[0003] In the following description, numerous specific details are set forth in order to provide a more thorough understanding of at least one embodiment. However, one skilled in the art will appreciate that the inventive concepts may be practiced without one or more of these specific details.
[0004] In at least one embodiment, a neural network is trained using one or more of the operations and techniques described herein. In at least one embodiment, a neural network is trained in parallel using a distributed system having a plurality of processors on a multiprocessor system. In at least one embodiment, input data to a neural network is partitioned and distributed among a plurality of processors such that a processor can perform one or more operations to train a neural network in parallel using the partitioned input data. In at least one embodiment, activations of partitioned input data are also partitioned and distributed among a plurality of processors based at least in part on the partitioning of the input data.In at least one embodiment, communication between the processors is concurrent, such that, for example, a first processor performs a first computation on a first partition of input data and, concurrently while the first processor performs a second computation, the results of the first computation are indicated, sent, or otherwise provided to a second processor. In at least one embodiment, a processor comprising one or more circuits is used to distribute the inferencing of two or more related portions of information only between two or more respective processing cores or otherwise perform operations described herein below in connection with. Fig. 1-6 are described.
[0005] In at least one embodiment, a processor causes the derivation of two or more contiguous portions of information distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more final portions of the information. In at least one embodiment, a portion of the information is a portion of a training data subset, as described in connection with Fig. 5. In at least one embodiment, a processing core is a processor such as those described in connection with Fig. 1 described processors 502A and 502B.
[0006] In at least one embodiment, portions of the training data subset are distributed among the multiple processors, starting from both ends of the subset until all portions are distributed. For example, in at least one embodiment, when dividing a training data subset into 16 portions to be distributed among 4 processors, a first portion and a last portion are distributed among a first processor, a second portion and a second-to-last portion are distributed among a second processor, a third portion and a third-to-last portion are distributed among a third processor, a fourth portion and a second-to-last portion are distributed among a fourth processor, and so on in a similar pattern for the remaining portions until the 16 portions are evenly distributed among the 4 processors.In at least one embodiment, in the above example, an eighth section and the eighth-last section (ninth section) are final sections and are distributed to the fourth processor.
[0007] In at least one embodiment, continuing the example above, the workloads from computing attentions across these bins vary during training based on their positions within the training data subset. In at least one embodiment, with this distribution approach, each processor receives an equal number of bins, with half requiring a larger workload and half requiring a smaller workload, resulting in an evenly distributed workload from attention computations.
[0008] Fig. 1 is a block diagram 100 illustrating a computer system in which partitioned training data and activations are used to train a neural network, according to at least one embodiment. In at least one embodiment, a processor 122 of a computer system 108 includes one or more circuits for training a neural network 110. In at least one embodiment, the processor 122 of the client computer system 108 includes one or more circuits for partitioning training data and activations of the neural network using one or more operations as described herein. In at least one embodiment, the client computer system 108 includes one or more client devices as described herein. In at least one embodiment, Fig. 1, the client computer system 108 includes one or more devices that are clients of a cloud computing environment as described herein. In at least one embodiment, the processor 122 is a processor as described below. In at least one embodiment, the processor 122 is a central processing unit (CPU), a graphics processing unit (GPU), a parallel processing unit (PPU), a general-purpose graphics processing unit (GPGPU), a compute cluster, and / or a combination of these and / or other such processors. In at least one embodiment, the processor 122 is part of a computer system as described herein.
[0009] In at least one embodiment, one or more circuits of the processor 122 are used to perform one or more operations to train the neural network 110. In at least one embodiment, one or more circuits of the processor 122 are used to perform one or more operations to train the neural network 110 using systems and methods as described herein at least in connection with Fig. 10. In at least one embodiment described in Fig. 1, one or more circuits of the processor 122 are used to perform one or more operations to train the neural network 110 by executing one or more application programming interfaces (APIs).
[0010] In at least one embodiment, a neural network trained using the neural network 110 is a neural network as described herein at least in connection with Fig. 9A-43. In at least one embodiment, a neural network trained using neural network 110 is a portion of a neural network as described herein. In at least one embodiment, a neural network trained using train neural network 110 is a large language model such as large language model 4312 described herein at least in connection with Fig. 43. In at least one embodiment, a neural network trained using train neural network 110 is a diffusion model (e.g., an image generation model that uses one or more neural networks to generate images based on an understanding of latent structures of a dataset, such as a dataset of images).
[0011] In at least one embodiment, one or more operations for training the neural network 110 are specified, sent, or otherwise communicated to a computer system 102. In at least one embodiment, Fig. 1, one or more operations for training the neural network 110 are specified, sent, or otherwise made available to a computer system 102 using one or more APIs.
[0012] In at least one embodiment, computer system 102 is a high-performance computer system. In at least one embodiment, computer system 102 is a distributed computer system. In at least one embodiment, computer system 102 is a deep learning computer system. In at least one embodiment, operations to train neural network 110 are performed by computer system 102 using the systems, methods, operations, and / or techniques described herein. In at least one embodiment, computer system 102 is a cloud computing environment as described herein. In at least one embodiment, computer system 102 includes one or more processors 104. In at least one embodiment described in Fig. 1, the computer system 102 includes one or more graphics processors. In at least one embodiment, the processors 104 include one or more processors as described herein. In at least one embodiment, the graphics processors of the computer system 102 include one or more graphics processors as described herein. In at least one embodiment, the processors 104 of the computer system 102 include one or more central processing units (CPUs), graphics processing units (GPUs), parallel processing units (PPUs), general-purpose graphics processing units (GPGPUs), compute clusters, and / or a combination of these and / or other such processors as described herein. In at least one embodiment, Fig. 1, one or more processors 104 of computer system 102 include a plurality of processors as described herein (e.g., processor 104A includes a plurality of processors, processor 104B includes a plurality of processors, etc.).
[0013] In at least one embodiment described in Fig. 1, one or more processors of processors 104 are interconnected using systems and methods as described herein. For example, in at least one embodiment, at least some processors of processors 104 are interconnected as one or more compute clusters (or GPU clusters), as described herein at least in connection with Fig. 20B are described.
[0014] In at least one embodiment, one or more circuits of processor 122 are used to perform one or more operations for partitioning training data 112. In at least one embodiment, training data 106 is provided, sent, or otherwise provided to processor 122 of customer computer system 108. In at least one embodiment, training data 106 includes data such as documents, images, videos, and / or other such data that can be used by processor 122 to train neural network 110, as described herein. In at least one embodiment, training data 106 is data of a training data set 1002 used by a training framework 1004 to generate a trained neural network 1008, as described herein at least in connection with Fig. 10 described.
[0015] In at least one embodiment, one or more training data partitioning operations 112 comprise one or more operations to generate one or more subsets of training data 106. In at least one embodiment, for example, when the training data 106 comprises a document, one or more training data partitioning operations 112 generate subsets of training data, where each subset of training data comprises one or more pages of the document. In at least one embodiment, for example, when the training data 106 comprises a document, one or more training data partitioning operations 112 generate subsets of training data, where each subset of training data comprises one or more portions of pages of the document (e.g., half pages, paragraphs, sentences, words, etc.).In at least one embodiment, for example, when the training data 106 comprises a plurality of images (e.g., a set of images and / or a video stream of images), one or more training data partitioning operations 112 generate subsets of training data, where each subset of the training data comprises one or more images. In at least one embodiment, for example, when the training data 106 comprises one or more images, one or more training data partitioning operations 112 generate subsets of training data, where each subset of the training data comprises one or more portions of images (e.g., portions of images, sets of pixels, etc.).
[0016] In at least one embodiment, one or more training data partitioning operations 112 generate one or more training data subsets and corresponding activations (e.g., activations of neural network layers associated with the one or more training data subsets). In at least one embodiment, one or more training data partitioning operations 112 generate one or more training data subsets and activations, such as training data subset and activations 114, training data subset and activations 116, training data subset and activations 118, and / or training data subset and activations 120. In at least one embodiment, training data subset and activations 114 includes a training data subset and one or more activations corresponding to the training data subset.In at least one embodiment, training data subset and activations 116 include a training data subset and one or more activations corresponding to the training data subset. In at least one embodiment, training data subset and activations 118 include a training data subset and one or more activations corresponding to the training data subset. In at least one embodiment, training data subset and activations 120 include a training data subset and one or more activations corresponding to the training data subset.
[0017] In at least one embodiment, training data and activations 114 are indicated, sent, or otherwise provided to computer system 102 by processor 122. In at least one embodiment, one or more of the processors 104 of computer system 102 use training data and activations 114 to train neural network 110, as described herein. In at least one embodiment, training data and activations 116 are indicated, sent, or otherwise provided to computer system 102 by processor 122. In at least one embodiment, one or more of the processors 104 of computer system 102 use training data and activations 116 to train neural network 110, as described herein. In at least one embodiment, training data and activations 118 are indicated, sent, or otherwise provided to computer system 102 by processor 122.In at least one embodiment, one or more of the processors 104 of the computer system 102 use training data and activations 118 to train the neural network 110, as described herein. In at least one embodiment, training data and activations 120 are indicated, sent, or otherwise provided to the computer system 102 by processor 122. In at least one embodiment, one or more processors 104 of the computer system 102 use the training data subset and activations 120 to train the neural network 110, as described herein.
[0018] In at least one embodiment, a portion of the training data subset and activations 114 and a portion of the training data subset and activations 116 are used by processor 104A to train neural network 110. In at least one embodiment, a portion of the training data subset and activations 114 and a portion of the training data subset and activations 116 are used by processor 104B to train neural network 110. In at least one embodiment, the portion of the training data subset and activations 114 used by processor 104A to train neural network 110 includes at least a portion of the training data 106 that is not included in the portion of the training data subset and activations 114 used by processor 104B to train neural network 110.In at least one embodiment, the portion of the training data subset and activations 116 used by processor 104A to train neural network 110 includes at least a portion of training data 106 that is not included in the portion of the training data subset and activations 116 used by processor 104B to train neural network 110.
[0019] In at least one embodiment, a first portion of the training data subset and activations 114 used by processor 104A to train neural network 110 is contiguous with a second portion of the training data subset and activations 114 used by processor 104B to train neural network 110 (e.g., the first portion and the second portion occupy a contiguous physical or virtual space in computer system memory of client computer system 108).In at least one embodiment, a first portion of the training data subset and activations 116 used by processor 104A to train neural network 110 is contiguous with a second portion of the training data subset and activations 116 used by processor 104B to train neural network 110 (e.g., the first portion and the second portion occupy a contiguous physical or virtual space in computer system memory of client computer system 108).
[0020] In at least one embodiment (in Fig. 1 not shown), one or more of the processors 104 include two or more processing cores. In at least one embodiment, a processor, such as one or more of the processors 104 that include two or more processing cores, includes one or more circuits to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores, such that a first of the two or more contiguous portions of information is distributed to a first processing core of the two or more respective processing cores and a second of the two or more contiguous portions of information is distributed to a second processing core of the two or more respective processing cores.In at least one embodiment, one or more circuits of a processor of processors 104 (e.g., processor 104A) cause the derivation of a first portion of training data subset and activations (e.g., training data subset and activations 114) contiguous with a second portion of training data subset and activations (e.g., training data subset and activations 114) to be distributed only between two or more respective processing cores of the processor of processors 104.In at least one embodiment, the one or more circuits of the processor of the processors 104 are to cause the derivation of a first portion of the training data subset and activations (e.g., training data subset and activations 116) associated with a second portion of the training data subset and activations (e.g., training data subset and activations 116) to be distributed only between the two or more respective processing cores of the processor of the processors 104.
[0021] In at least one embodiment, a first portion of the training data subset and activations 114 used by processor 104A to train neural network 110 is not contiguous with a second portion of the training data subset and activations 114 used by processor 104B to train neural network 110 (e.g., the first portion and the second portion do not occupy contiguous physical or virtual space in computer system memory of client computer system 108).In at least one embodiment, where a first portion of the training data subset and activations 114 used by processor 104A to train neural network 110 is not related to a second portion of the training data subset and activations 114 used by processor 104B to train neural network 110, one or more portions between the first portion and the second portion (e.g., portions in physical or virtual memory between the first portion and the second portion) are not used to train neural network 110.
[0022] In at least one embodiment, a first portion of the training data subset and activations 116 used by processor 104A to train neural network 110 is not contiguous with a second portion of the training data subset and activations 116 used by processor 104B to train neural network 110 (e.g., the first portion and the second portion do not occupy contiguous physical or virtual space in computer system memory of client computer system 108).In at least one embodiment, where a first portion of the training data subset and activations 116 used by processor 104A to train neural network 110 is not related to a second portion of the training data subset and activations 116 used by processor 104B to train neural network 110, one or more portions between the first portion and the second portion (e.g., portions in physical or virtual memory between the first portion and the second portion) are not used to train neural network 110.
[0023] In at least one embodiment, a portion of the training data subset and activations 118 and a portion of the training data subset and activations 120 are used by processor 104C to train neural network 110. In at least one embodiment, a portion of the training data subset and activations 118 and a portion of the training data subset and activations 120 are used by processor 104D to train neural network 110. In at least one embodiment, the portion of the training data subset and activations 118 used by processor 104C to train neural network 110 includes at least a portion of training data 106 that is not included in the portion of the training data subset and activations 118 used by processor 104D to train neural network 110.In at least one embodiment, the portion of the training data subset and activations 120 used by processor 104C to train neural network 110 includes at least a portion of training data 106 that is not included in the portion of the training data subset and activations 120 used by processor 104D to train neural network 110.
[0024] In at least one embodiment, a first portion of the training data subset and activations 118 used by processor 104C to train neural network 110 is contiguous with a second portion of the training data subset and activations 118 used by processor 104D to train neural network 110 (e.g., the first portion and the second portion occupy a contiguous physical or virtual space in computer system memory of client computer system 108).In at least one embodiment, a first portion of training data and activations 120 used by processor 104C to train neural network 110 is contiguous with a second portion of training data and activations 120 used by processor 104D to train neural network 110 (e.g., the first portion and the second portion occupy a contiguous physical or virtual space in computer system memory of client computer system 108).
[0025] In at least one embodiment, a first portion of the training data subset and activations 118 used by processor 104C to train neural network 110 is not contiguous with a second portion of the training data subset and activations 118 used by processor 104D to train neural network 110 (e.g., the first portion and the second portion do not occupy contiguous physical or virtual space in computer system memory of client computer system 108).In at least one embodiment, where a first portion of the training data subset and activations 118 used by processor 104C to train neural network 110 is not related to a second portion of the training data subset and activations 118 used by processor 104D to train neural network 110, one or more portions between the first portion and the second portion (e.g., portions in physical or virtual memory between the first portion and the second portion) are not used to train neural network 110.
[0026] In at least one embodiment, a first portion of the training data subset and activations 120 used by processor 104C to train neural network 110 is not contiguous with a second portion of the training data subset and activations 120 used by processor 104D to train neural network 110 (e.g., the first portion and the second portion do not occupy contiguous physical or virtual space in computer system memory of client computer system 108).In at least one embodiment, where a first portion of the training data subset and activations 120 used by processor 104C to train neural network 110 is not related to a second portion of the training data subset and activations 120 used by processor 104D to train neural network 110, one or more portions between the first portion and the second portion (e.g., portions in physical or virtual memory between the first portion and the second portion) are not used to train neural network 110.
[0027] In at least one embodiment, processor 104A has training data 106 comprising a portion of the training data subset and activations 114 and a portion of the training data subset and activations 116. In at least one embodiment, processor 104B has training data 106 comprising a portion of the training data subset and activations 114 and a portion of the training data subset and activations 116. In at least one embodiment, processor 104A and processor 104B 124 communicate intermediate results (e.g., forward propagation, back propagation, etc.) of performing neural network training to train neural network 110, as described herein.
[0028] In at least one embodiment, processor 104A communicates one or more intermediate training results, as described herein, to processor 104B using systems and methods as described herein. In at least one embodiment, processor 104A communicates one or more intermediate derivation results, as described herein, to processor 104B using systems and methods as described herein. In at least one embodiment, processor 104B communicates one or more intermediate training results, as described herein, to processor 104A using systems and methods as described herein. In at least one embodiment, processor 104B communicates one or more intermediate derivation results, such as described herein, to processor 104A using systems and methods as described herein.In at least one embodiment described in . Fig. 1, processor 104A and / or processor 104B 124 transmit one or more intermediate results of the training and / or derivation to one or more other processors of computer system 102.
[0029] In at least one embodiment, processor 104C has training data 106 comprising a portion of the training data subset and activations 118 and a portion of the training data subset and activations 120. In at least one embodiment, processor 104D has training data 106 comprising a portion of the training data subset and activations 118 and a portion of the training data subset and activations 120. In at least one embodiment, processor 104C and processor 104D 126 communicate intermediate results (e.g., forward propagation, back propagation, etc.) of performing neural network training to train neural network 110, as described herein.
[0030] In at least one embodiment, processor 104C 126 communicates one or more intermediate training results as described herein to processor 104D using systems and methods as described herein. In at least one embodiment, processor 104C 126 communicates one or more intermediate derivation results as described herein to processor 104D using systems and methods as described herein. In at least one embodiment, processor 104D 126 communicates one or more intermediate training results as described herein to processor 104C using systems and methods as described herein.In at least one embodiment, processor 104D communicates one or more intermediate derivation results, such as those described herein, to processor 104C using systems and methods as described herein. In at least one embodiment, described in . Fig. 1, processor 104C and / or processor 104D 126 transmit one or more intermediate results of the training and / or derivation to one or more other processors of computer system 102.
[0031] In at least one embodiment, one or more layers of a neural network, such as those described herein, compute an output. In at least one embodiment, the output is stored in a location addressable by one or more processors, such as processors 104. In at least one embodiment, the output is provided to the one or more processors, which use the output to compute an output using one or more layers of the neural network (e.g., a next layer, a previous layer, or a same layer). In at least one embodiment, partitioning training data and activations (e.g., as described herein) reduces the need for a processor to provide data to processors, for example, by executing a portion of the neural network on the processor.In at least one embodiment, partitioning training data and activations (e.g., as described herein) reduces bottlenecks that can occur when certain processors have more important activations (e.g., outputs from a layer) than others, and those more important activations are needed by other processors. In at least one embodiment, these bottlenecks can lead to memory problems when a processor needs to recompute activations that it does not have in memory. In at least one embodiment, larger sequences of training data have a large number of activations that cannot all be stored on a single processor. In at least one embodiment, the processor may receive repeated requests to provide activations (e.g., to other processors) that it does not have, requiring frequent recomputing.
[0032] In at least one embodiment, one or more processors (e.g., processor 122, one or more of processors 104, and / or other processors such as those described herein) include one or more circuits to perform operations or instructions described herein, such as one or more circuits to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores. In at least one embodiment, one or more processors include one or more circuits to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores to train a neural network such as those described herein.In at least one embodiment, one or more processors include one or more circuits for distributing the derivation of two or more contiguous pieces of information only between two or more respective processing cores to execute a neural network as described herein. In at least one embodiment, one or more processors include one or more circuits for distributing the derivation of two or more contiguous pieces of information only between two or more respective processing cores to train and / or execute a large language model (LLM) as described herein.In at least one embodiment, one or more processors include one or more circuits for causing the derivation of two or more contiguous portions of information that are distributed only between two or more respective processing cores to train and / or execute a diffusion model, such as those described herein. In at least one embodiment, one or more processors include one or more circuits for performing the operations or instructions described herein, such as one or more circuits that cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence.In at least one embodiment, one or more processors include one or more circuits to perform operations or instructions described herein, such as one or more circuits to cause the derivation of two or more contiguous portions of information that are distributed only between two or more respective processing cores to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence. In at least one embodiment, described in . Fig. 1 is not shown, a non-transferable, machine-readable medium has stored therein an instruction set which, when executed by one or more processors, which are shown here at least in conjunction with Fig. 1-12, such as operations that cause the derivation of two or more contiguous pieces of information to be distributed only between two or more respective processing cores.
[0033] Fig. 2 is a block diagram 200 illustrating a computer system in which data and activations are partitioned according to partitioning criteria, in accordance with at least one embodiment. In at least one embodiment, a processor 222 of a computer system 208 executes instructions to partition training data 212, as described herein. In at least one embodiment, the processor 222 is a processor such as the processor 122, described here at least in conjunction with Fig. 1. In at least one embodiment, the computer system 208 is a computer system such as the client computer system 108, which is described here at least in connection with Fig. 1. In at least one embodiment, one or more training data partitioning instructions 212 comprise instructions such as training data partitioning instructions 112, herein at least in conjunction with Fig. 1 described.
[0034] In at least one embodiment, training data partitioning instructions 212 executed by processor 222 generate training data subsets from training data 206, as described herein. In at least one embodiment, training data 206 is training data such as training data 106, described herein at least in connection with Fig. 1. In at least one embodiment, training data 212 partitioning instructions executed by processor 222 generate a training data subset 216A-training data subset 216N (e.g., a total of "N" training data subsets). In at least one embodiment, for example, when training data 206 comprises a ten-page document, training data 212 partitioning instructions executed by processor 222 generate ten training data subsets (e.g., "N" equals ten), each comprising one page. In at least one embodiment, for example, when training data 206 comprises a ten-page document, training data 212 partitioning instructions executed by processor 222 generate twenty training data subsets (e.g., "N" equals twenty), each comprising half a page.In at least one embodiment, for example, when training data 206 comprises a ten-page document, training data partitioning instructions 212 executed by processor 222 generate five training data subsets (e.g., "N" equals five), each comprising two pages. In at least one embodiment, other values of "N" (e.g., a number of training data subsets) may be generated by training data partitioning instructions 212 executed by processor 222.
[0035] In at least one embodiment, training data 212 partitioning instructions executed by processor 222 generate one or more activations based at least in part on training data subset 216A-training data subset 216N. In at least one embodiment, training data 212 partitioning instructions executed by processor 222 generate an activation corresponding to each training data subset (e.g., training data subset 216A-training data subset 216N). For example, in at least one embodiment, training data 212 partitioning instructions executed by processor 222 generate activations 218A (e.g., comprising a single activation) corresponding to training data subset 216A, activations 218N (e.g., comprising a single activation) corresponding to training data subset 216N, and activations (not included in Fig. 2) corresponding to each other training data subset (also not shown in Fig. 2).
[0036] In at least one embodiment described in Fig. 2, instructions for partitioning training data 212 executed by processor 222 generate a plurality of activations corresponding to each training data subset (e.g., training data subset 216A-training data subset 216N). For example, in at least one embodiment, instructions for partitioning training data 212 executed by processor 222 generate a plurality of activations 218A corresponding to training data subset 216A, a plurality of activations 218N corresponding to training data subset 216N, and a plurality of activations (in Fig. 2 not shown) according to other training data subsets (also in Fig. 2 not shown). In at least one embodiment, the training data partitioning instructions 212 executed by processor 222 generate both single activations and a plurality of training data set activations.
[0037] In at least one embodiment, instructions for partitioning training data 212 executed by processor 222 are executed based at least in part on processor resources 214. In at least one embodiment, processor resources 214 are indications of resources of processors 204 of a high-performance computer system 202. In at least one embodiment, processors 204 are processors such as processors 104, which are described here at least in connection with Fig. 1. In at least one embodiment, the high-performance computer system 202 is a high-performance computer system such as the computer system 102, described here at least in connection with Fig. 1 is described.
[0038] In at least one embodiment, processor resources 214 include an indication of the resources of high-performance computing system 202 to be used to train the neural network using partitioned training data and activations 210. In at least one embodiment, processor 222 uses the information indicated by processor resources 214 to determine how to partition training data 212 to generate training data subsets (e.g., training data subset 216A-training data subset 216N) and to determine activations (e.g., activations 218A-activations 218N). In at least one embodiment, processor resources 214 include an indication of a number of processors 204 available for training the neural network using partitioned training data and activations 210.In at least one embodiment, the processing resources 214 include an indication of one or more specifications of the processors 204 (e.g., type of processor, memory capacity of the processor, other processes running on the processor, etc.) available for training the neural network using partitioned training data and activations 210. In at least one embodiment, the processor resources 214 include an indication of the memory bandwidth between the processors available for training the neural network using partitioned training data and activations 210.In at least one embodiment, the processor resources 214 include an indication of one or more architectures of the high-performance computing system 202 (e.g., an indication of the relationships between the processors 204 of the high-performance computing system 202) available for training the neural network using partitioned training data and activations 210. In at least one embodiment, at least a portion of the computational resources 214 are stored in the high-performance computing system 202. In at least one embodiment, as shown in FIG. Fig. 2, at least a portion of the processor resources 214 are stored in the computer system 208. In at least one embodiment, the processor resources 214 are indicated, sent, or otherwise provided to the processor 222 via one or more application programming interfaces (APIs) as described herein.
[0039] In at least one embodiment, instructions for partitioning training data 212 executed by processor 222 are executed based at least in part on partition criteria 220. In at least one embodiment, partition criteria 220 includes an indication of a number of partitions of training data 206 to be used to train a neural network using partitioned training data and activations 210. In at least one embodiment, partition criteria 220 includes an indication of the size of the partitions of training data 206 to be used to train the neural network using partitioned training data and activations 210. In at least one embodiment, the indications in partition criteria 220 are based at least in part on one or more properties of training data 206 (e.g., size of the data, type of data, location of the data, etc.).). In at least one embodiment, partition criteria 220 are specified, sent, or otherwise provided to processor 222 using one or more application programming interfaces (APIs) as described herein.
[0040] In at least one embodiment, processor 222 uses training data subsets (e.g., training data subset 216A-training data subset 216N) and activations (e.g., activations 218A-activations 218N) to cause processors 204 of high-performance computing system 202 to train a neural network using partitioned training data and activations 210, as described herein.In at least one embodiment, processor 222 uses training data subsets (e.g., training data subset 216A-training data subset 216N) and activations (e.g., activations 218A-activations 218N) to cause processors 204 of high-performance computing system 202 to train a neural network using partitioned training data and activations 210, to cause derivation of two or more contiguous portions of information that are distributed only between two or more corresponding processing cores, and / or to perform other operations as described herein.
[0041] Fig. 3 is a block diagram 300 illustrating a process for partitioning training data and activations of a neural network, according to at least one embodiment. In at least one embodiment, a processor such as processor 122 and / or one or more processors 104 (as described herein at least in connection with Fig. 1) performs one or more steps of a process for partitioning training data and activations of a neural network, as shown in block diagram 300 (or any other processes described herein, or variations and / or combinations thereof), using systems, methods, operations, and techniques as described herein. In at least one embodiment, a process for partitioning training data and activations of a neural network, as shown in block diagram 300 (or other processes described herein, or variations and / or combinations thereof), is performed in whole or in part under the control of one or more computer systems, as described in Fig. 9A-43, configured with computer-executable instructions and implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more applications) that execute collectively on one or more processors, by hardware, software, or combinations thereof. In at least one embodiment, the code is stored on a computer-readable storage medium in the form of a computer program that includes a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium.
[0042] In at least one embodiment, in step 302 of a process for splitting training data and neural network activations, as shown in block diagram 300, a processor executes instructions to receive training data. In at least one embodiment, the training data received in step 302 is training data such as the training data 106, which is described here at least in connection with Fig. 1. In at least one embodiment, in step 302, the received training data is received by a processor such as processor 122, which is used here at least in conjunction with Fig. 1. In at least one embodiment, in step 302, the received training data is indicated, transmitted, or otherwise provided using one or more operations as described herein. In at least one embodiment, after step 302, a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 304.
[0043] In at least one embodiment, in step 304 of a process for partitioning training data and activations of a neural network, as shown in block diagram 300, a processor executes instructions to receive processor resources. In at least one embodiment at step 304, the received processor resources are processor resources such as processor resources 214, which are described here at least in connection with Fig. 2. In at least one embodiment, in step 304, the received processor resources are received by a processor such as processor 122, which is described here at least in conjunction with Fig. 1. In at least one embodiment, in step 304, the received computing resources are used by a high-performance computer system such as computer system 102, described here at least in conjunction with Fig. 1. In at least one embodiment, after step 304, a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 306.
[0044] In at least one embodiment, in step 306 of a process for partitioning training data and neural network activations, as shown in block diagram 300, a processor executes instructions to receive partition criteria. In at least one embodiment, the partition criteria received in step 306 are partition criteria such as the partition criteria 220, which are described herein at least in connection with Fig. 2. In at least one embodiment, in step 306, the received partition criteria are received by a processor such as processor 122, which is described here at least in connection with Fig. 1. In at least one embodiment, in step 306, the received partition criteria are specified, sent, or otherwise provided using one or more operations as described herein. In at least one embodiment, after step 306, a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 308.
[0045] In at least one embodiment, in step 308 of a process for partitioning training data and neural network activations, as shown in block diagram 300, a processor executes instructions to determine a data subset size. In at least one embodiment, in step 308, the data subset size is an indication of a size and / or number of training data subsets (e.g., training data subset 216A-training data subset 216N, herein at least in connection with Fig. 2). In at least one embodiment, in step 308, the data subset size is determined based at least in part on the training data received in step 302, as described herein. In at least one embodiment, in step 308, the data subset size is determined based at least in part on the processor resources received in step 304, as described herein. In at least one embodiment, in step 308, the data subset size is determined based at least in part on the partition criteria received in step 306, as described herein. In at least one embodiment, after step 308, a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 310.
[0046] In at least one embodiment, in step 310 of a process for partitioning training data and neural network activations, as shown in block diagram 300, a processor executes instructions to select a first partition. In at least one embodiment, in step 310, a partition is selected from training data received in step 302. In at least one embodiment, in step 310, a partition is selected based at least in part on a data subset size determined in step 308. In at least one embodiment, a first partition is selected from the training data received in step 302 from an initial or starting location of the training data received in step 302.In at least one embodiment, a next partition is selected from the training data received in step 302 from a position of partitions previously selected from the training data received in step 302. In at least one embodiment, a partition is randomly selected from the training data received in step 302. In at least one embodiment, a partition is selected from the training data received in step 302 according to one or more other criteria (e.g., criteria specified in the partitioning criteria received in step 306). In at least one embodiment, after step 310, a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 312.
[0047] In at least one embodiment, in step 312 of a process for partitioning training data and neural network activations, as shown in block diagram 300, a processor executes instructions for partitioning training data to generate a training data subset, as described herein. In at least one embodiment, in step 312, the instructions for partitioning training data to generate a training data subset are based at least in part on a data subset size determined in step 308. In at least one embodiment, after step 302, a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 314.
[0048] In at least one embodiment, in step 314 of a process for partitioning training data and activations of a neural network, as shown in block diagram 300, a processor executes instructions for determining one or more activations of a training data subset, as described herein at least in connection with Fig. 1. In at least one embodiment, the instructions for determining one or more activations of a training data subset in step 314 include instructions for determining one or more activations of a training data subset partitioned in step 312. In at least one embodiment, after step 314, a process for partitioning training data and activations of a neural network, as shown in block diagram 300, continues in step 316.
[0049] In at least one embodiment, at step 316 of a process for partitioning training data and neural network activations, as shown in block diagram 300, a processor executes instructions to determine whether to select a next partition. In at least one embodiment, at step 316, the determination of whether to select a next partition is based at least in part on training data received in step 302, processor resources received in step 304, and / or partition criteria received in step 306, as described herein. In at least one embodiment, at step 316, the decision of whether to select a next partition is based at least in part on a data subset size determined in step 308.In at least one embodiment, in step 316, it is determined whether to select a next partition based at least in part on one or more previously selected partitions (e.g., in step 310). In at least one embodiment, if it is determined in step 316 that a next partition should be selected (branch "YES"), a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 310 to select a next partition. In at least one embodiment, in step 316, if it is determined that a next partition should not be selected ("NO" branch), a process for partitioning training data and neural network activations, as shown in block diagram 300, continues in step 318.
[0050] In at least one embodiment, at step 318 of a process for partitioning training data and activations of a neural network, as shown in block diagram 300, a processor executes instructions to train a neural network with partitioned training data and activations. In at least one embodiment, the instructions for training a neural network with partitioned training data and activations in step 318 include instructions such as instructions for training a neural network with partitioned training data and activations 210, which are described here at least in connection with Fig. 2. In at least one embodiment, after step 318, a process for partitioning training data and activations of the neural network ends, as shown in block diagram 300. In at least one embodiment described in Fig. 3, after step 318, a process for partitioning training data and activations of the neural network, which is illustrated in block diagram 300, continues in step 302 to obtain new training data. In at least one embodiment, which is illustrated in Fig. 3, after step 318, a process for partitioning training data and activations of the neural network, as shown in block diagram 300, continues in step 304 to obtain new processor resources. In at least one embodiment, as shown in Fig. 3 is not shown, after step 318, a process for partitioning training data and activations of the neural network, as shown in block diagram 300, continues in step 306 to obtain new partition criteria. In at least one in Fig. 3, after step 318, a process for partitioning training data and activations of the neural network, as shown in block diagram 300, continues in step 308 to determine a new data subset size.
[0051] In at least one embodiment, operations of a process for partitioning training data and activations of a neural network shown in block diagram 300 are performed in a different order than in Fig. 3. In at least one embodiment, the operations of a process for partitioning training data and neural network activations, as shown in block diagram 300, are performed concurrently or in parallel. In at least one embodiment, operations of a process for partitioning training data and neural network activations, shown in block diagram 300, that are not dependent on each other (e.g., are independent of order), are performed concurrently or in parallel. In at least one embodiment, operations of a process for partitioning training data and neural network activations, shown in block diagram 300, are performed by a plurality of threads executing on a processor such as those described herein.
[0052] Fig. 4 is a block diagram 400 illustrating forward and backward propagation in a neural network according to at least one embodiment. In at least one embodiment, forward propagation 402 is used to calculate output values 414 of a neural network. In at least one embodiment, forward propagation 402 is used to pass input data (e.g., from an input layer 404) through one or more layers of a neural network (e.g., hidden layers 406) to produce an output (e.g., an output layer 408). In at least one embodiment, a processor as described herein includes code and / or data storage (e.g., code and / or data storage 905, described herein at least in connection with Fig. 9A and Fig. 9B) to store feedforward and / or output weights and / or input / output data and / or other parameters to configure neurons or layers of a neural network that is trained and / or used in inferences in aspects of one or more embodiments. In at least one embodiment, logic 915 includes and / or is coupled to code and / or data storage 905 to store graph code or other software for controlling the timing and / or order in which information about weights and / or other parameters is to be loaded to configure logic 915, as described herein at least in connection with Fig. 9A and Fig. 9B.) In at least one embodiment, code (e.g., graph code) loads information about weights or other parameters into processors based on a neural network architecture to which such code corresponds. In at least one embodiment, the code and / or data store 905 stores weight parameters and / or layers of input / output data of a neural network being trained using aspects of one or more embodiments or used in connection with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference.
[0053] In at least one embodiment, the forward propagation 402 is performed by one or more processors, such as those described herein at least in connection with Fig. 1. In at least one embodiment, forward propagation 402 is a step in training a neural network (e.g., training neural network 110 as described herein at least in connection with Fig. 1). In at least one embodiment, calculating output values 414 of a neural network consists of applying biases and weights to nodes (or neurons) of a neural network and applying one or more activations (or activation functions) to generate one or more output values, as described herein.
[0054] In at least one embodiment, forward propagation 402 is performed using one or more layers of a neural network as described herein. In at least one embodiment, a layer of the neural network is an input layer, such as input layer 404, a hidden layer, such as one of hidden layers 406, and / or an output layer, such as output layer 408. In at least one embodiment described in Fig. 4 is not shown, a layer of a neural network is a combination of this and / or other such layers of a neural network.
[0055] In at least one embodiment, the neural network layers used to perform forward propagation 402 include an input layer 404. In at least one embodiment, input layer 404 is a neural network layer that includes input data for a neural network. In at least one embodiment, Fig. 4, the neural network layers used to perform forward propagation 402 include a plurality of input layers such as input layer 404.
[0056] In at least one embodiment, the neural network layers used to perform forward propagation 402 include one or more hidden layers 406. In at least one embodiment, hidden layers 406 are inner layers of a neural network used to calculate output values 414. In at least one embodiment, Fig. 4, the hidden layers 406 are empty (i.e., they do not include activations, thresholds, weights, etc.). In at least one embodiment, data from the input layer 404 is indicated, sent, or otherwise provided to one or more layers of the hidden layers 406, as described herein. In at least one embodiment, data from one or more layers of the hidden layers 406 is indicated, sent, or otherwise provided to one or more other layers of the hidden layers 406, as described herein.
[0057] In at least one embodiment, the neural network layers used to perform forward propagation 402 include an output layer 408. In at least one embodiment, the output layer 408 is a neural network layer that includes output data from a neural network. In at least one embodiment, Fig. 4, the neural network layers used to perform forward propagation 402 include a plurality of output layers such as output layer 408. In at least one embodiment, data from one or more of the hidden layers 406 is indicated, sent, or otherwise provided to the output layers 408 as described herein.
[0058] In at least one embodiment, the data of output layer 408 is evaluated to determine the accuracy of one or more parameters and / or hyperparameters of a neural network. In at least one embodiment, the data of output layer 408 is compared to one or more expected values 410. In at least one embodiment, the expected values 410 are values that a neural network is expected to generate based, at least in part, on input layer 404. In at least one embodiment, for example, if a neural network is trained to recognize dogs in an image and an input image is received that includes one or more dogs, the expected values 410 include indications that dogs are present in the input image.
[0059] In at least one embodiment, the data from output layer 408 and expected values 410 are used to determine one or more loss values 412. In at least one embodiment, loss values 412 are differences between the values in output layer 408 and expected values 410. In at least one embodiment, for example, when training a neural network to detect dogs in an image, receiving an input image containing one or more dogs, and expected values 410 do not include any indication of the presence of dogs in the input image, then loss values 412 include indications of where dogs were not found but were expected to be found. In at least one embodiment, loss values 412 are indicated, sent, or otherwise provided to backpropagation 416 (also referred to herein as "backpropagation").
[0060] In at least one embodiment, backpropagation 416 (or backpropagation) is used to update biases and weights 426 of a neural network. In at least one embodiment, backpropagation 416 (or backpropagation) is used to backpropagate losses (e.g., loss values 412) through a neural network to determine how much of the loss each node (or neuron) is responsible for, and to update weights and biases of a neural network to minimize the loss. In at least one embodiment, a processor as described herein includes code and / or data storage (e.g., code and / or data storage 905, described herein at least in connection with Fig. 9A and Fig. 9B) to store backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network being trained and / or used to infer in aspects of one or more embodiments as described herein. In at least one embodiment, logic 915 includes and / or is coupled to a code and / or data store 905 for storing graph code or other software for controlling the timing and / or order in which weights and / or other parameter information are to be loaded to configure logic 915, as described herein at least in connection with Fig. 9A and Fig. 9B.) In at least one embodiment, code (e.g., graph code) loads information about weights or other parameters into processors based on a neural network architecture to which such code corresponds. In at least one embodiment, the code and / or the data store stores weight parameters and / or input / output data of layers of a neural network being trained using aspects of one or more embodiments or used in connection with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inference.
[0061] In at least one embodiment, the backward propagation 416 (or backpropagation) is performed by one or more processors such as those described herein at least in connection with Fig. 1. In at least one embodiment, the backpropagation 416 is a step in training a neural network (e.g., training the neural network 110 as described herein at least in connection with Fig. 1). In at least one embodiment, updating biases and weights 426 of a neural network is a process of adjusting weights and biases of a neural network using backpropagation to minimize loss, as described herein. In at least one embodiment, loss values 418 are used to update biases and weights 426 of a neural network. In at least one embodiment, loss values 418 are identical to loss values 412. In at least one embodiment, loss values 418 are altered versions of loss values 412 (e.g., loss values 418 are generated by applying one or more transformations to loss values 412).
[0062] In at least one embodiment, backpropagation 416 is performed using one or more layers of a neural network, as described herein. In at least one embodiment, backpropagation 416 uses loss values 418 to update one or more weights and biases of an output layer 420, as described herein. In at least one embodiment, output layer 420 is identical to output layer 408 before backpropagation 416 uses loss values 418 to update one or more weights and biases of output layer 420. In at least one embodiment, output layer 420 is altered (e.g., has altered weights and biases) after backpropagation 416 uses loss values 418 to update one or more weights and biases of output layer 420.In at least one embodiment, the output layer 420 is not changed (e.g., has unchanged weights and biases) after the backpropagation 416 uses loss values 418 to update one or more weights and biases of the output layer 420.
[0063] In at least one embodiment, the updated weights and biases of output layer 420 are indicated, broadcast, or otherwise communicated to one or more of hidden layers 422. In at least one embodiment, backpropagation 416 uses loss values 418 to update one or more weights and biases of hidden layers 422, as described herein. In at least one embodiment, hidden layers 422 are identical to hidden layers 406 before backpropagation 416 uses loss values 418 to update one or more weights and biases of hidden layers 422.In at least one embodiment, one or more of the hidden layers 422 are altered (e.g., have altered weights and biases) after the backpropagation 416 uses loss values 418 to update one or more weights and biases of the hidden layers 422. In at least one embodiment, the hidden layers 422 are not altered (e.g., have unchanged weights and biases) after the backpropagation 416 uses loss values 418 to update one or more weights and biases of the hidden layers 422.
[0064] In at least one embodiment, the updated weights and biases of hidden layers 422 are indicated, sent, or otherwise communicated to an input layer 424. In at least one embodiment, backpropagation 416 uses loss values 418 to update one or more weights and biases of input layer 424, as described herein. In at least one embodiment, input layer 424 is identical to input layer 404 before backpropagation 416 uses loss values 418 to update one or more weights and biases of input layer 424. In at least one embodiment, input layer 424 is altered (e.g., has altered weights and biases) after backpropagation 416 uses loss values 418 to update one or more weights and biases of input layer 424.In at least one embodiment, the input layer 424 is not changed (e.g., has unchanged weights and biases) after the backpropagation 416 uses loss values 418 to update one or more weights and biases of the input layer 424.
[0065] In at least one embodiment described in Fig. 4, includes training a neural network using systems, methods, techniques, and operations as described herein (e.g., training the neural network 110 described herein at least in connection with Fig. 1), a plurality of forward propagations such as the forward propagation 402. In at least one in Fig. 4, includes training a neural network using systems, methods, techniques, and processes as described herein (e.g., training the neural network 110 described herein at least in connection with Fig. 1), a variety of backpropagations, such as backpropagation 416.
[0066] Fig. 5 is a block diagram 500 illustrating a first portion of a computer system in which partitioned training data and neural network activations are used to train a neural network according to at least one embodiment. In at least one embodiment, a processor 502A performs one or more operations to use partitioned training data and neural network activations to train a neural network, as described herein at least in connection with Fig. 1-3. In at least one embodiment, processor 502A is a processor, such as one or more of the processors 104 of a high-performance computer system, such as computer system 102, both as described herein at least in connection with Fig. 1. In at least one embodiment described in Fig. 5, processor 502A includes a plurality of processors that perform one or more operations to use partitioned training data and neural network activations to train a neural network, as described herein at least in connection with Fig. 1-3. In at least one embodiment, a processor 502B performs one or more operations to use partitioned training data and neural network activations as described herein at least in connection with Fig. 1-3. In at least one embodiment, processor 502B is a processor, such as one or more of the processors 104 of a high-performance computer system, such as computer system 102, both as described herein at least in connection with Fig. 1. In at least one embodiment described in Fig. 5, processor 502B includes a plurality of processors that perform one or more operations to use partitioned training data and neural network activations to train a neural network, as described herein at least in connection with Fig. 1-3. In at least one embodiment, processor 502A and processor 502B together perform one or more operations to use partitioned training data and neural network activations, as described herein at least in connection with Fig. 1-3. In at least one embodiment described in Fig. 5, one or more additional processors of a high-performance computer system such as computer system 102 (as described herein at least in connection with Fig. 1) executes one or more instructions to use partitioned training data and neural network activations to train a neural network, as shown in block diagram 500.
[0067] In at least one embodiment, processor 502A receives input 504A. In at least one embodiment, input 504A includes input data "IN0" and input data "IN1." In at least one embodiment, "IN0" of input 504A includes a first portion of training data and activations, such as a portion of training data subset and activations 114, described herein at least in connection with Fig. 1. In at least one embodiment, “IN1” of input 504A includes a first portion of training data and activations, such as a portion of training data subset and activations 116 described herein at least in connection with Fig. 1. In at least one embodiment, processor 502B receives input 504B. In at least one embodiment, input 504B includes input data "IN2" and input data "IN3." In at least one embodiment, "IN2" of input 504B includes a second portion of training data and activations, such as a portion of training data subset and activations 114 described herein at least in connection with Fig. 1. In at least one embodiment, “IN3” of input 504B includes a second portion of training data and activations, such as a portion of training data subset and activations 116 described herein at least in connection with Fig. 1. In at least one embodiment, "IN0" and "IN2" together comprise a training data subset and activations as described herein, and "IN1" and "IN3" together comprise a training data subset and activations as described herein.
[0068] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to perform all gathers 506. In at least one embodiment, "all gather 506" is an operation wherein processors used to train a neural network and / or perform one or more inferences using a trained neural network aggregate values into an output. For example, in at least one embodiment, if there are N processors and an "all gather" such as "all gather 506" is used to gather K values from each of the N processors, an output of dimension N*K is produced (e.g., K values from each of the N processors).
[0069] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to perform reduce scatter 508. In at least one embodiment, reduction scatter 508 is an operation for performing reductions on data (e.g., a sum operation, a max operation, a min operation, etc.) and distributing (or scattering) the results of those reductions to processors. In at least one embodiment, for example, when there are K values from each of N processors (e.g., as described above), a reduce scatter with a reduce operation that determines a maximum value for each set of K values scatters a maximum value of the K values to each of the N processors.
[0070] In at least one embodiment, processor 502A executes one or more instructions to generate a QKV tensor 510A (indicated as "GENERATE Q,K,V" in block diagram 500). In at least one embodiment, processor 502B executes one or more instructions to generate a QKV tensor 510B (indicated as "GENERATE Q,K,V" in block diagram 500). In at least one embodiment, generating QKV tensor 510A and generating QKV tensor 510B generates a tensor that includes a Q vector, a K vector, and a V vector. In at least one embodiment, a Q vector is a linear output layer vector based at least in part on what is encoded by a layer (e.g., an output of an encoder layer or an output of a decoder layer).In at least one embodiment, a K-vector is also a linear output layer vector based at least in part on an input (e.g., input 504A or input 504B). In at least one embodiment, a V-vector is a learned vector based at least in part on one or more computations and also based at least in part on inputs (e.g., input 504A or input 504B). In at least one embodiment, processor 502A executes one or more instructions to generate QKV tensor 510A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to generate QKV tensor 510B.
[0071] In at least one embodiment, processor 502A executes one or more instructions to perform an all gather / reduce scatter 512A (indicated as "AG / RS" in block diagram 500). In at least one embodiment, processor 502B executes one or more instructions to perform an all gather / reduce scatter 512B (indicated as "AG / RS" in block diagram 500). In at least one embodiment, all gather / reduce scatter 512A and all gather / reduce scatter 512B are joint operations that perform an all gather on forward propagation, as described herein, and a reduce scatter on backward propagation, also as described herein.In at least one embodiment, processor 502A executes one or more instructions to perform all of the gather / reduce scatters 512A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to perform all of the gather / reduce scatters 512B.
[0072] In at least one embodiment, processor 502A executes one or more instructions to perform attention 514A (indicated as "ATTENTION" in block diagram 500). In at least one embodiment, processor 502B executes one or more instructions to perform attention 514B (indicated as "ATTENTION" in block diagram 500). In at least one embodiment, attention 514A and attention 514B enhance some portions of the input data (e.g., input 504A or input 504B) and de-emphasize other portions of the input data. In at least one embodiment, attention 514A and attention 514B enable a neural network to focus more on portions of data (e.g., specific sentences, words, pixels, etc.) during training.In at least one embodiment, processor 502A executes one or more attention execution instructions 514A concurrently (e.g., at the same time as) processor 502B executing one or more attention execution instructions 514B.
[0073] In at least one embodiment, processor 502A executes one or more instructions to perform a no-op / AII gather 516A (indicated as "NO-OP / AG" in block diagram 500). In at least one embodiment, processor 502B executes one or more instructions to perform a no-op / AII gather 516B (indicated as "NO-OP / AG" in block diagram 500). In at least one embodiment, no-op / AII gather 516A and no-op / AII gather 516B are joint operations that perform no operations on forward propagation (e.g., perform no operation and / or perform a null operation) and perform an all gather on backpropagation, as described herein.In at least one embodiment, processor 502A executes one or more instructions to perform no-op / AII gather 516A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to perform no-op / AII gather 516B.
[0074] In at least one embodiment, processor 502A executes one or more instructions to perform an attention output 518A (indicated as "(ATTENTION OUTPUT)" in block diagram 500). In at least one embodiment, processor 502B executes one or more instructions to perform an attention output 518B (indicated as "ATTENTION OUTPUT" in block diagram 500). In at least one embodiment, attention output 518A and attention output 518B generate results of the application of attention 514A or attention 514B for further processing. In at least one embodiment, processor 502A executes one or more instructions to perform attention output 518A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to perform attention output 518B.
[0075] In at least one embodiment, processor 502A proceeds to 520A, using partitioned training data and neural network activations to train a neural network, as shown in block diagram 600. In at least one embodiment, processor 502B proceeds to 520B, using partitioned training data and neural network activations to train a neural network, as shown in block diagram 600.
[0076] In at least one embodiment, the Fig. 5 are performed in the order indicated in block diagram 500 (e.g., first All-Gather 506, then Reduce Scatter 508, then Generate QKV Tensor 510A and Generate QKV Tensor 510B, then All-Gather / Reduce Scatter 512A and All-Gather / Reduce Scatter 512B, etc.). In at least one embodiment, the operations described in Fig. 5 are performed in a different order than that indicated in block diagram 500. In at least one embodiment, the operations described in Fig. 5 are performed simultaneously or in parallel. In at least one embodiment, the operations described in Fig. 5, which are not dependent on each other (e.g., independent of the order), are executed simultaneously or in parallel. In at least one embodiment, the operations described in Fig. 5 are performed by a plurality of threads executing on a processor such as those described here.
[0077] Fig. 6 is a block diagram 600 illustrating a second portion of a computer system in which partitioned training data and neural network activations are used to train a neural network, according to at least one embodiment. In at least one embodiment, processor 502A (e.g., as described herein in connection with Fig. 5) proceeds to 604A, using partitioned training data and neural network activations to train a neural network, as shown in block diagram 500. In at least one embodiment, processor 502B (e.g., as described herein in connection with Fig. 5) proceeds to 604B, using partitioned training data and neural network activations to train a neural network as shown in block diagram 500. In at least one embodiment described in Fig. 6, one or more additional processors of a high-performance computer system such as computer system 102 (as described here at least in connection with Fig. 1) executes one or more instructions to use partitioned training data and neural network activations to train a neural network, as shown in block diagram 600.
[0078] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to reduce scatter 606. In at least one embodiment, reduced scatter 606 is a reduced scatter such as reduced scatter 508, described herein at least in connection with Fig. 5 is described.
[0079] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to perform all-gather 608. In at least one embodiment, all-gather 608 is an all-gather like all-gather 506, herein at least in conjunction with Fig. 5 described.
[0080] In at least one embodiment, processor 502A executes one or more instructions to perform a dropout 610A (indicated as "DROPOUT" in block diagram 600). In at least one embodiment, processor 502B executes one or more instructions to perform a dropout 610B (indicated as "DROPOUT" in block diagram 600). In at least one embodiment, dropouts 610A and dropouts 610B are performed by a processor (e.g., processor 502A and / or processor 502B) to remove data and / or noise from a neural network during training. In at least one embodiment, dropout 610A and dropout 610B randomly remove data and / or noise from a neural network to improve training and processing times. In at least one embodiment, dropout 610A and dropout 610B remove a certain portion of data and / or noise from a neural network (e.g., one-third).In at least one embodiment, one or more enhancements (e.g., increasing weight) are made to data and / or noise that has not been removed from a neural network. In at least one embodiment, processor 502A executes one or more instructions to perform dropout 610A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to perform dropout 610B.
[0081] In at least one embodiment, processor 502A executes one or more instructions to perform layer normalization 612A (indicated as "LAYER-NORM" in block diagram 600). In at least one embodiment, processor 502B executes one or more instructions to perform layer normalization 612B (indicated as "LAYER-NORM" in block diagram 600). In at least one embodiment, layer normalization 612A and layer normalization 612B are processes for estimating normalization statistics from summed inputs within a hidden layer such that normalization does not introduce new dependencies between training cases (e.g., with different input data) when training a neural network.In at least one embodiment, when performing layer normalization, the hidden layers share normalization terms, but different training cases (e.g., with different input data) have different normalization terms. In at least one embodiment, processor 502A executes one or more instructions to perform layer normalization 612A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to perform layer normalization 612B.
[0082] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to perform all-gather 614. In at least one embodiment, all-gather 614 is an all-gather such as all-gather 506, which is described herein at least in conjunction with Fig. 5 is described.
[0083] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to reduce scatter 616. In at least one embodiment, reduced scatter 616 is a reduced scatter such as reduced scatter 508, described herein at least in connection with Fig. 5 is described.
[0084] In at least one embodiment, processor 502A executes one or more instructions to create a fully connected layer 1 618A (indicated as "FC1" in block diagram 600). In at least one embodiment, processor 502B executes one or more instructions to create a fully connected layer 618B (indicated as "FC1" in block diagram 600). In at least one embodiment, creating a fully connected layer 1 618A and creating a fully connected layer 1 618B are layer transformations that cause each neuron of one layer (e.g., a first hidden layer) to be connected to each neuron of another layer (e.g., a second hidden layer).In at least one embodiment, as a result of generating a fully connected layer 1 618A and generating a fully connected layer 1 618B, each input of an input vector affects each output of an output vector. In at least one embodiment, processor 502A executes one or more instructions to generate a fully connected layer 1 618A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to generate a fully connected layer 1 618B.
[0085] In at least one embodiment, processor 502A executes one or more instructions to execute a linear Gaussian error unit 620A (indicated as "GeLU" in block diagram 600). In at least one embodiment, processor 502B executes one or more instructions to execute a linear Gaussian error unit 620B (indicated as "GeLU" in block diagram 600). In at least one embodiment, linear Gaussian error unit 620A and linear Gaussian error unit 620B are activation functions that weight inputs according to a Gaussian distribution function (e.g., a percentile of the inputs). In at least one embodiment, processor 502A executes one or more instructions to execute linear Gaussian error unit 620A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to execute linear Gaussian error unit 620B.
[0086] In at least one embodiment, processor 502A executes one or more instructions to create a fully connected layer two 622A (indicated as "FC2" in block diagram 600). In at least one embodiment, processor 502B executes one or more instructions to create a fully connected layer two 622B (indicated as "FC2" in block diagram 500). In at least one embodiment, the creation of a fully connected layer two 622A and the creation of a fully connected layer two 622B are as described above in connection with the creation of a fully connected layer one 618A and the creation of a fully connected layer one 618B.In at least one embodiment, processor 502A executes one or more instructions to generate a fully connected layer two 622A concurrently (e.g., at the same time as) processor 502B executing one or more instructions to generate a fully connected layer two 622B.
[0087] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to reduce scatter 624. In at least one embodiment, reduce scatter 624 is a reduce scatter such as reduce scatter 508, described here at least in connection with Fig. 5 is described.
[0088] In at least one embodiment, processor 502A and processor 502B execute one or more instructions to perform all-gather 626. In at least one embodiment, all-gather 626 is an all-gather such as all-gather 506, which is described herein at least in conjunction with Fig. 5 is described.
[0089] In at least one embodiment, the Fig. 6 are performed in an order indicated in block diagram 600 (e.g., first Reduce Scatter 606, then All-Gather 608, then Dropout 610A and Dropout 610B, then Layer Normalization 612A and Layer Normalization 612B, etc.). In at least one embodiment, the operations described in Fig. 6 are performed in a different order than that indicated in block diagram 600. In at least one embodiment, the operations described in Fig. 6 are performed simultaneously or in parallel. In at least one embodiment, the operations described in Fig. 6, which are not dependent on each other (e.g., independent of the order), are executed simultaneously or in parallel. In at least one embodiment, the operations described in Fig. 6 are performed by a plurality of threads executing on a processor such as those described herein.
[0090] Fig. 7 is a block diagram 700 illustrating a system for partitioning a training data subset 716 among multiple processors to balance the workload among the processors, according to at least one embodiment. In at least one embodiment, the processors may be a group of processors, each receiving one or more portions of the training data subset 716. In at least one embodiment, the activations of the input data are also partitioned and distributed to the processor based at least in part on the partitioning of the data subset. In at least one embodiment, for illustrative purposes, Fig. 7 uses two processors 502A and 502B.
[0091] In at least one embodiment, a client computer system, such as that described herein in connection with Fig. 1, is used to divide the training data subset 716. In at least one embodiment, the client machine divides the training data subset 716 into a number of chunks, where the number of chunks is an even multiple of a number of processors. In at least one embodiment, since two processors 502A and 502B in the group receive the distribution of the training data subset 716, the client machine may divide the training data subset 716 into four chunks, eight chunks, 12 chunks, 16 chunks, and so on, as long as the number of chunks is equal to 2mn, where m is any integer and n is a number of processors (such as processors 502A and 502B) that each receive distributions of the training data subset 716 from the client machine.
[0092] In at least one embodiment, each section of the training data subset 716 includes an equal number of tokens. In at least one embodiment, a token represents a word or part of a word in text. In at least one embodiment, a token represents characters, subword units (e.g., syllables or parts of words). In at least one embodiment, a token represents entire sentences or paragraphs, depending on the granularity of a tokenization process. In at least one embodiment, each section of the eight sections of the training data subset 716 includes an equal number of tokens.In at least one embodiment, if the training data subset 716 does not include tokens that allow an equal number of tokens to be distributed to each of processors 502A and 502B, padding techniques are used to standardize the training data subset 716 to remove or add tokens therein so that each processor can receive an equal number of tokens. In at least one embodiment, the training data subset 716 is divided into eight consecutive sections labeled "IN0," "IN1," "IN2," "IN3," "IN4," "IN5," "IN7," and "IN8." In at least one embodiment, the section "IN0" is at the beginning and is a first section of the training data subset 716, and the section "IN7" is at the end and is a last section of the training data subset 716.
[0093] In at least one embodiment, the client machine distributes portions of the training data subset 716 from opposite ends of the training data subset 716 to the processors 502A and 502B. For example, in at least one embodiment, the client machine distributes portions "IN0" and "IN7" to processor 502A, portions "IN1" and "IN6" to processor 502B, portions "IN2" and "IN5" to processor 502A, and portions "IN3" and "IN4" to processor 502B, resulting in each processor 502A and 502B having four portions. In at least one embodiment, as in Fig. 7, processor 502A receives sections “IN0”, “IN2”, “IN5”, and “IN7” as input, and processor 502B receives sections “IN1”, “IN3”, “IN4”, and “IN6” as input.
[0094] However, in at least one embodiment, each processor 502A and 502B requires the entirety of the training data subset 716 during training to compute self-attention. In at least one embodiment, token embeddings are exchanged between the processors 502A and 502B so that each processor 502A and 502B may have embeddings of all tokens in the training data subset 716. In at least one embodiment, a token embedding is a representation of the token in a numerical format, such as a vector in a high-dimensional space. In at least one embodiment, such a vector captures one or more semantic and syntactic properties of the token. For example, in at least one embodiment, words with similar meanings or that appear in similar contexts may have embeddings that are close to each other in a high-dimensional space.
[0095] In at least one embodiment, self-attention enables a neural network to weigh the importance of different tokens of an input sequence (e.g., an input sequence in training data subset 716) when processing a particular token of that input sequence. In at least one embodiment, the particular token is a current token of the input sequence. In at least one embodiment, during self-attention, the current token (e.g., the word) is "paid attention" to itself and to each preceding token in the input sequence to determine which tokens are more and less important for predicting a next token. In at least one embodiment, paying attention to a token means focusing on that token and giving it importance.In at least one embodiment, self-attention enables a neural network to recognize internal structure and relationships within a given input sequence (e.g., input sequence in training data subset 716). In at least one embodiment, such an approach, in which future tokens in an input sequence are masked and a neural network sees only current and past information, is referred to as "causal masking," which can force the trained neural network to learn how to predict a next token without knowing what the next token is.
[0096] In at least one embodiment, by distributing portions from opposite ends of the training data subset 716, the workload resulting from calculating self-attention is balanced across the training data subset 716. In at least one embodiment, the workload resulting from calculating self-attention for a token is determined by a position of the token in the training data subset 716. In at least one embodiment, a token located close to the beginning of the training data subset 716 may require less work than a token not as close to the beginning due to the "causal masking effect" described above.
[0097] Fig. Figure 8 is a block diagram illustrating how workload balancing is achieved in conjunction with self-attention computations according to at least one embodiment. In at least one embodiment, Fig. 8 that in Fig. 7, in which the training data subset 716, which is divided into eight sections, is distributed from opposite ends of the input data sequence to the processors 502A and 502B, resulting in an equal workload of self-attention computations on each processor 502A and 502B.
[0098] In at least one embodiment, a workload of self-attention calculations for a current token is determined by a number of tokens from a beginning of the training data subset 716 to that current token. In at least one embodiment, each of the eight bins "IN0", "IN1", "IN2", "IN3", "IN4", "IN5", "IN6", "IN7" comprises an equal number (n) of tokens. Therefore, in at least one embodiment, the number of bins in each of the columns 801-817 is proportional to the number of tokens in that column. For example, in at least one embodiment, column 801 has n tokens, column 809 has 2n tokens, and so on.
[0099] In at least one embodiment, the height of a Fig. 3, measured by the number of vertically stacked sections, indicates a self-attention calculation workload for a section located at the bottom of that column. In at least one embodiment, a combined total height of columns 801-807 is equal to a combined total height of columns 809-817. This equality means that processors 502A and 502B have an equivalent self-attention calculation workload for the respective sections of training data subset 716 assigned to each processor.
[0100] Fig. Figure 9 is a block diagram illustrating a process for calculating self-attention for an input sequence, according to at least one embodiment. In at least one embodiment, given an input sequence (e.g., a sequence of tokens in training data subset 716), processor 502A performs attention calculation 514, which includes a series of operations described below.
[0101] In at least one embodiment, an input sequence with eight tokens [T1, T2, T3, T4, T5, T6, T7, T8] is used to demonstrate a self-attention operation in a causal masking manner. In at least one embodiment, as described in connection with Fig. As described in Section 7, the calculation of self-attention in a causal masking manner involves a process in which each brand is only allowed to pay attention to itself and to brands that precede it.
[0102] In at least one embodiment, in block 901, each token (T1, T2, ..., T8) is transformed into a high-dimensional vector. In at least one embodiment, these vectors are learned representations that capture semantic and syntactic properties of the tokens. In at least one embodiment, with an embedding dimension of 128, each token would be a vector of 128 numbers. For example, in at least one embodiment, T1 would be encoded into a 128-dimensional vector V1, T2 into a 128-dimensional vector V2, and so on, until eight embedded vectors with 128 dimensions (V1 to V8) are created.
[0103] In at least one embodiment, an 8x8 matrix is created that dictates which tokens can pay attention to which others to prevent a token from being influenced by tokens that come after it in the stated input order. In at least one embodiment, a diagonal and a lower triangle of the matrix are filled with 1s and an upper triangle with 0s, where 1s indicate allowed attention and 0s prevent attention. In at least one embodiment, other masking techniques may also be used. In at least one embodiment, the matrix of 0s and 1s is referred to as a causal mask.
[0104] In at least one embodiment, in block 903, each of the embedded vectors (V1 through V8) is transformed into three new sets of vectors using three different weight matrices that are part of the architecture of a trained neural network. In at least one embodiment, for each token in the input sequence, a set of query (Q), key (K), and value (V) vectors is generated by multiplying the token embedding by these weight matrices. In at least one embodiment, a parallel process is used to generate Q, K, and V vectors for all tokens in an input sequence at once. In at least one embodiment, a Q vector for the token is a transformed version of the token embedding that quantifies the relevance of all other tokens to the token in the input sequence.In at least one embodiment, a K-vector for a token is a transformed version of the token embedding that facilitates determining the relevance of that token to every other token in an input sequence. In at least one embodiment, a V-vector for a token is a transformed representation of the token embedding that contains information to be used for a weighted output. In at least one embodiment, after the relevance of each token has been determined by the Q-vector and the K-vector, the V-vector provides the actual content, which is aggregated to form a weighted output that reflects contextual relationships within the input sequence.
[0105] Continuing the example above, in at least one embodiment, at block 905, processor 502A calculates for each of the eight tokens how much attention it pays to itself and each previous token, for example, by taking a dot product of the token's Q vector with a K vector of each other token, yielding an 8x8 matrix of scores, where each element (i, j) represents an attention score from token Ti to token Tj. In at least one embodiment, processor 502A generates an attention score matrix from these attention scores, where a series of scores for each token indicates its relevance to every other token in an input sequence. In at least one embodiment, processor 502A modifies the attention score matrix using the causal mask generated above to generate a masked attention score matrix in which attention scores corresponding to a "0" in the causal mask are blocked.
[0106] In at least one embodiment, in block 907, processor 502A applies a softmax function to each row of the masked attention score matrix to generate a matrix of attention weights, which is a matrix of probabilities, where each row corresponds to a different token in the input sequence.
[0107] In block 909, processor 502A uses these probabilities in the attention weight matrix to weight the value vectors of the unmasked tokens and then sums these weighted vectors to produce a final output.
[0108] In at least one embodiment, the process of calculating self-attention may be used to calculate self-attention for each token in a portion (e.g., IN0, IN2, IN5, and IN7) of the training data subset 716 associated with the processor 502A and for each token in a portion (e.g., IN1, IN3, IN4, and IN6) of the training data subset 716 associated with the processor 502B.
[0109] Fig. 10 is a visual representation of the workload across different processors as a result of distributing equal-sized portions from opposite ends of the input sequence, according to at least one embodiment. In this example, an input sequence is divided along a sequence dimension into four portions, each containing the same number of tokens. In at least one embodiment, GPU0 receives two portions from both ends of said input sequence, while GPU1 receives the two middle portions. In at least one embodiment, each of GPU0 and GPU1 is a processor, such as processor 502A or 502B, used in conjunction with Fig. 1 was described.
[0110] In at least one embodiment, as in Fig. 10, shaded areas 1009-1015 represent self-attention computation workloads. In at least one embodiment, a self-attention computation workload includes the operations and computations described in blocks 901-909. In at least one embodiment, the workload may be measured by computer runtime, which refers to a period of time during which instructions for self-attention computations are executed. In at least one embodiment, this period includes the time required to execute the instructions and manage system resources.
[0111] In at least one embodiment, shaded area 1009 represents the workload of GPU0 in calculating self-attention for a first section, which may require the least workload for this purpose. In at least one embodiment, shaded area 1015 represents a workload of GPU0 in calculating attention for a last section, which may require the greatest workload for this purpose. In at least one embodiment, these shaded areas 1009 and 1015 taken together represent a total workload of GPU0 in calculating self-attention for the first and last sections. In at least one embodiment, shaded areas 1011 and 1013 added together represent a total workload of GPU1 from self-attention calculations for the two middle sections assigned to that GPU.In at least one embodiment, the total workload of GPU1 is equal to the total workload of GPU0.
[0112] Fig. 11 is a block diagram 1100 illustrating the use of partitioned training data and neural network activations to train a neural network according to at least one embodiment. In at least one embodiment, a processor 1116 (e.g., a processor such as one or more of the processors described herein at least in connection with Fig. 1 described processors 104) training data 1102 (for example training data 106, here also at least in connection with Fig. 1). In at least one embodiment, the processor 1116 first divides the training data into training data subsets 1104, as described herein in connection with Fig. 1-3. In at least one embodiment, the processor 1116 then determines the activations for the training data subsets, as described herein in connection with the Fig. 1-3. In at least one embodiment, the processor 1116 then determines activations corresponding to the training data subsets 1106, also as described herein in connection with Fig. 1-3 described.
[0113] In at least one embodiment, the high-performance computer system 1108 uses training data subsets and activations to train the neural network 1110 (e.g., the neural network 110 described here at least in connection with Fig. 1), as described here in conjunction with Fig. 1-3. In at least one embodiment, the high-performance computer system 1108 uses one or more processors (e.g., one or more of the processors described herein at least in connection with Fig. 1 described processors 104) to use training data subsets and activations to train the neural network 1110, using here at least in conjunction with Fig. 5 and Fig. 6. In at least one embodiment, an untrained neural network 1112 (e.g., untrained neural network 1006) is trained using trained neural network 1110 to generate a trained neural network 1114 (e.g., trained neural network 1008) using operations as described herein at least in connection with Fig. 10. In at least one embodiment, the trained neural network 1114 is a large language model such as the large language model 4312 described here at least in connection with Fig. 43. In at least one embodiment, the trained neural network 1114 is a diffusion model (e.g., an image generation model that uses one or more neural networks to generate images based on an understanding of latent structures of a dataset, such as a dataset of images). In at least one embodiment, the trained neural network 1114 is another neural network as described herein.
[0114] Fig. 12 is a block diagram 1200 illustrating a processor and modules according to at least one embodiment. In at least one embodiment, a processor 1202 performs one or more processes as described herein to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores. In at least one embodiment, the processor 1202 performs one or more processes as described herein to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores using systems, methods, operations, and techniques associated with Fig. 1-12 are distributed.
[0115] In at least one embodiment, the processor 1202 comprises one or more processors as described in connection with the Fig. 9A-43. In at least one embodiment, processor 1202 is a processor such as processor 122 and / or one or more of the processors 104 described herein at least in connection with Fig. 1. In at least one embodiment, processor 1202 is any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, GPGPUs, PPUs, and / or variations thereof. In at least one embodiment, processor 1202 includes or has access to a client computing module 1204, a high-performance computing module 1206, a data partitioning module 1208, a neural network training module 1210, an activation module 1212, and a neural network inference module 1214. In at least one embodiment, one or more of client compute module 1204, high performance compute module 1206, data partitioning module 1208, neural network training module 1210, activation module 1212, and / or neural network inference module 1214 are part of processor 1202 and / or one or more other processors as described herein.In at least one embodiment, the client compute module 1204, the high-performance compute module 1206, the data partitioning module 1208, the neural network training module 1210, the activation module 1212, and the neural network inference module 1214 are distributed across multiple processors that communicate via a bus, a network, by writing to shared memory, and / or any suitable communication method such as those described herein.
[0116] In at least one embodiment, a module, as used in an implementation described herein, unless the context otherwise indicates or expressly states otherwise, refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functions described herein. In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions, and "hardware," such as that used by a processor in any embodiment described herein, may include, for example, individually or in any combination, hardwired circuitry, programmable circuitry, state machine circuitry, fixed function circuitry, execution units, and / or firmware storing instructions executed by programmable circuitry.In at least one embodiment, modules may be implemented collectively or individually as circuits that are part of a larger system, such as an integrated circuit (IC), a system-on-chip (SoC), and so forth. In at least one embodiment, a module performs one or more processes in conjunction with any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, GPGPUs, PPUs, and / or variations thereof.
[0117] In at least one embodiment, processor 1202 uses client compute module 1204 to execute or otherwise implement one or more client environments as described herein. In at least one embodiment, processor 1202 uses client compute module 1204 to perform one or more training data partitioning and neural network activation operations as described herein. In at least one embodiment, processor 1202 uses client compute module 1204 to perform one or more training data partitioning and neural network activation operations in conjunction with one or more of the following: high-performance compute module 1206, data partitioning module 1208, neural network training module 1210, activation module 1212, and / or neural network inference module 1214.In at least one embodiment, processor 1202 uses client computing module 1204 to perform operations of client computer system 108, described herein at least in connection with. Fig. 1. In at least one embodiment, processor 1202 uses client computing module 1204 to perform one or more processes as described herein, at least by including or otherwise encoding instructions that cause the one or more processes to be performed or that can otherwise be used to perform the one or more processes (e.g., by processor 1202). In at least one embodiment, processor 1202 uses client computing module 1204 to execute or otherwise implement one or more computing environments using systems, methods, operations, and techniques described herein, at least in connection with Fig. 1-12. In at least one embodiment, processor 1202 uses client compute module 1204 to perform one or more operations to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores.
[0118] In at least one embodiment, processor 1202 uses high-performance computing module 1206 to perform or otherwise implement one or more high-performance computing environments, such as those described herein. In at least one embodiment, processor 1202 uses high-performance computing module 1206 to perform one or more neural network training data and activation partitioning operations, as described herein. In at least one embodiment, processor 1202 uses high-performance computing module 1206 to perform one or more neural network training data and activation partitioning operations in conjunction with one or more of the following: client computing module 1204, data partitioning module 1208, neural network training module 1210, activation module 1212, and / or neural network inference module 1214.In at least one embodiment, processor 1202 uses high performance computing module 1206 to perform operations of computer system 102, described herein at least in connection with. Fig. 1. In at least one embodiment, the processor 1202 uses the high-performance computing module 1206 to perform one or more processes as described herein, at least by including or otherwise encoding instructions that cause the one or more processes to be performed or that can otherwise be used to perform the one or more processes (e.g., by the processor 1202). In at least one embodiment, the processor 1202 uses the high-performance computing module 1206 to execute or otherwise implement one or more computing environments using systems, methods, operations, and techniques described herein at least in connection with Fig. 1-12. In at least one embodiment, processor 1202 uses high-performance computing module 1206 to perform one or more operations to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores.
[0119] In at least one embodiment, the processor 1202 uses the data partitioning module 1208 to partition the neural network training data. In at least one embodiment, the processor 1202 uses the data partitioning module 1208 to partition the neural network training data, as described herein at least in connection with Fig. 1-3. In at least one embodiment, processor 1202 uses data partitioning module 1208 to perform one or more processes as described herein, at least by including or otherwise encoding instructions that cause the one or more processes to be performed or that can otherwise be used to perform the one or more processes (e.g., by processor 1202). In at least one embodiment, processor 1202 uses data partitioning module 1208 to perform one or more operations for partitioning training data and neural network activations as described herein.In at least one embodiment, processor 1202 uses data partitioning module 1208 to perform one or more training data partitioning operations and neural network activations in conjunction with one or more of the following: client compute module 1204, high performance compute module 1206, neural network training module 1210, activation module 1212, and / or neural network derivation module 1214. In at least one embodiment, processor 1202 uses data partitioning module 1208 to partition one or more environments using systems, methods, operations, and techniques described herein, at least in conjunction with. Fig. 1-12. In at least one embodiment, processor 1202 uses data partitioning module 1208 to perform one or more operations to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores.
[0120] In at least one embodiment, the processor 1202 uses the neural network training module 1210 to train a neural network. In at least one embodiment, the processor 1202 uses the neural network training module 1210 to train a neural network 110, as described herein at least in connection with Fig. 1. In at least one embodiment, the processor 1202 uses the neural network training module 1210 to perform one or more operations to partition training data and neural network activations, as described herein. In at least one embodiment, the processor 1202 uses the neural network training module 1210 to perform one or more operations to partition training data and neural network activations in conjunction with one or more of the following: client compute module 1204, high-performance compute module 1206, data partitioning module 1208, activation module 1212, and / or neural network derivation module 1214. In at least one embodiment, the processor 1202 uses the neural network training module 1210 to train one or more environments using systems, methods, operations, and techniques described herein, at least in conjunction with Fig. 1-12. In at least one embodiment, processor 1202 uses neural network training module 1210 to perform one or more operations to cause the derivation of two or more contiguous pieces of information to be distributed only between two or more respective processing cores.
[0121] In at least one embodiment, the processor 1202 uses the activation module 1212 to partition and distribute the neural network activations. In at least one embodiment, the processor 1202 uses the activation module 1212 to partition and distribute the neural network activations, as described herein at least in connection with Fig. 1-3. In at least one embodiment, processor 1202 uses activation module 1212 to perform one or more processes as described herein, at least by including or otherwise encoding instructions that cause the one or more processes to be performed or that can otherwise be used to perform the one or more processes (e.g., by processor 1202). In at least one embodiment, processor 1202 uses activation module 1212 to perform one or more operations for partitioning training data and neural network activations, as described herein.In at least one embodiment, processor 1202 uses activation module 1212 to perform one or more training data partitioning and neural network activation operations in conjunction with one or more of the following: client compute module 1204, high performance compute module 1206, data partitioning module 1208, neural network training module 1210, and / or neural network derivation module 1214. In at least one embodiment, processor 1202 uses activation module 1212 to execute or otherwise implement one or more environments using systems, methods, operations, and techniques described herein, at least in conjunction with. Fig. 1-12. In at least one embodiment, processor 1202 uses enablement module 1212 to perform one or more operations to cause the derivation of two or more contiguous portions of information to be distributed only between two or more respective processing cores.
[0122] In at least one embodiment, the processor 1202 uses the neural network derivation module 1214 to execute a trained neural network. In at least one embodiment, the processor 1202 uses the neural network derivation module 1214 to execute a trained neural network trained by the neural network 110, as described herein at least in connection with Fig. 1. In at least one embodiment, processor 1202 uses neural network derivation module 1214 to perform one or more processes as described herein, at least by including or otherwise encoding instructions that cause the one or more processes to be performed or that can otherwise be used to perform the one or more processes (e.g., by processor 1202). In at least one embodiment, processor 1202 uses neural network derivation module 1214 to perform one or more operations for partitioning training data and neural network activations, as described herein.In at least one embodiment, processor 1202 uses neural network derivation module 1214 to perform one or more operations to partition training data and neural network activations in conjunction with one or more of the following: client compute module 1204, high performance compute module 1206, data partitioning module 1212, neural network training module 1210, and / or activation module 1212. In at least one embodiment, processor 1202 uses neural network derivation module 1214 to perform one or more environments using systems, methods, operations, and techniques described herein, at least in conjunction with. Fig. 1-12. In at least one embodiment, processor 1202 uses neural network derivation module 1214 to perform one or more operations to cause the derivation of two or more contiguous portions of information that are distributed only between two or more respective processing cores.
[0123] In at least one embodiment, the processor 1202 includes circuitry to cause one or more circuits of the processor 1202 to cause the derivation of two or more contiguous pieces of information that are shared only between two or more respective processing cores using one or more of client compute modules 1204, high performance compute modules 1206, data partitioning modules 1208, neural network training modules 1210, activation modules 1212, and / or neural network derivation modules 1214 using systems, methods, operations, and / or techniques described herein at least in connection with Fig. 1-12 are distributed. LOGIC
[0124] Fig. 13A shows logic 1315, which, as described elsewhere herein, may be used in one or more devices to perform operations such as those discussed herein in accordance with at least one embodiment. In at least one embodiment, logic 1315 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, logic 1315 is inference and / or training logic. Details of logic 1315 are described below in connection with Fig. 13A and / or 13B. In at least one embodiment, logic refers to any combination of software logic, hardware logic, and / or firmware logic to provide the functions or operations described herein, where the logic may be embodied, in whole or in part, as circuitry that forms part of a larger system, such as an integrated circuit (IC), a system-on-chip (SoC), or one or more processors (e.g., CPU, GPU).
[0125] In at least one embodiment, logic 1315 may include, without limitation, code and / or data storage 1301 to store feedforward and / or output weights and / or input / output data and / or other parameters to configure neurons or layers of a neural network trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, logic 1315 may include or be coupled to graph code and / or data storage 1301 to store graph code or other software that controls the timing and / or order in which information about weights and / or other parameters is loaded to configure logic, including integer and / or floating-point units (collectively referred to as arithmetic logic units (ALUs)).In at least one embodiment, code, such as graph code, loads weights or other parameter information into processor ALUs based on a neural network architecture to which such code corresponds. In at least one embodiment, code and / or data storage 1301 stores weight parameters and / or input / output data of each layer of a neural network being trained using aspects of one or more embodiments or used in connection with one or more embodiments during forward propagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments. In at least one embodiment, each portion of code and / or data storage 1301 may include other on-chip or off-chip data stores, including a processor's L1, L2, or L3 cache or system memory.
[0126] In at least one embodiment, each portion of code and / or data storage 1301 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, code and / or data storage 1301 may be cache memory, dynamic random addressable memory ("DRAM"), static random addressable memory ("SRAM"), non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether the code and / or code and / or data storage 1301 is, for example, internal or external to a processor or comprises DRAM, SRAM, flash, or another type of memory may depend on the available on-chip versus off-chip memory, the latency requirements of the derivation and / or inference functions performed, the batch size of the data used in the derivation and / or training of a neural network, or a combination of these factors.
[0127] In at least one embodiment, logic 1315 may include, without limitation, a code and / or data storage 1305 to store backward and / or output weights and / or input / output data corresponding to neurons or layers of a neural network being trained and / or used for inference in aspects of one or more embodiments. In at least one embodiment, code and / or data storage 1305 stores weight parameters and / or input / output data of each layer of a neural network being trained or used in connection with one or more embodiments during backpropagation of input / output data and / or weight parameters during training and / or inference using aspects of one or more embodiments.In at least one embodiment, logic 1315 may include or be coupled to code and / or data memory 1305 to store graph code or other software that controls the timing and / or order in which information about weights and / or other parameters is loaded to configure logic, including integer and / or floating point units (collectively: arithmetic logic units (ALUs)).
[0128] In at least one embodiment, code, such as graph code, causes information about weights or other parameters to be loaded into processor ALUs based on a neural network architecture to which such code corresponds. In at least one embodiment, each portion of code and / or data storage 1305 may comprise other on-chip or off-chip data storage, including the L1, L2, or L3 cache or system memory of a processor. In at least one embodiment, each portion of code and / or data storage 1305 may be internal or external to one or more processors or other hardware logic devices or circuitry. In at least one embodiment, code and / or data storage 1305 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory.In at least one embodiment, the choice of whether the code and / or data memory 1305 is, for example, internal or external to a processor, or comprises DRAM, SRAM, flash memory, or another type of memory, may depend on the available memory on-chip versus off-chip, the latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors.
[0129] In at least one embodiment, code and / or data memory 1301 and code and / or data memory 1305 may be separate memory structures. In at least one embodiment, code and / or data memory 1301 and code and / or data memory 1305 may be a combined memory structure. In at least one embodiment, code and / or data memory 1301 and code and / or data memory 1305 may be partially combined and partially separate. In at least one embodiment, each portion of code and / or data memory 1301 and code and / or data memory 1305 may include other on-chip or off-chip data storage, including a processor's L1, L2, or L3 cache or system memory.
[0130] In at least one embodiment, logic 1315 may include, without limitation, one or more arithmetic logic unit(s) ("ALU(s)") 1310, including integer and / or floating point units, to perform logical and / or mathematical operations based at least in part on or specified by training and / or inference code (e.g., graph code), the result of which may produce activations stored in activation memory 1320 (e.g., output values of layers or neurons within a neural network) that are functions of input / output and / or weighting parameter data stored in code and / or data memory 1301 and / or code and / or data memory 1305.In at least one embodiment, activations stored in activation memory 1320 are generated according to linear algebraic and / or matrix-based mathematics performed by ALU(s) 1310 in response to execution instructions or other code, using weight values stored in code and / or data memory 1305 and / or data memory 1301 as operands along with other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, any or all of which may be stored in code and / or data memory 1305 or code and / or data memory 1301 or other memory on or off-chip.
[0131] In at least one embodiment, ALU(s) 1310 are included in one or more processors or other logical hardware devices or circuits, while in another embodiment, ALU(s) 1310 may be external to a processor or other logical hardware device or circuit that uses them (e.g., a co-processor). In at least one embodiment, ALU(s) 1310 may be included in the execution units of a processor or otherwise in a bank of ALUs that can be accessed by the execution units of a processor, either within the same processor or distributed among different processors of different types (e.g., central processing units, graphics processing units, fixed function units, etc.).In at least one embodiment, code and / or data memory 1301, code and / or data memory 1305, and enable memory 1320 may share a processor or other logical hardware device or circuitry, while in another embodiment, they may be located in different processors or other logical hardware devices or circuitry, or in a combination of the same and different processors or other logical hardware devices or circuitry. In at least one embodiment, each portion of enable memory 1320 may include other on-chip or off-chip data stores, including a processor's L1, L2, or L3 cache or system memory.In addition, derivation and / or training code may be stored along with other code accessible by a processor or other hardware logic or circuitry, and retrieved and / or processed using a processor's derivation, decoding, scheduling, execution, elimination, and / or other logic circuitry.
[0132] In at least one embodiment, activation memory 1320 may be a cache memory, a DRAM, an SRAM, a non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, activation memory 1320 may be located entirely or partially inside or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether activation memory 1320 is, for example, inside or outside a processor, or includes DRAM, SRAM, flash memory, or another type of memory, may depend on the available on-chip versus off-chip memory, the latency requirements of the training and / or inference functions performed, the batch size of the data used in the inference and / or training of a neural network, or a combination of these factors.
[0133] In at least one embodiment, the Fig. 13A may be used in conjunction with an application-specific integrated circuit (“ASIC”), such as a TensorFlow® Processing Unit from Google, an inference processing unit (IPU) from Graphcore™, or a Nervana® (e.g., “Lake Crest”) processor from Intel Corp. In at least one embodiment, the logic 1315 shown in Fig. 13A may be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate arrays (“FPGAs”).
[0134] Fig. 13B shows logic 1315 according to at least one embodiment. In at least one embodiment, logic 1315 is inference and / or training logic. In at least one embodiment, logic 1315 may include, without limitation, hardware logic in which computational resources associated with weight values or other information corresponding to one or more layers of neurons within a neural network are dedicated or otherwise exclusively used. In at least one embodiment, the logic shown in Fig. 13B may be used in conjunction with an application-specific integrated circuit (ASIC), such as Google's TensorFlow® Processing Unit, a Graphcore™ Inference Processing Unit (IPU), or a Nervana® (e.g., "Lake Crest") processor from Intel Corp. In at least one embodiment, the logic 1315 shown in Fig. 13B may be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU), or other hardware, such as field-programmable gate arrays (FPGAs). In at least one embodiment, logic 1315 includes, without limitation, code and / or data memory 1301 and code and / or data memory 1305, which may be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, pulse values, and / or other parameter or hyperparameter information. In at least one embodiment, Fig. 13B, each code and / or data memory 1301 and each code and / or data memory 1305 is coupled to a dedicated computing resource, such as computer hardware 1302 and computer hardware 1306, respectively. In at least one embodiment, each of computer hardware 1302 and computer hardware 1306 includes one or more ALUs that perform mathematical functions, such as linear algebraic functions, only on information stored in code and / or data memory 1301 and code and / or data memory 1305, respectively, the result of which is stored in activation memory 1320.
[0135] In at least one embodiment, each of the code and / or data memories 1301 and 1305 and the corresponding computer hardware 1302 and 1306 correspond to different layers of a neural network, such that the resulting activation from one memory / computing pair 1301 / 1302 of code and / or data memory 1301 and computer hardware 1302 is provided as input to a next memory / computing pair 1305 / 1306 of code and / or data memory 1305 and computer hardware 1306 to reflect a conceptual organization of a neural network. In at least one embodiment, each of the memory / computing pairs 1301 / 1302 and 1305 / 1306 may correspond to more than one layer of the neural network. In at least one embodiment, additional memory / compute pairs (not shown) may be included in logic 1315 subsequent to or in parallel with memory / compute pairs 1301 / 1302 and 1305 / 1306. TRAINING AND DEPLOYMENT OF A NEURAL NETWORK
[0136] Fig. 14 illustrates the training and deployment of a deep neural network according to at least one embodiment. In at least one embodiment, the untrained neural network 1406 is trained using a training dataset 1402. In at least one embodiment, the training framework 1404 is a PyTorch framework, while in other embodiments, the training framework 1404 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or another training framework. In at least one embodiment, the training framework 1404 trains an untrained neural network 1406 and facilitates its training using the processing resources described herein to produce a trained neural network 1408. In at least one embodiment, the weights may be selected randomly or by pre-training using a deep belief network.In at least one embodiment, the training may be performed in either a supervised, semi-supervised, or unsupervised manner.
[0137] In at least one embodiment, the untrained neural network 1406 is trained using supervised learning, where the training data set 1402 includes an input paired with a desired output for an input, or where the training data set 1402 includes an input with a known output and an output of the untrained neural network 1406 is manually ranked. In at least one embodiment, the untrained neural network 1406 is trained in a supervised manner and processes inputs from the training data set 1402 and compares the resulting outputs to a set of expected or desired outputs. In at least one embodiment, the errors are then backtracked through the untrained neural network 1406. In at least one embodiment, the training framework 1404 adjusts the weights that control the untrained neural network 1406.In at least one embodiment, the training framework 1404 includes tools for monitoring the convergence of the untrained neural network 1406 toward a model, e.g., the trained neural network 1408, that can generate correct answers, e.g., in the output 1414, based on input data, e.g., a new data set 1412. In at least one embodiment, the training framework 1404 repeatedly trains the untrained neural network 1406, adjusting the weights to refine an output of the untrained neural network 1406 using a loss function and an adaptation algorithm, such as stochastic gradient descent. In at least one embodiment, the training framework 1404 trains the untrained neural network 1406 until the untrained neural network 1406 achieves the desired accuracy.In at least one embodiment, the trained neural network 1408 may then be used to implement any number of machine learning operations.
[0138] In at least one embodiment, the untrained neural network 1406 is trained using unsupervised learning, where the untrained neural network 1406 attempts to train itself using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 1402 comprises input data without associated output or ground truth data. In at least one embodiment, the untrained neural network 1406 can learn groupings within the training dataset 1402 and determine how individual inputs are related to the untrained dataset 1402. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in a trained neural network 1408 capable of performing operations useful in reducing the dimensionality of the new dataset 1412.In at least one embodiment, unsupervised training may also be used for anomaly detection, enabling the identification of data points in the new data set 1412 that deviate from normal patterns of the new data set 1412.
[0139] In at least one embodiment, semi-supervised learning may be used, i.e., a technique in which the training dataset 1402 comprises a mixture of labeled and unlabeled data. In at least one embodiment, the training framework 1404 may be used to perform incremental learning, for example, through transfer learning techniques. In at least one embodiment, incremental learning allows the trained neural network 1408 to adapt to a new dataset 1412 without forgetting the knowledge instilled in the trained neural network 1408 during initial training.
[0140] In at least one embodiment, the training framework 1404 is a framework processed in conjunction with a software development toolkit such as OpenVINO (Open Visual Inference and Neural Network Optimization). In at least one embodiment, an OpenVINO toolkit is a toolkit such as that developed by Intel Corporation of Santa Clara, CA. In at least one embodiment, OpenVINO includes logic 1315 or uses logic 1315 to perform the operations described herein. In at least one embodiment, an SoC, integrated circuit, or processor uses OpenVINO to perform the operations described herein.
[0141] In at least one embodiment, OpenVINO is a toolkit for facilitating the development of applications, in particular neural network applications, for various tasks and operations, such as human vision emulation, speech recognition, natural language processing, recommender systems, and / or variations thereof. In at least one embodiment, OpenVINO supports neural networks such as convolutional neural networks (CNNs), recurrent and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries such as OpenCV, OpenCL, and / or variants thereof.
[0142] In at least one embodiment, OpenVINO supports neural network models for various tasks and operations, such as classification, segmentation, object detection, face recognition, speech recognition, pose estimation (e.g., of humans and / or objects), monocular depth estimation, image colorization, style transfer, action recognition, colorization, and / or variations thereof.
[0143] In at least one embodiment, OpenVINO comprises one or more software modules and / or tools for model optimization, also referred to as model optimizers. In at least one embodiment, a model optimizer is a command-line tool that facilitates the transitions between training and deployment of neural network models. In at least one embodiment, a model optimizer optimizes neural network models for execution on different devices and / or processing units, such as GPU, CPU, PPU, GPGPU, and / or variations thereof. In at least one embodiment, a model optimizer generates an internal representation of a model and optimizes the model to generate an intermediate representation. In at least one embodiment, a model optimizer reduces a number of layers of a model. In at least one embodiment, a model optimizer removes layers of a model used for training.In at least one embodiment, a model optimizer performs various operations of a neural network, such as changing the inputs to a model (e.g., changing the size of the inputs to a model), changing the size of the inputs of a model (e.g., changing the batch size of a model), changing a model structure (e.g., modifying layers of a model), normalization, standardization, quantization (e.g., converting weights of a model from a first representation, e.g., floating point, to a second representation, e.g., integer), and / or variations thereof.
[0144] In at least one embodiment, OpenVINO includes one or more software libraries for inference, also referred to as an inference engine. In at least one embodiment, an inference engine is a C++ library or other suitable library in a programming language. In at least one embodiment, an inference engine is used to infer input data. In at least one embodiment, an inference engine implements various classes for inferring input data and producing one or more results. In at least one embodiment, an inference engine implements one or more API functions to process an intermediate representation, specify input and / or output formats, and / or execute a model on one or more devices.
[0145] In at least one embodiment, OpenVINO provides various capabilities for heterogeneously executing one or more neural network models. In at least one embodiment, heterogeneous execution or heterogeneous computing refers to one or more computing processes and / or systems using one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software capabilities for executing a program on one or more devices. In at least one embodiment, OpenVINO provides various software capabilities for executing a program and / or portions of a program on different devices. In at least one embodiment, OpenVINO provides various software capabilities, for example, to execute a first portion of the code on a CPU and a second portion of the code on a GPU and / or FPGA.In at least one embodiment, OpenVINO provides various software functions to execute one or more layers of a neural network on one or more devices (e.g., a first set of layers on a first device, such as a GPU, and a second set of layers on a second device, such as a CPU).
[0146] In at least one embodiment, OpenVINO includes various functionalities similar to those associated with a CUDA programming model, such as various neural network model operations associated with frameworks such as TensorFlow, PyTorch, and / or variants thereof. In at least one embodiment, one or more CUDA programming model operations are performed with OpenVINO. In at least one embodiment, various systems, methods, and / or techniques described herein are implemented using OpenVINO. DATA CENTER
[0147] Fig. 15 shows an exemplary data center 1500 in which at least one embodiment may be used. In at least one embodiment, the data center 1500 includes a data center infrastructure layer 1510, a framework layer 1520, a software layer 1530, and an application layer 1540.
[0148] In at least one embodiment, as in Fig. 15, the data center infrastructure layer 1510 may include a resource orchestrator 1512, clustered compute resources 1514, and node compute resources (“Node CRs”) 1516(1)-1516(N), where “N” represents a positive integer (which may be a different integer “N” than used in other figures). In at least one embodiment, the node CRs 1516(1)-1516(N) may include any number of central processing units (“CPUs”) or other processors (including accelerators, field programmable gate arrays (FPGAs), graphics processors, etc.), storage devices 1518(1)-1518(N) (e.g., dynamic read-only memory, solid-state storage, or hard disk drives), network input / output devices (“NW I / O”), network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more node CRs among the node CRs 1516(1)-1516(N) may be a server that has one or more of the computing resources listed above.
[0149] In at least one embodiment, the grouped computing resources 1514 may include separate groupings of node CRs housed in one or more racks (not shown) or multiple racks housed in data centers in different geographic locations (also not shown). In at least one embodiment, separate groupings of node CRs within the grouped computing resources 1514 may include grouped computing, networking, memory, or storage resources that may be configured or allocated to support one or more workloads. In at least one embodiment, multiple node CRs, including CPUs or processors, may be grouped in one or more racks to provide computing resources to support one or more workloads.In at least one embodiment, one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0150] In at least one embodiment, resource orchestrator 1512 may configure or otherwise control one or more node CRs 1516(1)-1516(N) and / or clustered computing resources 1514. In at least one embodiment, resource orchestrator 1512 may include a software design infrastructure ("SDI") management entity for data center 1500. In at least one embodiment, resource orchestrator 1512 may include hardware, software, or a combination thereof.
[0151] In at least one embodiment, as in Fig. 15, the framework layer 1520 includes a job scheduler 1522, a configuration manager 1524, a resource manager 1526, and a distributed file system 1528. In at least one embodiment, the framework layer 1520 may include a framework for supporting the software 1532 of the software layer 1530 and / or one or more applications 1542 of the application layer 1540. In at least one embodiment, the software 1532 or the application(s) 1542 may include web-based service software or applications, such as those provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, the framework layer 1520 may be some type of free and open source software web application framework such as Apache Spark™ (hereinafter "Spark"), which may utilize a distributed file system 1528 for processing large amounts of data (e.g., "Big Data"), but is not limited thereto.In at least one embodiment, job scheduler 1522 may include a Spark driver to facilitate scheduling workloads supported by different layers of data center 1500. In at least one embodiment, configuration manager 1524 may be capable of configuring different layers, such as software layer 1530 and framework layer 1520, which include Spark and distributed file system 1528 to support processing large amounts of data. In at least one embodiment, resource manager 1526 may be capable of managing clustered or grouped compute resources allocated to support distributed file system 1528 and job scheduler 1522. In at least one embodiment, the clustered or grouped compute resources may include grouped compute resources 1514 in data center infrastructure layer 1510.In at least one embodiment, the resource manager 1526 may be coordinated with the resource orchestrator 1512 to manage these allocated or assigned computing resources.
[0152] In at least one embodiment, the software 1532 included in software layer 1530 may include software used by at least portions of node CRs 1516(1)-1516(N), clustered computer systems 1514, and / or distributed file systems 1528 of framework layer 1520. In at least one embodiment, one or more types of software may include, but are not limited to, Internet website search software, email virus scanning software, database software, and streaming video content software.
[0153] In at least one embodiment, the application(s) 1542 included in application layer 1540 may include one or more types of applications used by at least portions of node CRs 1516(1)-1516(N), clustered computing resources 1514, and / or distributed file systems 1528 of framework layer 1520. In at least one embodiment, one or more types of applications may include any number of a genome application, a cognitive computing application, and a machine learning application, including, but not limited to, training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in connection with one or more embodiments.
[0154] In at least one embodiment, configuration manager 1524, resource manager 1526, and resource orchestrator 1512 may implement any number and type of self-modifying actions based on any amount and type of data collected in any technically feasible manner. In at least one embodiment, self-modifying actions may relieve an operator of a data center 1500 from potentially making poor configuration decisions and potentially avoiding underutilized and / or poorly performing portions of a data center.
[0155] In at least one embodiment, data center 1500 may include tools, services, software, or other resources to train one or more machine learning models or to predict or infer information using one or more machine learning models according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model may be trained by calculating weighting parameters according to a neural network architecture using software and computational resources described above with respect to data center 1500.In at least one embodiment, trained machine learning models corresponding to one or more neural networks may be used to infer or predict information using the resources described above with respect to the data center 1500 by using weighting parameters calculated by one or more training techniques described herein.
[0156] In at least one embodiment, the data center may use CPUs, application-specific integrated circuits (ASICs), GPUs, FPGAs, or other hardware to perform training and / or inference using the resources described above. Furthermore, one or more of the software and / or hardware resources described above may be configured as a service to enable users to train or infer information, such as image recognition, speech recognition, or other artificial intelligence services.
[0157] Logic 1315 is used to perform inference and / or training operations in connection with one or more embodiments. Details of logic 1315 are described herein in connection with Fig. 13A and / or 13B. In at least one embodiment, logic 1315 may be used in data center 1500 for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0158] In at least one embodiment, at least one Fig. 15 shown or described component is used to control the device used in conjunction with Fig. 1-12. In at least one embodiment, at least one of Fig. 15 is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one component shown in Fig. 15, is used to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and to cause the derivation of two or more contiguous portions of input data to a neural network to be distributed between two or more respective processing cores based at least in part on positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information. AUTONOMOUS VEHICLE
[0159] Fig. 16A shows an example of an autonomous vehicle 1600 according to at least one embodiment. In at least one embodiment, the autonomous vehicle 1600 (alternatively referred to herein as "vehicle 1600") may be, without limitation, a passenger vehicle, such as a car, a truck, a bus, and / or another type of vehicle that accommodates one or more passengers. In at least one embodiment, the vehicle 1600 may be a semi-trailer truck used for transporting goods. In at least one embodiment, the vehicle 1600 may be an aircraft, a robotic vehicle, or another type of vehicle.
[0160] Autonomous vehicles may be described in terms of automation levels defined by the National Highway Traffic Safety Administration ("NHTSA"), a division of the U.S. Department of Transportation, and the Society of Automotive Engineers ("SAE") "Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles" (e.g., Standard No. J3016-201806, published June 15, 2018, Standard No. J3016-201609, published September 30, 2016, and prior and future versions of this standard). In at least one embodiment, the vehicle 1600 may be capable of performing functionality according to one or more of Levels 1 through Level 5 of autonomous driving. For example, in at least one embodiment, the vehicle 1600 may be capable of conditionally automated (Level 3), highly automated (Level 4), and / or fully automated (Level 5), depending on the embodiment.
[0161] In at least one embodiment, vehicle 1600 may include, without limitation, components such as a chassis, a vehicle body, wheels (e.g., 2, 4, 6, 8, 18, etc.), tires, axles, and other components of a vehicle. In at least one embodiment, vehicle 1600 may include, without limitation, a propulsion system 1650, such as an internal combustion engine, a hybrid electric power plant, an all-electric motor, and / or another type of propulsion system. In at least one embodiment, propulsion system 1650 may be connected to a drivetrain of vehicle 1600, which may include, without limitation, a transmission, to facilitate propulsion of vehicle 1600. In at least one embodiment, propulsion system 1650 may be controlled in response to receiving signals from gas pedal / accelerator(s) 1652.
[0162] In at least one embodiment, a steering system 1654, which may include, without limitation, a steering wheel, is used to steer the vehicle 1600 (e.g., along a desired path or route) when the propulsion system 1650 is operating (e.g., when the vehicle 1600 is in motion). In at least one embodiment, the steering system 1654 may receive signals from the steering actuator(s) 1656. In at least one embodiment, a steering wheel may be optional for full automation functionality (Level 5). In at least one embodiment, a brake sensor system 1646 may be used to apply the vehicle brakes in response to receiving signals from brake actuator(s) 1648 and / or brake sensors.
[0163] In at least one embodiment, the controller(s) 1636, which may include, without limitation, one or more system-on-chips (“SoCs”) (in Fig. 16A not shown) and / or graphics processing unit(s) ("GPU(s)"), send signals (e.g., representative of commands) to one or more components and / or systems of the vehicle 1600. For example, in at least one embodiment, the controller(s) 1636 may send signals to actuate the vehicle brakes via the brake actuator(s) 1648, to actuate the steering system 1654 via the steering actuator(s) 1656, to actuate the propulsion system 1650 via the accelerator pedal(s) 1652. In at least one embodiment, the controller(s) 1636 may include one or more built-in (e.g., integrated) computing devices that process sensor signals and issue operational commands (e.g., signals representing commands) to enable autonomous driving and / or to assist a human driver in operating the vehicle 1600.In at least one embodiment, controller(s) 1636 may include a first controller for autonomous driving functions, a second controller for functional safety functions, a third controller for artificial intelligence functions (e.g., computer vision), a fourth controller for infotainment functions, a fifth controller for emergency redundancy, and / or other controllers. In at least one embodiment, a single controller may perform two or more of the above functions, two or more controllers may perform a single function, and / or any combination thereof.
[0164] In at least one embodiment, the controller(s) 1636 provide signals to control one or more components and / or systems of the vehicle 1600 in response to sensor data received from one or more sensors (e.g., sensor inputs). In at least one embodiment, sensor data may be obtained, for example and without limitation, from Global Navigation Satellite System ("GNSS") sensor(s) 1658 (e.g., Global Positioning System sensor(s)), RADAR sensor(s) 1660, ultrasonic sensor(s) 1662, LIDAR sensor(s) 1664, Inertial Measurement Unit ("IMU") sensor(s) 1666 (e.g., accelerometer(s), gyroscope(s), magnetic compass(es), magnetometer(s), etc.), microphone(s) 1696, stereo camera(s) 1668, wide-angle camera(s) 1670 (e.g., fisheye cameras), infrared camera(s) 1672, environmental camera(s) 1674 (e.g., 360-degree cameras), long-range cameras (in Fig. 16A not shown), mid-range camera(s) (in Fig. 16A not shown), speed sensor(s) 1644 (e.g., for measuring the speed of the vehicle 1600), vibration sensor(s) 1642, steering sensor(s) 1640, brake sensor(s) (e.g., as part of the brake sensor system 1646), and / or other types of sensors.
[0165] In at least one embodiment, one or more of the controller(s) 1636 may receive inputs (e.g., represented by input data) from an instrument cluster 1632 of the vehicle 1600 and provide outputs (e.g., represented by output data, display data, etc.) via a human-machine interface (“HMI”) display 1634, an audible annunciator, a speaker, and / or via other components of the vehicle 1600. In at least one embodiment, the outputs may include information such as vehicle speed, RPM, time, map information (e.g., a high-resolution map (in Fig. 16A not shown)), location data (e.g., the location of the vehicle 1600, as on a map), direction, location of other vehicles (e.g., an occupancy grid), information about objects and the status of objects as perceived by the controller(s) 1636, etc. For example, in at least one embodiment, the HMI display 1634 may include information about the presence of one or more objects (e.g., a road sign, a warning sign, a changing traffic light, etc.) and / or information about maneuvers the vehicle has performed, is currently performing, or will perform (e.g., lane change now, exit 34B in two miles, etc.).
[0166] In at least one embodiment, the vehicle 1600 further includes a network interface 1624 that may utilize wireless antenna(s) 1626 and / or modem(s) to communicate over one or more networks. For example, in at least one embodiment, the network interface 1624 may be capable of communicating over Long-Term Evolution ("LTE"), Wideband Code Division Multiple Access ("WCDMA"), Universal Mobile Telecommunications System ("UMTS"), Global System for Mobile Communication ("GSM"), IMT-CDMA Multi-Carrier ("CDMA2000") networks, etc. In at least one embodiment, the wireless antenna(s) 1626 may also enable communication between objects in the environment (e.g., vehicles, mobile devices, etc.) using local area networks such as Bluetooth, Bluetooth Low Energy ("LE"), Z-Wave, ZigBee, etc., and / or low-power wide area networks ("LPWANs") such as LoRaWAN, SigFox, etc.
[0167] Logic 1315 is used to perform inference and / or training operations in connection with one or more embodiments. Details of logic 1315 are described herein in connection with Fig. 13A and / or 13B. In at least one embodiment, logic 1315 in vehicle 1600 may be used for inference or prediction operations based at least in part on weight parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0168] In at least one embodiment, at least one Fig. 20A is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig. 20A is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 20A to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and cause the derivation of two or more contiguous portions of input data to a neural network to be distributed among two or more respective processing cores based at least in part on positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0169] Fig. 16B shows an example of camera positions and fields of view for the autonomous vehicle 1600 of Fig. 16A according to at least one embodiment. In at least one embodiment, the cameras and the respective fields of view represent an exemplary embodiment and are not to be considered limiting. For example, in at least one embodiment, additional and / or alternative cameras may be included and / or cameras may be arranged at different locations on the vehicle 1600.
[0170] In at least one embodiment, the camera types for cameras may include, but are not limited to, digital cameras that may be adapted for use with components and / or systems of the vehicle 1600. In at least one embodiment, the camera(s) may operate at Automotive Safety Integrity Level ("ASIL") B and / or another ASIL. In at least one embodiment, the camera types may achieve any image capture rate, such as 60 frames per second (fps), 1220 fps, 240 fps, etc., depending on the embodiment. In at least one embodiment, the cameras may use rolling shutter, global shutter, another shutter type, or a combination thereof.In at least one embodiment, the color filter array may comprise a red-clear-clear color filter array ("RCCC"), a red-clear-blue color filter array ("RCCB"), a red-blue-green clear color filter array ("RBGC"), a Foveon X3 color filter array, a Bayer sensor color filter array ("RGGB"), a monochrome sensor color filter array, and / or another type of color filter array. In at least one embodiment, clear-pixel cameras, such as cameras with an RCCC, an RCCB, and / or an RBGC color filter array, may be used to increase light sensitivity.
[0171] In at least one embodiment, one or more cameras may be used to implement advanced driver assistance systems ("ADAS") (e.g., as part of a redundant or fail-safe design). For example, in at least one embodiment, a multi-function mono camera may be installed, including functions such as lane departure warning, traffic sign assist, and intelligent headlight control. In at least one embodiment, one or more of the cameras (e.g., all cameras) may simultaneously record and provide image data (e.g., video).
[0172] In at least one embodiment, one or more cameras may be mounted in a mounting fixture, such as a custom-designed (three-dimensional ("3D") printed) fixture, to eliminate stray light and reflections within the vehicle 1600 (e.g., reflections from the dashboard reflected in the windshield mirrors) that may impair the camera's ability to capture images. With respect to the mounting of exterior mirrors, in at least one embodiment, exterior mirror assemblies may be custom 3D printed so that a camera mounting plate conforms to the shape of an exterior mirror. In at least one embodiment, camera(s) may be integrated into the exterior mirrors. In at least one embodiment, for side-facing cameras, the camera(s) may also be integrated into four pillars at each corner of the cabin.
[0173] In at least one embodiment, cameras with a field of view encompassing portions of an environment in front of the vehicle 1600 (e.g., forward-facing cameras) may be used for the environmental view to help identify forward paths and obstacles, and to provide information critical to establishing an occupancy grid and / or determining preferred vehicle paths with the aid of one or more controllers 1636 and / or control SoCs. In at least one embodiment, forward-facing cameras may be used to perform many similar ADAS functions as LIDAR, including, without limitation, emergency braking, pedestrian detection, and collision avoidance.In at least one embodiment, forward-facing cameras may also be used for ADAS features and systems, including, without limitation, lane departure warnings (“LDW”), autonomous cruise control (“ACC”), and / or other features such as traffic sign recognition.
[0174] In at least one embodiment, a plurality of cameras may be used in a forward-facing configuration, including, for example, a monocular camera platform comprising a complementary metal oxide semiconductor (CMOS) color imager. In at least one embodiment, a wide-angle camera 1670 may be used to detect objects entering view from the periphery (e.g., pedestrians, crossing traffic, or bicycles). Although in Fig. 16B shows only one wide-angle camera 1670, in other embodiments, the vehicle 1600 may include any number (including zero) of wide-angle cameras. In at least one embodiment, any number of long-range camera(s) 1698 (e.g., a long-range stereo camera pair) may be used for depth-based object detection, particularly for objects for which a neural network has not yet been trained. In at least one embodiment, the long-range camera(s) 1698 may also be used for object detection and classification, as well as basic object tracking.
[0175] In at least one embodiment, any number of stereo cameras 1668 may also be in a forward-facing configuration. In at least one embodiment, one or more of the stereo cameras 1668 may include an integrated control unit comprising a scalable processing unit that may provide a programmable logic ("FPGA") and a multi-core microprocessor with an integrated network interface ("CAN") or Ethernet interface on a single chip. In at least one embodiment, such a unit may be used to create a 3D map of the environment of the vehicle 1600 that includes a distance estimate for all points in an image.In at least one embodiment, one or more of the stereo camera(s) 1668 may comprise, without limitation, compact stereo vision sensors, which may comprise, without limitation, two camera lenses (one on the left and one on the right) and an image processing chip that can measure the distance between the vehicle 1600 and the target object and use the generated information (e.g., metadata) to activate autonomous emergency braking and lane departure warning functions. In at least one embodiment, other types of stereo camera(s) 1668 may be used in addition to or alternatively to those described herein.
[0176] In at least one embodiment, cameras with a field of view that includes portions of the environment on the sides of the vehicle 1600 (e.g., side cameras) may be used for the environment view and provide information used to create and update an occupancy grid and to generate side impact warnings. In at least one embodiment, the environment camera(s) 1674 (e.g., four environment cameras, as in Fig.16B) may be positioned on the vehicle 1600. In at least one embodiment, the surround camera(s) 1674 may include, without limitation, any number and combination of wide-angle cameras, fisheye cameras, 360-degree cameras, and / or similar cameras. For example, in at least one embodiment, four fisheye cameras may be positioned on a front, a rear, and the sides of the vehicle 1600. In at least one embodiment, the vehicle 1600 may utilize three surround camera(s) 1674 (e.g., left, right, and rear) and may utilize one or more other cameras (e.g., a forward-facing camera) as a fourth surround camera.
[0177] In at least one embodiment, cameras with a field of view that includes portions of an environment behind the vehicle 1600 (e.g., rearview cameras) may be used for parking assistance, surround view, rear collision warnings, and the creation and updating of an occupancy grid. In at least one embodiment, a variety of cameras may be used, including, but not limited to, cameras that are also suitable as forward-facing cameras (e.g., long-range cameras 1698 and / or mid-range camera(s) 1676, stereo camera(s) 1668, infrared camera(s) 1672, etc.), as described herein.
[0178] In at least one embodiment, at least one Fig. 20B is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig.20B is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 20B to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and cause the derivation of two or more contiguous portions of input data to a neural network to be distributed among two or more respective processing cores based at least in part on positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0179] Fig. 16C is a block diagram illustrating an example system architecture for the autonomous vehicle 1600 of Fig. 16A according to at least one embodiment. In at least one embodiment, each of the components, features, and systems of the vehicle 1600 is Fig.16C as connected via a bus 1602. In at least one embodiment, bus 1602 may include, without limitation, a CAN data interface (alternatively referred to herein as a "CAN bus"). In at least one embodiment, a CAN may be a network within vehicle 1600 used to support the control of various features and functions of vehicle 1600, such as brake application, acceleration, braking, steering, windshield wipers, etc. In at least one embodiment, bus 1602 may be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., a CAN ID). In at least one embodiment, bus 1602 may be read to determine steering wheel angle, vehicle speed, engine speed, button positions, and / or other indications of vehicle status.In at least one embodiment, bus 1602 may be a CAN bus that is ASIL B compliant.
[0180] In at least one embodiment, FlexRay and / or Ethernet protocols may be used in addition to or as an alternative to CAN. In at least one embodiment, there may be any number of buses that make up bus 1602, which may include, without limitation, zero or more CAN buses, zero or more FlexRay buses, zero or more Ethernet buses, and / or zero or more other types of buses with different protocols. In at least one embodiment, two or more buses may be used to perform different functions and / or may be used for redundancy. For example, a first bus may be used for collision avoidance functionality and a second bus may be used for actuation control.In at least one embodiment, each bus of bus 1602 may communicate with any components of vehicle 1600, and two or more buses of bus 1602 may communicate with corresponding components. In at least one embodiment, each of any number of system(s) on chip(s) ("SoC(s)") 1604 (such as SoC 1604(A) and SoC 1604(B)), each of controller(s) 1636, and / or each computer within the vehicle may have access to the same input data (e.g., inputs from sensors of vehicle 1600) and be connected to a common bus, such as a CAN bus.
[0181] In at least one embodiment, the vehicle 1600 may include one or more controllers 1636 as described herein with reference to Fig.16A. In at least one embodiment, the controller(s) 1636 may be used for a variety of functions. In at least one embodiment, the controller(s) 1636 may be coupled to any of various other components and systems of the vehicle 1600 and may be used for control of the vehicle 1600, the artificial intelligence of the vehicle 1600, the infotainment for the vehicle 1600, and / or other functions.
[0182] In at least one embodiment, vehicle 1600 may include any number of SoCs 1604. In at least one embodiment, each of SoCs 1604 may include, without limitation, central processing units ("CPU(s)") 1606, graphics processing units ("GPU(s)") 1608, processor(s) 1610, cache(s) 1612, accelerators 1614, data storage 1616, and / or other components and features not shown. In at least one embodiment, SoC(s) 1604 may be used to control vehicle 1600 in a variety of platforms and systems. For example, in at least one embodiment, SoC(s) 1604 in a system (e.g., system of vehicle 1600) may be combined with a high-definition (“HD”) map 1622 that may receive map refreshes and / or updates via network interface 1624 from one or more servers (in Fig. 16C not shown).
[0183] In at least one embodiment, the CPU(s) 1606 may comprise a CPU cluster or CPU complex (alternatively referred to herein as a "CCPLEX"). In at least one embodiment, the CPU(s) 1606 may comprise multiple cores and / or level two ("L2") caches. For example, in at least one embodiment, the CPU(s) 1606 may comprise eight cores in a coherent multiprocessor configuration. In at least one embodiment, the CPU(s) 1606 may comprise four dual-core clusters, each cluster having a dedicated L2 cache (e.g., a 2 megabyte (MB) L2 cache). In at least one embodiment, the CPU(s) 1606 (e.g., CCPLEX) may be configured to support concurrent cluster operations such that any combination of clusters of CPU(s) 1606 may be active at any given time.
[0184] In at least one embodiment, one or more of the CPU(s) 1606 may implement power management features, including, without limitation, one or more of the following features: individual hardware blocks may be automatically clocked when idle to conserve dynamic power; each core clock may be clocked when such core is not actively executing instructions due to the execution of Wait for Interrupt ("WFI") / Wait for Event ("WFE") instructions; each core may be independently power-driven; each core cluster may be independently clock-driven if all cores are clock-driven or power-driven; and / or each core cluster may be independently power-driven if all cores are power-driven.In at least one embodiment, CPU(s) 1606 may further implement an enhanced power state management algorithm, where allowable power states and expected wake-up times are specified, and the hardware / microcode determines which power state is best for the core, cluster, and CCPLEX. In at least one embodiment, processing cores may support simplified sequences for entering power states in software, offloading the work to microcode.
[0185] In at least one embodiment, GPU(s) 1608 may comprise an integrated GPU (alternatively referred to herein as an "iGPU"). In at least one embodiment, GPU(s) 1608 may be programmable and efficient for parallel workloads. In at least one embodiment, GPU(s) 1608 may utilize an extended Tensor instruction set. In at least one embodiment, GPU(s) 1608 may comprise one or more streaming microprocessors, where each streaming microprocessor may comprise a Level 1 ("L1") cache (e.g., an L1 cache with a memory capacity of at least 96 KB), and two or more streaming microprocessors may share an L2 cache (e.g., an L2 cache with a memory capacity of 512 KB). In at least one embodiment, GPU(s) 1608 may comprise at least eight streaming microprocessors.In at least one embodiment, the GPU(s) 1608 may use one or more application programming interfaces (API(s)) for computation. In at least one embodiment, the GPU(s) 1608 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA model).
[0186] In at least one embodiment, one or more of the GPU(s) 1608 may be power-optimized for best performance in automotive and embedded use cases. For example, in at least one embodiment, the GPU(s) 1608 may be fabricated on fin field-effect transistor ("FinFET") circuits. In at least one embodiment, each streaming microprocessor may include a number of mixed-precision processing cores divided into multiple blocks. For example, 64 PF32 cores and 32 FP64 cores could be divided into four processing blocks. In at least one embodiment, each processing block could be assigned 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two NVIDIA mixed-precision Tensor cores for deep learning matrix arithmetic, a level zero ("L0") instruction cache, a scheduler (e.g., warp scheduler) or sequencer, a dispatch unit, and / or a 64 KB register file.In at least one embodiment, streaming microprocessors may include independent parallel integer and floating-point datapaths to enable efficient execution of workloads with a mix of computations and addressing calculations. In at least one embodiment, streaming microprocessors may include independent thread scheduling capability to enable finer-grained synchronization and collaboration between parallel threads. In at least one embodiment, streaming microprocessors may include a combined L1 data cache and shared memory unit to improve performance while simplifying programming.
[0187] In at least one embodiment, one or more of the GPU(s) 1608 may include high-bandwidth memory ("HBM") and / or a 16 GB HBM2 memory subsystem to provide, in some examples, a peak memory bandwidth of approximately 900 GB / second. In at least one embodiment, in addition to or alternatively to the HBM memory, a synchronous graphics random-access memory ("SGRAM") may be used, such as a synchronous graphics double data rate random-access memory type 5 ("GDDR5").
[0188] In at least one embodiment, the GPU(s) 1608 may comprise a unified memory technology. In at least one embodiment, address translation services ("ATS") support may be used to allow the GPU(s) 1608 to directly access page tables of the CPU(s) 1606. In at least one embodiment, an address translation request may be communicated to the CPU(s) 1606 when a GPU of the GPU(s) 1608 memory management unit ("MMU") encounters a fault. In response, the CPU of the CPU(s) 1606 may look up a virtual-physical mapping for an address in its page tables and transmit the translation back to the GPU(s) 1608, in at least one embodiment.In at least one embodiment, unified memory technology may enable a single unified virtual address space for the memory of both the CPU(s) 1606 and the GPU(s) 1608, simplifying programming of the GPU(s) 1608 and porting of applications to the GPU(s) 1608.
[0189] In at least one embodiment, the GPU(s) 1608 may include any number of access counters that may track the frequency of access by the GPU(s) 1608 to the memory of other processors. In at least one embodiment, access counters may help ensure that memory pages are moved to the physical memory of a processor that accesses pages most frequently, thereby improving the efficiency of memory regions shared between processors.
[0190] In at least one embodiment, one or more of the SoC(s) 1604 may include any number of cache(s) 1612, including those described herein. For example, in at least one embodiment, the cache(s) 1612 could include a Level 3 ("L3") cache available to both the CPU(s) 1606 and the GPU(s) 1608 (e.g., connected to the CPU(s) 1606 and the GPU(s) 1608). In at least one embodiment, the cache(s) 1612 may include a write-back cache that can track the states of lines, for example, using a cache coherence protocol (e.g., MEI, MESI, MSI, etc.). In at least one embodiment, an L3 cache may include 4 MB of memory or more, depending on the embodiment, although smaller cache sizes may also be used.
[0191] In at least one embodiment, one or more of the SoC(s) 1604 may include one or more accelerators 1614 (e.g., hardware accelerators, software accelerators, or a combination thereof). In at least one embodiment, the SoC(s) 1604 may include a hardware acceleration cluster, which may include optimized hardware accelerators and / or large on-chip memory. In at least one embodiment, large on-chip memory (e.g., 4 MB of SRAM) may enable a hardware acceleration cluster to accelerate neural networks and other computations. In at least one embodiment, a hardware acceleration cluster may be used to supplement the GPU(s) 1608 and offload some tasks from the GPU(s) 1608 (e.g., to free up more cycles of the GPU(s) 1608 to perform other tasks).In at least one embodiment, the accelerator(s) 1614 could be used for targeted workloads (e.g., perception, convolutional neural networks ("CNNs"), recurrent neural networks ("RNNs"), etc.) that are robust enough to be suitable for acceleration. In at least one embodiment, a CNN may include a region-based or regional neural network ("RCNN") and a fast RCNN (e.g., for object detection), or another type of CNN.
[0192] In at least one embodiment, the accelerator(s) 1614 (e.g., hardware acceleration clusters) may include one or more deep learning accelerators ("DLAs"). In at least one embodiment, the DLA(s) may include, without limitation, one or more tensor processing units ("TPUs") that may be configured to provide an additional tens of trillion operations per second for deep learning applications and inferences. In at least one embodiment, TPUs may be accelerators configured and optimized to perform image processing functions (e.g., for CNNs, RCNNs, etc.). In at least one embodiment, the DLA(s) may be further optimized for a particular set of neural network types and floating-point operations, as well as for inferences.In at least one embodiment, the design of DLA(s) can provide more performance per millimeter than a typical general-purpose GPU, typically far exceeding the performance of a CPU. In at least one embodiment, the TPU(s) can perform multiple functions, including a single-instance convolution function, supporting, for example, INT8, INT16, and FP16 data types for both features and weights, as well as post-processing functions.In at least one embodiment, DLA(s) can quickly and efficiently execute neural networks, particularly CNNs, on processed or unprocessed data for a variety of functions, including, for example and without limitation: a CNN for object identification and recognition using data from camera sensors; a CNN for distance estimation using data from camera sensors; a CNN for emergency vehicle detection and identification and recognition using data from microphones; a CNN for facial recognition and vehicle owner identification using data from camera sensors; and / or a CNN for safety-relevant and / or security-related events.
[0193] In at least one embodiment, DLA(s) may perform any function of GPU(s) 1608, and by using an inference accelerator, for example, a developer may dedicate either DLA(s) or GPU(s) 1608 to any function. For example, in at least one embodiment, a developer may focus the processing of CNNs and floating-point operations on DLA(s) and leave other functions to GPU(s) 1608 and / or accelerator(s) 1614.
[0194] In at least one embodiment, the accelerator(s) 1614 may comprise a programmable image processing accelerator ("PVA"), which may also be referred to herein as a computer vision accelerator. In at least one embodiment, the PVA may be designed and configured to accelerate image processing algorithms for advanced driver assistance systems ("ADAS") 1638, autonomous driving, augmented reality ("AR"), and / or virtual reality ("VR") applications. In at least one embodiment, the PVA may provide a balance between performance and flexibility. For example, in at least one embodiment, each PVA may comprise, without limitation, any number of reduced instruction set ("RISC") cores, direct memory access ("DMA") cores, and / or any number of vector processors.
[0195] In at least one embodiment, RISC cores may interact with image sensors (e.g., image sensors of cameras described herein), image signal processor(s), etc. In at least one embodiment, each RISC core may include any amount of memory. In at least one embodiment, the RISC cores may use any number of protocols, depending on the embodiment. In at least one embodiment, RISC cores may execute a real-time operating system ("RTOS"). In at least one embodiment, RISC cores may be implemented with one or more integrated circuits, application-specific integrated circuits ("ASICs"), and / or memory devices. In at least one embodiment, RISC cores could include, for example, an instruction cache and / or tightly coupled RAM.
[0196] In at least one embodiment, DMA may enable components of the PVA to access system memory independently of the CPU(s) 1606. In at least one embodiment, DMA may support any number of features designed to optimize a PVA, including, but not limited to, support for multi-dimensional addressing and / or circular addressing. In at least one embodiment, DMA may support up to six or more dimensions of addressing, which may include, without limitation, block width, block height, block depth, horizontal block stepping, vertical block stepping, and / or depth stepping.
[0197] In at least one embodiment, vector processors may be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing functions. In at least one embodiment, a PVA may include a PVA core and two vector processing subsystem partitions. In at least one embodiment, a PVA core may include a processor subsystem, DMA engine(s) (e.g., two DMA engines), and / or other peripherals. In at least one embodiment, a vector processing subsystem may operate as the primary engine of a PVA and include a vector processing unit (“VPU”), an instruction cache, and / or a vector memory (e.g., “VMEM”).In at least one embodiment, the VPU core may include a digital signal processor, such as a single instruction multiple data ("SIMD") and very long instruction word ("VLIW") digital signal processor. In at least one embodiment, a combination of SIMD and VLIW may increase throughput and speed.
[0198] In at least one embodiment, each of the vector processors may include an instruction cache and be connected to dedicated memory. Consequently, in at least one embodiment, each vector processor may be configured to operate independently of other vector processors. In at least one embodiment, vector processors included in a particular PVA may be configured to use data parallelism. For example, in at least one embodiment, a plurality of vector processors included in a single PVA may execute a common computer vision algorithm, but on different regions of an image.In at least one embodiment, the vector processors included in a particular PVA may simultaneously execute different image processing algorithms on an image, or even different algorithms on successive images or portions of an image. In at least one embodiment, among other things, any number of PVAs may be included in a hardware acceleration cluster, and each PVA may include any number of vector processors. In at least one embodiment, the PVA may include additional error-correcting code ("ECC") memory to increase the security of the overall system.
[0199] In at least one embodiment, the accelerator(s) 1614 may include an on-chip computer vision network and static random access memory ("SRAM") to provide high-bandwidth, low-latency SRAM for the accelerator(s) 1614. In at least one embodiment, the on-chip memory may include at least 4 MB of SRAM, including, for example and without limitation, eight field-configurable memory blocks accessible by both a PVA and a DLA. In at least one embodiment, each pair of memory blocks may include an extended peripheral bus ("APB") interface, configuration circuitry, a controller, and a multiplexer. In at least one embodiment, any type of memory may be used.In at least one embodiment, a PVA and a DLA may access the memory via a backbone that provides high-speed access to the memory for a PVA and a DLA. In at least one embodiment, a backbone may include an on-chip computer vision network that interconnects a PVA and a DLA to the memory (e.g., using APB).
[0200] In at least one embodiment, an on-chip computer vision network may include an interface that determines that both a PVA and a DLA are providing ready and valid signals before transmitting control signals / addresses / data. In at least one embodiment, an interface may provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst communication for continuous data transmission. In at least one embodiment, an interface may conform to International Organization for Standardization ("ISO") 26262 or International Electrotechnical Commission ("EC") 61508 standards, although other standards and protocols may be used.
[0201] In at least one embodiment, one or more of the SoC(s) 1604 may include a hardware accelerator for real-time ray tracing. In at least one embodiment, the real-time ray tracing hardware accelerator may be used for quickly and efficiently determining positions and extents of objects (e.g., within a world model), generating real-time visualization simulations, radar signal interpretation, sound propagation synthesis and / or analysis, simulating sonar systems, general wave propagation simulation, comparing with lidar data for localization and / or other functions, and / or for other purposes.
[0202] In at least one embodiment, the accelerator(s) 1614 may have a wide range of uses for autonomous driving. In at least one embodiment, a PVA may be used for critical processing steps in ADAS and autonomous vehicles. In at least one embodiment, the capabilities of a PVA are well suited to algorithmic domains that require predictable, low-power, and low-latency processing. In other words, a PVA is well suited for semi-dense or dense regular computations, even on small datasets, that require predictable, low-latency, and low-power runtimes. In at least one embodiment, such as in vehicle 1600, PVAs could be designed to execute classical computer vision algorithms because they can be efficient at object detection and integer math processing.
[0203] For example, according to at least one embodiment of the technology, a PVA is used to perform computer stereo vision. In at least one embodiment, a semi-global matching-based algorithm may be used in some examples, although this is not intended to be limiting. In at least one embodiment, Level 3-5 autonomous driving applications utilize motion estimation / stereo matching while driving (e.g., structure from motion, pedestrian detection, lane detection, etc.). In at least one embodiment, a PVA may perform computer stereo vision functions on inputs from two monocular cameras.
[0204] In at least one embodiment, a PVA may be used to perform dense optical flow. In at least one embodiment, a PVA could, for example, process raw radar data (e.g., using a 4D Fast Fourier Transform) to provide processed radar data. In at least one embodiment, a PVA is used for depth-of-flight processing, processing raw time-of-flight data to provide, for example, processed time-of-flight data.
[0205] In at least one embodiment, a DLA may be used to power any type of network to improve control and driving safety, including, for example and without limitation, a neural network that outputs a confidence measure for each object detection. In at least one embodiment, confidence may be represented or interpreted as a probability, or as the relative "weight" of each detection compared to other detections. In at least one embodiment, a confidence measure allows the system to make further decisions about which detections should be considered true positives and which should be considered false positives. In at least one embodiment, a system may set a threshold for the confidence measure and consider only detections that exceed the threshold to be true positives.In an embodiment using an automatic emergency braking ("AEB") system, false positive detections would cause the vehicle to automatically perform emergency braking, which is clearly undesirable. In at least one embodiment, highly confident detections can be considered as triggers for AEB. In at least one embodiment, a DLA can employ a neural network to regress the confidence value.In at least one embodiment, the neural network may use as input at least a subset of parameters, such as the dimensions of the bounding box, the ground plane estimate obtained (e.g., from another subsystem), the output of the IMU sensor(s) 1666 correlated with the orientation of the vehicle 1600, the range, the 3D position estimates of the object obtained from the neural network and / or other sensors (e.g., LIDAR sensor(s) 1664 or RADAR sensor(s) 1660), and others.
[0206] In at least one embodiment, one or more SoC(s) 1604 may include one or more data stores 1616 (e.g., memories). In at least one embodiment, the data store(s) 1616 may be on-chip memory of the SoC(s) 1604, which may store neural networks to be executed on the GPU(s) 1608 and / or a DLA. In at least one embodiment, the capacity of the data store(s) 1616 may be large enough to store multiple neural network instances for redundancy and security. In at least one embodiment, the data store(s) 1616 may include L2 or L3 cache(s).
[0207] In at least one embodiment, one or more of the SoC(s) 1604 may include any number of processor(s) 1610 (e.g., embedded processors). In at least one embodiment, the processor(s) 1610 may include a boot and power management processor, which may be a dedicated processor and subsystem to handle boot power and management functions and associated security enforcement. In at least one embodiment, a boot and power management processor may be part of a boot sequence of SoC(s) 1604 and provide runtime power management services.In at least one embodiment, a boot power and management processor may provide clock and voltage programming, assist with low-power state transitions, manage the thermal and temperature sensors of the SoC(s) 1604, and / or manage the power states of the SoC(s) 1604. In at least one embodiment, each temperature sensor may be implemented as a ring oscillator whose output frequency is proportional to temperature, and SoC(s) 1604 may use ring oscillators to sense the temperatures of the CPU(s) 1606, GPU(s) 1608, and / or accelerator(s) 1614.In at least one embodiment, if temperatures are determined to exceed a threshold, a boot and power management processor may enter a temperature fault routine and place SoC(s) 1604 into a lower power state and / or place vehicle 1600 into a chauffeur-to-safe stop mode (e.g., bring vehicle 1600 to a safe stop).
[0208] In at least one embodiment, processor(s) 1610 may further comprise a set of embedded processors that may serve as an audio processing engine, which may be an audio subsystem enabling full hardware support for multi-channel audio across multiple interfaces and a wide and flexible range of audio I / O interfaces. In at least one embodiment, an audio processing engine is a dedicated processing core including a digital signal processor with dedicated RAM.
[0209] In at least one embodiment, the processor(s) 1610 may further include an "always on" processor engine that may provide the necessary hardware functions to support low-power sensor management and wake use cases. In at least one embodiment, an "always on" processor engine may include, without limitation, a processor core, tightly coupled memory, supporting peripherals (e.g., timers and interrupt controllers), various I / O control peripherals, and routing logic.
[0210] In at least one embodiment, the processor(s) 1610 may further comprise a Safety Cluster Engine, which may comprise, without limitation, a dedicated processor subsystem for handling safety management for automotive applications. In at least one embodiment, a Safety Cluster Engine may comprise, without limitation, two or more processor cores, tightly coupled memory, supporting peripherals (e.g., timers, an interrupt controller, etc.), and / or routing logic. In a safety mode, in at least one embodiment, two or more cores may operate in a lockstep mode, functioning as a single core with comparison logic to detect any differences between their operations.In at least one embodiment, processor(s) 1610 may further comprise a real-time camera engine, which may, without limitation, comprise a dedicated processor subsystem for handling real-time camera management. In at least one embodiment, processor(s) 1610 may further comprise a high dynamic range signal processor, which may, without limitation, comprise an image signal processor that is a hardware engine that is part of a camera processing pipeline.
[0211] In at least one embodiment, processor(s) 1610 may include a video image compositor, which may be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to generate a final image for a player window. In at least one embodiment, a video image compositor may perform lens distortion correction on the wide-angle camera(s) 1670, the surround camera(s) 1674, and / or the sensors of the in-cabin surveillance camera(s). In at least one embodiment, the sensor(s) of the in-cabin surveillance camera(s) is / are preferably monitored by a neural network running on another instance of SoC 1604 and configured to detect and respond to events in the cabin.In at least one embodiment, a system within the vehicle may perform lip reading without limitation to activate cellular service and place a call, dictate emails, change a vehicle's destination, activate or change a vehicle's infotainment system and settings, or enable voice-activated web browsing. In at least one embodiment, certain functions are available to the driver when the vehicle is operating in an autonomous mode and are disabled otherwise.
[0212] In at least one embodiment, a video image compositor may include enhanced temporal noise reduction for both spatial and temporal noise reduction. For example, in at least one embodiment where motion occurs in a video, the noise reduction appropriately weights spatial information and reduces the weight of information provided by neighboring frames. In at least one embodiment where an image or a portion of an image does not include motion, the temporal noise reduction performed by the video image compositor may use information from a previous image to reduce noise in the current image.
[0213] In at least one embodiment, a video image compositor may also be configured to perform stereo rectification on the input stereo lens images. In at least one embodiment, a video image compositor may further be used for interface composition when an operating system desktop is in use and the GPU(s) 1608 are not required to continuously render new surfaces. In at least one embodiment, a video image compositor may be used to offload the GPU(s) 1608 to improve performance and responsiveness when the GPU(s) 1608 are turned on and active and performing 3D rendering.
[0214] In at least one embodiment, one or more of SoC(s) 1604 may further include a Mobile Industrial Processor Serial Interface ("MIPI") for receiving video and input from cameras, a high-speed interface, and / or a video input block that may be used for a camera and associated pixel input functions. In at least one embodiment, one or more of SoC(s) 1604 may further include one or more input / output controllers that may be controlled by software and may be used to receive I / O signals that are not tied to a specific role.
[0215] In at least one embodiment, one or more of SoC(s) 1604 may further include a wide range of peripheral interfaces to enable communication with peripherals, audio encoders / decoders ("codecs"), power management, and / or other devices. In at least one embodiment, SoC(s) 1604 may be used to receive data from cameras (e.g., via Gigabit Multimedia Serial Link and Ethernet channels), sensors (e.g., LIDAR sensor(s) 1664, RADAR sensor(s) 1660, etc., which may be connected via Ethernet channels), data from bus 1602 (e.g., speed of vehicle 1600, steering wheel position, etc.), data from GNSS sensor(s) 1658 (e.g., connected via an Ethernet bus or a CAN bus), etc.In at least one embodiment, one or more of SoC(s) 1604 may further include dedicated high-performance mass storage controllers, which may include their own DMA engines and which may be used to free CPU(s) 1606 from routine data management tasks.
[0216] In at least one embodiment, the SoC(s) 1604 may be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture, leveraging computer vision and ADAS techniques for diversity and redundancy, and providing a platform for a flexible, reliable driving software stack along with deep learning tools. In at least one embodiment, the SoC(s) 1604 may be faster, more reliable, and even more power and space efficient than conventional systems. For example, in at least one embodiment, the accelerator(s) 1614, in combination with the CPU(s) 1606, the GPU(s) 1608, and the data memory(s) 1616, may provide a fast, efficient platform for Level 3-5 autonomous vehicles.
[0217] In at least one embodiment, computer vision algorithms may be executed on CPUs that can be configured using a high-level programming language, such as C, to execute a variety of processing algorithms on a wide variety of visual data. However, in at least one embodiment, CPUs are often unable to meet the performance requirements of many image processing applications, such as execution time and power consumption. In at least one embodiment, many CPUs are unable to execute complex object detection algorithms in real time, which are used in in-vehicle ADAS applications and in practical Level 3-5 autonomous vehicles.
[0218] The embodiments described herein enable multiple neural networks to be executed simultaneously and / or sequentially and the results to be combined to enable Level 3-5 autonomous driving functionality. For example, in at least one embodiment, a CNN executing on a DLA or a discrete GPU (e.g., GPU(s) 1620) may include text and word recognition that enables the reading and understanding of traffic signs, including signs for which a neural network has not been specifically trained. In at least one embodiment, a DLA may further include a neural network capable of identifying, interpreting, and semantically understanding a sign and passing this semantic understanding to path planning modules running on a CPU complex.
[0219] In at least one embodiment, multiple neural networks may be operated simultaneously, such as in Level 3, 4, or 5 driving. For example, in at least one embodiment, a warning sign stating "Caution: Flashing lights indicate icing" along with an electric light may be interpreted independently or jointly by multiple neural networks. In at least one embodiment, such a warning sign may itself be identified as a traffic sign by a first deployed neural network (e.g., a trained neural network), and the text "Flashing lights indicate icy conditions" may be interpreted by a second deployed neural network, which informs a vehicle's path planning software (preferably executing on a CPU complex) that, when flashing lights are detected, icy conditions are present.In at least one embodiment, a turn signal may be identified by operating a third neural network over multiple frames, which informs a vehicle's path planning software of the presence (or absence) of turn signals. In at least one embodiment, all three neural networks may run concurrently, for example, within a DLA and / or on GPU(s) 1608.
[0220] In at least one embodiment, a facial recognition and vehicle owner identification CNN may use data from camera sensors to identify the presence of an authorized driver and / or owner of the vehicle 1600. In at least one embodiment, an "always on" sensor processing engine may be used to unlock a vehicle when an owner approaches a driver's door and turns on the lights, and to disable such a vehicle in a security mode when an owner exits such a vehicle. In this way, the SoC(s) 1604 provide security against theft and / or carjacking.
[0221] In at least one embodiment, a CNN for detecting and identifying emergency vehicles may use data from microphones 1696 to detect and identify emergency vehicle sirens. In at least one embodiment, the SoC(s) 1604 use a CNN to classify environmental and urban sounds, as well as to classify visual data. In at least one embodiment, a CNN running on a DLA is trained to detect a relative approach speed of an emergency vehicle (e.g., using a Doppler effect). In at least one embodiment, a CNN may also be trained to identify emergency vehicles specific to a local area in which a vehicle is traveling, as identified by GNSS sensor(s) 1658.In at least one embodiment, when deployed in Europe, a CNN will attempt to detect European sirens, and when deployed in North America, a CNN will attempt to identify only North American sirens. In at least one embodiment, once an emergency vehicle is detected, a control program may be used to execute an emergency vehicle safety routine, decelerate a vehicle, pull over to the side of the road, park a vehicle, and / or idle a vehicle using ultrasonic sensor(s) 1662 until emergency vehicles pass by.
[0222] In at least one embodiment, vehicle 1600 may include CPU(s) 1618 (e.g., discrete CPU(s) or dCPU(s)) that may be coupled to SoC(s) 1604 via a high-speed interconnect (e.g., PCIe). In at least one embodiment, CPU(s) 1618 may include, for example, an x86 processor. CPU(s) 1618 may be used to perform a variety of functions, including reconciling potentially conflicting results between ADAS sensors and SoC(s) 1604 and / or monitoring the status and health of controller(s) 1636 and / or an infotainment system on a chip (“Infotainment SoC”) 1630, for example. In at least one embodiment, SoC(s) 1604 includes one or more interconnects, and an interconnect may include a Peripheral Component Interconnect Express (PCIe).
[0223] In at least one embodiment, vehicle 1600 may include GPU(s) 1620 (e.g., discrete GPU(s) or dGPU(s)) that may be coupled to SoC(s) 1604 via a high-speed interconnect (e.g., NVIDIA's NVLINK channel). In at least one embodiment, GPU(s) 1620 may provide additional artificial intelligence functionality, e.g., by executing redundant and / or distinct neural networks, and may be used to train and / or update neural networks based at least in part on inputs (e.g., sensor data) from sensors of vehicle 1600.
[0224] In at least one embodiment, vehicle 1600 may further include a network interface 1624, which may include, without limitation, one or more wireless antennas 1626 (e.g., one or more wireless antennas for different communication protocols, such as a cellular antenna, a Bluetooth antenna, etc.). In at least one embodiment, network interface 1624 may be used to enable wireless connection to internet cloud services (e.g., to server(s) and / or other network devices), to other vehicles, and / or to computing devices (e.g., passenger client devices). In at least one embodiment, communication with other vehicles may involve a direct connection between vehicle 1600 and another vehicle and / or an indirect connection (e.g., via networks and the internet).In at least one embodiment, direct connections may be established via a vehicle-to-vehicle communication link. In at least one embodiment, a vehicle-to-vehicle communication link may provide information to vehicle 1600 about vehicles in the vicinity of vehicle 1600 (e.g., vehicles in front of, beside, and / or behind vehicle 1600). In at least one embodiment, such aforementioned functionality may be part of a cooperative adaptive cruise control function of vehicle 1600.
[0225] In at least one embodiment, the network interface 1624 may include an SoC that provides modulation and demodulation functionality and enables the controller(s) 1636 to communicate over wireless networks. In at least one embodiment, the network interface 1624 may include a radio frequency front-end for upconverting from baseband to radio frequency and downconverting from radio frequency to baseband. In at least one embodiment, the frequency conversions may be performed in any technically feasible manner. For example, frequency conversions may be performed by known methods and / or using superheterodyne techniques. In at least one embodiment, the radio frequency front-end functionality may be provided by a separate chip.In at least one embodiment, the network interfaces may include wireless capabilities for communication via LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.
[0226] In at least one embodiment, the vehicle 1600 may further include one or more data stores 1628, which may include, without limitation, off-chip memory (e.g., off-SoC(s) 1604). In at least one embodiment, the data store(s) 1628 may include, without limitation, one or more memory elements, including RAM, SRAM, dynamic random access memory ("DRAM"), video random access memory ("VRAM"), flash memory, hard drives, and / or other components and / or devices capable of storing at least one bit of data.
[0227] In at least one embodiment, the vehicle 1600 may further include GNSS sensor(s) 1658 (e.g., GPS and / or assisted GPS sensors) to assist with mapping, sensing, occupancy grid generation, and / or path planning. In at least one embodiment, any number of GNSS sensor(s) 1658 may be used, including, for example, and without limitation, a GPS with a USB port with an Ethernet-to-serial bridge (e.g., RS-232).
[0228] In at least one embodiment, vehicle 1600 may further include RADAR sensor(s) 1660. In at least one embodiment, RADAR sensor(s) 1660 may be used by vehicle 1600 for long-range vehicle detection, even in darkness and / or adverse weather conditions. In at least one embodiment, RADAR sensor(s) 1660 may use a CAN bus and / or bus 1602 (e.g., for transmitting data generated by RADAR sensor(s) 1660) for control and access to object tracking data, with raw data being accessed via Ethernet channels in some examples. In at least one embodiment, a wide range of RADAR sensors may be used. For example, and without limitation, RADAR sensor(s) 1660 may be suitable for use as front, rear, and side RADAR.In at least one embodiment, one or more sensors of the RADAR sensor(s) 1660 is a pulse Doppler RADAR sensor.
[0229] In at least one embodiment, the RADAR sensor(s) 1660 may include different configurations, such as long-range with a narrow field of view, short-range with a wide field of view, short-range side coverage, etc. In at least one embodiment, long-range RADAR may be used for the adaptive cruise control function. In at least one embodiment, long-range RADAR systems may provide a wide field of view realized by two or more independent scans, for example, within a range of 250 m (meters). In at least one embodiment, RADAR sensor(s) 1660 may assist in distinguishing between static and moving objects and may be used by the ADAS system 1638 for emergency braking assistance and forward collision warning.In at least one embodiment, the sensor(s) 1660 included in a long-range RADAR system may comprise, without limitation, a monostatic multimodal RADAR with multiple (e.g., six or more) fixed RADAR antennas and a high-speed CAN and FlexRay interface. In at least one embodiment with six antennas, four antennas in the center may create a focused beam pattern useful for detecting the vehicle's surroundings at higher speeds with minimal interference from traffic in adjacent lanes. In at least one embodiment, two additional antennas may expand the field of view, allowing for rapid detection of vehicles entering or exiting a lane of vehicle 1600.
[0230] For example, in at least one embodiment, medium-range radar systems may include a range of up to 160 m (front) or 80 m (rear) and a field of view of up to 42 degrees (front) or 150 degrees (rear). In at least one embodiment, short-range radar systems may include, without limitation, any number of radar sensors 1660 that may be installed at either end of a rear bumper. In at least one embodiment, a radar sensor system, when installed at either end of a rear bumper, may create two beams that continuously monitor blind spots to the rear and to the side of a vehicle. In at least one embodiment, short-range radar systems may be used in ADAS system 1638 for blind spot detection and / or lane change assistance.
[0231] In at least one embodiment, the vehicle 1600 may further include ultrasonic sensor(s) 1662. In at least one embodiment, the ultrasonic sensor(s) 1662, which may be arranged at a front, rear, and / or side location of the vehicle 1600, may be used for parking assistance and / or for creating and updating an occupancy grid. In at least one embodiment, a plurality of ultrasonic sensor(s) 1662 may be used, and different ultrasonic sensors 1662 may be used for different detection ranges (e.g., 2.5 m, 4 m). In at least one embodiment, the ultrasonic sensor(s) 1662 may operate at functional safety levels of ASIL B.
[0232] In at least one embodiment, the vehicle 1600 may include the LIDAR sensor(s) 1664. In at least one embodiment, the LIDAR sensor(s) 1664 may be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. In at least one embodiment, the LIDAR sensor(s) 1664 may operate at the ASIL B functional safety level. In at least one embodiment, the vehicle 1600 may include multiple LIDAR sensors 1664 (e.g., two, four, six, etc.) that may use an Ethernet channel (e.g., to provide data to a Gigabit Ethernet switch).
[0233] In at least one embodiment, the LIDAR sensor(s) 1664 may be capable of providing a list of objects and their distances for a 360-degree field of view. In at least one embodiment, the commercially available LIDAR sensor(s) 1664 may have an advertised range of approximately 100 m, with an accuracy of 2 cm to 3 cm, and with support for a 100 Mbps Ethernet connection, for example. In at least one embodiment, one or more non-protruding LIDAR sensors may be used. In such an embodiment, the LIDAR sensor(s) 1664 may comprise a small device that can be embedded in a front, rear, side, and / or corner position of the vehicle 1600.In at least one embodiment, the LIDAR sensor(s) 1664 in such an embodiment may provide a horizontal field of view of up to 120 degrees and a vertical field of view of up to 35 degrees with a range of 200 m, even for objects with low reflectivity. In at least one embodiment, the front-mounted LIDAR sensor(s) 1664 may be configured for a horizontal field of view between 45 degrees and 135 degrees.
[0234] In at least one embodiment, LIDAR technologies, such as 3D flash LIDAR, may also be used. In at least one embodiment, 3D flash LIDAR uses a laser flash as a transmission source to illuminate the surroundings of the vehicle 1600 up to approximately 200 m. In at least one embodiment, a flash LIDAR unit includes, without limitation, a receptor that records the time of flight of the laser pulse and the reflected light at each pixel, which in turn corresponds to a distance from the vehicle 1600 to objects. In at least one embodiment, flash LIDAR may enable highly accurate and distortion-free images of the surroundings to be generated with each laser flash. In at least one embodiment, four flash LIDAR sensors may be deployed, one on each side of the vehicle 1600.In at least one embodiment, 3D flash LIDAR systems include, without limitation, a solid-state 3D star array LIDAR camera with no moving parts other than a fan (e.g., a non-scanning LIDAR device). In at least one embodiment, the flash LIDAR device may use a 5-nanosecond Class I (eye-safe) laser pulse per image and collect the reflected laser light as a 3D range point cloud and co-registered intensity data.
[0235] In at least one embodiment, the vehicle 1600 may further include one or more IMU sensors 1666. In at least one embodiment, the IMU sensor(s) 1666 may be located in the center of a rear axle of the vehicle 1600. In at least one embodiment, the IMU sensor(s) 1666 may include, for example, and without limitation, accelerometers, magnetometers, gyroscopes, a magnetic compass, magnetic compasses, and / or other types of sensors. In at least one embodiment, such as in six-axis applications, the IMU sensor(s) 1666 may include, without limitation, accelerometers and gyroscopes. In at least one embodiment, such as in nine-axis applications, the IMU sensor(s) 1666 may include, without limitation, accelerometers, gyroscopes, and magnetometers.
[0236] In at least one embodiment, the IMU sensor(s) 1666 may be implemented as a miniaturized, high-performance GPS-based inertial navigation system ("GPS / INS") that combines microelectromechanical systems ("MEMS") inertial sensors, a high-sensitivity GPS receiver, and advanced Kalman filter algorithms to provide estimates of position, velocity, and attitude. In at least one embodiment, the IMU sensor(s) 1666 may enable the vehicle 1600 to estimate its heading without requiring input from a magnetic sensor by directly observing velocity changes from a GPS and correlating them with the IMU sensor(s) 1666. In at least one embodiment, the IMU sensor(s) 1666 and the GNSS sensor(s) 1658 may be combined into a single integrated unit.
[0237] In at least one embodiment, the vehicle 1600 may include the microphone(s) 1696 disposed in and / or around the vehicle 1600. In at least one embodiment, the microphone(s) 1696 may be used, among other things, for detecting and identifying emergency vehicles.
[0238] In at least one embodiment, the vehicle 1600 may further include any number of camera types, including stereo camera(s) 1668, wide-angle camera(s) 1670, infrared camera(s) 1672, surround camera(s) 1674, long-range camera(s) 1698, medium-range camera(s) 1676, and / or other camera types. In at least one embodiment, cameras may be used to capture image data around the entire perimeter of the vehicle 1600. In at least one embodiment, the types of cameras used depend on the vehicle 1600. In at least one embodiment, any combination of camera types may be used to provide the required coverage around the vehicle 1600. In at least one embodiment, the number of cameras employed may vary depending on the embodiment.For example, in at least one embodiment, the vehicle 1600 could include six cameras, seven cameras, ten cameras, twelve cameras, or any other number of cameras. In at least one embodiment, the cameras can support, for example, and without limitation, Gigabit Multimedia Serial Link ("GMSL") and / or Gigabit Ethernet communication. In at least one embodiment, each camera can be configured as previously described in . Fig. 16A and Fig. 16B will be described in more detail.
[0239] In at least one embodiment, the vehicle 1600 may further include vibration sensor(s) 1642. In at least one embodiment, the vibration sensor(s) 1642 may measure vibrations of components of the vehicle 1600, such as the axle(s). For example, in at least one embodiment, changes in vibrations may indicate a change in the road surface. In at least one embodiment, when two or more vibration sensors 1642 are used, differences between vibrations may be used to determine friction or slippage of the road surface (e.g., when there is a difference in vibration between a driven axle and a free-spinning axle).
[0240] In at least one embodiment, the vehicle 1600 may include an ADAS system 1638. In at least one embodiment, the ADAS system 1638 may include, in some examples, without limitation, an SoC. In at least one embodiment, the ADAS system 1638 may include, without limitation, any number and combination of an autonomous / adaptive / automatic cruise control (“ACC”) system, a cooperative adaptive cruise control (“CACC”) system, a forward crash warning (“FCW”) system, an automatic emergency braking (“AEB”) system, a lane departure warning (“LDW”) system, a lane keep assist (“LKA”) system, a blind spot warning (“BSW”) system, a rear cross traffic warning (“RCTW”) system, a forward collision warning (“CW”) system, a lane centering (“LC”) system, and / or other systems, features, and / or functions.
[0241] In at least one embodiment, the ACC system may use RADAR sensor(s) 1660, LIDAR sensor(s) 1664, and / or any number of cameras. In at least one embodiment, the ACC system may include a longitudinal ACC system and / or a lateral ACC system. In at least one embodiment, a longitudinal ACC system monitors and controls the distance to another vehicle immediately in front of the vehicle 1600 and automatically adjusts the speed of the vehicle 1600 to maintain a safe distance from preceding vehicles. In at least one embodiment, a lateral ACC system performs follow-through and advises the vehicle 1600 to change lanes when necessary. In at least one embodiment, a lateral ACC system is connected to other ADAS applications, such as LC and CW.
[0242] In at least one embodiment, a CACC system utilizes information from other vehicles, which may be received via a network interface 1624 and / or wireless antenna(s) 1626 from other vehicles over a wireless connection or indirectly via a network connection (e.g., over the Internet). In at least one embodiment, direct connections may be provided by a vehicle-to-vehicle ("V2V") communication link, while indirect connections may be provided by an infrastructure-to-vehicle ("I2V") communication link. Generally, V2V communication provides information about immediately preceding vehicles (e.g., vehicles immediately in front of and in the same lane as vehicle 1600), while I2V communication provides information about further ahead traffic.In at least one embodiment, a CACC system may include either one or both I2V and V2V information sources. In at least one embodiment, a CACC system may be more reliable given information about vehicles ahead of vehicle 1600 and has the potential to improve traffic flow and reduce congestion on the road.
[0243] In at least one embodiment, an FCW system is configured to alert a driver of a hazard so that the driver can take corrective action. In at least one embodiment, an FCW system utilizes a forward-facing camera and / or radar sensor(s) 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibrating component. In at least one embodiment, an FCW system may provide a warning, such as a sound, a visual warning, a vibration, and / or a rapid braking pulse.
[0244] In at least one embodiment, an AEB system detects an impending forward collision with another vehicle or other object and may automatically apply the brakes if a driver does not take corrective action within a specified time or distance parameter. In at least one embodiment, the AEB system may utilize forward-facing camera(s) and / or RADAR sensor(s) 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. In at least one embodiment, when an AEB system detects a hazard, it will typically first alert a driver to take corrective action to avoid a collision, and if that driver does not take corrective action, the AEB system may automatically apply the brakes to prevent or at least mitigate the effects of a predicted collision.In at least one embodiment, an AEB system may include techniques such as dynamic brake assistance and / or crash-imminent braking.
[0245] In at least one embodiment, an LDW system provides visual, audible, and / or tactile warnings, such as steering wheel or seat vibrations, to alert the driver when the vehicle 1600 crosses lane markings. In at least one embodiment, an LDW system is not activated when a driver indicates an intentional lane departure, for example, by activating a turn signal. In at least one embodiment, an LDW system may utilize forward-facing cameras coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to provide feedback to the driver, for example, via a display, speaker, and / or vibrating component. In at least one embodiment, an LKA system is a variation of an LDW system.In at least one embodiment, an LKA system provides a steering input or braking to correct the vehicle 1600 when the vehicle 1600 begins to depart from its lane.
[0246] In at least one embodiment, a BSW system detects and warns a driver of vehicles in an automobile's blind spot. In at least one embodiment, a BSW system may provide a visual, audible, and / or tactile warning to indicate that merging or changing lanes is unsafe. In at least one embodiment, a BSW system may provide an additional warning when a driver activates a turn signal. In at least one embodiment, a BSW system may utilize rear-facing camera(s) and / or RADAR sensor(s) 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to driver feedback, such as a display, speaker, and / or vibrating component.
[0247] In at least one embodiment, an RCTW system may provide visual, audible, and / or tactile notification when an object is detected outside the range of the rearview camera when the vehicle 1600 is reversing. In at least one embodiment, an RCTW system includes an AEB system to ensure that the vehicle brakes are applied to avoid a crash. In at least one embodiment, an RCTW system may utilize one or more rear-facing RADAR sensors 1660 coupled to a dedicated processor, DSP, FPGA, and / or ASIC electrically coupled to provide feedback to the driver, such as a display, speaker, and / or vibrating component.
[0248] In at least one embodiment, conventional ADAS systems may be prone to false positive results, which may be annoying and distracting for a driver, but are typically not catastrophic because conventional ADAS systems alert a driver and allow that driver to decide whether a safety condition truly exists and act accordingly. In at least one embodiment, in the event of conflicting results, the vehicle 1600 itself decides whether to consider the result of a primary computer or a secondary computer (e.g., a first controller or a second controller of the controllers 1636). In at least one embodiment, the ADAS system 1638 may, for example, be a backup and / or secondary computer that provides perception information to a rationality module of the backup computer.In at least one embodiment, a backup computer rationality monitor may execute redundant, diverse software on hardware components to detect errors in perception and dynamic driving tasks. In at least one embodiment, the outputs of the ADAS system 1638 may be forwarded to a supervisory MCU. In at least one embodiment, if the outputs of a primary computer and the outputs of a secondary computer conflict, a supervisory MCU determines how to resolve the conflict to ensure safe operation.
[0249] In at least one embodiment, a primary computer may be configured to provide a score to a supervising MCU indicating the primary computer's confidence in a selected outcome. In at least one embodiment, the supervising MCU may follow the primary computer's instruction if that confidence score exceeds a threshold, regardless of whether the secondary computer provides a conflicting or inconsistent outcome. In at least one embodiment, in cases where a confidence score does not meet a threshold and where primary and secondary computers indicate different outcomes (e.g., a conflict), a supervising MCU may arbitrate between the computers to determine an appropriate outcome.
[0250] In at least one embodiment, a monitoring MCU may be configured to execute a neural network(s) trained and configured to determine, based at least in part on the outputs of a primary computer and the outputs of a secondary computer, the conditions under which the secondary computer provides false alarms. In at least one embodiment, the neural network(s) in a monitoring MCU may learn when the output of a secondary computer can and cannot be trusted. In at least one embodiment, when the secondary computer is a RADAR-based FCW system, a neural network(s) mayNeural networks in this monitoring MCU can learn when an FCW system identifies metallic objects that are not actually hazards, such as a drain grate or manhole cover that triggers an alarm. In at least one embodiment, when a secondary computer is a camera-based LDW system, a neural network in a monitoring MCU can learn to override the LDW system when cyclists or pedestrians are present and leaving the lane is actually the safest maneuver. In at least one embodiment, a monitoring MCU can include at least one DLA or GPU suitable for executing neural networks with associated memory. In at least one embodiment, a monitoring MCU can comprise and / or be included as a component of the SoC(s) 1604.
[0251] In at least one embodiment, ADAS system 1638 may include a secondary computer that performs ADAS functions using conventional computer vision rules. In at least one embodiment, this secondary computer may use classic computer vision (if-then) rules, and the presence of a neural network(s) in a higher-level MCU may improve reliability, safety, and performance. In at least one embodiment, the different implementation and intentional non-identity make the overall system more fault-tolerant, particularly against errors caused by software functions (or software-hardware interfaces).For example, in at least one embodiment, if a software error occurs in the software running on a primary computer and non-identical software code runs on a secondary computer that produces a consistent overall result, then a supervising MCU may have greater confidence that an overall result is correct and an error in the software or hardware on that primary computer does not cause a significant error.
[0252] In at least one embodiment, an output of the ADAS system 1638 may be fed to the perception block of a primary computer and / or the dynamic driving task block of a primary computer. For example, in at least one embodiment, if the ADAS system 1638 indicates a forward crash warning due to an object immediately ahead, a perception block may use this information in identifying objects. In at least one embodiment, a secondary computer may have its own neural network trained, thus reducing the risk of false alarms, as described herein.
[0253] In at least one embodiment, vehicle 1600 may further include an infotainment SoC 1630 (e.g., an in-vehicle infotainment system (IVI)). Although shown and described as an SoC, in at least one embodiment, infotainment system SoC 1630 may not be an SoC and may include, without limitation, two or more discrete components. In at least one embodiment, infotainment SoC 1630 may include, without limitation, a combination of hardware and software that may be used to provide audio (e.g., music, a personal digital assistant, navigation commands, news, radio, etc.), video (e.g., television, movies, streaming, etc.), phone (e.g., hands-free calling), network connectivity (e.g., LTE, Wi-Fi, etc.), and other functions.) and / or information services (e.g., navigation systems, rear parking assistance, a radio data system, vehicle-related information such as fuel level, total distance traveled, brake fuel level, oil level, door open / close, air filter information, etc.) to the vehicle 1600. The infotainment SoC 1630 could include, for example, radios, record players, navigation systems, video players, USB and Bluetooth connectivity, car computers, in-car entertainment, WiFi, steering wheel audio controls, hands-free calling, a heads-up display (“HUD”), an HMI display 1634, a telematics device, a control panel (e.g., for controlling and / or interacting with various components, functions, and / or systems), and / or other components.In at least one embodiment, the infotainment SoC 1630 may be further used to provide information (e.g., visual and / or audible) to the user(s) of the vehicle 1600, such as information from the ADAS system 1638, autonomous driving information such as planned vehicle maneuvers, trajectories, environmental information (e.g., intersection information, vehicle information, road information, etc.), and / or other information.
[0254] In at least one embodiment, the infotainment SoC 1630 may include any amount and type of GPU functionality. In at least one embodiment, the infotainment SoC 1630 may communicate with other devices, systems, and / or components of the vehicle 1600 via the bus 1602. In at least one embodiment, the infotainment SoC 1630 may be coupled to a supervisory MCU so that a GPU of an infotainment system may perform some self-driving functions if the primary controller(s) 1636 (e.g., primary and / or backup computers of the vehicle 1600) fail. In at least one embodiment, the infotainment SoC 1630 may place the vehicle 1600 into a chauffeur-to-safe-stop mode, as described herein.
[0255] In at least one embodiment, the vehicle 1600 may further include an instrument cluster 1632 (e.g., a digital dashboard, an electronic instrument cluster, a digital instrument panel, etc.). In at least one embodiment, the instrument cluster 1632 may include, without limitation, a controller and / or a supercomputer (e.g., a discrete controller or a supercomputer). In at least one embodiment, the instrument cluster 1632 may include, without limitation, any number and combination of instruments, such as, but not limited to, speedometer, fuel level, oil pressure, tachometer, odometer, turn signals, shift position indicator, seat belt warning light(s), parking brake warning light(s), engine malfunction light(s), supplemental restraint system information (e.g., airbags), lighting controls, safety system controls, navigation information, etc.In some examples, information may be displayed and / or shared between infotainment SoC 1630 and instrument cluster 1632. In at least one embodiment, instrument cluster 1632 may comprise a portion of infotainment SoC 1630, or vice versa.
[0256] In at least one embodiment, at least one Fig. 20C is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig. 20C is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig.20C to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and cause the derivation of two or more contiguous portions of input data to a neural network to be distributed among two or more respective processing cores based at least in part on positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0257] Fig. 16D is a diagram of a system for communication between the cloud-based server(s) and the autonomous vehicle 1600 of Fig.16A, according to at least one embodiment. In at least one embodiment, the system may include, without limitation, server(s) 1678, network(s) 1690, and any number and type of vehicles, including vehicle 1600. In at least one embodiment, server(s) 1678 may include, without limitation, a plurality of GPUs 1684(A)-1684(H) (collectively referred to herein as GPUs 1684), PCIe switches 1682(A)-1682(D) (collectively referred to herein as PCIe switches 1682), and / or CPUs 1680(A)-1680(B) (collectively referred to herein as CPUs 1680). In at least one embodiment, the GPUs 1684, the CPUs 1680, and the PCIe switches 1682 may be interconnected with high-speed interconnects, such as, for example, and without limitation, the NVLink interfaces 1688 and / or PCIe interconnects 1686 developed by NVIDIA.In at least one embodiment, the GPUs 1684 are connected via an NVLink and / or NVSwitch SoC, and the GPUs 1684 and PCIe switches 1682 are connected via PCIe interconnects. Although eight GPUs 1684, two CPUs 1680, and four PCIe switches 1682 are shown, this is not intended to be limiting. In at least one embodiment, each of the servers 1678 may include, without limitation, any number of GPUs 1684, CPUs 1680, and / or PCIe switches 1682 in any combination. For example, in at least one embodiment, the server(s) 1678 could each include eight, sixteen, thirty-two, and / or more GPUs 1684.
[0258] In at least one embodiment, the server(s) 1678 may receive, via the network(s) 1690 and from vehicles, image data representative of images depicting unexpected or changed road conditions, such as recently commenced roadwork. In at least one embodiment, the server(s) 1678 may transmit, via the network(s) 1690 and to the vehicles, updated or other neural network 1692 and / or map information 1694 including, among other things, information about traffic and road conditions. In at least one embodiment, the updates to the map information 1694 may include, without limitation, updates to the HD map 1622, such as information about construction, potholes, detours, flooding, and / or other obstacles.In at least one embodiment, the neural networks 1692 and / or the map information 1694 may result from recent training and / or experience represented in data received from any number of vehicles in an environment and / or may be based at least in part on training performed in a data center (e.g., using server(s) 1678 and / or other servers).
[0259] In at least one embodiment, the server(s) 1678 may be used to train machine learning models (e.g., neural networks) based at least in part on training data. In at least one embodiment, the training data may be generated by vehicles and / or generated in a simulation (e.g., using a game engine). In at least one embodiment, any amount of training data is tagged (e.g., if the associated neural network benefits from supervised learning) and / or subjected to other preprocessing. In at least one embodiment, any amount of training data is untagged and / or subjected to preprocessing (e.g., if the associated neural network does not require supervised learning).In at least one embodiment, once trained, the machine learning models may be used by the vehicles (e.g., transmitted to the vehicles via network(s) 1690), and / or the machine learning models may be used by server(s) 1678 to remotely control the vehicles.
[0260] In at least one embodiment, the server(s) 1678 may receive data from vehicles and apply data to current neural networks for intelligent inference in real time. In at least one embodiment, the server(s) 1678 may include deep learning supercomputers and / or dedicated AI computers powered by GPU(s) 1684, such as DGX and DGX Station machines developed by NVIDIA. However, in at least one embodiment, the server(s) 1678 may also include a deep learning infrastructure using CPU-powered data centers.
[0261] In at least one embodiment, the deep learning infrastructure of server(s) 1678 may be capable of fast, real-time inference and may utilize this capability to assess and verify the health of processors, software, and / or associated hardware in vehicle 1600. For example, in at least one embodiment, the deep learning infrastructure may receive periodic updates from vehicle 1600, such as a sequence of images and / or objects that vehicle 1600 has located in that sequence of images (e.g., via computer vision and / or other machine learning object classification techniques).In at least one embodiment, the deep learning infrastructure may run its own neural network to identify objects and compare them to objects identified by the vehicle 1600, and if the results do not match and the deep learning infrastructure concludes that the AI in the vehicle 1600 is malfunctioning, then the server(s) 1678 may send a signal to the vehicle 1600 instructing a fail-safe computer of the vehicle 1600 to take over control, notify the passengers, and perform a safe parking maneuver.
[0262] In at least one embodiment, the server(s) 1678 may include GPU(s) 1684 and one or more programmable inference accelerators (e.g., NVIDIA's TensorRT 3 devices). In at least one embodiment, a combination of GPU-driven servers and inference accelerators may enable real-time responsiveness. In at least one embodiment, for example, where performance is less critical, servers with CPUs, FPGAs, and other processors may also be used for inference. In at least one embodiment, hardware structure(s) 1315 are used to perform one or more embodiments. Details of the hardware structure(s) 1315 are described herein in connection with Fig. 13A and / or 13B. COMPUTER SYSTEMS
[0263] Fig.17 is a block diagram illustrating an exemplary computer system, which may be a system of interconnected devices and components, a system-on-a-chip (SOC), or a combination thereof, formed with a processor that may include execution units for executing an instruction, according to at least one embodiment. In at least one embodiment, a computer system 1700 may include, without limitation, a component, such as a processor 1702, for employing execution units including logic for performing algorithms for processing data in accordance with the present disclosure, as in the embodiment described herein.In at least one embodiment, computer system 1700 may include processors such as the PENTIUM® processor family, Xeon™, Itanium®, XScale™ and / or StrongARM™, Intel® Core™, or Intel® Nervana™ microprocessors available from Intel Corporation of Santa Clara, California, although other systems (including personal computers with other microprocessors, technical workstations, set-top boxes, and the like) may be used. In at least one embodiment, computer system 1700 may run a version of the WINDOWS operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (e.g., UNIX and Linux), embedded software, and / or graphical interfaces may also be used.
[0264] Embodiments may be used in other devices such as handheld devices and embedded applications. Some examples of portable devices include cellular phones, Internet Protocol devices, digital cameras, personal digital assistants ("PDAs"), and handheld PCs. In at least one embodiment, embedded applications may include a microcontroller, a digital signal processor ("DSP"), a system on a chip, network computers ("NetPCs"), set-top boxes, network hubs, wide area network ("WAN") switches, or any other system capable of executing one or more instructions in accordance with at least one embodiment.
[0265] In at least one embodiment, computer system 1700 may include, without limitation, a processor 1702, which may include, without limitation, one or more execution units 1708 to perform machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, computer system 1700 is a desktop or server system having a processor, but in another embodiment, computer system 1700 may be a multiprocessor system. In at least one embodiment, processor 1702 may include, without limitation, a complex instruction set computer (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor implementing a combination of instruction sets, or any other device, such as a digital signal processor.In at least one embodiment, the processor 1702 may be connected to a processor bus 1710 that may transmit data signals between the processor 1702 and other components of the computer system 1700.
[0266] In at least one embodiment, processor 1702 may include, without limitation, an internal Level 1 ("L1") cache ("cache") 1704. In at least one embodiment, processor 1702 may have a single internal cache or multiple levels of internal cache. In at least one embodiment, the cache may be external to processor 1702. Other embodiments may also include a combination of internal and external caches, depending on the particular implementation and needs. In at least one embodiment, a register file 1706 may store different data types in various registers, including, without limitation, integer registers, floating-point registers, status registers, and an instruction pointer register.
[0267] In at least one embodiment, execution unit 1708, which includes, without limitation, logic for performing integer and floating-point operations, is also located in processor 1702. In at least one embodiment, processor 1702 may also include microcode read-only memory ("ROM") ("ucode") that stores microcode for certain macroinstructions. In at least one embodiment, execution unit 1708 may include logic for handling a packed instruction set 1709. In at least one embodiment, by including packed instruction set 1709 in the instruction set of a general-purpose processor, along with associated instruction execution circuitry, operations used by many multimedia applications may be performed using packed data in processor 1702.In at least one embodiment, many multimedia applications can be accelerated and executed more efficiently by using the entire width of a processor's data bus to perform operations on packed data, thereby eliminating the need to transfer smaller units of data across the processor's data bus to perform one or more operations on one data element at a time.
[0268] In at least one embodiment, execution unit 1708 may also be used in microcontrollers, embedded processors, graphics devices, DSPs, and other types of logic circuits. In at least one embodiment, computer system 1700 may include, without limitation, a memory 1720. In at least one embodiment, memory 1720 may be a dynamic random access memory (DRAM) device, a static random access memory (SRAM) device, a flash memory device, or other storage device. In at least one embodiment, memory 1720 may store instruction(s) 1719 and / or data 1721 represented by data signals that may be executed by processor 1702.
[0269] In at least one embodiment, a system logic chip may be connected to the processor bus 1710 and the memory 1720. In at least one embodiment, a system logic chip may include, without limitation, a memory control hub ("MCH") 1716, and the processor 1702 may communicate with the MCH 1716 via the processor bus 1710. In at least one embodiment, the MCH 1716 may provide a high-bandwidth memory path 1718 to the memory 1720 for instruction and data storage, as well as for graphics command, data, and texture storage. In at least one embodiment, the MCH 1716 may route data signals between the processor 1702, the memory 1720, and other components in the computer system 1700, and may bridge data signals between the processor bus 1710, the memory 1720, and a system I / O interface 1722.In at least one embodiment, a system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 1716 may be coupled to memory 1720 via a high-bandwidth memory path 1718, and a graphics / video card 1712 may be coupled to MCH 1716 via an Accelerated Graphics Port ("AGP") interconnect 1714.
[0270] In at least one embodiment, computer system 1700 may use system I / O interface 1722 as a proprietary hub interface bus to connect MCH 1716 to an I / O control hub ("ICH") 1730. In at least one embodiment, ICH 1730 may provide direct connections to some I / O devices via a local I / O bus. In at least one embodiment, a local I / O bus may include, without limitation, a high-speed I / O bus for connecting peripherals to memory 1720, a chipset, and processor 1702. Examples may include, without limitation, an audio controller 1729, a firmware hub (“Flash BIOS”) 1728, a wireless transceiver 1726, a data store 1724, a legacy I / O controller 1723 with user input and keyboard interfaces 1725, a serial expansion port 1727, such as a Universal Serial Bus (“USB”) interface, and a network controller 1734.In at least one embodiment, data storage 1724 may comprise a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0271] In at least one embodiment, Fig. 17 a system comprising interconnected hardware devices or “chips”, while in other embodiments Fig. 17 may show an exemplary SoC. In at least one embodiment, the Fig. 17 may be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of computer system 1700 are interconnected using Compute Express Link (CXL) interconnects.
[0272] Logic 1315 is used to perform inference and / or training operations associated with one or more embodiments. Details of logic 1315 are described herein in connection with Fig. 13A and / or 13B. In at least one embodiment, logic 1315 in computer system 1700 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0273] In at least one embodiment, at least one Fig. 13 is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig.13 is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 13 is used to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and to cause the derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0274] Fig. 18 is a block diagram illustrating an electronic device 1800 for using a processor 1810 according to at least one embodiment. In at least one embodiment, the electronic device 1800 may be, for example and without limitation, a notebook, a tower server, a rack server, a blade server, a laptop, a desktop, a tablet, a mobile device, a phone, an embedded computer, or any other suitable electronic device.
[0275] In at least one embodiment, the electronic device 1800 may include, without limitation, a processor 1810 communicatively connected to any number or type of components, peripherals, modules, or devices. In at least one embodiment, the processor 1810 is coupled via a bus or interface, such as an I2C bus, a System Management Bus ("SMBus"), a Low Pin Count (LPC) bus, a Serial Peripheral Interface ("SPI"), a High Definition Audio (HDA) bus, a Serial Advance Technology Attachment (SATA) bus, a Universal Serial Bus ("USB") (versions 1, 2, 3, etc.), or a Universal Asynchronous Receiver / Transmitter (UART) bus. In at least one embodiment, Fig. 18 a system comprising interconnected hardware devices or “chips”, while in other embodiments Fig.18 may show an exemplary SoC. In at least one embodiment, the Fig. 18 may be interconnected with proprietary interconnects, standardized interconnects (e.g., PCIe), or a combination thereof. In at least one embodiment, one or more components of Fig. 18 interconnected using CXL (Compute Express Link) connections.
[0276] At least in one embodiment, Fig.18 a display 1824, a touchscreen 1825, a touchpad 1830, a Near Field Communications unit (“NFC”) 1845, a sensor hub 1840, a thermal sensor 1846, an Express Chipset (“EC”) 1835, a Trusted Platform Module (“TPM”) 1838, BIOS / Firmware / Flash Memory (“BIOS, FW Flash”) 1822, a DSP 1860, a drive 1820 such as a Solid State Disk (“SSD”) or a Hard Drive (“HDD”), a Wireless Local Area Network unit (“WLAN”) 1850, a Bluetooth unit 1852, a Wireless Wide Area Network unit (“WWAN”) 1856, a Global Positioning System (GPS) unit 1855, a camera (“USB 3.0 Camera”) 1854, such as a USB 3.0 Camera, and / or a Low Power Double Data Rate ("LPDDR") memory unit ("LPDDR3") 1815, implemented, for example, according to an LPDDR3 standard. These components may each be implemented in any suitable manner.
[0277] In at least one embodiment, other components may be communicatively coupled to the processor 1810 via components described herein. In at least one embodiment, an accelerometer 1841, an ambient light sensor ("ALS") 1842, a compass 1843, and a gyroscope 1844 may be communicatively coupled to the sensor hub 1840. In at least one embodiment, a thermal sensor 1839, a fan 1837, a keyboard 1836, and a touchpad 1830 may be communicatively coupled to the EC 1835. In at least one embodiment, speakers 1863, headphones 1864, and a microphone ("mic") 1865 may be communicatively coupled to an audio unit ("audio codec and class D amp") 1862, which in turn may be communicatively coupled to the DSP 1860. In at least one embodiment, the audio unit 1862 may include, for example and without limitation, an audio encoder / decoder ("codec") and a Class D amplifier.In at least one embodiment, a SIM card ("SIM") 1857 may be communicatively coupled to the WWAN unit 1856. In at least one embodiment, components such as the WLAN unit 1850 and the Bluetooth unit 1852, as well as the WWAN unit 1856, may be implemented in a Next Generation Form Factor ("NGFF").
[0278] Logic 1315 is used to perform inference and / or training operations in connection with one or more embodiments. Details of logic 1315 are described herein in connection with Fig.13A and / or 13B. In at least one embodiment, logic 1315 in electronic device 1800 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0279] In at least one embodiment, at least one Fig. 18 is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig.18 is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 1 to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and to cause the derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0280] Fig. Figure 19 illustrates a computer system 1900 according to at least one embodiment. In at least one embodiment, the computer system 1900 is configured to implement various processes and methods described in this disclosure.
[0281] In at least one embodiment, computer system 1900 includes, without limitation, at least one central processing unit ("CPU") 1902 connected to a communications bus 1910 implemented using any suitable protocol, such as PCI ("Peripheral Component Interconnect"), Peripheral Component Interconnect Express ("PCI-Express"), AGP ("Accelerated Graphics Port"), HyperTransport, or other bus or point-to-point communications protocol. In at least one embodiment, computer system 1900 includes, without limitation, main memory 1904 and control logic (e.g., implemented as hardware, software, or a combination thereof), and data is stored in main memory 1904, which may take the form of random access memory ("RAM").In at least one embodiment, a network interface subsystem (“network interface”) 1922 provides an interface to other computing devices and networks to receive and transmit data from and to other systems with computer system 1900.
[0282] In at least one embodiment, computer system 1900 includes, without limitation, input devices 1908, a parallel processing system 1912, and display devices 1906, which may be implemented using a conventional cathode ray tube ("CRT"), a liquid crystal display ("LCD"), a light-emitting diode ("LED"), a plasma display, or other suitable display technology. In at least one embodiment, user input is provided via input devices 1908 such as a keyboard, mouse, touchpad, microphone, etc. In at least one embodiment, each module described herein may be packaged on a single semiconductor platform to form a processing system.
[0283] Logic 1315 is used to perform inference and / or training operations in connection with one or more embodiments. Details of inference and / or training logic 1315 are described herein in connection with Fig.13A and / or 13B. In at least one embodiment, logic 1315 in computer system 1900 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0284] In at least one embodiment, at least one Fig. 19 is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig.19 is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 1 to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and to cause the derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0285] Fig. 20 shows a computer system 2000 according to at least one embodiment. In at least one embodiment, the computer system 2000 includes, without limitation, a computer 2010 and a USB flash drive 2020. In at least one embodiment, the computer 2010 may include, without limitation, any number and type of processor(s) (not shown) and memory (not shown). In at least one embodiment, the computer 2010 includes, without limitation, a server, a cloud instance, a laptop, and a desktop computer.
[0286] In at least one embodiment, the USB flash drive 2020 includes, without limitation, a processing unit 2030, a USB interface 2040, and USB interface logic 2050. In at least one embodiment, the processing unit 2030 may be any instruction execution system, apparatus, or device capable of executing instructions. In at least one embodiment, the processing unit 2030 may include, without limitation, any number and type of processing cores (not shown). In at least one embodiment, the processing unit 2030 includes an application-specific integrated circuit ("ASIC") optimized to perform any number and type of machine learning-related operations.For example, in at least one embodiment, processing unit 2030 is a tensor processing unit ("TPC") optimized for performing machine learning operations. In at least one embodiment, processing unit 2030 is a video processing unit ("VPU") optimized for performing machine vision and machine learning operations.
[0287] In at least one embodiment, USB interface 2040 may be any type of USB plug or receptacle. For example, in at least one embodiment, USB interface 2040 is a USB 3.0 Type-C receptacle for data and power. In at least one embodiment, USB interface 2040 is a USB 3.0 Type-A plug. In at least one embodiment, the logic of USB interface 2050 may include any amount and type of logic that enables processing unit 2030 to communicate with devices (e.g., computer 2010) via USB port 2040.
[0288] Logic 1315 is used to perform inference and / or training operations in connection with one or more embodiments. Details of logic 1315 are described herein in connection with Fig.13A and / or 13B. In at least one embodiment, logic 1315 in computer system 2000 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0289] In at least one embodiment, at least one Fig. 20 is used to perform techniques and / or functions associated with Fig. 1-12. In at least one embodiment, at least one Fig.20 is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 1 to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and to cause the derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0290] Fig. 21A illustrates an example architecture in which a plurality of GPUs 2110(1)-2110(N) are communicatively coupled to a plurality of multi-core processors 2105(1)-2105(M) via high-speed interconnects 2140(1)-2140(N) (e.g., buses, point-to-point interconnects, etc.). In at least one embodiment, the high-speed interconnects 2140(1)-2140(N) support communication throughput of 4 GB / s, 30 GB / s, 80 GB / s, or greater. In at least one embodiment, various interconnect protocols may be used, including, but not limited to, PCIe 4.0 or 5.0 and NVLink 2.0. In various figures, "N" and "M" represent positive integers, the values of which may vary from figure to figure. In at least one embodiment, one or more GPUs in a plurality of GPUs 2110(1)-2110(N) includes one or more graphics cores (also referred to simply as “cores”) 2400, as shown in the Fig. 24A and Fig. 24B. In at least one embodiment, one or more graphics cores 2400 may be referred to as streaming multiprocessors ("SMs"), stream processors ("SPs"), stream processing units ("SPUs"), compute units ("CUs"), execution units ("EUs"), and / or slices, where a slice in this context may refer to a portion of processing resources in a processing unit (e.g., 16 cores, a ray tracing unit, a thread director, or scheduler).
[0291] Additionally, and in at least one embodiment, two or more GPUs 2110 are interconnected via high-speed interconnects 2129(1)-2129(2), which may be implemented using similar or different protocols / connections than those used for high-speed interconnects 2140(1)-2140(N). Similarly, two or more multi-core processors 2105 may be interconnected via a high-speed interconnect 2128, which may be symmetric multiprocessor (SMP) buses operating at 20 GB / s, 30 GB / s, 120 GB / s, or higher. Alternatively, all communication between the various Fig. 21A using similar protocols / connections (e.g., via a common connection structure).
[0292] In at least one embodiment, each multi-core processor 2105 is communicatively connected to a processor memory 2101(1)-2101(M) via memory interconnects 2126(1)-2126(M), and each GPU 2110(1)-2110(N) is communicatively connected to GPU memory 2120(1)-2120(N) via GPU memory interconnects 2150(1)-2150(N). In at least one embodiment, memory interconnects 2126 and 2150 may use similar or different memory access technologies. For example, the processor memories 2101(1)-2101(M) and the GPU memories 2120 may be volatile memories, such as dynamic random access memories (DRAMs) (including stacked DRAMs), graphics DDR SDRAM (GDDR) (e.g., GDDR5, GDDR6), or high bandwidth memory (HBM), and / or they may be non-volatile memories, such as 3D XPoint or Nano-Ram.In at least one embodiment, a portion of the processor memory 2101 may be volatile memory and another portion may be non-volatile memory (e.g., using a two-level memory (2LM) hierarchy).
[0293] As described herein, various multi-core processors 2105 and GPUs 2110 may be physically connected to a particular memory 2101 or 2120, respectively, and / or a unified memory architecture may be implemented in which a virtual system address space (also referred to as "effective address space") is distributed across different physical memories. For example, processor memories 2101(1)-2101(M) may each comprise 64 GB of system address space, and GPU memories 2120(1)-2120(N) may each comprise 32 GB of system address space, resulting in a total of 256 GB of addressable memory when M=2 and N=4. Other values for N and M are possible.
[0294] Fig.Figure 21B shows additional details for an interconnect between a multi-core processor 2107 and a graphics acceleration module 2146 according to an example embodiment. In at least one embodiment, the graphics acceleration module 2146 may include one or more GPU chips integrated on a line card connected to the processor 2107 via a high-speed interconnect 2140 (e.g., a PCIe bus, NVLink, etc.). Alternatively, in at least one embodiment, the graphics acceleration module 2146 may be integrated on a package or die with the processor 2107.
[0295] In at least one embodiment, processor 2107 includes a plurality of cores 2160A-2160D (which may be referred to as "execution units"), each having a translation lookaside buffer ("TLB") 2161A-2161D and one or more caches 2162A-2162D. In at least one embodiment, cores 2160A-2160D may include various other components for executing instructions and processing data, not shown. In at least one embodiment, caches 2162A-2162D may include Level 1 (L1) and Level 2 (L2) caches. In addition, one or more shared caches 2156 may be included in caches 2162A-2162D and shared by groups of cores 2160A-2160D. For example, one embodiment of processor 2107 includes 24 cores, each with its own L1 cache, twelve shared L2 caches, and twelve shared L3 caches.In this embodiment, one or more L2 and L3 caches are shared between two adjacent cores. In at least one embodiment, processor 2107 and graphics acceleration module 2146 are coupled to system memory 2114, which includes processor memories 2101(1)-2101(M) of FIG. Fig. 21A may include.
[0296] In at least one embodiment, coherency for data and instructions stored in various caches 2162A-2162D, 2156, and system memory 2114 is maintained via inter-core communication over a coherency bus 2164. For example, in at least one embodiment, each cache may have cache coherency logic / circuitry coupled to it to communicate over the coherency bus 2164 in response to detected reads or writes to particular cache lines. In at least one embodiment, a cache coherency protocol is implemented over the coherency bus 2164 to sniff out cache accesses.
[0297] In at least one embodiment, a proxy circuit 2125 communicatively couples the graphics acceleration module 2146 to the coherence bus 2164, allowing the graphics acceleration module 2146 to participate in a cache coherence protocol as a peer of the cores 2160A-2160D. Specifically, in at least one embodiment, an interface 2135 provides connectivity to the proxy circuit 2125 via the high-speed interconnect 2140, and an interface 2137 connects the graphics acceleration module 2146 to the high-speed interconnect 2140.
[0298] In at least one embodiment, an accelerator integration circuit 2136 provides cache management, memory access, context management, and interrupt management services on behalf of a plurality of graphics processing engines 2131(1)-2131(N) of the graphics acceleration module 2146. In at least one embodiment, the graphics processing engines 2131(1)-2131(N) may each include a separate graphics processing unit (GPU). In at least one embodiment, a plurality of graphics processing engines 2131(1)-2131(N) of the graphics acceleration module 2146 include one or more graphics cores 2400, as described in connection with the Fig. 24A and Fig.24B. In at least one embodiment, the graphics processing engines 2131(1)-2131(N) may alternatively comprise different types of graphics processing engines within a GPU, such as graphics execution units, media processing engines (e.g., video encoders / decoders), samplers, and blit engines. In at least one embodiment, the graphics acceleration module 2146 may be a GPU with a plurality of graphics processing engines 2131(1)-2131(N), or the graphics processing engines 2131(1)-2131(N) may be individual GPUs integrated on a common package, line card, or die.
[0299] In at least one embodiment, accelerator integration circuitry 2136 includes a memory management unit (MMU) 2139 for performing various memory management functions, such as virtual-to-physical memory translations (also referred to as effective-to-real memory translations) and memory access protocols for accessing system memory 2114. In at least one embodiment, MMU 2139 may also include a translation lookaside buffer (TLB) (not shown) to cache virtual / effective to physical / real address translations. In at least one embodiment, a cache 2138 may store instructions and data for efficient access by graphics processing engines 2131(1)-2131(N).In at least one embodiment, the data stored in cache 2138 and graphics memories 2133(1)-2133(M) is kept coherent with core caches 2162A-2162D, 2156, and system memory 2114, possibly using a fetch unit 2144. As mentioned, this may be done via a proxy circuit 2125 on behalf of cache 2138 and memories 2133(1)-2133(M) (e.g., sending updates to cache 2138 regarding changes / accesses to cache lines in processor caches 2162A-2162D, 2156 and receiving updates from cache 2138).
[0300] In at least one embodiment, a set of registers 2145 stores context data for threads executed by graphics processing engines 2131(1)-2131(N), and a context management circuit 2148 manages thread contexts. For example, context management circuit 2148 may perform save and restore operations to save and restore the contexts of various threads during context switches (e.g., when a first thread is saved and a second thread is saved to allow a second thread to be executed by a graphics processing engine). For example, upon a context switch, context management circuit 2148 may save the current register values to a specific region of memory (e.g., identified by a context pointer). The register values may then be restored upon return to a context.In at least one embodiment, an interrupt management circuit 2147 receives and processes interrupts received from system devices.
[0301] In at least one embodiment, virtual / effective addresses from a graphics processing engine 2131 are translated by the MMU 2139 into real / physical addresses in system memory 2114. In at least one embodiment, the accelerator integration circuit 2136 supports multiple (e.g., 4, 8, 16) graphics acceleration modules 2146 and / or other accelerator devices. In at least one embodiment, the graphics acceleration module 2146 may be dedicated to a single application executing on the processor 2107 or may be shared among multiple applications. In at least one embodiment, a virtualized graphics execution environment is illustrated in which the resources of the graphics processing engines 2131(1)-2131(N) are shared among multiple applications or virtual machines (VMs).In at least one embodiment, the resources may be divided into "slices" that are allocated to different VMs and / or applications based on processing requirements and priorities associated with VMs and / or applications.
[0302] In at least one embodiment, accelerator integration circuitry 2136 acts as a bridge to a system for graphics acceleration module 2146 and provides address translation and system memory caching services. Furthermore, in at least one embodiment, accelerator integration circuitry 2136 may provide virtualization facilities to a host processor to manage the virtualization of graphics processing engines 2131(1)-2131(N), interrupts, and memory management.
[0303] Because, in at least one embodiment, the hardware resources of graphics processing engines 2131(1)-2131(N) are explicitly mapped to a real address space seen by host processor 2107, each host processor can directly address these resources via an effective address value. In at least one embodiment, a function of accelerator integration circuitry 2136 is to physically separate graphics processing engines 2131(1)-2131(N) so that they appear to a system as independent entities.
[0304] In at least one embodiment, one or more graphics memories 2133(1)-2133(M) are coupled to each of the graphics processing engines 2131(1)-2131(N), where N=M. In at least one embodiment, the graphics memories 2133(1)-2133(M) store instructions and data processed by each of the graphics processing engines 2131(1)-2131(N). In at least one embodiment, the graphics memories 2133(1)-2133(M) may be volatile memories such as DRAMs (including stacked DRAMs), GDDR memories (e.g., GDDR5, GDDR6), or HBM, and / or they may be non-volatile memories such as 3D XPoint or Nano-Ram.
[0305] In at least one embodiment, to reduce data traffic over high-speed interconnect 2140, bias techniques may be used to ensure that the data stored in graphics memories 2133(1)-2133(M) is data most frequently used by graphics processing engines 2131(1)-2131(N) and preferably not used (at least not frequently) by cores 2160A-2160D. Similarly, in at least one embodiment, a bias mechanism attempts to keep data needed by cores (and preferably not by graphics processing engines 2131(1)-2131(N)) in caches 2162A-2162D, 2162A-2162D, and system memory 2114.
[0306] Fig.21C shows another exemplary embodiment in which accelerator integration circuitry 2136 is integrated with processor 2107. In this embodiment, graphics processing engines 2131(1)-2131(N) communicate directly over high-speed interconnect 2140 with accelerator integration circuitry 2136 via interface 2137 and interface 2135 (which may again be any form of bus or interface protocol). In at least one embodiment, accelerator integration circuitry 2136 may perform operations similar to those described in Fig.21B, but possibly with higher throughput due to its proximity to the coherence bus 2164 and caches 2162A-2162D, 2156. In at least one embodiment, an accelerator integration circuit supports different programming models, including a dedicated process programming model (without graphics acceleration module virtualization) and shared programming models (with virtualization), which may include programming models controlled by the accelerator integration circuit 2136 and programming models controlled by the graphics acceleration module 2146.
[0307] In at least one embodiment, the graphics processing engines 2131(1)-2131(N) are dedicated to a single application or process under a single operating system. In at least one embodiment, a single application can forward other application requests to the graphics processing engines 2131(1)-2131(N), thus enabling virtualization within a VM / partition.
[0308] In at least one embodiment, graphics processing engines 2131(1)-2131(N) may be shared between multiple VM / application partitions. In at least one embodiment, shared models may use a system hypervisor to virtualize the graphics processing engines 2131(1)-2131(N) to allow access by any operating system. In at least one embodiment, for systems with a single partition without a hypervisor, the graphics processing engines 2131(1)-2131(N) are owned by an operating system. In at least one embodiment, an operating system may virtualize the graphics processing engines 2131(1)-2131(N) to allow access to any process or application.
[0309] In at least one embodiment, the graphics acceleration module 2146 or an individual graphics processing engine 2131(1)-2131(N) selects a process element using a process handle. In at least one embodiment, process elements are stored in system memory 2114 and are addressable using an effective address to real address translation technique described herein. In at least one embodiment, a process handle may be an implementation-specific value provided to a host process when it registers its context with the graphics processing engine 2131(1)-2131(N) (i.e., when it calls system software to add a process element to a linked process element list). In at least one embodiment, the lower 16 bits of a process handle may be an offset of a process element within a process element list.
[0310] Fig.21D shows an exemplary accelerator integration slice 2190. In at least one embodiment, a "slice" comprises a particular portion of the processing resources of accelerator integration circuitry 2136. In at least one embodiment, an application stores process elements 2183 in an effective address space 2182 within system memory 2114. In at least one embodiment, process elements 2183 are stored in response to GPU calls 2181 from applications 2180 executing on processor 2107. In at least one embodiment, a process element 2183 contains the state of the corresponding application 2180. In at least one embodiment, a work description (WD) 2184 contained in process element 2183 may be a single job requested by an application or may contain a pointer to a queue of jobs.In at least one embodiment, the WD 2184 is a pointer to a job request queue in the effective address space 2182 of an application.
[0311] In at least one embodiment, the graphics acceleration module 2146 and / or individual graphics processing engines 2131(1)-2131(N) may be shared by all or a subset of the processes in a system. In at least one embodiment, an infrastructure for establishing process states and sending a WD 2184 to a graphics acceleration module 2146 to start a job in a virtualized environment may be included.
[0312] In at least one embodiment, a dedicated process programming model is implementation-specific. In at least one embodiment, in this model, a single process owns the graphics acceleration module 2146 or an individual graphics processing engine 2131. In at least one embodiment, when the graphics acceleration module 2146 is owned by a single process, a hypervisor initializes the accelerator integration circuit 2136 for an owning partition, and an operating system initializes the accelerator integration circuit 2136 for an owning process when the graphics acceleration module 2146 is allocated.
[0313] In operation, in at least one embodiment, a WD fetch unit 2191 in accelerator integration slice 2190 fetches the next WD 2184, which includes an indication of the work to be performed by one or more graphics processing engines of graphics acceleration module 2146. In at least one embodiment, the data of WD 2184 may be stored in registers 2145 and used by MMU 2139, interrupt management circuitry 2147, and / or context management circuitry 2148, as shown. For example, one embodiment of MMU 2139 includes segment / page walkup circuitry for accessing segment / page tables 2186 within a virtual address space 2185 of the operating system. In at least one embodiment, circuitry 2147 may process interrupt events 2192 received from graphics acceleration module 2146.In at least one embodiment, when performing graphics operations, an effective address 2193 generated by a graphics processing engine 2131(1)-2131(N) is translated into a real address by the MMU 2139.
[0314] In at least one embodiment, registers 2145 are duplicated for each graphics processing engine 2131(1)-2131(N) and / or each graphics acceleration module 2146 and may be initialized by a hypervisor or an operating system. In at least one embodiment, each of these duplicated registers may be included in an accelerator integration slice 2190. Example registers that may be initialized by a hypervisor are shown in Table 1. Table 1 - Initialized hypervisor registers Register # Description 1 Slice control register (slice control register) 2 Real Address (RA) pointer for the Scheduled Processes area 3 Authority mask override register 4 Interrupt vector table entry offset 5 Interrupt vector table entry boundary 6 Condition register 7 Logical partition ID 8 Real Address (RA) Hypervisor Accelerator Workload Set Pointer 9 Memory description register
[0315] Example registers that can be initialized by an operating system are listed in Table 2. Table 2 - Initialized operating system registers Register # Description 1 Process and thread identification 2 Effective Address (EA) Context Store / Restore Pointer 3 Virtual Address (VA) Accelerator Workload Set Pointer 4 Virtual Address (VA) Pointer to memory segment table 5 Authority mask 6 Job description
[0316] In at least one embodiment, each WD 2184 is specific to a particular graphics acceleration module 2146 and / or graphics processing engines 2131(1)-2131(N). In at least one embodiment, it contains all the information needed by a graphics processing engine 2131(1)-2131(N) to perform work, or it may be a pointer to a memory location where an application has established a command queue of work to be performed.
[0317] Fig.Figure 21E shows additional details for an exemplary embodiment of a joint model. This embodiment includes a real hypervisor address space 2198 in which a process element list 2199 is stored. In at least one embodiment, the real hypervisor address space 2198 is accessible via a hypervisor 2196 that virtualizes graphics acceleration engine engines for the operating system 2195.
[0318] In at least one embodiment, shared programming models allow all or a subset of processes from all or a subset of partitions in a system to utilize a graphics acceleration module 2146. In at least one embodiment, there are two programming models in which the graphics acceleration module 2146 is shared among multiple processes and partitions: time-slice sharing and graphics sharing.
[0319] In at least one embodiment, in this model, the system hypervisor 2196 has the graphics acceleration module 2146 and makes its functionality available to all operating systems 2195. In at least one embodiment, a graphics acceleration module 2146 may meet certain requirements to support virtualization by the system hypervisor 2196, such as (1) the job request of an application must be autonomous (i.e.the state does not need to be maintained between jobs), or the graphics acceleration module 2146 must provide a mechanism for saving and restoring the context, (2) the graphics acceleration module 2146 guarantees that an application's job request will be completed in a specified amount of time, including any translation errors, or the graphics acceleration module 2146 provides the ability to preempt processing of a job, and (3) the graphics acceleration module 2146 must be guaranteed fairness between processes when operating in a directed joint programming model.
[0320] In at least one embodiment, application 2180 must make a system call to operating system 2195 with a graphics acceleration module type, a work description (WD), an authority mask register (AMR) value, and a context save / restore pointer (CSRP). In at least one embodiment, the graphics acceleration module type describes a targeted acceleration function for a system call. In at least one embodiment, the graphics acceleration module type may be a system-specific value. In at least one embodiment, WD is formatted specifically for graphics acceleration module 2146 and may be in the form of a graphics acceleration module instruction, a pointer to an effective address of a user-defined structure, a pointer to an effective address of an instruction queue, or another data structure describing the work to be performed by graphics acceleration module 2146.
[0321] In at least one embodiment, an AMR value is an AMR state to be used for a current process. In at least one embodiment, a value passed to an operating system is comparable to an application setting an AMR. In at least one embodiment, if accelerator integration circuitry 2136 (not shown) and graphics acceleration module 2146 do not support an authority mask override register (UAMOR), an operating system may apply a current UAMOR value to an AMR value before passing an AMR in a hypervisor call. In at least one embodiment, hypervisor 2196 may optionally apply a current authority mask override register (AMOR) value before placing an AMR in process element 2183.In at least one embodiment, CSRP is one of the registers 2145 that contain an effective address of a region in an application's effective address space 2182 for the graphics acceleration module 2146 to save and restore state. In at least one embodiment, this pointer is optional if no state needs to be saved between jobs or if a job aborts prematurely. In at least one embodiment, the context save / restore region may be pinned to system memory.
[0322] Upon receiving a system call, the operating system 2195 may verify whether the application 2180 has and has been granted permission to use the graphics acceleration module 2146. In at least one embodiment, the operating system 2195 then invokes the hypervisor 2196 with the information shown in Table 3. Table 3 - Parameters for calling the operating system to the hypervisor Parameters # Description 1 A job description (WD) 2 An Authority Mask Register (AMR) value (potentially masked) 3 An effective address (EA) Context save / restore pointer (CSRP) 4 A process ID (PID) and optionally a thread ID (TID) 5 A virtual address (VA) accelerator utilization set pointer (AURP) 6 Virtual address of the pointer to the memory segment table (SSTP) 7 A logical interrupt service number (LISN)
[0323] In at least one embodiment, upon receiving a hypervisor call, the hypervisor 2196 checks whether the operating system 2195 has and has been granted permission to use the graphics acceleration module 2146. In at least one embodiment, the hypervisor 2196 then places the process element 2183 in a process element list for a corresponding type of graphics acceleration module 2146. In at least one embodiment, a process element may include the information shown in Table 4. Table 4 - Process element information Item # Description 1 A job description (WD) 2 An Authority Mask Register (AMR) value (possibly masked). 3 An effective address (EA) Context save / restore pointer (CSRP) 4 A process ID (PID) and optionally a thread ID (TID) 5 A virtual address (VA) accelerator utilization set pointer (AURP) 6 Virtual address of the pointer to the memory segment table (SSTP) 7 A logical interrupt service number (LISN) 8 Interrupt vector table derived from hypervisor call parameters 9 A status register value (SR) 10 A logical partition ID (LPID) 11 Ein Zeiger für den Beschleunigerauslastungssatz des Hypervisors mit Realer Adresse (RA) 12 Speicher-Deskriptor-Register (SDR)
[0324] In at least one embodiment, the hypervisor initializes a plurality of registers 2145 for accelerator integration slice 2190.
[0325] As in Fig. 21F, in at least one embodiment, a unified memory addressable via a common virtual memory address space is used to access physical processor memories 2101(1)-2101(N) and GPU memories 2120(1)-2120(N). In this implementation, operations performed on GPUs 2110(1)-2110(N) use the same virtual / effective address space to access processor memories 2101(1)-2101(M) and vice versa, simplifying programmability. In at least one embodiment, a first portion of a virtual / effective address space is assigned to processor memory 2101(1), a second portion is assigned to second processor memory 2101(N), a third portion is assigned to GPU memory 2120(1), and so on.In at least one embodiment, this distributes an entire virtual / effective memory space (sometimes referred to as effective address space) across each of the processor memories 2101 and GPU memories 2120, such that each processor or GPU can access each physical memory with a virtual address associated with that memory.
[0326] In at least one embodiment, the bias / coherence management circuitry 2194A-2194E within one or more MMUs 2139A-2139E ensures cache coherence between the caches of one or more host processors (e.g., 2105) and GPUs 2110 and implements bias techniques that indicate in which physical memories certain data types should be stored. In at least one embodiment, while multiple instances of the bias / coherence management circuitry 2194A-2194E in Fig. 21F, the bias / coherence circuitry may be implemented within an MMU of one or more host processors 2105 and / or within the accelerator integration circuitry 2136.
[0327] In one embodiment, GPU memories 2120 may be mapped as part of system memory and accessed using shared virtual memory (SVM) technology without the performance penalties associated with full system cache coherence. In at least one embodiment, the ability to access GPU memories 2120 as system memory without burdensome cache coherence overhead provides a favorable operating environment for GPU offload. In at least one embodiment, this arrangement allows host processor 2105 software to set operands and access computation results without the overhead of traditional I / O DMA data copies. In at least one embodiment, such traditional copies involve driver calls, interrupts, and memory-mapped I / O (MMIO) accesses, all of which are inefficient compared to simple memory accesses.In at least one embodiment, the ability to access GPU memory 2120 without cache coherence overheads may be critical to the execution time of an offloaded computation. For example, in at least one embodiment, cache coherence overhead may significantly reduce the effective write bandwidth of a graphics processor 2110 in cases with significant streaming write memory traffic. In at least one embodiment, operand construction efficiency, result access efficiency, and GPU computation efficiency may play a role in determining the effectiveness of a GPU offload.
[0328] In at least one embodiment, the selection of the GPU bias and the host processor bias is controlled by a bias tracker data structure. For example, in at least one embodiment, a bias table may be used, which may be a page-granular structure (e.g., controlled at the granularity of a memory page) comprising 1 or 2 bits per GPU-attached memory page. In at least one embodiment, a bias table may be implemented in a stolen memory region of one or more GPU memories 2120, with or without a bias cache in a GPU 2110 (e.g., to cache frequently / recently used entries of a bias table). Alternatively, in at least one embodiment, an entire bias table may be maintained in a GPU.
[0329] In at least one embodiment, a bias table entry associated with each access to GPU-attached memory 2120 is accessed prior to the actual GPU memory access, causing the following operations. In at least one embodiment, local requests from a GPU 2110 that find their page in GPU-biased are forwarded directly to a corresponding GPU memory 2120. In at least one embodiment, local requests from a GPU that find its page in the host's bias are forwarded to processor 2105 (e.g., over a high-speed connection, as described herein). In at least one embodiment, requests from processor 2105 that find a requested page in the host processor's bias complete a request like a normal memory read. Alternatively, requests directed to a GPU-biased page may be forwarded to GPU 2110.In at least one embodiment, a GPU may forward a page to a host processor bias when it is not currently using the page. In at least one embodiment, a page's bias state may be changed either by a software-based mechanism, a hardware-assisted software-based mechanism, or, for a limited number of cases, a purely hardware-based mechanism.
[0330] In at least one embodiment, a mechanism for changing bias state uses an API call (e.g., OpenCL), which in turn invokes a graphics processor's device driver, which in turn sends a message to a graphics processor (or queues a command descriptor) to instruct it to change a bias state and, on some transitions, perform a cache flush operation in a host. In at least one embodiment, a cache flush operation is used for a transition from the host processor 2105 bias to the GPU bias, but not for an opposite transition.
[0331] In at least one embodiment, cache coherence is maintained by temporarily making GPU-biased pages uncacheable by the host processor 2105. In at least one embodiment, to access these pages, the processor 2105 may request access from the GPU 2110, which may or may not grant access immediately. Therefore, in at least one embodiment, to reduce communication between the processor 2105 and the GPU 2110, it is advantageous to ensure that GPU-biased pages are those required by a GPU but not by the host processor 2105, and vice versa.
[0332] Hardware structure(s) 1315 are used to carry out one or more embodiments. Details of a hardware structure (or more hardware structures) 1315 may be described herein in connection with Fig. 13A and / or 13B must be specified.
[0333] Fig. Figure 22 illustrates exemplary integrated circuits and associated graphics processors that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to the illustrated embodiments, other logic and circuitry may be included in at least one embodiment, including additional graphics processors / cores, peripheral interface controllers, or general-purpose processor cores.
[0334] Fig. 22 is a block diagram illustrating an exemplary integrated circuit 2200 that may be fabricated using one or more IP cores according to at least one embodiment. In at least one embodiment, the integrated circuit 2200 includes one or more application processors 2205 (e.g., CPUs), at least one graphics processor 2210, and may additionally include an image processor 2215 and / or a video processor 2220, each of which may be a modular IP core. In at least one embodiment, the integrated circuit 2200 includes peripheral or bus logic, including a USB controller 2225, a UART controller 2230, an SPI / SDIO controller 2235, and an I22S / I22C controller 2240.In at least one embodiment, integrated circuit 2200 may include a display device 2245 coupled to one or more of the following interfaces: a High-Definition Multimedia Interface (HDMI) controller 2250 and a Mobile Industry Processor Interface (MIPI) interface 2255. In at least one embodiment, memory may be provided by a flash memory subsystem 2260, which includes flash memory and a flash memory controller. In at least one embodiment, a memory interface may be provided via a memory controller 2265 for accessing SDRAM or SRAM memory devices. In at least one embodiment, some integrated circuits additionally include an embedded security engine 2270.
[0335] Logic 1315 is used to perform inference and / or training operations in connection with one or more embodiments. Details of logic 1315 are described herein in connection with Fig. 13A and / or 13B. In at least one embodiment, logic 1315 in integrated circuit 2200 may be used for inference or prediction operations based at least in part on weighting parameters calculated using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0336] In at least one embodiment, at least one Fig. 21A-22 is used to control the component shown or described in connection with Fig. 1-12. In at least one embodiment, at least one of Fig. 21A-22 is used to distribute the derivation of two or more contiguous sections of information only between two or more respective processing cores. In at least one embodiment, at least one in Fig. 1 to cause two or more other processors to each perform one or more operations of one or more neural networks based at least in part on two or more non-contiguous portions of an input sequence, and to cause the derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information.
[0337] Fig. 23A and Fig. 23B illustrate example integrated circuits and associated graphics processors that may be fabricated using one or more IP cores according to various embodiments described herein. In addition to the embodiments shown, other logic and circuitry may also be included in at least one embodiment, including additional graphics processors / cores,...
Claims
[1] Processor comprising: one or more circuits for causing the inferencing of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information. [2] The processor of claim 1, wherein the one or more circuits further cause the derivation of an equal number of portions of the information to be distributed to each of the two or more respective processing cores. [3] The processor of claim 1, wherein the information is divided into a number of sections, the number of sections being an even multiple of a number of processing cores each receiving one or more sections of the information. [4] The processor of claim 1, wherein at least two of the two or more final portions of information are distributed to the same processing core. [5] The processor of claim 1, wherein at least a portion of the information at the beginning of the information and at least a portion of the information at the end of the information are distributed to the same processing core. [6] The processor of claim 1, wherein each of the two or more contiguous portions of information comprises an equal number of tokens. [7] The processor of claim 1, wherein one or more activations based at least in part on the two or more contiguous portions of information are to be distributed to the two or more respective processing cores. [8] Method comprising: Causing the derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on the positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information. [9] Method according to claim 8, wherein the derivation is to be performed using a large language model. [10] The method of claim 8, wherein the derivation is used to train a large language model. [11] The method of claim 8, wherein each processing core of the two or more respective processing cores receives at least two portions of the information. [12] The method of claim 8, wherein each processing core of the two or more respective processing cores receives an even number of portions of the information. [13] The method of claim 8, wherein each processing core of the two or more respective processing cores receives an equal number of portions of the information. [14] The method of claim 8, wherein the two or more respective processing cores exchange token embeddings such that each of the two or more respective processing cores has a token embedding for each token in the information. [15] Computer system comprising: one or more processors and a memory storing instructions that, when executed by the one or more processors, are to cause derivation of two or more contiguous portions of information to be distributed between two or more respective processing cores based at least in part on positions of the two or more contiguous portions within the information relative to one or more terminating portions of the information. [16] The computer system of claim 15, wherein the instructions, when executed by the one or more processors, further cause the derivation of an even number of portions of the information to be distributed to each of the two or more respective processing cores. [17] The computer system of claim 15, wherein each of the one or more respective processing cores has an equal workload resulting from processing portions of the information distributed to that respective processing core. [18] The computer system of claim 15, wherein at least two of the two or more final portions of the information are distributed to the same processing core of the two or more processing cores. [19] The computer system of claim 15, wherein a first portion of the information and a last portion of the information are distributed to the same processing core of the two or more processing cores. [20] The computer of claim 15, wherein each of the processing cores of the two or more respective processing cores has an embedding for each token in the information.
Citation Information
Patent Citations
3016-201609
3016-201806
Cited By
Method for establishing transportation and diffusion long-term model of seabed leaked carbon dioxide in seawater
CN120524866A