Efficient SOFTMAX computation
By using methods to reduce precision and decomposition operations in Softmax computation, the problems of low memory utilization and high computational cost in traditional Softmax computation are solved, achieving more efficient neural network computation.
Patent Information
- Application Number
- CN202111007923.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-04
- Filing Date
- 2021-08-27
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2041-08-27
AI Technical Summary
Traditional Softmax computation suffers from low memory utilization and high computational cost in neural networks, especially in transformation neural networks and deep learning applications, which affects the performance and efficiency of the system.
A method with reduced precision is adopted, which calculates the exponent by using 2x instead of ex and decomposes Softmax into unnormalized and normalized operations, combined with the calculation of the maximum integer value, thus optimizing the Softmax calculation process.
It improves the efficiency of Softmax computation, reduces computational costs and memory requirements, and enhances the performance and energy efficiency of neural networks.
Smart Images

Figure CN114118354B_ABST
Abstract
Description
[0001] Cross Reference to Related Applications
[0002] This application claims priority to and the benefit of U.S. Application Serial No. 62 / 817,413, filed March 12, 2019, under 35 U.S.C. 119(e), the contents of which are incorporated herein in their entirety. BACKGROUND
[0003] Softmax computations are commonly used in various types of neural networks and deep learning applications. Examples of neural networks that utilize Softmax are recurrent neural networks, convolutional neural networks, and Transformer neural networks.
[0004] Conventional computations of Softmax suffer from low memory utilization and high computational cost in some aspects. As a result, neural networks and deep learning applications can benefit from more efficient Softmax computations.
[0005] Transformer neural networks show particularly promising results in conversational artificial intelligence (AI) applications. Transformer networks use attention mechanisms that utilize Softmax in encoder and decoder stages, and can particularly benefit from more efficient Softmax computations.
[0006] Deep neural networks (DNNs) are a class of neural networks that have become a key method for solving complex problems in various technical fields, particularly those involving deep machine learning. Applications of DNNs have different performance, accuracy, and power requirements depending on the implementation. Due to high design complexity and manufacturing challenges, the cost of building a specialized DNN for the requirements of a particular implementation can be prohibitive. Deep neural networks also tend to use Softmax computations heavily, and can therefore also benefit from more efficient Softmax computations.
[0007] In some aspects, a system includes one or more processors. The system includes logic that, when applied to the one or more processors, computes an unnormalized Softmax vector from an input vector by raising elements of the input vector to a power of 2 and computing an integer vector maximum of the input vector. The system also includes logic that, when applied to the one or more processors, converts the unnormalized Softmax vector to a normalized Softmax vector.
[0008] In other aspects, the artificial neural network includes one or more feedforward layers, and one or more Softmax layers coupled to the one or more feedforward layers. The artificial neural network includes at least one Softmax layer configured to compute an unnormalized Softmax vector from an input vector by raising elements of the input vector to a power of 2 and computing an integer vector maximum of the input vector.
[0009] In other aspects, the transformer artificial neural network includes a self-attention layer and an encoder-decoder attention layer. Each of the self-attention layer and the encoder-decoder attention layer includes a Softmax layer configured to generate an unnormalized Softmax vector from an input vector by raising elements of the input vector to a power of 2 and computing an integer vector maximum of the input vector. BRIEF DESCRIPTION OF DRAWINGS
[0010] For ease of reference to discussions of any particular element or act, one or more most significant digits of a reference number are indicative of the figure number in which the element was first introduced.
[0011] Figure 1 An exemplary system 100 utilizing an artificial neural network is depicted.
[0012] Figure 2 A deep learning system 202 according to one embodiment is depicted.
[0013] Figure 3 A transformer neural network 302 according to one embodiment is depicted.
[0014] Figure 4 An encoder 402 according to one embodiment is depicted.
[0015] Figure 5 A decoder 502 according to one embodiment is depicted.
[0016] Figure 6 An attention layer 602 according to one embodiment is depicted.
[0017] Figures 7A-7C A Softmax algorithm 700 according to one embodiment is depicted.
[0018] Figures 8A-8D Softmax computation logic 800 in one embodiment is depicted.
[0019] Figure 9 A distributed computing system 900 for Softmax computation according to one embodiment is depicted.
[0020] Figure 10 A multi-die package 1012 according to one embodiment is depicted.
[0021] Figure 11 A neural network processor 1100 implemented on a single chip according to one embodiment is depicted.
[0022] Figure 12 A local processing element 1200 according to one embodiment is depicted.
[0023] Figure 13 A local processing element 1300 according to one embodiment is depicted in more detail.
[0024] Figure 14 Details of a post-processor 1212 according to one embodiment are depicted.
[0025] Figure 15 A global processing element 1522 according to one embodiment is depicted.
[0026] Figure 16 A parallel processing unit 2008b according to one embodiment is depicted.
[0027] Figure 17 A general processing cluster 1700 according to one embodiment is depicted.
[0028] Figure 18 A memory partition unit 1800 according to one embodiment is depicted.
[0029] Figure 19 A streaming multiprocessor 1900 according to one embodiment is depicted.
[0030] Figure 20 A processing system 2000 according to one embodiment is depicted.
[0031] Figure 21 An exemplary processing system 2100 according to another embodiment is depicted.
[0032] Figure 22 A graphics processing pipeline 2200 according to one embodiment is depicted. DETAILED DESCRIPTION
[0033] In many deep learning applications, inference is often performed on trained models using less accurate data representations to improve performance (increase throughput or latency per inference) and reduce the computational energy consumed per inference. These models can be applied to tensor cores on programmable graphics processing units (GPUs) or specialized deep learning accelerators. Some solutions focus on improving the performance of neural network layers by implementing layers such as batched matrix multiplication operations in GPUs. A "neural network" refers to an algorithm or computing system based on a collection of connected units or nodes called artificial neurons that loosely model the neurons in a biological system. Each connection between neurons, like a synapse in a biological brain, can transmit a signal (activation) from one artificial neuron to another. An artificial neuron that receives a signal (input activation) can process it and then send out signals (output activation) to other artificial neurons connected to it. "Input activation" refers to the activation received by a neuron in a neural network. "Output activation" refers to the activation output by a neuron in a neural network. Output activation is typically computed based on the input activation of the neuron and weights applied to the input activation. "Weights" refer to values that are multiplied with activations to increase or decrease the impact of the activation value in an activation function. "Activation" refers to the output value of a neuron in a neural network, computed based at least in part on the weights input to the neuron and an activation function of the neuron. Activation is also referred to as "activation value."
[0034] As subsequent generations of GPU hardware have emerged, the computational performance of core matrix multiplication computations has improved. Other aspects of deep learning applications have therefore become bottlenecks. For example, in many conversational artificial intelligence workloads, such as transformer-based neural networks, Softmax computations can become a bottleneck.
[0035] Conversational AI implementations that utilize transformer neural networks can be particularly impacted by poor performance of Softmax. At a high level, transformer neural network structures include encoding components, decoding components, and connections between these components. The encoding components can include a stack of multiple encoding stages, and the decoding components can include a stack of multiple decoding stages, typically the same number as the encoding stages. The encoding stages (referred to simply as "encoders") are neural networks, which can typically be identical in structure to each other unless they can be differentiated during training (e.g., trained to have different weights from each other). Likewise, the decoding stages (referred to simply as "decoders") can all typically have the same structure except for differentiations obtained during training. The encoders and decoders can include "layers" that perform operations on vector inputs to generate vector or scalar outputs. These vectors can be multi-dimensional (typically NxMx...P, N, M,...P > 1) and nested, often referred to as tensors.
[0036] Traditional Softmax computation typically involves the following operations: (i) computing the maximum value in the input vector, (ii) applying the exponential to floating point or fixed point numbers, (iii) performing a sum of the exponential values, and (iv) dividing the exponential values by the sum. The formula for traditional Softmax operation is
[0037]
[0038] Traditional Softmax equation
[0039] One algorithm for computing traditional Softmax is:
[0040] 1 : m0← -∞
[0041] 2: for k ← 1, V do
[0042] 3: m k ← max(m k-1 , x k )
[0043] 4: end for
[0044] 5: d0← 0
[0045] 6: for j ← 1, V do
[0046] 7:
[0047] 8: end for
[0048] 9: for i ← 1, V do
[0049] 10:
[0050] 11: end for
[0051] Traditional Softmax algorithm
[0052] The algorithm involves multiple accesses to memory and exhibits low operation number reuse, sometimes leading to poor performance. Loop 2: - 4: over vector V (finding the maximum member m v of vector V) involves reading a vector from memory; loop 6: - 8: computing the sum of the exponentials (involving another) and loop 9: - 11: normalizing V (involving yet another).
[0053] Implementing exponent and reciprocal functions in hardware (for speed) can result in high design overhead (circuit area and / or power consumption). For example, the exponent and reciprocal functions can be performed in special function units (SFUs) of GPUs configured with look-up tables (LUTs) having 32-bit floating point precision. The high circuit area overhead of these components can make it too costly to replicate SFU units to achieve high throughput.
[0054] Embodiments are disclosed herein that improve the efficiency of Softmax computation. These solutions can be used to implement fast and efficient deep learning inference in transformer neural networks and other neural networks. The disclosed Softmax computation includes reduced precision implementations of various operations, substitution of 2x for ex to reduce instruction overhead associated with computing ex, and substitution of integer max computation for floating point max computation of vector elements. An extensible implementation decomposes Softmax into separate UnNormalized Softmax and Normalization operations.
[0055] The disclosed methods compute Softmax by formulating a 2 x vector of value elements. The expression“2 x vector of value elements” refers to a vector of elements each raised to the power of 2, where the exponent of the power of 2 is computed using an input value x from an input vector of elements. It should be understood that when referring to a“2 x vector of value elements,” the actual exponent of the power of 2 can in fact not be x, but a value derived from x (e.g., x-x max ), where x max is a running computation maximum of the input vector elements.
[0056] Embodiments are also described herein of efficient, tiled DNN processors that utilize an extensible design. These embodiments can benefit from the disclosed improvements to Softmax computation. The disclosed embodiments include beneficial features including: 1) a fully distributed, tile-based architecture, 2) flexible and efficient weight and activation tiling at the processing element (PE) level, chip level, and in some embodiments package level, improving data locality and reducing communication costs; 3) multi-level dataflow, improving data reuse and energy efficiency.
[0057] DNN processor embodiments utilize a data path designed to address the low compute-to-memory ratio of neural network layers. In some implementations, the data path includes local and global processing elements. Each local processing element includes logic to perform local multiply-accumulates of weights and input activations, as well as post-processing such as ReLu, MaxPool, Softmax, etc. By "logic" is meant machine memory circuitry and non-transitory machine-readable media, including machine-executable instructions (software and firmware) and / or circuitry (hardware) through which material and / or material-energy configurations including control and / or program signals, and / or settings and values (e.g., resistances, impedances, capacitances, inductances, current / voltage ratings, etc.) are used to affect the operation of a device. Magnetic media, electronic circuitry, electrical and optical storage (volatile and non-volatile), and firmware are all examples of logic. Logic specifically excludes pure signals or software by itself (but not machine memory containing software and thereby forming a material configuration).
[0058] Memory buffers in the form of collectors and register files can be arranged within processing elements and / or in the data path between processing elements. By "buffer" is meant a memory that stores values as inputs to a computation or as results of a computation. By "collector" is meant a buffer that is disposed between another buffer and an input or output of a data processor, such as a multiply-accumulate unit. By "multiply-accumulate unit" is meant a data processing circuit that performs multiply-accumulate operations, in which the product of two numbers is computed and added to an accumulator. Multiply-accumulate units can be referred to herein by their acronym, MAC, or MAC unit. A multiply-accumulate unit performs a computation of the form a <- a + (b * c). A vector multiply-accumulate unit computes the product of two vectors using an array of multipliers, then performs a reduction operation by adding all of the outputs of the multipliers to produce a partial sum, which is then added to an accumulator. By "partial sum" is meant an intermediate multiply-accumulate result in a dot-product accumulation computation. By "dot-product accumulation" is meant the computation of a dot product. A dot product is the sum of the products of corresponding entries of two sequences of numbers (vectors). A dot product is effectively computed using a vector multiply-accumulate unit.
[0059] DNN processor embodiments provide a multi-level memory and compute hierarchy that utilizes local storage of weights and output activations to improve the energy efficiency of neural network execution. Conventional neural network accelerator designs only exploit reuse opportunities at the innermost level of execution (e.g., a loop), whereas the disclosed architecture provides a multi-level memory and processing hierarchy to exploit data reuse opportunities across multiple loop levels, enabling a diverse set of energy-saving data flows. For example, instead of capturing temporal reuse only for weights or outputs, multi-level data flows can be implemented using weight and partial sum reuse during execution.
[0060] To efficiently implement a specific data flow, each local processing unit can use one or more collectors (e.g., small register files): one before the weight buffer, another before the accumulation buffer, and another before the input activation buffer. An "activation buffer" is a memory buffer used to store activation values (activations) used in neural network computations. Activations are computed by each neuron in a neural network layer using an activation function, sometimes also called a "transfer function." Activations can be simple binary values (e.g., "1" or "0" representing "ON" or "OFF"), or they may take a range of values from certain activation functions. These collectors filter out (reduce) expensive reads and writes to weight and partial summation buffers (e.g., SRAM), thereby improving overall energy efficiency. Global processing elements and / or chips can provide additional storage (e.g., global or shared register files) and processing power along the data path of neural network computations.
[0061] The disclosed DNN processor embodiments provide a heterogeneous block-based computing platform for different types of neural network computations. In addition to dense convolutions, many neural networks perform element-level computations and depth-level convolutions. To facilitate this computation, the architecture includes two general-purpose types of processing elements. The first type, called local processing elements, is specifically designed to perform dense convolutions with significant data reuse. The second type, called global processing elements, provides secondary storage for local processing elements during dense convolutions. Furthermore, global processing elements can perform element-wise operations and depth-wise convolutions with a low computation-to-memory ratio without transferring large amounts of data through neural network layers.
[0062] Figure 1 An exemplary system 100 utilizing artificial neural networks is depicted. Neural networks are widely used in applications such as speech-to-text conversion, natural language processing, language translation, image recognition and classification, and search.
[0063] In the specific example depicted, person 102 speaks into the microphone 110 of digital device 118, for example, to communicate with a mobile phone or home automation device (e.g., This can be used for voice assistant interaction on devices such as IoT devices 106 and / or cloud computing systems 104, for example, via a local area network 108 and / or a wide area network 112. The conversion of human speech 102 into text and / or commands understood by the IoT device 106 and / or cloud computing system 104 can be performed by one or more neural networks 114 utilizing one or more Softmax layers 116. Examples of neural networks 114 that can be used for these purposes include transform neural networks, recurrent neural networks, convolutional neural networks, and mixtures of these types, as well as other types known in the art.
[0064] Figure 2 An exemplary scenario of applying a neural network in a deep learning system 202 that utilizes Softmax computation is depicted in accordance with some embodiments. Deep learning system 202 can be used in computing systems 204, vehicles 206, and robots 208, to name a few examples. Deep learning system 202 can include one or more neural networks that provide image recognition and classification (machine vision), conversational AI, control systems for autonomous vehicles and robots, and so on.
[0065] Figure 3 A transformer neural network 302 in one embodiment is depicted. As previously noted, transformer neural networks can utilize Softmax computation extensively in attention layers. Transformer neural network 302 receives an input sequence 304 at a first encoder 306 of an encoder stack 308. Encoder 306 performs encoding on input sequence 304 and passes the result to an encoder 310, which performs additional encoding and passes the result to an encoder 312. Although three encoders are depicted in encoder stack 308, any manageable number can in fact be present.
[0066] The result of the last encoder 312 in encoder stack 308 is provided to a decoder stack 314. Decoder stack 314 as depicted includes three decoders (decoder 316, decoder 318, and decoder 320), but any manageable number can in fact be present. The encoded result of final encoder 312 is provided to a first decoder 316 of decoder stack 314, and the attention result of final encoder 312 can be fully connected to the encoder-decoder attention layer 504 of each encoder in decoder stack 314, in one embodiment. Decoder stack 314 operates on the result provided by encoder stack 308 to generate an output sequence 322 transformation of input sequence 304. There can typically be a linear layer and a Softmax layer (not depicted) at the output of the final decoder 320 level to produce output sequence 322.
[0067] Generally, attention vectors from any encoder self-attention layer can be provided to any decoder encoder-decoder attention layer. Moreover, the attention layers can be“multi-headed” as known in the art.
[0068] Figure 4An encoder 402 in one embodiment is depicted. The encoder 402 receives an input vector at a self-attention layer 404, which transforms the input vector before passing the input vector to a feed-forward neural network 406. The result of the feed-forward neural network 406 is passed to the next encoder stage (if there is one), and / or to a decoder (if the encoder 402 is the final encoder stage). Depending on the implementation, the result of the self-attention layer 404 can also be passed to one or more decoder stages (e.g., if the encoder 402 is the final encoder stage). There can typically be a sum and normalization layer (not depicted) after each of the self-attention layer 404 and the feed-forward neural network 406.
[0069] Figure 5 A decoder 502 in one embodiment is depicted. The decoder 502 receives an input at a self-attention layer 506 (either from a preceding decoder stage or from an encoder stage). The result of the self-attention layer 506 is passed to an encoder-decoder attention layer 504, which can also receive attention input from one or more self-attention layers 404 of the encoder stack 308. The encoder-decoder attention layer 504 helps the decoder 502 focus on more relevant parts of the input sequence at particular positions in the input sequence (similar to attention in seq2seq models). The encoder-decoder attention layer 504 is followed by a feed-forward neural network 508. There can typically be a sum and normalization layer (not depicted) after each of the self-attention layer 506, the encoder-decoder attention layer 504, and the feed-forward neural network 508.
[0070] The result of the encoder-decoder attention layer 504 is passed to the feed-forward neural network 508, which generates an output to the next decoder stage or a final output result (possibly after additional processing by a linear and Softmax layer).
[0071] Figure 6 An attention layer 602 in one embodiment is depicted. A matrix multiplication 604 is performed on an input vector to the attention layer 602 to form a query vector 606, a key vector 608, and a value vector 610. The matrices applied in the matrix multiplication 604 are derived by training the neural network that includes the attention layer 602.
[0072] Next, a score vector 612 is derived by performing a dot product 614 of the query vector 606 and the key vector 608. The element values in the score vector 612 determine how much attention to place on other parts (e.g., tokens) of the input vector when processing a particular token of the input vector. The score vector 612 is then processed with a Softmax 616 algorithm to normalize the scores so that they are all positive and add up to one. The Softmax scores determine the degree to which each token of the input sequence is expressed at a particular input sequence token position.
[0073] The weighted value vector 610 is then multiplied 618 by the Softmax score vector values and the weighted value vector 610 is summed (vector sum 620).
[0074] Figure 7A , Figure 7B , and Figure 7C An embodiment of a Softmax algorithm 700 is depicted that can improve some of the shortcomings of traditional Softmax. In Figure 7A the 7th row computation of the exponent is replaced by a more efficient power of 2 computation. The power of 2 is used for the normalization operation in the 10th row instead of the exponent.
[0075] In Figure 7B the computation of the maximum element in the vector V is combined with the computation of the sum of the vector elements (rows 3-8) to eliminate one of the three passes over the vector elements V. This results in the unnormalized Softmax vector V being renormalized by V in rows 9-11.
[0076] In Figure 7C the computation of the maximum element is computed in integer precision (row 4), which makes it Figure 7B the computationally expensive multiplication operation in row 5 is replaced by Figure 7C the computationally less expensive right shift operation in row 5 in Figure 7B the multiplication operation in the renormalization in row 10 is also Figure 7C replaced by the less expensive right shift in row 10.
[0077] The architecture of an embodiment of the unnormalized Softmax unit 828 and the normalization unit 830 is depicted in Figures 8A to 8DThe units can cooperate to implement, for example, the Softmax algorithm 700. Those skilled in the art will appreciate that the units can be implemented as hardware (e.g., in special function units 1912), firmware, software, or a combination thereof (i.e., "logic"). For example, aspects of the units can be implemented in hardware, with certain functions (e.g., linear piecewise approximation) micro-coded as extended ISA (Instruction Set Architecture) instructions executed by a processor. Some embodiments can implement many or all components of the units in software on a high performance computing platform.
[0078] The overall input vector for Softmax can be broken into smaller vectors that are fed to unnormalized Softmax units 828. These smaller portions of the overall vector can be processed in parallel by multiple processing elements (see, e.g., FIG. 8B), or they can be input and processed sequentially by the unnormalized Softmax units 828. Figure 9
[0079] The vector integer max unit 802 receives an input vector and computes the integer maximum (LocalMax) among the vector elements. Each element of the input vector is rounded to an integer, and the maximum value element after rounding and comparison (max comparator 834) is selected as the local maximum (LocalMax). If the input vector is a small portion of the Softmax overall vector, then the maximum value element of the vector is the "local" maximum. This local maximum can be shared among other processing elements that are processing other segments of the overall vector for comparison and determination of the global maximum of the vector. In one embodiment, a central processor / controller that coordinates execution of the Softmax algorithm across processing elements also collects the LocalMax values and determines the global maximum (GlobalMax) of the vector. A "controller" refers to any logic that controls the operation of other logic. When the controller is implemented in hardware, it can be, for example, many of the well-known microprocessor models, a graphics processing unit, or a custom controller implemented using an application specific integrated circuit (ASIC), a system on a chip (SOC), or one of many other ways known in the art. The controller can also be implemented in software or firmware, as computer instructions stored in volatile or non-volatile memory. The controller is typically used to coordinate the operation of one or more other components in a system, such as providing signals to other components to start and stop their operation, or instructing other components to perform using specific commands.
[0080] The 2 power computation unit 804 receives the input vector and the LocalMax value computed by the vector integer max unit 802. The 2 power computation unit 804 subtracts the LocalMax from each element value x of the input vector, and then computes 2 (x-LocalMax) Low-precision micro-architecture can be implemented to improve the computational efficiency (i.e., reduce the computational complexity) in the 2-power computation unit 804. In one embodiment, the input vector elements and LocalMax can be implemented in low-precision (meaning lower precision than typical floating-point or long integer) fixed-point representation (with six integer bits and two fractional bits). The linear piecewise computation unit 822 can utilize a fixed-point fractional divider 824 to direct the fractional bits to the lookup table 832 and the integer bits to the left-shift 826 logic to generate the binary power.
[0081] The linear piecewise computation unit 822 can be implemented using a lookup table 832, which in one embodiment, each includes four entries and ten bits. The use of an integer maximum (IntMax) can simplify the 2-power computation unit 804 and the normalization unit 830 in terms of circuit area, power consumption, and computation speed. This can simplify the floating-point subtraction operation to integer in the 2-power computation unit 804 and avoid the need for 2 x power computation (LPW) in the linear piecewise computation unit 822.
[0082] The un-normalized Softmax values generated by the 2-power computation unit 804 can be sequentially reduced in the reduction unit 806 to compute the power sum (PowSum). It is within the ordinary skill of those in the art to sequentially perform the reduction operation of the PowSum across the entire Softmax vector in the reduction unit 806 using the vector element adder 836, the power sum selector 816 (selecting the power sum from the vector element adder 836 or from another processing element), the right shifter 818, and the adder 820.
[0083] Similarly, the reduction operation of the LocalMax values can be performed using the maximum selector 808 and the maximum comparator 810 and the results stored in the memory buffer 812. When the sub-vectors of the Softmax computation are spatially distributed across multiple processing elements, the cross-processing element (PE) reduction operation of the local PowSum and IntMax values can be performed and shared among the processing elements (via the memory buffer 812 and the power sum buffer 814) to determine the global maximum (GlobalMax) and the global power sum (GlobalPowSum).
[0084] The normalization unit 830 can receive UnnormedSoftmax (un-normalized Softmax vector), LocalMax, GlobalMax, and GlobalPowSum values as inputs and perform the normalization operation by first calculating the inverse of GlobalPowSum using a low-precision LPW reciprocal unit, which can be implemented as a 10-byte size LUT in one embodiment. The final Softmax vector elements can be calculated by right-shifting the un-normalized Softmax (UnnormedSoftMax) vector elements and multiplying by the inverse of GlobalPowSum.
[0085] By using reduced bit-width operands and reduced LUTs, the implementation cost (in area, power, and / or speed) of the Softmax computation unit can be significantly reduced from floating-point SFUs used to perform similar functions on traditional GPUs. Figure 8D Reduced bit-width values Q(n, m) are depicted for one embodiment for various factors in a Softmax computation, where n is the number of integer bits and m is the number of fractional bits.
[0086] Figure 9 A distributed computing system 900 that can be configured to implement a scalable neural network processor in one embodiment is depicted. The distributed computing system 900 includes multiple processing elements 904 that pass values between each other using local router interfaces 902 to perform distributed neural network computations. The weights of a deep neural network are tiled across the local memory spaces of the processing elements 904. The processing elements 904 can be distributed across multiple chips in a single package / device / printed circuit board, or across multiple packages / devices / printed circuit boards.
[0087] The overall deep neural network distributed computation is coordinated by a controller 906 with the intermediate values of the computation stored in local memories of the processing elements 904 or a global global memory buffer 910. A "global memory buffer" refers to a buffer available for use by all or at least multiple processing elements on a chip. Tensors, weights, and other values of the computation can also be read and written from memory 912 (e.g., a larger but slower DRAM device) at least initially.
[0088] In one embodiment, the controller 906 is configured to coordinate the processing elements 904 to perform un-normalized Softmax, which is then normalized by a normalization unit 908.
[0089] The requirements for deep neural network application can vary greatly. For example, a typical data center inference application (such as image recognition) can prioritize performance and scalability at low latency and can be willing to sacrifice classification accuracy, while an inference for an autonomous driving workload can prioritize energy efficiency within real-time constraints while maintaining the best achievable network accuracy. Distributed computing system 900 is a general-purpose architecture that can be configured as a specialized inference accelerator with performance and power advantages over general-purpose solutions.
[0090] A multi-die package 1012 embodiment for implementing a DNN accelerator is depicted in Figure 10 The multi-die package 1012 can be a semiconductor package that includes multiple dies 1018 (chips). Each die 1018 includes multiple processing elements 1014, a global buffer 1002, and a controller 1004 (e.g., an open-source RISC-V processor). The elements of each chip / die communicate via on-chip network routers 1008. The multiple chips in the package communicate with each other via a package network router 1016, and can also communicate with a host 1020 system that includes DRAM 1006 or other memory, via a field programmable gate array (FPGA 1010), joint test action group (JTAG) logic, or other interface technology known in the art.
[0091] Some or all of the processing elements are local processing elements that include a weight buffer for receiving and storing weight values of a deep neural network. A “weight buffer” refers to a buffer that stores weight values. The local processing elements include an activation buffer for receiving activation values of the deep neural network. The weight buffer and the activation buffer can be separate elements within each processing element. The local processing elements also include multiple multiply-accumulate units to combine the weight values and the activation values in parallel to generate partial sums.
[0092] The multi-die package 1012 can be configured to distribute the weight values and the activation values among the local processing elements spatially and temporally (over time). The global memory buffer of each chip can act as a secondary buffer for activation values during computation. A “secondary buffer” refers to a memory that stores and retrieves values when the values are needed for computation but are not available in a primary buffer. Here, the chip global buffer can act as a secondary buffer to the primary activation buffer of the chip processing elements. The distribution of weights and activations during computation can be performed by the controller 1004 of the chip. The controller 1004 or a local controller of any of the processing elements 1014 can be configured by instructions stored in memory to perform the various data flows described below. The memory configured in this way can be referred to simply as “logic” herein. The location of this logic is a design choice. The memory storing these instructions can be any of the memories depicted in the figure, or a different memory not depicted.
[0093] Figure 11 A neural network processor 1100 embodied on a single chip is depicted. The neural network processor 1100 can utilize fixed-point data paths between multiple processing elements 1014. The neural network processor 1100 also includes the aforementioned global buffer 1002 and controller 1004, which can be a RISC-V processor, for example. The processing elements 1014 and global buffer 1002 communicate via on-chip network routers 1008 or other interconnect technology (see GPU implementation, described further below). If routers are used, they can be implemented as centralized or in a distributed fashion on each processing element 1014. The processing elements 1014 communicate with processing elements on the same package using the routers / interconnect, or in some embodiments across packages via network package routers 1016.
[0094] Figure 12 An exemplary local processing element 1200 is depicted at a high level. The processing element 1200 includes multiple vector multiply-accumulate units 1202, a weight buffer 1204, an activation buffer 1206, a router 1208, a controller 1214, an accumulation memory buffer 1210, and a post-processor 1212. The “accumulation memory buffer” refers to a memory buffer used to store the results of computations by one or more multiply-accumulate units. The “post-processor” refers to logic applied after multiplication and accumulation in neural network computations. In one embodiment, the activation buffer 1206 can be implemented as a dual-port SRAM to receive activation values from the global buffer 1002 or from other local or global processing elements via the router 1208 or other interconnect. The router 1208 can be a component of the distributed on-chip network routers 1008, including a serializer / deserializer, a packetizer, an arbitrator, a high-level extensible interface, and other components known in the art, in one embodiment.
[0095] In one embodiment, the weight buffer 1204 can be implemented as a single-port SRAM storing weight values. The weight values used by the vector multiply-accumulate units 1202 can be “weight stationary,” meaning that they are not updated at every clock cycle, but rather once an output activation value is computed for a particular layer of a deep neural network.
[0096] The accumulation memory buffer 1210 can include one or more SRAM devices to store output activations computed by the vector multiply-accumulate units 1202. The router 1208 communicates these output activations and control signals from the processing element 1200 to other processing elements.
[0097] The processing elements 1200 can efficiently perform all operations of the convolutional and fully connected layers of a DNN, including multiply-accumulate, clipping, scaling, bias addition, ReLU, and pooling (the last five in the post-processor 1212). "Bias addition" refers to the inclusion of a bias (e.g., a fixed output value or an increment to an output value) of one or more neurons of a neural network layer. Bias addition is a technique used to ensure that at least one neuron of the layer produces a non-zero activation to the next layer when the layer detects no features in its input. The vector multiply-accumulate units 1202 can operate on the same input using different filters. In one embodiment, each vector multiply-accumulate unit 1202 performs a dot product of eight input channels at each clock cycle and accumulates the result into an accumulation memory buffer 1210. The weights stored in the weight buffer 1204 do not change until the entire computation of the output activations is complete. Each processing element 1200 reads the input activations in the activation buffer 1206, performs a multiply-accumulate operation, and writes the output activations to the accumulation memory buffer 1210 at each clock cycle. The frequency of accessing the weight buffer 1204 depends on the input activation matrix dimension and the number of filters used.
[0098] The vector multiply-accumulate units 1202 of each processing element 1200 compute a portion of the wide dot-product-accumulate as a partial result and forward the partial result to a neighboring processing element. A "neighboring processing element" refers to a processing element that is one hop distance away from another processing element on a data communication network fabric, such as an on-chip network or an on-package network.
[0099] The post-processor 1212 converts the partial results into final results and transfers to the global buffer 1002. The global buffer 1002 acts as a staging area for the final multiply-accumulate results between deep neural network layers.
[0100] The accumulation memory buffer 1210 receives the output from the vector multiply-accumulate units 1202. The central controller 1004 allocates the weight values and activation values among the processing elements and utilizes the global memory buffer as a secondary buffer for the activation values. When processing an image, the controller 1004 configures the processing of the deep neural network layer spatially across the processing elements by the input / output channel size and temporally by the image height / width.
[0101] The global buffer 1002 stores both input activations and output activations from the processing elements 1014 for distribution by the aforementioned transceivers to the processing elements via multicasting. "Multicasting" refers to a group communication mechanism whereby a data transmission is addressed to a group of target devices (e.g., processing elements) simultaneously. Multicasting can enable one-to-many or many-to-many distribution. In one embodiment, each processing element 1014 includes a router 1208 to transfer 64 bits of data input and 64 bits of data output per clock cycle. This enables accumulation of partial sums of wide dot products whose computations are spatially tiled across the processing elements 1014.
[0102] Figure 13 An exemplary local processing element 1300 is depicted in more detail. The processing element 1300 includes the aforementioned vector multiply-accumulate unit 1202, weight buffer 1204, activation buffer 1206, router 1208, controller 1214, accumulation memory buffer 1210, and post-processor 1212 (e.g., post-processor 1212). Also depicted are a weight collector 1320 interposed between the weight buffer 1204 and the vector multiply-accumulate unit 1202, and an accumulation collector 1322 interposed between the vector multiply-accumulate unit 1202 and the accumulation memory buffer 1210. The accumulation collector 1322 can also be referred to herein as an "output collector." Also depicted are various memory buffer managers that can be used (e.g., weight memory buffer manager 1310, activation memory buffer manager 1312, and accumulation memory buffer manager 1316). A "memory buffer manager" refers to logic for managing the contents of a memory buffer, e.g., managing the availability of certain data (e.g., weights, activations) in the buffer when requested by a processing element.
[0103] The processing element 1300 includes vector multiply-accumulate units 1202, the number of which, N, for a given dataflow, is operable. Each vector multiply-accumulate unit 1324 performs V multiplications and additions per clock cycle. Thus, in each clock cycle, the processing element 1300 can multiply a weight matrix of dimension NxV with an input activation vector of size V to generate a partial sum vector of size N. In other words, each vector multiply-accumulate unit 1202 can perform a V-width dot product computation per clock cycle. One or both of N and V can be configured at the controller 1004.
[0104] Input activation buffer 1206 has an operational size IA and weight buffer 1204 has an operational size W. "Operational size" refers to a pool of resources available for performing computations during device operation, which can be smaller than the total number or maximum size of the pool of resources. The operational size can be configured using registers or other settings, for example, for higher performance or lower power consumption. One or both of W and IA can be configured at controller 1004. Accumulation memory buffer 1210 has an operational size of A.
[0105] Each vector multiply-accumulate unit 1202 includes a weight collector 1320 buffer, which has a configurable depth (e.g., number of different registers or addresses in a register file used by the vector multiply-accumulate unit 1202 during computation) WD and a width V x N x WP (WP also referred to as weight precision). Input activations have a width IAP. Each vector multiply-accumulate unit 1202 also includes an accumulation collector 1322, which has a configurable operational depth AD and a width NxAP (AP also referred to as accumulator precision). V-wide dot products and N-sized partial sum vectors can thus be computed by each vector multiply-accumulate unit 1324 at mixed precision. Some or all of WD, WP, IAP, AD, and AP can be configured by controller 1004.
[0106] The weight buffer 1204 read (output) port is WP x N x V bits wide and is capable of providing different weight vectors to different vector multiply-accumulate units 1202. The activation buffer 1206 is IAP x V bits wide because the same IA vector is provided to all N vector multiply-accumulate units 1202 in parallel.
[0107] For example, the values of V and N can be adjusted to achieve a certain amount of computation parallelism and weight reuse. Depending on the configuration of N and V, other parameters such as W, IA, A, etc. can be adjusted to ensure that the vector multiply-accumulate units 1202 remain busy during convolution computation.
[0108] The weight buffer 1204 and the activation buffer 1206 each have an associated address generator (address generator 1314 and address generator 1318, respectively) that generates an address each cycle. "Address generator" refers to logic that computes an address value in memory to read or write data from that address. The order of operations performed by the vector multiply-accumulate units 1202 is controlled by these address generators, which can be configured to support temporal reuse of weight or result values produced in a clock cycle in the accumulation collector 1322 for different types of data streams. The depth WD of the weight collector 1320 can be configured to achieve different amounts of temporal reuse of partial sum values depending on the requirements of the data stream. Likewise, the depth AD of the accumulation collector 1322 can be configured to achieve different amounts of temporal reuse of weight values depending on the requirements of the data stream.
[0109] The processing element 1300 can also include an input gatherer 1328 disposed between the activation buffer 1206 and the vector multiply-accumulate unit 1202. The operational depth IC of the input gatherer 1328 can be configured to set different levels of input activation fixed data flow, as further described below.
[0110] Each of the weight buffer 1204 and the activation buffer 1206 also has a buffer manager (weight memory buffer manager 1310 and activation memory buffer manager 1312, respectively) responsive to the controller 1214 and determines availability of data to the vector multiply-accumulate unit 1202. In some embodiments, the dimensionality of the address generator and the granularity of data movement from the weight buffer 1204 and the activation buffer 1206 to the vector multiply-accumulate unit 1202 can be configured at the controller 1004.
[0111] The accumulation memory buffer 1210 stores the partial sums from all N vector multiply-accumulate units 1202 and can be optimized to perform read-modify-write operations per cycle. The partial sums from the N vector multiply-accumulate units 1202 are packed into vectors of width AP x N and stored in the accumulation memory buffer 1210. From there, they can be sent directly to another processing element for cross-processing element reduction or to the post-processor 1212 to produce the final output activations. The post-processor 1212 can provide scaling and quantization operations, as well as additional ReLU and pooling operations to implement layer fusion.
[0112] Input weights 1302 arrive through the router 1208 and are stored in the weight buffer 1204. Input activations 1304 also arrive through the router 1208 and are stored in the activation buffer 1206. The computed output activations 1326 (after post-processing by the post-processor 1212) or partial sums 1306 from the accumulation memory buffer 1210 are output via the router 1208 to the global buffer 1002 or to an adjacent processing element, respectively. Cross-processing element reductions 1308 from the adjacent processing element can be received by the router 1208 and accumulated in the accumulation memory buffer 1210. By “cross-processing element reduction” is meant reducing the partial computation results of a first processing element to the final or more complete computation results of one or more other processing elements.
[0113] Figure 14A post-processor 1212 in one embodiment is depicted. The post-processor 1212 can include logic (e.g., special function units 1912) for common neural network operations such as pooling, ReLu activation, bias addition, rounding, and scaling. In some embodiments, the unnormalized Softmax unit 828 can be implemented as a special function unit in the post-processor 1212 of processing elements 1014.
[0114] Figure 15 A global processing element 1522 in one embodiment is shown. The global processing element 1522 includes a global memory buffer 1502 with an arbitrated memory bank 1520 (e.g., a "scratchpad"), a controller 1504 that computes on data in the arbitrated memory bank 1520, and an active address generator 1506 and a destination address generator 1510 that generate source and destination addresses, respectively, for computation. A "memory bank" refers to a logical unit of memory storage. The memory bank can be determined by a memory controller along with the physical organization of the hardware memory interface. In a typical synchronous dynamic random access memory (SDRAM) or double data rate synchronous dynamic random access memory (DDR SDRAM), the memory bank includes rows and columns of memory cells, possibly distributed over multiple memory chips. The global processing element 1522 communicates with other processing elements via the router 1208.
[0115] A data path 1508 to and from the global memory buffer 1502 includes a register file 1512 that can act as a collector for one or more of input activations 1518, output activations 1514, and partial sums 1516 to and from local processing elements, as required by the dataflow.
[0116] Many neural networks utilize computations such as element-wise computation and deep convolution to improve overall accuracy. Local processing elements are specialized to perform dense convolutions with significant data reuse. The global buffer 1002 can be used by local processing elements as a secondary data store during dense convolution, and can also perform computations for element-wise operations and deep convolutions. Global processing elements perform computations locally at a low compute-to-memory ratio without the need to transfer data through layers (and chips) of the neural network.
[0117] The controller 1504 can be local to each global processing element 1522 or can be implemented by a chip master controller (controller 1004). Likewise, the global memory buffer 1502 can be local to the global processing element 1522 or implemented by the global buffer 1002.
[0118] The algorithms and techniques disclosed herein can be executed by computing devices that utilize at least one graphics processing unit (GPU) and / or general purpose data processors (e.g.,“central processing units or CPUs”). Exemplary architectures are described that can be configured to execute the techniques disclosed herein on such devices.
[0119] The following description can make use of certain acronyms and abbreviations, which are as follows:
[0120] •“DPC” refers to“Data Processing Cluster”;
[0121] •“GPC” refers to“General Purpose Cluster”;
[0122] •“I / O” refers to“Input / Output”;
[0123] •“L1 cache” refers to“Level 1 Cache”;
[0124] •“L2 cache” refers to“Level 2 Cache”;
[0125] •“LSU” refers to“Load / Store Unit”;
[0126] •“MMU” refers to“Memory Management Unit”;
[0127] •“MPC” refers to“M Pipe Controller”;
[0128] •“PPU” refers to“Parallel Processing Unit”;
[0129] •“PROP” refers to“Pre-Raster Operations Unit”;
[0130] •“ROP” refers to“Raster Operations”;
[0131] •“SFU” refers to“Special Function Unit”;
[0132] •“SM” refers to“Streaming Multi-Processor”;
[0133] •“Viewport SCC” refers to“Viewport Scaling, Clipping, and Culling”;
[0134] •“WDX” refers to“Work Distribution Crossbar”;
[0135] •“XBar” refers to“Crossbar”.
[0136] Parallel processing unit
[0137] Figure 16A parallel processing unit 2008b according to one embodiment is shown. In one embodiment, the parallel processing unit 2008b is a multi-threaded processor that is implemented on at least one integrated circuit device. The parallel processing unit 2008b is a latency hiding architecture designed to process many threads in parallel. A thread (i.e., an execution thread) is an instance of a set of instructions configured to be executed by the parallel processing unit 2008b. In one embodiment, the parallel processing unit 2008b is a graphics processing unit (GPU) that is configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data in order to generate two-dimensional (2D) image data for display on a display device, such as a liquid crystal display (LCD) device. In other embodiments, the parallel processing unit 2008b can be used to perform general purpose computations. Although one exemplary parallel processor is provided herein for illustrative purposes, it should be noted that this processor is set forth for illustrative purposes and any processor can be used in addition to and / or in place of the processors described herein.
[0138] The at least one parallel processing unit 2008b module can be configured to accelerate thousands of high performance computing (HPC), datacenter, and machine learning applications. The parallel processing unit 2008b can be configured to accelerate numerous deep learning systems and applications, including autonomous vehicle platforms, deep learning, high-accuracy speech, image, and text recognition systems, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.
[0139] As shown in FIG. 17A, the parallel processing unit 2008b includes an input / output (I / O) unit 1602, a front-end unit 1604, a scheduler unit 1608, a work distribution unit 1610, a hub 1606, a crossbar (Xbar) 1614, at least one general processing cluster (GPC) 1700 module, and at least one memory partition unit 1800 module. The parallel processing unit 2008b can be connected to a host processor and other parallel processing units 2008b via an NVLink 1616 interconnect. The parallel processing unit 2008b can be connected to a host processor and other peripheral devices via an interconnect 1618. The parallel processing unit 2008b can also be connected to a local memory, which can include a number of memory devices 1612. In one embodiment, the local memory can include a number of dynamic random access memory (DRAM) devices. The DRAM devices can be configured as a high-bandwidth memory (HBM) subsystem, with multiple DRAM dies stacked Figure 16 within each device. The memory 1612 can include logic to configure the parallel processing unit 2008b to perform aspects of the techniques disclosed herein.
[0140] The NVLink 1616 interconnect enables system scaling and includes at least one parallel processing unit 2008b module in conjunction with at least one CPU, supports cache coherency between the parallel processing unit 2008b module and the CPU, and is CPU master. Data and / or commands can be sent by the NVLink 1616 to other units of the parallel processing unit 2008b or from them, such as at least one copy engine, a video encoder, a video decoder, a power management unit, etc. (not explicitly shown), via the hub 1606. In conjunction with each other, the parallel processing unit 2008b, and the CPU, facilitate a scalable and efficient means for processing of multimedia data. Figure 20 The NVLink 1616 is described in more detail.
[0141] The I / O unit 1602 is configured to transmit and receive communications (e.g., commands, data, etc.) from a host processor (not shown) over the interconnect 1618. The I / O unit 1602 can communicate directly with the host processor via the interconnect 1618 or can do so by
[0142] The I / O unit 1602 decodes data packets received via the interconnect 1618. In one embodiment, the data packets represent commands configured to cause the parallel processing unit 2008b to perform various operations. The I / O unit 1602 transmits the decoded commands, as specified by the commands, to various other units of the parallel processing unit 2008b. For example, some commands can be transmitted to the front end unit 1604. Other commands can be transmitted to the hub 1606 or to other units of the parallel processing unit 2008b, such as at least one copy engine, a video encoder, a video decoder, a power management unit, etc. (not explicitly shown). In other words, the I / O unit 1602 is configured to route communications between and among various logical units of the parallel processing unit 2008b.
[0143] In one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload for processing to the parallel processing unit 2008b. The workload can comprise a number of instructions and data to be processed by those instructions. The buffer is a region of memory that is accessible (e.g., read / write) by both the host processor and the parallel processing unit 2008b. For example, the I / O unit 1602 can be configured to access the buffer in system memory connected to the interconnect 1618 via memory requests transmitted over the interconnect 1618. In one embodiment, the host processor writes the command stream into the buffer and then sends a pointer to the start of the command stream to the parallel processing unit 2008b. The front-end unit 1604 receives the pointer to the at least one command stream. The front-end unit 1604 manages the at least one stream, reading commands from the stream and forwarding the commands to the various units of the parallel processing unit 2008b.
[0144] The front-end unit 1604 is coupled to a scheduler unit 1608, which is configured to schedule tasks for execution on the various general processing clusters 1700. The scheduler unit 1608 is configured to track state information related to various tasks managed by the scheduler unit 1608. The state can indicate which general processing cluster 1700 a task is assigned to, whether the task is active or inactive, the priority associated with the task, etc. The scheduler unit 1608 manages execution of multiple tasks on the at least one general processing cluster 1700 module.
[0145] The scheduler unit 1608 is coupled to a work distribution unit 1610, which is configured to distribute tasks to the general processing clusters 1700 for execution. The work distribution unit 1610 can track a number of scheduled tasks received from the scheduler unit 1608. In one embodiment, the work distribution unit 1610 manages a pending task pool and an active task pool for each general processing cluster 1700 module. The pending task pool can include a number of slots (e.g., 32 slots) that contain tasks assigned to be processed by a particular general processing cluster 1700. The active task pool can include a number of slots (e.g., 4 slots) for tasks that are currently being actively processed by a general processing cluster 1700 module. When a general processing cluster 1700 completes processing of a task, the task is evicted from the active task pool for that general processing cluster 1700, and one of the other tasks from the pending task pool is selected and scheduled for execution on the general processing cluster 1700. If the active task on a general processing cluster 1700 has idled, e.g., while waiting for a data dependency to be resolved, then the active task can be evicted from the general processing cluster 1700 and returned to the pending task pool, and another task from the pending task pool is selected and scheduled for execution on the general processing cluster 1700.
[0146] The work distribution unit 1610 communicates with at least one general processing cluster 1700 module via an XBar (crossbar) 1614. The XBar 1614 is an interconnect network that couples many of the processing units 2008b to other processing units 2008b. For example, the XBar 1614 can be configured to couple the work distribution unit 1610 to a particular general processing cluster 1700. Although not explicitly shown, at least one other processing unit 2008b can also be connected to the XBar 1614 via the hub 1606.
[0147] Tasks are managed by a scheduler unit 1608 and dispatched by a work distribution unit 1610 to the general processing clusters 1700. The general processing clusters 1700 are configured to process tasks and generate results. The results can be consumed by other tasks within the general processing clusters 1700, routed via the XBar 1614 to a different general processing cluster 1700, or stored in the memory 1612. The results can be written to the memory 1612 via a memory partition unit 1800 module, which implements a memory interface for reading from and writing to the memory 1612. The results can be sent over the NVLink 1616 to another parallel processing unit 2008b or CPU. In one embodiment, the parallel processing unit 2008b includes a number U of memory partition units 1800 modules, equal to the number of independent and distinct memory devices 1612 coupled to the parallel processing unit 2008b. The memory partition unit 1800 is described in greater detail below in conjunction with FIG. 18. Figure 18 The memory partition unit 1800 is described in greater detail below.
[0148] In one embodiment, the host processor executes a driver kernel that implements an application programming interface (API) that enables an application to execute on the host processor to schedule operations for execution on the parallel processing unit 2008b. In one embodiment, multiple compute applications are executed simultaneously by the parallel processing unit 2008b, and the parallel processing unit 2008b provides isolation, quality of service (QoS), and independent address spaces for the multiple compute applications. An application can generate instructions (e.g., API calls) that cause the driver kernel to generate tasks for execution by the parallel processing unit 2008b. The driver kernel outputs the tasks to a stream that is being processed by the parallel processing unit 2008b. Each task can include at least one related thread group, referred to herein as a warp. In one embodiment, a warp includes 32 related threads that can be executed in parallel. Cooperative threads can refer to a plurality of threads that execute instructions of a task and can exchange data through shared memory. Cooperative threads are described in greater detail below in conjunction with FIG. 19. Figure 19 Threads and cooperative threads are described in greater detail below.
[0149] Figure 17 An embodiment is shown. Figure 16 The 2008b parallel processing unit is part of the general-purpose processing cluster 1700 module. For example... Figure 17 As shown, each general-purpose processing cluster 1700 module includes multiple hardware units for processing tasks. In one embodiment, each general-purpose processing cluster 1700 module includes a pipeline manager 1702, a pre-raster operation unit 1704, a raster engine 1708, a job allocation crossbar switch 1714, a memory management unit 1716, and at least one data processing cluster 1706. It should be understood that... Figure 17 The general-purpose processing cluster 1700 can include alternatives Figure 17 Other hardware units of the unit shown or excluding Figure 17 Other hardware units besides the unit shown.
[0150] In one embodiment, the operation of the general-purpose processing cluster 1700 is controlled by a pipeline manager 1702. The pipeline manager 1702 manages the configuration of at least one data processing cluster 1706 module for processing tasks assigned to the general-purpose processing cluster 1700. In one embodiment, the pipeline manager 1702 may configure at least one of the data processing cluster 1706 modules to implement at least a portion of the graphics rendering pipeline. For example, the data processing cluster 1706 may be configured to execute vertex shaders on a programmable streaming multiprocessor (SM) 1900. The pipeline manager 1702 may also be configured to route packets received from the job allocation unit 1610 to appropriate logical units within the general-purpose processing cluster 1700. For example, some packets may be routed to fixed-function hardware units in the pre-raster operation unit 1704 and / or the raster engine 1708, while other packets may be routed to the data processing cluster 1706 modules for processing by the primitive engine 1712 or the streaming multiprocessor 1900. In one embodiment, pipeline manager 1702 can configure at least one of the data processing cluster 1706 modules to implement neural network models and / or computation pipelines.
[0151] The pre-raster operation unit 1704 is configured to route data generated by the raster engine 1708 and the data processing cluster 1706 module to the raster operation (ROP) unit, in conjunction with... Figure 18 For a more detailed description, the pre-raster operation unit 1704 can also be configured to perform improved color mixing operations, organize pixel data, perform address translation, etc.
[0152] The raster engine 1708 includes several fixed-function hardware units configured to perform various raster operations. In one embodiment, the raster engine 1708 includes a setup engine, a coarse raster engine, a cull engine, a clip engine, a fine raster engine, and a tessellation engine. The setup engine receives transformed vertices and generates plane equations associated with the geometric primitives defined by the vertices. The plane equations are sent to the coarse raster engine to generate coverage information for the primitives (e.g., x, y coverage masks for tiles). The output of the coarse raster engine is sent to the cull engine, where fragments associated with primitives that fail a z-test are culled, and to the clip engine, where fragments that are outside a view frustum are clipped. Those fragments that survive clipping and culling can be passed to the fine raster engine to generate attributes for pixel fragments based on the plane equations generated by the setup engine. The output of the raster engine 1708 includes, for example, fragments to be processed by a fragment shader implemented within the data processing cluster 1706.
[0153] Each data processing cluster 1706 included in the general processing cluster 1700 includes an M-pipe controller 1710, a primitive engine 1712, and at least one streaming multiprocessor 1900 module. The M-pipe controller 1710 controls the operation of the data processing cluster 1706, routing data packets received from the pipeline manager 1702 to the appropriate units in the data processing cluster 1706. For example, data packets associated with vertices can be routed to the primitive engine 1712, which is configured to fetch vertex attributes associated with the vertices from the memory 1612. In contrast, data packets associated with a shader program can be sent to the streaming multiprocessor 1900.
[0154] Streaming multiprocessor 1900 includes a programmable streaming processor configured to process tasks represented by multiple threads. Each streaming multiprocessor 1900 is multithreaded and configured to concurrently execute multiple threads (e.g., 32 threads) from a particular thread group. In one embodiment, streaming multiprocessor 1900 implements a single-instruction, multiple-data (SIMD) architecture, where each thread in a thread group (e.g., warp) is configured to process a different dataset based on the same instruction set. All threads in the thread group execute the same instructions. In another embodiment, streaming multiprocessor 1900 implements a single-instruction, multiple-thread (SIMT) architecture, where each thread in a thread group is configured to process a different dataset based on the same instruction set, but where individual threads in the thread group are allowed to diverge during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each thread bundle, enabling concurrency between the thread bundle and serial execution within the thread bundle when threads within the thread bundle diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby achieving equal concurrency across all threads within and between thread bundles. When maintaining execution state for each individual thread, threads executing the same instructions can converge and execute in parallel for maximum efficiency. The following is in conjunction with... Figure 19 A more detailed description of the Streaming Multiprocessor 1900.
[0155] The memory management unit 1716 provides an interface between the general-purpose processing cluster 1700 and the memory partitioning unit 1800. The memory management unit 1716 can provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, the memory management unit 1716 provides one or more translation back buffers (TLBs) for performing translations from virtual addresses to physical addresses in memory 1612.
[0156] Figure 18 An embodiment is shown. Figure 16 The parallel processing unit 2008b has a memory partition unit 1800. For example... Figure 18As shown, the memory memory partition unit 1800 includes a raster operations unit 1802, a cache memory 1804, and a memory interface 1806. The memory interface 1806 is coupled to the memory 1612. The memory interface 1806 can implement a 32, 64, 128, 1024-bit data bus for high-speed data transfer. In one embodiment, the parallel processing unit 2008b incorporates U memory interface 1806 modules, one for each pair of memory partition units 1800, with each pair of memory partition units 1800 modules connected to a corresponding memory device 1612. For example, the parallel processing unit 2008b can be connected to up to Y memory devices 1612, such as high bandwidth memory stacks or graphics double data rate version 5 synchronous dynamic random access memory or other types of persistent storage.
[0157] In one embodiment, the memory interface 1806 implements an HBM2 memory interface and Y is equal to half of U. In one embodiment, the HBM2 memory stacks are located on the same physical package as the parallel processing unit 2008b, providing significant power and area savings compared to a traditional GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies and Y is equal to 4, with the HBM2 stack including two 128-bit channels per die for a total of 8 channels and a data bus width of 1024 bits.
[0158] In one embodiment, the memory 1612 supports single error correction double error detection (SECDED) error correcting code (ECC) to protect data. For compute applications that are sensitive to data corruption, the ECC provides improved reliability. In a large cluster computing environment, where the parallel processing unit 2008b processes large data sets and / or runs applications for an extended period, the reliability is particularly important.
[0159] In one embodiment, the parallel processing unit 2008b implements a multi-level memory hierarchy. In one embodiment, the memory memory partition unit 1800 supports unified memory to provide a single, unified virtual address space for the CPU and the parallel processing unit 2008b memory, enabling virtual memory systems to share data tightly and efficiently. In one embodiment, the frequency of access to memory locations by the parallel processing unit 2008b is tracked, with the results used to identify memory pages that are frequently accessed by the parallel processing unit 2008b. In one embodiment, the NVLink 1616 supports an address translation service, which allows the parallel processing unit 2008b to access page tables of the CPU directly and provide complete access to CPU memory by the parallel processing unit 2008b.
[0160] In one embodiment, the copy engine transfers data between multiple parallel processing unit 2008b modules or between a parallel processing unit 2008b module and the CPU. The copy engine can generate a page fault for an address that is not mapped to a page table. The memory memory partition unit 1800 can then service the page fault, map the address into a page table, after which the copy engine can perform the transfer. In a conventional system, memory is fixed (e.g., non-paged) for multiple copy engine operations between multiple processors, which significantly reduces the available memory. Due to the hardware page fault, an address can be passed to the copy engine without worrying about whether a memory page is resident, and whether the copy process is transparent.
[0161] Data from memory 1612 or other system memory can be retrieved and stored in the secondary cache 1804, which is on-chip with a respective general processing cluster 1700 module and shared among various general processing cluster 1700 modules. As shown, each memory memory partition unit 1800 includes a portion of the secondary cache 1804 that is associated with a corresponding memory device 1612. Lower level caches can then be implemented within the various units of a general processing cluster 1700 module. For example, each streaming multi-processor 1900 can implement a LI cache. The LI cache is a dedicated memory that is only accessible by a particular streaming multi-processor 1900. Data from the secondary cache 1804 can be fetched and stored into each LI cache for processing by the functional units of the streaming multi-processor 1900 module. The secondary cache 1804 is coupled to the memory interface 1806 and the crossbar 1614.
[0162] The raster operations unit 1802 performs graphics raster operations related to pixel colors such as color compression, pixel blending, etc. The raster operations unit 1802 also implements depth testing with the raster engine 1708, receiving a depth for a sample location associated with a pixel fragment from the culling engine of the raster engine 1708. The depth for the sample location associated with the fragment is tested against a corresponding depth in a depth buffer. If the fragment passes the depth test for the sample location, the raster operations unit 1802 updates the depth buffer and sends the results of the depth test to the raster engine 1708. It will be appreciated that the number of memory memory partition units 1800 modules can be different from the number of general processing cluster 1700 modules, and thus each raster operations unit 1802 can be coupled to each general processing cluster 1700 module. The raster operations unit 1802 tracks data packets received from different general processing cluster 1700 modules and determines to which general processing cluster 1700 module results generated by the raster operations unit 1802 are routed through the crossbar 1614. Although in the embodiment shown in FIG. 17, the raster operations unit 1802 is shown as a separate unit from the general processing cluster 1700, it will be appreciated that the raster operations unit 1802 can be implemented within the general processing cluster 1700. Figure 18The raster operation unit 1802 is included within the memory partition unit 1800, but in other embodiments, the raster operation unit 1802 may be located outside the memory partition unit 1800. For example, the raster operation unit 1802 may reside in the general-purpose processing cluster 1700 or another unit.
[0163] Figure 19 An embodiment is shown. Figure 17 The streaming multiprocessor 1900. For example... Figure 19 As shown, the streaming multiprocessor 1900 includes an instruction cache 1902, one or more scheduler units 1904 modules (e.g., scheduler unit 1608), a register file 1908, one or more processing cores 1910 modules, one or more special function units 1912, one or more load / store units 1914 modules, an interconnect network 1916, and a shared memory / L1 cache 1918.
[0164] As described above, the work allocation unit 1610 schedules tasks to execute on the general-purpose processing cluster 1700 module of the parallel processing unit 2008b. Tasks are assigned to a specific data processing cluster 1706 within the general-purpose processing cluster 1700, and if the task is associated with a shader program, it can be assigned to the streaming multiprocessor 1900. The scheduler unit 1608 receives tasks from the work allocation unit 1610 and manages the instruction scheduling of one or more thread blocks assigned to the streaming multiprocessor 1900. The scheduler unit 1904 schedules thread blocks to execute as thread bundles of parallel threads, wherein each thread block is assigned at least one thread bundle. In one embodiment, each thread bundle executes 32 threads. The scheduler unit 1904 can manage multiple different thread blocks, assign thread bundles to different thread blocks, and then dispatch instructions from multiple different cooperative groups to various functional units (i.e., the core 1910 module, the special function unit 1912 module, and the load / store unit 1914 module) during each clock cycle.
[0165] Cooperative groups are a programming model for organizing groups of communicating threads that allow developers to express the granularity at which threads are communicating, enabling richer, more efficient parallel decomposition. Cooperative launch APIs support synchronization between thread blocks to execute parallel algorithms. Traditional programming models provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often want to define thread groups at a granularity smaller than a thread block and synchronize within the defined groups to enable higher performance, design flexibility, and software reuse in the form of collective group-wide function interfaces.
[0166] Cooperative groups enable programmers to explicitly define thread groups at sub-block (e.g., as small as a single thread) and multi-block granularities and perform collective operations, such as synchronization across threads in a cooperative group. The programming model supports clean composition across software boundaries so that libraries and utility functions can safely synchronize in their local environment without making assumptions about convergence. Cooperative group primitives enable new patterns of cooperative parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across an entire thread block grid.
[0167] The dispatch unit 1906 is configured within the scheduler unit 1904 to transmit instructions to one or more functional units. In one embodiment, the scheduler unit 1904 includes two dispatch units 1906 that enable two different instructions from the same thread bundle to be dispatched during each clock cycle. In alternative embodiments, each scheduler unit 1904 can include a single dispatch unit 1906 or additional dispatch units 1906.
[0168] Each streaming multiprocessor 1900 includes a register file 1908 that provides a set of registers for functional units of the streaming multiprocessor 1900. In one embodiment, the register file 1908 is partitioned between each of the functional units such that each functional unit is allocated a dedicated portion of the register file 1908. In another embodiment, the register file 1908 is partitioned between different thread bundles executed by the streaming multiprocessor 1900. The register file 1908 provides temporary storage for operands of the data
[0169] Each streaming multiprocessor 1900 includes L processing core 1910 modules. In one embodiment, the streaming multiprocessor 1900 includes a large number (e.g., 128, etc.) of distinct processing core 1910 modules. Each core 1910 can include a fully pipelined, single-precision, double-precision, and / or mixed-precision processing unit that includes floating-point and integer arithmetic logic units. In one embodiment, the floating-point arithmetic logic units implement the IEEE 754-2008 standard for floating-point operations. In one embodiment, the core 1910 module includes 64 single-precision (32-bit) floating-point cores, 64 integer cores, 32 double-precision (64-bit) floating-point cores, and 8 tensor cores.
[0170] The tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are included in the core 1910 module. Specifically, the tensor cores are configured to perform deep learning matrix operations, such as convolution operations for neural network training and inference. In one embodiment, each tensor core operates on 4x4 matrices and performs matrix multiply and accumulate operations D = A' B + C, where A, B, C, and D are 4x4 matrices.
[0171] In one embodiment, the matrix multiply inputs A and B are 16-bit floating-point matrices, while the accumulate matrices C and D can be 16-bit floating-point or 32-bit floating-point matrices. The tensor cores operate on 16-bit floating-point input data and 32-bit floating-point accumulation. The 16-bit floating-point multiplication does 64 operations to produce a full-precision product, which is then accumulated using 32-bit floating-point addition with other intermediate products of the 4x4x4 matrix multiplication. In practice, the tensor cores are used to perform larger two-dimensional or higher-dimensional matrix operations built up from these smaller elements. APIs, such as the CUDA 9 C++ API, expose specialized matrix load, matrix multiply and accumulate, and matrix store operations to efficiently use the tensor cores from a CUDA-C++ program. At the CUDA level, the warp-level interface assumes 16x16 size matrices across all 32 threads of a warp.
[0172] Each streaming multiprocessor 1900 also includes M special-function units 1912 modules that perform special functions (e.g., attribute evaluations, inverse square root, etc.). In one embodiment, the special-function units 1912 modules can include a tree traversal unit configured to traverse a hierarchical tree data structure. In one embodiment, the special-function units 1912 modules can include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to load a texture map (e.g., 2D array of texture pixels) from memory 1612 and sample the texture map to produce sampled texture values for use in a shader program executed by the streaming multiprocessor 1900. In one embodiment, the texture map is stored in shared memory / L1 cache 1918. The texture unit implements texture operations, such as filtering operations using mip maps (i.e., texture maps of varying levels of detail). In one embodiment, each streaming multiprocessor 1900 includes two texture units.
[0173] Each streaming multiprocessor 1900 also includes N load / store units 1914 modules that implement load and store operations between shared memory / L1 cache 1918 and register file 1908. Each streaming multiprocessor 1900 includes an interconnect network 1916 that connects each functional unit to the register file 1908, as well as connects the load / store units 1914 to the register file 1908, shared memory / L1 cache 1918. In one embodiment, the interconnect network 1916 is a crossbar that can be configured to connect any functional unit to any register in the register file 1908, as well as connect the load / store units 1914 modules to memory locations in the register file 1908 and shared memory / L1 cache 1918.
[0174] Shared memory / L1 cache 1918 is an on-chip memory array that allows data storage and communication between the streaming multiprocessor 1900 and the primitive engine 1712, as well as between threads in the streaming multiprocessor 1900. In one embodiment, shared memory / L1 cache 1918 includes 128 KB of storage capacity and is in the path from the streaming multiprocessor 1900 to the memory partition unit 1800. Shared memory / L1 cache 1918 can be used for cache reads and writes. One or more of shared memory / L1 cache 1918, level two cache 1804, and memory 1612 are backing stores.
[0175] Combining data cache and shared memory functionality into a single memory block provides the best overall performance for both types of memory accesses. The capacity can be used by a program as a cache that does not use shared memory. For example, if the shared memory is configured to use half the capacity, then texture and load / store operations can use the remaining capacity. The integration within shared memory / L1 cache 1918 causes shared memory / L1 cache 1918 to function as a high-throughput pipeline for streaming data and, at the same time, provide high-bandwidth and low-latency access to frequently-reused data.
[0176] When configured for general-purpose parallel computation, a simpler configuration can be used compared to graphics processing. Specifically, Figure 16 The illustrated fixed-function graphics processing units are bypassed, creating a simpler programming model. In the general-purpose parallel computation configuration, work distribution unit 1610 assigns and dispatches thread blocks directly to data processing cluster 1706 modules. The threads in a block execute the same program, using the unique thread ID in the computation to ensure each thread generates a unique result, using streaming multiprocessor 1900 to execute the program and perform the computation, using shared memory / L1 cache 1918 to communicate between threads, and using load / store unit 1914 to read from and write to global memory through shared memory / L1 cache 1918 and memory memory partition unit 1800. When configured for general-purpose parallel computation, streaming multiprocessor 1900 can also write to scheduler unit 1608 commands that are usable to launch new work on data processing cluster 1706 modules.
[0177] Parallel processing unit 2008b can be included in a desktop computer, laptop computer, tablet computer, server computer, supercomputer, smart- phone (e.g., wireless, hand-held device), personal digital assistant (PDA), digital camera, vehicle, head-mounted display, hand-held electronic device, etc. In one embodiment, parallel processing unit 2008b is contained on a single semiconductor die. In another embodiment, parallel processing unit 2008b is included on a system-on-a-chip (SoC) along with one or more other devices, such as additional parallel processing unit 2008b modules, memory 1612, reduced instruction set computer (RISC) CPU, memory management unit (MMU), digital-to-analog converter (DAC), etc.
[0178] In one embodiment, parallel processing unit 2008b can be included on a graphics card that includes one or more memory devices. The graphics card can be configured to interface with a PCIe slot on a motherboard of a desktop computer. In yet another embodiment, parallel processing unit 2008b can be an integrated graphics processing unit (iGPU) or parallel processor contained in a chipset of a motherboard.
[0179] Exemplary computing system
[0180] Systems with multiple GPUs and CPUs are being used across various industries as developers expose to and leverage greater parallelism in applications such as artificial intelligence computing. High-performance GPU-accelerated systems with tens to thousands of compute nodes are being deployed in data centers, research institutions, and supercomputers to tackle larger problems. As the number of processing devices within high-performance systems increases, communication and data transmission mechanisms need to be scaled to support this increased bandwidth.
[0181] Figure 20 This is based on the use of one embodiment. Figure 16 A conceptual diagram of a processing system 2000 implemented using parallel processing units 2008b is provided. The processing system 2000 includes a central processing unit 2006, a switch 2004, and each of multiple parallel processing unit 2008b modules, as well as corresponding memory modules 1612. An NVLink 1616 provides a high-speed communication link between each parallel processing unit 2008b module. Although... Figure 20 A specific number of NVLink 1616 and interconnect 1618 connections are shown, but the number of connections to each parallel processing unit 2008b and central processing unit 2006 can vary. Switch 2004 interfaces between interconnect 1618 and central processing unit 2006. The parallel processing unit 2008b module, memory 1612 module, and NVLink 1616 connections can reside on a single semiconductor platform to form parallel processing module 2002. In one embodiment, switch 2004 supports two or more protocols that interface between various different connections and / or links.
[0182] In another embodiment (not shown), NVLinks 1616 provide one or more high-speed communication links between each parallel processing unit module (parallel processing unit 2008a, parallel processing unit 2008b, parallel processing unit 2008c,... parallel processing unit 2008d) and central processing unit 2006, and a switch 2004 interfaces between interconnect 1618 and each parallel processing unit module. The parallel processing unit modules, memory 1612 modules, and interconnect 1618 can be located on a single semiconductor platform to form a parallel processing module 2002. In yet another embodiment (not shown), interconnect 1618 provides one or more communication links between each parallel processing unit module and central processing unit 2006, and a switch 2004 uses NVLinks 1616 to interface between each parallel processing unit module to provide one or more high-speed communication links between parallel processing unit modules. In another embodiment (not shown), NVLinks 1616 provide one or more high-speed communication links between parallel processing unit module B and central processing unit 2006 through switch 2004. In yet another embodiment (not shown), interconnect 1618 directly provides one or more communication links between each parallel processing unit module. One or more NVLink 1616 high-speed communication links can be implemented as physical NVLink interconnects or on-chip or off-chip interconnects using the same protocol as NVLinks 1616.
[0183] In the context of this specification, a single semiconductor platform can refer to a sole unitary semiconductor-based integrated circuit that is fabricated in a single fabrication operation. It should be noted that the term single semiconductor platform can also refer to multi-chip modules with increased connectivity which simulate on-chip operation, and make substantial improvements over utilizing a conventional bus implementation. Of course, the various circuits or devices can alternatively be formed on a single semiconductor platform depending on the desires of a user. Optionally, parallel processing module 2002 can be implemented as a circuit board substrate, and each of the parallel processing unit modules and / or memory 1612 modules can be packaged devices. In one embodiment, central processing unit 2006, switch 2004, and parallel processing module 2002 are on a single semiconductor platform.
[0184] In one embodiment, the signaling rate of each NVLink 1616 is 20 to 25 gigabits per second, and each parallel processing unit module includes six NVLink 1616 interfaces (as Figure 20As shown, each parallel processing unit module includes five NVLink 1616 interfaces. Each NVLink 1616 provides a 25 gigabit / second data transfer rate in each direction, with six lanes providing 300 gigabit / second. The NVLink 1616 can be used exclusively for communications between the central processing unit 2006 and the parallel processing system module 2002, or some combination of PPU-to-PPU communications, PPU-to-PPU, and PPU-to-CPU. Figure 20
[0185] In one embodiment, the NVLink 1616 allows direct load / store / atomic operations from the central processing unit 2006 to the memory 1612 of each parallel processing unit module. In one embodiment, the NVLink 1616 supports coherency operations, allowing data read from the memory 1612 module to be stored in the cache hierarchy of the central processing unit 2006, reducing cache access latency for the central processing unit 2006. In one embodiment, the NVLink 1616 includes support for address translation services (ATS), allowing the parallel processing unit module to directly access page tables within the central processing unit 2006. At least one NVLink 1616 can also be configured to operate in a low power mode.
[0186] Figure 21 An exemplary system 2100 is shown in which various previously described embodiments of various architectures and / or functionality can be implemented. As shown, a system 2100 is provided that includes at least one central processing unit 2006 connected to an interconnect bus 2110. The interconnect bus 2110 can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol(s). The exemplary processing system 2100 also includes main memory 2102. Control logic (software) and data are stored in the main memory 2102, which can take the form of random access memory (RAM).
[0187] The exemplary processing system 2100 also includes input device(s) 2108, parallel processing system module 2002, and display device 2106, such as a conventional CRT (cathode ray tube), LCD (liquid crystal display), LED (light-emitting diode), plasma display, or the like. User input can be received from input device(s) 2108, such as a keyboard, mouse, touchpad, microphone, or the like. Each of the foregoing modules and / or devices can even be located on a single semiconductor platform, as shown in the exemplary processing system 2100. Alternatively, various modules can be located on different semiconductor platforms, which can be arranged to be in communication with one another over a communication medium.
[0188] Furthermore, exemplary processing system 2100 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, a cable network, etc.) through network interface 2104 for communication purposes.
[0189] Exemplary processing system 2100 can also include secondary storage (not shown). Secondary storage 610 includes, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a magnetic tape drive, a compact disk drive, digital versatile disk (DVD) drive, recording device, universal serial bus (USB) flash drive. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner.
[0190] Computer programs, or computer control logic algorithms, can be stored in main memory 2102 and / or secondary storage. Such computer programs, when executed, enable the exemplary processing system 2100 to perform various functions. The main memory 2102, storage and / or any other storage are possible examples of computer-readable media.
[0191] The architectures and / or functionalities of the various preceding figures can be implemented in the context of a general computer system, a circuit board system, a game console system dedicated for entertainment purposes, a special purpose system, and / or any other desired system. For example, exemplary processing system 2100 can take the form of a desktop computer, laptop computer, tablet computer, server computer, super computer, smart telephone (e.g., wireless, hand held device), personal digital assistant (PDA), digital camera, vehicle, head mounted display, hand held electronic device, mobile telephone device, television, workstation, game console, embedded system, and / or any other type of logic.
[0192] While various embodiments have been described above, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of an embodiment should not be limited by any of the above described exemplary embodiments, but should instead be defined in accordance with the following claims and their equivalents.
[0193] Graphics Processing Pipeline
[0194] Figure 22 is a graphics processing pipeline according to one embodiment Figure 16FIG. 22 illustrates a conceptual diagram of a graphics processing pipeline 2200 implemented by the parallel processing unit 2008b. In one embodiment, the parallel processing unit 2008b includes a graphics processing unit (GPU). The parallel processing unit 2008b is configured to receive commands defining a shading program to be used for processing graphics data. The graphics data can be defined as a set of primitives, such as points, lines, triangles, quads, triangle strips, etc. Typically, a primitive includes data specifying a number of vertices (e.g., in a model space coordinate system) and attributes associated with each vertex of the primitive. The parallel processing unit 2008b can be configured to process the primitives to generate a frame buffer (e.g., pixel data for each of the pixels in a display).
[0195] An application writes model data (e.g., a collection of vertices and attributes) for a scene to memory, such as system memory or memory 1612. The model data defines each of the objects that can be visible on a display. The application then makes an API call to a driver kernel, which requests the model data to be rendered and displayed. The driver kernel reads the model data and writes commands to at least one stream to perform operations to process the model data. These commands can reference different shading programs to be implemented on the streaming multi-processor 1900 modules of the parallel processing unit 2008b, including one or more of vertex shading, hull shading, domain shading, geometry shading, and pixel shading. For example, at least one of the streaming multi-processor 1900 modules can be configured to execute a vertex shading program that processes a number of vertices defined by the model data. In one embodiment, different streaming multi-processor 1900 modules can be configured to execute different shading programs at the same time. For example, a first subset of the streaming multi-processor 1900 modules can be configured to execute a vertex shading program while a second subset of the streaming multi-processor 1900 modules can be configured to execute a pixel shading program. The first subset of the streaming multi-processor 1900 modules processes the vertex data to produce processed vertex data and writes the processed vertex data to the L2 cache 1804 and / or memory 1612. After the processed vertex data is rasterized (e.g., converted from three-dimensional data into two-dimensional data in screen space) to produce fragment data, the second subset of the streaming multi-processor 1900 modules executes the pixel shading to produce processed fragment data, which is then blended with other processed fragment data and written to a frame buffer in memory 1612. The vertex shading program and the pixel shading program can be executed at the same time, processing different data from the same scene in a pipelined fashion until all of the model data for the scene has been rendered to the frame buffer. The contents of the frame buffer are then transmitted to a display controller for display on a display device.
[0196] The graphics processing pipeline 2200 is an abstract flowchart of processing steps implemented to generate 2D computer-generated images from 3D geometric data. It is well known that pipeline architectures can more efficiently perform long-latency operations by dividing operations into multiple stages, where the output of each stage is coupled to the input of the next successive stage. Therefore, the graphics processing pipeline 2200 receives input data 601 passed from one stage of the graphics processing pipeline 2200 to the next stage to generate output data 2204. In one embodiment, the graphics processing pipeline 2200 may represent a process... The graphics processing pipeline is defined by the API. Alternatively, the graphics processing pipeline 2200 can be implemented within the context of the functionality and architecture of the previous figures and / or one or more subsequent figures.
[0197] like Figure 22 As shown, the graphics processing pipeline 2200 includes a pipeline architecture comprising multiple stages. These stages include, but are not limited to, a data assembly stage 2206, a vertex shading stage 2208, a primitive assembly stage 2210, a geometry shading stage 2212, a viewport SCC stage 2214, a rasterization stage 2216, a fragment shading stage 2218, and a raster operation stage 2220. In one embodiment, input data 2202 includes commands that configure processing units to implement the stages of the graphics processing pipeline 2200 and configure geometric primitives (e.g., points, lines, triangles, quadrilaterals, triangular strips, or sectors, etc.) to be processed by these stages. Output data 2204 may include pixel data (i.e., color data), which is copied to a frame buffer or other type of surface data structure in memory.
[0198] The data assembly stage 2206 receives input data 2202, which specifies vertex data for higher-order surfaces, primitives, etc. The data assembly stage 2206 collects vertex data from temporary storage or a queue, such as by receiving a command from the host processor including a pointer to a buffer in memory and reading vertex data from that buffer. The vertex data is then passed to the vertex shading stage 2208 for processing.
[0199] The vertex shading 2208 stage processes vertex data by performing a set of operations (e.g., a vertex shader or program) on each of the vertices. A vertex can be specified, for example, as a 4-coordinate vector (e.g., <x, y, z, w>) associated with one or more vertex attributes (e.g., color, texture coordinates, surface normal, etc.). The vertex shading 2208 stage can manipulate individual vertex attributes, such as position, color, texture coordinates, etc. In other words, the vertex shading 2208 stage performs operations on vertex coordinates or other vertex attributes associated with a vertex. These operations typically include lighting operations (e.g., modifying a color attribute of a vertex) and transformation operations (e.g., modifying a coordinate space of a vertex). For example, a vertex can be specified using coordinates in an object coordinate space, which are transformed by multiplying the coordinates by a matrix that converts the coordinates from the object coordinate space to a world space or a normalized-device-coordinate (NDC) space. The vertex shading 2208 stage generates transformed vertex data that is passed to the primitive assembly 2210 stage.
[0200] The primitive assembly 2210 stage collects vertices output by the vertex shading 2208 stage and groups the vertices into geometric primitives for processing by the geometry shading 2212 stage. For example, the primitive assembly 2210 stage can be configured to group every three consecutive vertices into a geometric primitive (e.g., a triangle) for passing to the geometry shading 2212 stage. In some embodiments, particular vertices can be reused for consecutive geometric primitives (e.g., two consecutive triangles in a triangle strip can share two vertices). The primitive assembly 2210 stage passes geometric primitives (e.g., a set of associated vertices) to the geometry shading 2212 stage.
[0201] The geometry shading 2212 stage processes geometric primitives by performing a set of operations (e.g., a geometry shader or program) on the geometric primitives. Tessellation operations can generate one or more geometric primitives from each geometric primitive. In other words, the geometry shading 2212 stage can tessellate each geometric primitive into a finer mesh of two or more geometric primitives for processing by the remainder of the graphics processing pipeline 2200. The geometry shading 2212 stage passes geometric primitives to the viewport SCC 2214 stage.
[0202] In one embodiment, graphics processing pipeline 2200 can perform processing operations in sequence at stream multi-processor and vertex shading 2208 stage, primitive assembly 2210 stage, geometry shading 2212 stage, fragment shading 2218 stage, and / or hardware / software internal operations associated therewith. Once the sequential processing operations are complete, in one embodiment, viewport SCC 2214 stage can utilize the data. In one embodiment, primitive data processed by one or more of the stages in graphics processing pipeline 2200 can be written into a cache (e.g., LI cache, vertex cache, etc.). In this case, in one embodiment, viewport SCC 2214 stage can access the data in the cache. In one embodiment, viewport SCC 2214 stage and rasterization 2216 stage are implemented as fixed function circuitry.
[0203] Viewport SCC 2214 stage performs viewport scaling, culling, and clipping of geometric primitives. Each surface being rendered is associated with an abstract camera position. The camera position represents the position of a viewer who is looking at the scene and defines a view frustum that encloses the objects of the scene. The view frustum can include a viewing plane, a back plane, and four clipping planes. Any geometric primitive that is completely outside of the view frustum can be culled (e.g., discarded) because it will not contribute to the final rendered scene. Any geometric primitive that is partially inside the view frustum and partially outside the view frustum can be clipped (e.g., converted to new geometric primitives that are enclosed within the view frustum). In addition, each geometric primitive can be scaled based on the depth of the view frustum. All potentially visible geometric primitives are then passed to rasterization 2216 stage.
[0204] Rasterization 2216 stage converts 3D geometric primitives into 2D fragments (e.g., that can be used for display, etc.). Rasterization 2216 stage can be configured to set up a set of plane equations with the vertices of the geometric primitive from which various attributes can be interpolated. Rasterization 2216 stage can also compute a coverage mask for a plurality of pixels that indicates whether one or more sample locations of the pixel intercept the geometric primitive. In one embodiment, a z-test can also be performed to determine whether the geometric primitive is occluded by other geometric primitives that have already been rasterized. Rasterization 2216 stage generates fragment data (e.g., interpolated vertex attributes associated with particular sample locations of each covered pixel) that is passed to fragment shading 2218 stage.
[0205] The fragment shading 2218 stage processes the fragment data by performing a set of operations (e.g., a fragment shader or program) on each of the fragments. The fragment shading 2218 stage can generate pixel data (e.g., color values) for the fragments, such as by performing lighting operations or sampling texture maps using the interpolated texture coordinates for the fragments. The fragment shading 2218 stage generates pixel data, which is sent to the raster operations 2220 stage.
[0206] The raster operations 2220 stage can perform various operations on the pixel data, such as performing alpha tests, stencil tests, and blending the pixel data with other pixel data corresponding to other fragments associated with the pixel. When the raster operations 2220 stage has completed processing of the pixel data (e.g., output data 2204), the pixel data can be written to a render target, such as a frame buffer, color buffer, etc.
[0207] It should be appreciated that at least one additional stage can be included in the graphics processing pipeline 2200 in addition to or instead of at least one of the stages described above. Various implementations of an abstract graphics processing pipeline can implement different stages. Further, in some embodiments, at least one of the stages described above can be excluded from the graphics processing pipeline (such as the geometry shading 2212 stage). Other types of graphics processing pipelines are contemplated as being within the scope of the present disclosure. Further, any of the stages of the graphics processing pipeline 2200 can be implemented by at least one dedicated hardware unit within a graphics processor, such as the parallel processing unit 2008b. Other stages of the graphics processing pipeline 2200 can be implemented by programmable hardware units, such as the streaming multiprocessors 1900 of the parallel processing unit 2008b.
[0208] Graphics processing pipeline 2200 can be implemented via an application program executed by a host processor, such as a CPU. In one embodiment, a device driver can implement an application programming interface (API) that defines various functions that can be utilized by an application program to generate graphics data for display. The device driver is a software program that includes a plurality of instructions that control the operation of parallel processing unit 2008b. The API provides an abstraction for programmers that allows programmers to utilize specialized graphics hardware, such as parallel processing unit 2008b, to generate graphics data without requiring the programmer to utilize the specific instruction set of parallel processing unit 2008b. An application program can include API calls that are routed to the device driver of parallel processing unit 2008b. The device driver interprets the API calls and performs various operations in response to the API calls. In some cases, the device driver can perform operations by executing instructions on the CPU. In other cases, the device driver can perform operations at least partially by initiating operations on parallel processing unit 2008b utilizing an input / output interface between the CPU and parallel processing unit 2008b. In one embodiment, the device driver is configured to utilize the hardware of parallel processing unit 2008b to implement graphics processing pipeline 2200.
[0209] Various programs can be executed within parallel processing unit 2008b in order to implement the various stages of graphics processing pipeline 2200. For example, a device driver can launch a kernel on parallel processing unit 2008b to perform the vertex shading 2208 stage on one streaming multiprocessor 1900 (or multiple streaming multiprocessor 1900 modules). The device driver (or an initial kernel executed by PPU 400) can also launch further kernels on PPU 400 to perform other stages of graphics processing pipeline 2200, such as the geometry shading 2212 stage and the fragment shading 2218 stage. In addition, some of the stages of graphics processing pipeline 2200 can be implemented on fixed function hardware within PPU 400, such as a rasterizer or a data assembler. It will be appreciated that the results of a stage, before being processed by a subsequent stage on a streaming multiprocessor 1900, can be processed by one or more intermediate fixed-function hardware units.
[0210] Drawing element list
[0211] 100 system
[0212] 102 human
[0213] 104 cloud computer system
[0214] 106 internet of things device
[0215] 108 local area network
[0216] 110 microphone
[0217] 112 wide area network
[0218] 114 neural network
[0219] 116 Softmax layer
[0220] 118 digital device
[0221] 202 deep learning system
[0222] 204 computing system
[0223] 206 vehicle
[0224] 208 robot
[0225] 302 transform neural network
[0226] 304 input sequence
[0227] 306 encoder
[0228] 308 encoder stack
[0229] 310 encoder
[0230] 312 encoder
[0231] 314 decoder stack
[0232] 316 decoder
[0233] 318 decoder
[0234] 320 decoder
[0235] 322 output sequence
[0236] 402 encoder
[0237] 404 self-attention layer
[0238] 406 feedforward neural network
[0239] 502 decoder
[0240] 504 encoder-decoder attention layer
[0241] 506 self-attention layer
[0242] 508 feedforward neural network
[0243] 602 attention layer
[0244] 604 matrix multiplication
[0245] 606 query vector
[0246] 608 key vector
[0247] 610 value vector
[0248] 612 score vector
[0249] 614 dot product
[0250] 616 Softmax
[0251] 618 multiply
[0252] 620 vector sum
[0253] 700 Softmax algorithm
[0254] 800 Softmax computation logic
[0255] 802 vector integer maximum unit
[0256] 804 power of two computation unit
[0257] 806 reduction unit
[0258] 808 maximum selector
[0259] 810 maximum comparator
[0260] 812 memory buffer
[0261] 814 power and buffer
[0262] 816 power and selector
[0263] 818 right shift
[0264] 820 adder
[0265] 822 linear segment computation unit
[0266] 824 fixed point fractional divider
[0267] 826 left shift
[0268] 828 unnormalized Softmax unit
[0269] 830 normalization unit
[0270] 832 look-up table
[0271] 834 maximum comparator
[0272] 836 vector element adder
[0273] 900 distributed computing system
[0274] 902 router interface
[0275] 904 processing element
[0276] 906 controller
[0277] 908 normalization unit
[0278] 910 global global memory buffer
[0279] 912 memory
[0280] 1002 global buffer
[0281] 1004 controller
[0282] 1006 dynamic random access memory
[0283] 1008 network-on-a-chip router
[0284] 1010 FPGA
[0285] 1012 multi-chip package
[0286] 1014 processing element
[0287] 1016 network package router
[0288] 1018 die
[0289] 1020 host
[0290] 1100 neural network processor
[0291] 1200 processing element
[0292] 1202 vector multiply-accumulate unit
[0293] 1204 weight buffer
[0294] 1206 activation buffer
[0295] 1208 router
[0296] 1210 accumulation memory buffer
[0297] 1212 post-processor
[0298] 1214 controller
[0299] 1300 processing element
[0300] 1302 input weight
[0301] 1304 input activation
[0302] 1306 partial sum
[0303] 1308 cross-processing element reduction
[0304] 1310 weight memory buffer manager
[0305] 1312 activation memory buffer manager
[0306] 1314 address generator
[0307] 1316 accumulation memory buffer manager
[0308] 1318 address generator
[0309] 1320 weight collector
[0310] 1322 accumulation collector
[0311] 1324 vector multiply accumulate unit
[0312] 1326 output activations
[0313] 1328 input collector
[0314] 1502 global memory buffer
[0315] 1504 controller
[0316] 1506 activation address generator
[0317] 1508 data path
[0318] 1510 destination address generator
[0319] 1512 register file
[0320] 1514 output activations
[0321] 1516 partial sums
[0322] 1518 input activations
[0323] 1520 arbitration repository
[0324] 1522 global processing unit
[0325] 1602 input / output unit
[0326] 1604 front end unit
[0327] 1606 hub
[0328] 1608 scheduler unit
[0329] 1610 work distribution unit
[0330] 1612 memory
[0331] 1614 crossbar
[0332] 1616 NVLink
[0333] 1618 interconnect
[0334] 1700 general processing cluster
[0335] 1702 pipeline manager
[0336] 1704 pre-raster operations unit
[0337] 1706 data processing cluster
[0338] 1708 raster engine
[0339] 1710 M-pipe controller
[0340] 1712 primitive engine
[0341] 1714 work distribution crossbar
[0342] 1716 memory management unit
[0343] 1800 memory partition unit
[0344] 1802 raster operations unit
[0345] 1804 secondary cache
[0346] 1806 memory interface
[0347] 1900 streaming multiprocessor
[0348] 1902 instruction cache
[0349] 1904 scheduler unit
[0350] 1906 dispatch
[0351] 1908 register file
[0352] 1910 core
[0353] 1912 special function unit
[0354] 1914 load / store unit
[0355] 1916 interconnection network
[0356] 1918 shared memory / L1 cache
[0357] 2000 processing system
[0358] 2002 parallel processing module
[0359] 2004 switch
[0360] 2006 central processing unit
[0361] 2008a parallel processing unit
[0362] 2008b parallel processing unit
[0363] 2008c parallel processing unit
[0364] 2008d parallel processing unit
[0365] 2100 exemplary processing system
[0366] 2102 main memory
[0367] 2104 network interface
[0368] 2106 display device
[0369] 2108 input device
[0370] 2110 communication bus
[0371] 2200 graphics processing pipeline
[0372] 2202 input data
[0373] 2204 output data
[0374] 2206 data assembly
[0375] 2208 vertex shading
[0376] 2210 primitive assembly
[0377] 2212 geometry shading
[0378] 2214 viewport SCC
[0379] 2216 rasterization
[0380] 2218 fragment shading
[0381] 2220 raster operations
[0382] The various functional operations described herein can be implemented in logic that is embodied in tangible operation instructions. The operation instructions can describe actions to be taken by hardware of the machine. Use of the terms "operation" or "operations" can refer to one or more operations, and the terms "operation" or "operations" can be used interchangeably with the terms "step" or "steps."
[0383] In the present disclosure, different entities (which can variously be referred to as “units,” “circuits,” other components, etc.) can be described or claimed as “configured” to perform one or more tasks or operations. This formulation—[entity] configured to [perform one or more tasks]—is used herein to convey that a structure (i.e., hardware) has been made or manufactured to perform the task or tasks. As used in this formulation, “configured” is not construed as a
[0384] The term “configured to” does not mean “configurable to.” For example, an FPGA that is not programmed to perform a certain function would not be said to be “configured to” perform that function, even if it could be programmed to do so.
[0385] The recitation of a structure being “configured to” perform one or more tasks in the appended claims, means that the structure is arranged to perform the task(s). Such formulations may
[0386] As used herein, the term “based on” is used to describe one or more factors that affect a determination. This term does not foreclose the possibility that additional factors may affect a determination. That is, a determination may be solely based on specified factors or based on specified factors and other, unspecified factors. Consider the phrase “determine A based on B.” This phrase specifies that B is a factor that affects the determination of A. This phrase does not foreclose the possibility that the determination of A may also be based on some other factors, such as C. This phrase is also intended to cover an embodiment in which A is determined based only on B. As used herein, the phrase “based on” is synonymous with the phrase “based at least in part on.”
[0387] As used herein, the phrase “in response to” describes one or more factors that trigger an effect. The phrase does not exclude the possibility that other factors can influence or otherwise trigger the effect. That is, the effect can be responsive only to those factors, or can be responsive to the specified factors as well as other unspecified factors. Consider the phrase “perform A in response to B.” This phrase specifies that B is a factor that triggers performance of A. The phrase does not exclude the possibility that performance of A can also be responsive to some other factor, e.g., C. The phrase is also intended to cover instances where A is performed only in response to B.
[0388] As used herein, the terms “first,” “second,” and the like, are used as labels for nouns that they precede, and do not expressly imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless otherwise specified. For example, in a register file having eight registers, the terms “first register” and “second register” can be used to refer to any two of the eight registers, not, for example, specifically logical registers 0 and 1.
[0389] The term “or” when used in a claim, is used as an inclusive or in the sense of the alternative being present unless otherwise indicated (e.g., “x, y, or z” means any of x, y, and z individually, and x, y, and z combined). In addition, the term “one or more” as used herein means that “at least one” but can include more than one (e.g., one, two, three, four, or more; one through five, one through four, one, two, or three, etc.). When used in a claim, the term “comprises” or “comprising” is used in the sense of the term “includes” and / or “including” and not the sense of “consists” and / or “consisting.” Further, the term “comprises” or “comprising” is open-ended, meaning it includes the stated elements, but not excluding additional elements.
[0390] As used herein, recitations of “and / or” in reference to two or more elements should be interpreted as meaning either one element or a combination of elements. For example, “element A, element B, and / or element C” can include just element A, just element B, just element C, element A and element B, element A and element C, element B and element C, or elements A, B, and C. Further, “at least one of element A or element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B. Further still, “at least one of element A and element B” can include at least one of element A, at least one of element B, or at least one of element A and at least one of element B.
[0391] The subject matter of the present disclosure is described herein with specificity to meet statutory requirements. However, the description itself is not intended to limit the scope of this disclosure. Rather, the claimed subject matter might also be embodied in other ways to include different steps or combinations of steps similar to the ones described in this document, in conjunction with other present or future technologies. Moreover, although the terms “step” and / or “block” might be used herein to connote different elements of methods employed, the terms should not be interpreted as implying any particular order among or between various steps herein disclosed unless and except when the order of individual steps is explicitly described.
[0392] Having thus described in detail preferred embodiments of the present disclosure, it will be apparent to those skilled in the art that modifications and variations are possible without departing from the scope of the claimed disclosure. The scope of the disclosed subject matter is not to be limited to the described embodiments but is instead to be defined by the following claims.
Claims
1. A system comprising: one or more processors configured to: determine an unnormalized Softmax vector from an input vector by: converting elements of the input vector to powers of two; determining an integer maximum of elements of the input vector; determining exponents of the powers of two from elements of the input vector and the integer maximum; convert the unnormalized Softmax vector to a normalized Softmax vector; and perform image recognition and classification processing based at least on the normalized Softmax vector.
2. The system of claim 1, the one or more processors comprising a plurality of processing elements, the processors further configured to configure the plurality of processing elements to compute the unnormalized Softmax vector in a distributed computation.
3. The system of claim 2, the one or more processors further configured to: configure at least some of the plurality of processing elements to compute local integer maximums of elements of their respective input vectors; and configure at least some of the plurality of processing elements to perform a cross-processing element reduction operation of the local integer maximums to a global integer maximum.
4. The system of claim 2, the one or more processors further configured to: configure the one or more processors to compute a sum of powers of two.
5. The system of claim 4, the one or more processors further configured to: configure at least some of the plurality of processing elements to compute local sums of powers of two of their respective input vectors; and configure at least some of the plurality of processing elements to perform a cross-processing element reduction operation of the local sums of powers of two to a global sum of powers of two.
6. The system of claim 2, further comprising: a central normalization unit to convert the unnormalized Softmax vector to the normalized Softmax vector.
7. The system of claim 1, the one or more processors further configured to: raise elements of the input vector to powers of two and compute the integer maximum in a single execution loop.
8. The system of claim 7, the one or more processors further configured to: compute a sum of powers of two in the single execution loop.
9. A processor comprising an artificial neural network, the artificial neural network comprising: one or more feedforward layers; and one or more Softmax layers coupled to the one or more feedforward layers; at least one Softmax layer configured to compute an unnormalized Softmax vector from an input vector to yield a normalized Softmax vector by: converting elements of the input vector to powers of two; determining an integer maximum of elements of the input vector; and determining exponents of the powers of two from elements of the input vector and the integer maximum; the processor configured to perform image recognition and classification processing based at least on the normalized Softmax vector. 10. The processor of claim 9, the at least one Softmax layer further configured to: convert the unnormalized Softmax vector to a normalized Softmax vector.
11. The processor of claim 10, the at least one Softmax layer further configured to: convert the unnormalized Softmax vector to a normalized Softmax vector using shifts and reciprocal operations without performing multiplication operations.
12. The processor of claim 9, the at least one Softmax layer further configured to: compute the unnormalized Softmax vector using multiple processing elements in a distributed computation.
13. The processor of claim 12, the at least one Softmax layer further configured to: compute local integer maxima of elements of their respective input vectors using at least some of the multiple processing elements; and perform a cross-processing element reduction operation of the local integer maxima to a global integer maxima using at least some of the multiple processing elements.
14. The processor of claim 12, the at least one Softmax layer further configured to compute a sum of powers of two.
15. The processor of claim 14, the at least one Softmax layer further configured to: compute local sums of powers of two of their respective input vectors using at least some of the multiple processing elements; and perform a cross-processing element reduction operation of the local sums of powers of two to a global sum of powers of two using at least some of the multiple processing elements.
16. The processor of claim 12, the at least one Softmax layer further configured to: convert the unnormalized Softmax vector to a normalized Softmax vector.
17. The processor of claim 9, the at least one Softmax layer further configured to: raise elements of the input vector to a power of two and compute the integer maxima in a single execution loop.
18. The processor of claim 17, the at least one Softmax layer further configured to: compute a sum of powers of two in the single execution loop.
19. A processor comprising a transformed artificial neural network, the transformed artificial neural network comprising: a self-attention layer; and an encoder-decoder attention layer; each of the self-attention layer and the encoder-decoder attention layer comprising a Softmax layer configured to generate an unnormalized Softmax vector from an input vector to yield a normalized Softmax vector by: converting elements of the input vector to powers of two; determining an integer maximum of elements of the input vector; and determining exponents of the powers of two from elements of the input vector and the integer maximum; the processor configured to perform image recognition and classification processing based at least on the normalized Softmax vector.
20. The processor of claim 19, each Softmax layer is further configured to: compute a sum of powers of 2 in a single execution cycle; and convert the unnormalized Softmax vector to a normalized Softmax vector using shift and reciprocal operations without performing multiplication operations.
21. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to perform image recognition and classification processing by: In the Softmax computation, 2 is expressed by the formula x a vector of value elements, where each x is an element of an input vector of the Softmax computation; computing an integer maximum value of x in the input vector; and determining 2 x exponent from elements of the input vector and the integer maximum value.
22. The non-transitory computer-readable storage medium of claim 21, the computer- readable storage medium further comprising instructions that, when executed by the computer, cause the computer to: from the 2 x The vector of elements and the integer maximum generate an unnormalized Softmax vector.
23. The non-transitory computer-readable storage medium of claim 22, the computer- readable storage medium further comprising instructions that, when executed by the computer, cause the computer to: normalize the unnormalized Softmax vector.
24. The non-transitory computer-readable storage medium of claim 21, the computer- readable storage medium further comprising instructions that, when executed by the computer, cause the computer to: In the Softmax computation, the 2 x vector of elements, the integer maximum value of the x in the input vector is computed and the 2 x sum of the vector of elements is computed in a single execution loop.
25. A method comprising: executing a first machine instruction to represent in a Softmax computation 2 x a vector of value elements, where each x is an element of an input vector of the Softmax computation; executing a second instruction to determine an integer maximum value of x in the input vector and to determine an exponent of 2 from an element of the input vector and the integer maximum value; and x an exponent of 2 from an element of the input vector and the integer maximum value. performing image recognition and classification processing based at least on a result of executing a first machine instruction and executing a second instruction.
Citation Information
Patent Citations
Execution unit in processor
US20200293315A1
Machine Learning Computer
US20220051095A1
Efficient softmax computation
US20220067513A1