Hardware architecture for parallel computations

The hardware architecture addresses memory bandwidth limitations in ANNs by prioritizing data fetching and leveraging sparsity, resulting in improved performance and energy efficiency for AI computations.

WO2025123143A1PCT designated stage expired Publication Date: 2025-06-19HEPZIBAH INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/CA2024/051660
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The performance of artificial neural networks (ANNs) and other forms of artificial intelligence is limited by memory bandwidth, particularly due to the need to swap large volumes of data in and out of system memory during complex computations, leading to delays and increased energy consumption.

Method used

A hardware architecture that prioritizes the fetching of coefficients and data from memory for single instruction, multiple data (SIMD) operations, leveraging sparsity to reduce the volume of data fetched and implementing a prioritized fetch scheme to improve just-in-time delivery of required values while reducing memory interface idle time.

Benefits of technology

The proposed hardware architecture enhances the performance of ANNs and other AI applications by reducing memory bandwidth limitations, improving computation efficiency, and minimizing energy consumption through optimized data fetching and sparsity utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CA2024051660_19062025_PF_FP_ABST
    Figure CA2024051660_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A computing system includes arrays of processing tiles each comprising a plurality of processing elements for executing single instruction, multiple data (SIMD) operations. The processing tiles are interconnected by distinct control and payload networks to cooperatively execute operations on tensors and other arrays. The system and / or tiles are configured to issue requests to pre-fetch data from system memory for use in a later operation, and to issue prioritized read requests for data for use in a next operation. Prioritized read requests may be made for coefficients or other operands corresponding to sparse data entries generated by an earlier operation.
Need to check novelty before this filing date? Find Prior Art

Description

HARDWARE ARCHITECTURE FOR PARALLEL COMPUTATIONSTechnical Field[oooi] The present disclosure relates to a hardware architecture and related techniques for performing complex computations, in particular parallel operations on input data sets exhibiting sparsity.Technical Background

[0002] Artificial neural networks, and in particular generative large language models, have proven to be quite valuable. Implementing such systems is expensive, however, with costs per query that are orders of magnitude more than conventional search algorithms. Neural networks generally involve massive arrays of data that must be swapped into processor memory for each stage of a calculation, resulting in delays due to complex memory read processes and limited memory bandwidth.Brief Description of the Drawings

[0003] In drawings which illustrate by way of example only embodiments of the present invention,

[0004] FIG. 1 is a schematic of a portion of a neural network depicting a multi-stage computation.

[0005] FIG. 2 is a block diagram of an example computing system including tensor processing devices and processing tiles.

[0006] FIG. 3 is a block diagram of an example tensor processing device comprising a number of processing tiles.

[0007] FIG. 4 is a block diagram of an example processing tile comprising a number of hardware processing elements.

[0008] FIG. 5 is a block diagram of an example processing element.

[0009] FIG. 6 is a timing diagram depicting an order of execution of memory instructions.

[0010] FIG. 7 is a block diagram of a further example processing tile including sparsity converter units.Detailed Description[ooii] The present disclosure provides a hardware architecture for use in a computing system executing parallel computing tasks, in particular single instruction, multiple data (SIMD) operations as are typically encountered in artificial neural networks (ANNs) and other forms of artificial intelligence, and may be encountered in other tasks such as simulations, computer vision, linear algebra, and signal processing. As set out below, fetching coefficients or other data from memory for SIMD operations executed by arrays of processing elements is prioritized to improve just-in-time delivery of required values while reducing memory interface idle time. The present disclosure further leverages sparse output from various stages of a computation to assist in prioritization and to reduce the volume of data fetched from memory.

[0012] The problem of memory bandwidth can be illustrated with reference to FIG. 1 , which is a schematic of a well-known artificial neural network (ANN) structure. FIG. 1 depicts an example segment of generative large language model (G-LFM) architecture, specifically a transformer sub-block, including the self-attention mechanism are query, key, and valuematrices, Tis time complexity, and dkis the dimensionality of the keys. The result of the attention mechanism is a weighted sum of the value vectors, in which weights are determined by the softmax of the dot product of Q and K. As those skilled in the art recognize, the time complexity increases quadratically with the length of the input sequence to the mechanism, with the result that memory requirements for executing the self-attention mechanism increases quadratically with the length of the input sequence. The sub-block also includes feedforward networks (multilayer perceptrons MFP In and MFP Out), which likewise operate on matrices.

[0013] Due to the size of the matrices employed in ANNs, it is not economical to store all coefficients and state in processor memory; it is necessary to swap data in and out of system memory. Thus, performance of a model is limited by the memory bandwidth available forswapping coefficients and state values (e.g., outputs of K and F) used in these computations. For example, to produce a single token, currently a G-LLM may need to load more than 170 x 109coefficients, and matrices ranging from 10 to 100 GB in size. The available memory bandwidth may be effectively reduced due to queued requests, which can increase the memory idle time. Moreover, swapping large volumes of data increases both energy consumption and the time required to complete computations.

[0014] Various measures can be used to reduce energy and / or performance costs, such as batching queries, grouped queries, increasing available memory (e.g., DRAM, dynamic random access memory), using high-bandwidth memory, stacking DRAM vertically with processors to reduce the distance travelled by data, reducing model size, or using lower- precision arithmetic. While these measures can provide energy or time savings, some of these solutions are either more expensive to manufacture, or potentially reduce the accuracy of the model.

[0015] It has been recognized that complexity of an ANN may be reduced by introducing sparsity, by removing certain parameters or activations. In a sparse model, for example, a portion of its weights may be set to zero, thereby avoiding a number of multiplication operations and reducing memory requirements. However, there is a potential trade-off in accuracy of the sparse model. As another example, if an activation function results in a zero, it is possible to skip accumulation of values; however, this savings in computation time is not matched by a savings in memory bandwidth consumption, since in prior art systems, the coefficients corresponding to that activation are still fetched from memory. As noted above, memory bandwidth is a limiting factor in ANN performance.

[0016] Accordingly, the present disclosure provides a hardware architecture for use in a computing system executing parallel computing tasks, in particular single instruction, multiple data (SIMD) operations as are typically encountered in ANNs and other forms of artificial intelligence, and may be encountered in other tasks such as simulations, computer vision, linear algebra, and signal processing. As set out below, fetching coefficients or other data from memory for SIMD operations executed by arrays of processing elements is prioritized to improve just-in-time delivery of required values while reducing memory interface idle time. Priority of fetch requests may be determined during execution of acomplex computation task according to how quickly the data required for a given stage of the computation can be identified in advance. Once priority is determined, the priority ranking of a fetch request may be specified a number of ways: for example, different levels of priority may be assigned to discrete queues in a memory controller, or alternatively, certain memory address ranges may be associated with higher and lower priority requests.

[0017] Priority may be dependent on the context of the computation carried out by the computing device. For example, in the case of ANNs, which frequently employ nonlinear activation functions, or processes involving other operations that may output sparse vectors or matrices (arrays with a relatively significant number of zero or effectively-zero values) to be used as input in a subsequent SIMD operation such as a multiply-accumulate (MAC) operation, those subsequent individual operations taking a zero or effectively-zero value as an operand will produce a zero value as output, or possibly a sufficiently insignificant number such that the output can be treated as a zero value (i.e., effectively zero).Consequently, the sparsity of the input array may be exploited to reduce the number or volume of memory fetches required to retrieve the additional operands required for the SIMD operation, by retrieving only those values that will be combined with a non-zero operand from the previous operation. Since those specific values can only be identified after completion of the prior operation that produced the sparse array, the fetches of those specific values are assigned a higher priority than other fetches of data not required until a later step of the computation. Thus, it can be appreciated that the present solutions are particularly useful for deep learning and similar applications, but are also useful where computations involve sparse arrays or would benefit from selective and prioritized retrieval of array data from memory.

[0018] Arrays computed by the processing elements can be converted to a sparse representation — effectively, a compressed representation — for use in identifying the values to be fetched from memory. The conversion may be carried out by a controller associated with the processing elements executing the SIMD operations, or by a special-purpose converter unit connected to sets of processing elements (e.g., corresponding to rows or columns of arrays).

[0019] As will be seen below, these solutions may be implemented in a computing system employing arrays of processing elements configured to execute arithmetic operations, which are grouped on processing tiles managed by local controllers. One or more processing tiles, in turn, may be mounted on processing cards; a single computing system may comprise one or more such cards. The size of the processing element array, number of processing tiles per card, and number of cards used in a given computing system may be selected according to the intended purpose and the size of arrays to be processed by the system.

[0020] FIG. 2 is a block diagram of select components of a computing system 10 in which prioritization of fetch requests and the compressed representation may be utilized. A computing system 10 of this type may be constructed on multiple networking server trays, housed a conventional rackmount server chassis, and may include a general-purpose computer server and main processor(s), running a conventional operating system (e.g., Linux) for managing the operations and resources of the system 10, and communicate instructions to the tensor processing devices 100 described below. It will be appreciated by those skilled in the art that description is not intended to be limiting, and the solutions described herein may be implemented on any suitable computing system that is configurable to operate as described, whether or not this device is primarily intended for execution of ANNs or other uses. Further, it will be appreciated that FIG. 1 illustrates only those components of the computing system 10 most relevant to the present disclosure, and therefore omits several conventional components for ease of exposition, such as a motherboard, power supply, user interfaces, and the like. The computing system 10 may also include other connected resources, not shown, such as data storage attached at the motherboard, via a card connection, or over a network connection.

[0021] The computing system 10 includes several tensor processing devices (TPDs) 100 which carry the processing elements used to carry out operations. In this disclosure, “tensor” is used for convenience to denote an array of any dimensionality, including vectors and matrices, and may be considered interchangeable with the term “array”; those skilled in the art will appreciate whether a particular tensor or array in a given computation is a scalar, vector, matrix, or tensor. The tensor processing devices 100 of a given tray may be provided on a standard card format such as Peripheral Component Interconnect Express™(PCIe™) mounted to a motherboard or backplane (not shown), with an edge connector using a standard interface such as PCIe or Compute Express Link™ (CXL™), an open standard interconnect that maintains memory coherency across attached devices. Each tensor processing device 100 is also connected to tensor processing devices 100 on other trays, for example via optical links or a network interface card (NIC) 50, shown in FIG. 1. Further, each tensor processing device 100 may be directly connected to one or more other tensor processing devices 100 on the same tray via relay nodes, shown in FIG. 3. The system further includes a memory bank 30 for storing coefficients or other values for swapping. The memory bank 30 may comprise DRAM and preferably high-bandwidth memory. Other types of memory such as ferroelectric RAM, magnetic RAM, static randomaccess memory, or combinations of different types of memory may be employed. Fetch requests are received from the tensor processing devices 100 via a memory interface 35 from one or more memory controllers 40 from the tensor processing devices 100.

[0022] A tensor processing device 100 is shown in further detail in FIG. 3. Each device 100 comprises an array of processing tiles 200, each tile 200 comprising a plurality of hardware processing elements 300, as shown in FIGS. 4 and 5. These tiles 200 are configured to cooperate to execute a parallel processing scheme, such as SIMD operations as noted above, and are substantially identical. However, in some implementations one tile 200 may be designated as a coordinating controller that manages communications for a plurality of other tiles 200, and monitors the fulfilment of requests for data from the memory bank 30. A coordinating controller tile 200 may manage other tiles 200 on the same tensor processing device 100, or across multiple tensor processing devices. Thus, while the tiles 200 in the computing system 100 may all cooperate to execute the parallel processing scheme, one or more may fulfil different functions.

[0023] The processing tiles 200 on a tensor processing device 100 may be provided as a chiplet or module 110 mounted to the tensor processing device 100, as shown in FIG. 3, with input / output and other components provided separately. Alternatively, the module 110 may be a multi-chip module or system on chip including these other components.

[0024] As shown in FIG. 3, the processing tiles 200 are arranged in a grid, and each is connected to its immediate neighbor tile 200 by distinct payload and link level controlnetworks (indicated by solid and dotted lines, respectively), implemented with pipelined buses. For ease of reference, the rows and columns of the processing tiles 200, and thus the payload and control buses, which run in a single direction, can be conceptualized as running east and west, and north and south; however, it will be understood by those skilled in the art that the arrangement of processing tiles 200 is by no means limited to cardinal directions. The tiles 200 may be arranged in other geometries, and need not be connected to their nearest neighbors. The payload network delivers messages comprising a header containing metadata, and a payload. The payload size may reflect the number of processing elements 300 on a processing tile 200; for example, a 128-byte payload may be employed for tiles 200 comprising an array of 128 processing elements 300.

[0025] The payload and control networks are local to each tensor processing device 100, and may be implemented using any suitable protocol, such as the Bunch of Wires (BoW) specification published by the Open Compute Project Foundation Open Domain-Specific Architecture (ODSA) group. The device 100 is provided with additional input / output components, including, for example, a server interface 120 and card connector 130 to connect the tensor processing device 100 in the computing system 10. The module 110 further includes at least one network relay node 140 which provides an interface for connecting with other modules 110 on other tensor processing devices 100 on the same tray, and converts messages transmitted along a payload or control networks to an external communication protocol, and vice versa. Communication between modules 110 may be effected using any suitable networking protocol, such as a low-latency Ethernet derivative. Thus, each processing tile 200 on each tensor processing device 100 may be considered to be a node in a network of tiles 200 spanning all tensor processing devices 100 in the computing system 10.

[0026] While FIG. 2 depicts the memory bank 30 as being distinct from the tensor processing devices 100, in some implementations each tensor processing device 100 additionally comprises its own memory bank 30, memory interface 35, and controller 40; the memory interface 35 and controller 40 may form part of the module 110. Thus, the prioritized fetch scheme may be implemented on a single tensor processing device 100 without accessing memory located elsewhere in the computing system 100.

[0027] An example processing tile 200 is shown in FIG. 4. The processing tile 200 includes an array of processing elements 300, which are connected to a tile controller 210. The controller 210 may comprise a reduced instruction set computer (e.g., RISC-V) and finite state machines for managing the transmission of messages to and from, and the operations of, individual processing elements 300. As shown in FIG. 4, the processing elements 300 are interconnected at least in rows, but other topologies may be possible.

[0028] Each processing element 300 is substantially identical. As shown in FIG. 5, each processing element 300 includes an arithmetic logic unit (AEU) 310 for executing operations such as addition, multiplication, and the like, and control and payload registers 320 for state information and instructions, and for storing operands for ALU operations, respectively. Payloads are delivered directly to the payload registers 320 of the processing elements. When an operation is to be performed by the processing elements 300, requests for the required operands are transmitted to the memory controllers 40; the set of operands is fetched from the memory bank 30 via the interface 35, and transmitted from the controllers 40 in messages to the processing tiles 200 in a store-and-forward manner, each addressed to the respective processing element 300; and each tile controller 210 passes messages to the appropriate processing element 300.

[0029] As will be appreciated by those skilled in the art, during a complex computation such as the attention sub-block illustrated in FIG. 1 , it is necessary to swap in large sets of values and coefficients to perform tensor operations, resulting in delays due to limited memory bandwidth. Further, in a simple prior art swapping system employing a G-LLM chip, typically coefficients are fetched from DRAM and sent to SRAM on the chip, which carries out the required operations for a first calculation; when the G-LLM chip has completed the first calculation, it then requests the values required for the next calculation. Each fetch request is relayed to the DRAM controller, which then sends a read command to the DRAM, which then must select and latch a row address to read data into a buffer. Consequently, the DRAM interface is idle while the chip performs its operations, and while the DRAM executes the complicated task of reading data. In theory, pre-fetching data may mitigate the delay attributable to the DRAM interface. In practice, however, if the chip memory is already full of data for the first calculation, pre-fetching the next tensor’ s-worthof data is not possible. Similarly, if pre-fetched data for a subsequent calculation is occupying the chip memory, the chip cannot retrieve data needed for an earlier set of operations. Additional memory would be required.

[0030] Some bandwidth delay may be mitigated without increasing memory requirements by configuring the DRAM controller to prioritize certain fetch requests over others, such that pre-fetches of data required for a subsequent calculation are executed only if all fetches needed more urgently have been completed. Additionally, performance can be improved if the volume of data to be fetched can be reduced. In the context of ANNs and other processes that may produce sparse data sets (i.e., containing a significant amount of zeros, a value that is effectively zero, or, in some instances, a negative number) as output in a given stage of a computation, this sparsity can be leveraged to reduce the amount of data that must be fetched for the next stage of the computation, if that next stage involves an operation (typically, multiplication) that, when it takes a zero, an effectively zero amount, or a negative value as an operand, will also yield a zero or a negligible value that may be treated as zero. If it is known that one of the operands of an operation to be executed is a value that will yield a zero or negligible result, there is no need to retrieve the other operand(s) from memory; instead, a zero value may be inserted into the output.

[0031] Returning to the example of FIG. 1, the set of operations represented by the various elements of the attention sub-block can be considered to be a multi-stage computation: one in which several sets of operations on tensors (i.e., arrays) are executed sequentially, with the output of one stage of multistage computation providing input to a subsequent stage. For instance, computation of the dot product of Q and N is a first stage requiring fetching of the values of Q and K, swapping individual operands from O and K into the registers of a corresponding number of processing elements 300, and execution of a MAC operation. The output dot product is then input for a subsequent stage of the multi-stage computation, softmax. As another example, the sequence MLP In - NonLin - MLP Out is a sequence of three stages of the same multi-stage computation, in which the output of the first multilayer perceptron (MLP In) is provided as input to a non-linear activation function (NonLin), and the output of the activation function is then input to the second multilayer perceptron (MLPOut). Moreover, a multi-stage computation may be followed by another multi-stage computation.

[0032] As the order of operations in these multi-stage computations is known, as are some of their operands (e.g., Q, K, V, coefficients used for MLP In and MLP Out), they could be pre-fetched from the memory bank 30 well in advance of the stages of the multi-stage computation where they are required. On the other hand, sparse output resulting from a stage such as NonLin is not predictable, since the zero, effectively-zero, or negative elements of the output cannot be identified until the operations of that stage are completed. Thus, if sparsity is to be leveraged to reduce the amount of data to be fetched for the next stage of the computation, the data cannot be pre-fetched; a request to fetch the reduced set of operands corresponding to the non-zero values of the sparse output can only be transmitted to the memory controllers 40 once the sparse output has been identified. In the prior art system mentioned above, leveraging sparsity in this manner may negate the benefit of pre-fetching as well: there is a risk that the memory bank 30 and interface 35 may be serving a pre-fetch request (e.g., Q, K, V for a subsequent multi-stage computation) at the time the reduced set of operands is urgently needed.

[0033] Accordingly, a prioritized fetch scheme is implemented with the memory controllers 40 to enable the controllers to prioritize requests for “unpredictable” or urgently needed data, such as the reduced operand set described above, while still servicing requests to prefetch data for subsequent stages of the same multi-stage computation, or a subsequent multistage computation. The memory controllers 40 are provided with at least a high-priority and a low-priority queue; received fetch requests are entered into the appropriate queue upon receipt by a memory controller. The controllers 40 are configured to respond to any fetch requests in the high-priority queue prior to responding to fetch requests in the low-priority queue. Otherwise, requests may be processed on a first-in-first-out basis.

[0034] FIG. 6 is a timing diagram illustrating a sequence of requests for pre-fetching “predictable” data and fetching “unpredictable” data with this prioritized fetch scheme. Again referring to the attention sub-block illustrated in FIG. 1 as an example, the system 10 may be in the process of completing a first multi-stage computation. Message Ml may represent a request issued from a processing tile 200 to pre-fetch values for a latercomputation stage, such as Q, K, and / or V for a subsequent multi-stage computation. This message Ml is transmitted to the memory controller 40, where it is recognized as a low- priority request and entered in the low-priority queue 44. As the high-priority queue 42 is currently empty, the controller 40 proceeds with the request of message Ml, and sends read instructions R1 to the memory bank 30 (the memory interface 35 is not shown in FIG. 6 for simplicity). The memory bank 30 then proceeds with the read operation and fetches the requested data Fl.

[0035] While the memory bank 30 is busy with F 1 , a second message M2 with a low- priority fetch request is sent from a tile 200. The high-priority queue 42 is still empty; however, the memory bank 30 is still busy with Fl, so the controller does not transmit instructions to the memory bank 30 at this time.

[0036] Subsequently, a tile 200 sends a third message M3 to the memory controller 40 with a high-priority fetch request, such as a request for a reduced set of operands corresponding to non-zero values produced by an activation function. This request is received in the high- priority queue 42 before the controller 40 determines that the memory bank 30 has completed Fl.

[0037] Once the memory bank 30 has completed the read instruction of R1 and returned the data DI to the controller 40 (which then transmits the data back to the requesting tile 200), the controller 40 determines that a new instruction can now be sent to the memory bank 30. There is currently a request queued in both the high-priority and low-priority queues 42, 44; therefore, the controller 40 handles the high-priority request 42, and transmits a read instruction R3 to the memory bank 30. The memory bank 30 then proceeds with the read operation and fetches the requested data F3. Once that read is complete and the data D3 is sent to the requesting tile 200, the high-priority queue 42 is now empty, leaving a final request in the low-priority queue. The controller 40 may now transmit that read instruction R2 to the memory bank 30, which then performs a fetch F2 and returns the requested data D2. This prioritized fetch scheme thus enables the computing system 10 to take advantage of pre-fetching to reduce latency in computations, while permitting a queue of pre-fetch requests to be interrupted by more urgent requests.

[0038] Fetch requests may be designated by the requesting entity, e.g., tile 200, as either high- or low-priority by having the requesting entity address the request to the appropriate queue 42, 44. Alternatively, high-priority data may be distinguished from low-priority data based on the range of memory addresses used to store the data in the memory bank 30; for example, Q, K, and V values, which are “predictable” and can be pre-fetched, may be stored in a lower address range, while coefficients, which are identified for retrieval only after the activation function is computed, are stored in a higher address range. The controller 40 may be configured to prioritize requests to fetch data from the higher address range, which correlates to higher priority fetch requests. In this latter implementation, separate high- and low-priority queues may not be required; instead, a single priority queue may be implemented, in which the controller 40 pulls the request with the highest address value.

[0039] Additionally, the physical addresses at which data is stored in the memory bank 30 may be selected to optimize system performance. For example, the coefficients for the subsequent stage computations may be arranged so that fetching a row of coefficients involves consecutive reads to the same memory page. Alternatively, the coefficients may be stored in different pages so that the pages can be read in parallel. This latter approach can result in lower latency compared to the former approach, and reduces the likelihood of a sequence of sparse reads on a single memory page.

[0040] FIG. 6 also depicts control messages C1-C3 and Cl'-C3' transmitted between processing tiles 200 and a coordinating controller 350. As mentioned above, one tile 200 among the multiple tiles 200 in a given tensor processing device 100 or the computing system 10 as a whole may be designated as a coordinating controller 350 to manage the other tiles 200. However, in other embodiments, the computing system 100 may include a distinct controller unit (e.g., another processor) in communication with all tiles 200 that assumes this function, rather than employing a tile 200 as the coordinating controller 350.

[0041] The coordinating controller 350 may track fetch requests issued by tiles 200, or issue fetch requests on behalf of tiles 200. If tiles 200 issue their own messages (M1-M3) requesting data from the memory bank 30, then they may also transmit messages C1-C3 to the coordinating controller 350 containing information about the request, and confirmation messages Cl'-C3' indicating that the requested operand was received. The confirmationmessages Cl'-C3' may alternatively be short messages indicating that the operation to be carried out with the received operand (e.g., a MAC operation) was completed. Such messages enable the coordinating controller 350 to manage collection of output from the various processing elements 300, and / or to determine whether the processing elements 300 are ready for the next operation.

[0042] The foregoing example of the prioritized fetch scheme associates prioritized fetches with “unpredictable” fetch requests resulting from sparse output from a computation stage. It will be understood by those skilled in the art that other factors or criteria may be used to define high- and low-priority fetch requests; for example, features other than sparsity may result in an “unpredictable” fetch request. It will also be understood that the prioritized fetch scheme may also be used even if there is a low degree of sparsity in the output, although the time savings in fetching corresponding operands from the memory bank 30 may not be significant.

[0043] In those cases where sparsity is a feature of the output of a computation stage and is used to determine which operands to request from the memory controller 40 for a next stage, those sparse operands must be identified in order to make the fetch request.Therefore, when a processing tile 200 completes an operation, the results are converted to a sparse form. A sparse form is, effectively, a compressed representation of the output, in which the non-zero values are identified by index. Any suitable representation scheme may be used. For example, a sparse vector may be represented by a list of pairs (z, a of the index z of the non-zero value and its value a,. A variation of the list of pairs is a run-length encoding (z- iprev, a , in which the first element is the run length, i.e., the difference between the index of the non-zero value and the index iprevof the previous non-zero value. If a run length value is greater than the word size used to store the representation, the run-length encoding may be split by including an explicit zero entry: for example, (300, 1.2) is equivalent to (255, 0), (44, 1.2). A sparse representation of an zz-dimensional tensor may be represented in the same manner as an zz-tuple.

[0044] The conversion may be carried out by an additional unit in the processing tile 200. FIG. 7 illustrates a variant of the tile of FIG. 4, including a sparsity converter unit 330 associated with each row of processing elements 330. The sparsity converter unit 330comprises logic operable to detect zero values in the output of the processing elements 300, and to produce the sparse representation. This sparse representation is passed to the tile controller 210, which is configured to construct messages for the memory controller 40 (and the coordinating controller 350) requesting the specific coefficients or other values corresponding to the non-zero values identified in the sparse representation that are required for the next operation.

[0045] It may also be possible to predict which operations in a multi-stage computation are likely to yield a zero or effectively-zero value. Again referring to FIG. 1, consider that a negative input to the nonlinear activation function NonLin will generally provide a zero result, thus contributing to the sparsity of the NonLin output. The input to NonLin is the result of a matrix multiplication in MLP In that ordinarily requires retrieving an entire array of coefficients. These coefficients may be split into “coarse” and “fine” terms (e.g., most significant and least significant bits), and the tile controller 210 may request that only the coarse terms be fetched for use in executing MLP In. If a result of MLP In is sufficiently negative, it may be possible to project that the result would still be negative or zero even if the fine term was included. In that case, it can be concluded simply on the basis of the coarse term that the outcome of NonLin will likewise be zero. Thus, it may be possible to start identifying which coefficients will be required, and which will not, prior to the completion of MLP In or NonLin. Additionally, fetching only the coarse coefficient term for the initial MLP In computation may reduce the overall volume of data to be fetched from the memory bank 30, since the fine term need only be fetched for those processing elements 330 that did not produce a sufficiently negative result.

[0046] A similar consideration may be applied to the softmax computation; if a result of softmax (QK) is effectively zero, it is not necessary to fetch the corresponding element of V for the next operation since the output will again be zero or effectively zero.

[0047] When sparse fetching is utilized, the order in which results are received will likely be unpredictable, as the sequence of physical memory access requests will be unpredictable, and different pages of the memory will likely experience different queuing. Consequently, the order in which operands are received by individual processing elements and the order in which operations are performed are expected to be unpredictable as well. Variations in theorder of computations are likely irrelevant for multiply-accumulate operations, since these operations are associative. However, with conventional floating-point arithmetic, operations are effectively non-associative due to roundoff errors, potentially introducing some randomness into the results of the computations. Therefore, if exact reproducibility of results is required, true associative arithmetic should be employed.

[0048] The configuration of the tensor processing device 100, processing tiles 200, and processing elements 300 provides an architecture that is theoretically infinitely scalable to meet the requirements of any computation exercise. The number of tiles 200 on a tensor processing device 100, the number of processing elements 300 per processing tile 200, and the number of tensor processing devices 100 employed in the computing system 10 may be determined according to the particular application, such as the sizes of tensors used in computations, subject to practical constraints in fabrication. In one example embodiment, each tensor processing device 100 comprises an array of 4096 tiles (64 x 64), each tile containing 128 processing elements 300. The topology of the network of modules 110 may be defined arbitrarily, and there is effectively no difference between messages passed between processing tiles 200 located in a same module 110 on a tensor processing device 100, and processing tiles 200 located in different modules 110 on different tensor processing devices 100, except for latency.

[0049] The solutions presented in this disclosure are not limited to applications with ANNs or large language models. Some applications may not require such a large number of processing elements 300. For some implementations, a single tensor processing device 100, or only one or a few processing tiles 200, may be sufficient. The prioritized fetch scheme, with or without sparsity conversion, may still be implemented with a single processing tile 200 or small set of processing tiles 200.

[0050] Accordingly, there is provided a computing system comprising: a memory bank in communication with at least one memory controller; one or more processing tiles in communication with the at least one memory controller, each of the one or more processing tiles comprising a local controller and a plurality of processing elements, each processing element comprising logic to perform one or more operations on at least one received operand; a coordinating controller of one of the one or more processing tiles beingconfigured to execute logic for retrieval of operands from the memory bank for use in operations by the array of processing elements, comprising: before or during execution of one or more operations by the plurality of processing elements in a first stage of a first multistage computation, initiating a first request to the at least one memory controller for a first set of operands required for one or more operations to be executed subsequent to the first stage, the first request having a first priority; after completion of the first stage of the first multi-stage computation, identifying a second set of operands required for a subsequent stage of the first multi-stage computation, and initiating a second request to the at least one memory controller for the second set of operands, the second request having a higher priority than the first request.

[0051] In one aspect, the coordinating controller comprises a local controller of one of the one or more processing tiles.

[0052] In another aspect, initiating the first request comprises initiating transmission of the first request for receipt in a first queue maintained by the at least one memory controller, and initiating the second request comprises initiating transmission of the second request for receipt in a second queue maintained by the at least one memory controller, the second queue having a higher priority than the first queue.

[0053] In a further aspect, the first set of operands is stored at a first set of memory addresses in the memory bank and the second set of operands is stored at a second set of memory addresses in the memory bank, the memory controller configured to execute logic to prioritize read requests for the second set of memory addresses over read requests for the first set of memory addresses.

[0054] In still another aspect, the first and subsequent multi-stage computations comprise first and subsequent layers in a neural network model, each layer comprising at least one neuron, each neuron comprising at least a first function and a second function, the first stage of the first multi-stage computation comprising computation of the first function and the subsequent stage of the first multi-stage computation comprising computation of the second function.

[0055] In another aspect, the one or more operations executed during a given stage of a multi-stage computation are executed by a subset of processing elements of the computer system.

[0056] In another aspect, subset of processing elements is comprised in a single processing tile.

[0057] In a further aspect, the subset of processing elements is comprised in more than one processing tile.

[0058] In yet another aspect, the first set of operands is required for one or more operations in a subsequent stage of the first multi-stage computation, or for one or more operations in a second multi-stage computation to be executed after completion of the first multi-stage computation.

[0059] In still a further aspect, the first set of operands is required for one or more operations in a second multi-stage computation to be executed after completion of the first multi-stage computation.

[0060] In another aspect, the second multi-stage computation immediately follows the first multi-stage computation.

[0061] In another aspect, the subsequent stage of the first multi-stage computation immediately follows the first stage of the first multi-stage computation.

[0062] In still another aspect, the one or more processing tiles receive the first set of operands, and receive the second set of operands; the first set of operands is received after the second set of operands, or vice versa.

[0063] In another aspect, different processing tiles receive the first set of operands and the second set of operands.

[0064] In another aspect, processing elements of the one or more processing tiles receive the first set of operands and the second set of operands.

[0065] In a further aspect, different processing elements receive the first set of operands and the second set of operands.

[0066] In yet another aspect, the second set of operands comprise operands that can be identified only after completion of the first stage of the first multi-stage computation.

[0067] In still a further aspect, the subsequent stage of the first multi-stage computation comprises a computation on values of an array and the second set of operands, the second set of operands comprising values corresponding to non-zero values of the array.

[0068] In another aspect, the one or more processing tiles comprises a plurality of networked processing tiles.

[0069] In another aspect, each of the plurality of processing elements on the one or more processing tiles perform operations on discrete operands substantially in parallel.

[0070] It should be understood that this description is not intended to be limiting, and that the examples contemplated herein include all alternatives, modifications, and equivalents as would be appreciated by the person skilled in the art, and are included within the scope of the accompanying claims. Although the features and elements various examples or embodiments may be described as being in particular combinations, the person of ordinary skill in the art will appreciate which features or elements can be used alone, without the other features and elements of the embodiments, or in various combinations with or without other features and elements disclosed herein. Further, individual features or variations described in respect of one example or embodiment in this disclosure can be used with other examples or embodiments mentioned herein, as would be understood by the person skilled in the art.

[0071] The examples and embodiments are presented only by way of example and are not meant to limit the scope of the subject matter described herein. Each example embodiment presented above may be combined, in whole or in part, with the other examples. Further, variations of these examples will be apparent to those in the art and are considered to be within the scope of the subject matter described herein. Some steps or acts in a process or method may be reordered or omitted, and features and aspects described in respect of one embodiment may be incorporated into other described embodiments.

[0072] Hardware components, software modules, engines, functions, and data structures may be connected directly or indirectly to each other in order to allow the flow of dataneeded for their operations. Functional elements described herein may comprise VLSI circuits or gate arrays; field-programmable gate arrays; programmable array logic; programmable logic devices; commercially available logic chips, transistors, and other such components.

Claims

CLAIMS1. A computing system comprising: a memory bank in communication with at least one memory controller; one or more processing tiles in communication with the at least one memory controller, each of the one or more processing tiles comprising a local controller and a plurality of processing elements, each processing element comprising logic to perform one or more operations on at least one received operand; a coordinating controller of one of the one or more processing tiles being configured to execute logic for retrieval of operands from the memory bank for use in operations by the array of processing elements, comprising: before or during execution of one or more operations by the plurality of processing elements in a first stage of a first multi-stage computation, initiating a first request to the at least one memory controller for a first set of operands required for one or more operations to be executed subsequent to the first stage, the first request having a first priority; after completion of the first stage of the first multi-stage computation, identifying a second set of operands required for a subsequent stage of the first multistage computation, and initiating a second request to the at least one memory controller for the second set of operands, the second request having a higher priority than the first request.

2. The computing system of claim 1 , wherein the coordinating controller comprises a local controller of one of the one or more processing tiles.

3. The computing system of claim 1 , wherein initiating the first request comprises initiating transmission of the first request for receipt in a first queue maintained by the at least one memory controller, and initiating the second request comprises initiating transmission of the second request for receipt in a second queue maintained by the at least one memory controller, the second queue having a higher priority than the first queue.

4. The computing system of claim 1 , wherein the first set of operands is stored at a first set of memory addresses in the memory bank and the second set of operands is stored at a second set of memory addresses in the memory bank, the memory controller configured to execute logic to prioritize read requests for the second set of memory addresses over read requests for the first set of memory addresses.

5. The computing system of claim 1, wherein the first and subsequent multi-stage computations comprise first and subsequent layers in a neural network model, each layer comprising at least one neuron, each neuron comprising at least a first function and a second function, the first stage of the first multi-stage computation comprising computation of the first function and the subsequent stage of the first multi-stage computation comprising computation of the second function.

6. The computing system of claim 1, wherein the one or more operations executed during a given stage of a multi-stage computation are executed by a subset of processing elements of the computer system.

7. The computing system of claim 6, wherein the subset of processing elements is comprised in a single processing tile.

8. The computing system of claim 7, wherein the subset of processing elements is comprised in more than one processing tile.

9. The computing system of claim 1 , wherein the first set of operands is required for one or more operations in a subsequent stage of the first multi-stage computation, or for one or more operations in a second multi-stage computation to be executed after completion of the first multi-stage computation.

10. The computing system of claim 9, wherein the first set of operands is required for one or more operations in a second multi-stage computation to be executed after completion of the first multi-stage computation.

11. The computing system of claim 10, wherein the second multi-stage computation immediately follows the first multi-stage computation.

12. The computing system of claim 9, wherein the subsequent stage of the first multistage computation immediately follows the first stage of the first multi-stage computation13. The computing system of claim 1 , further comprising the one or more processing tiles: receiving the first set of operands, and receiving the second set of operands.

14. The computing system of claim 13, wherein the first set of operands is received after the second set of operands.

15. The computing system of claim 13, wherein different processing tiles receive the first set of operands and the second set of operands.

16. The computing system of claim 1 , wherein processing elements of the one or more processing tiles receive the first set of operands and the second set of operands.

17. The computing system of claim 1 , wherein different processing elements receive the first set of operands and the second set of operands.

18. The computing system of claim 1, wherein the second set of operands comprise operands that can be identified only after completion of the first stage of the first multi-stage computation.

19. The computing system of claim 1, wherein the subsequent stage of the first multistage computation comprises a computation on values of an array and the second set of operands, the second set of operands comprising values corresponding to non-zero values of the array.

20. The computing system of claim 1, wherein the one or more processing tiles comprises a plurality of networked processing tiles.

21. The computing system of claim 1 , wherein each of the plurality of processing elements on the one or more processing tiles perform operations on discrete operands substantially in parallel.

Citation Information

Patent Citations

  • SIMD parallel processor architecture

    US20110134131A1