Systolic ai processor compiler
Patent Information
- Application Number
- PCT/US2024/034736
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-19
- Filing Date
- 2024-06-20
- Publication Date
- 2025-07-31
AI Technical Summary
Current AI processing systems face challenges in achieving high computational throughput while minimizing power consumption and cost, particularly in large-scale AI applications that require high memory bandwidth and low memory access power.
The development of a systolic AI processor compiler that utilizes a parallelized computation architecture with multiple processing cores and memory devices, allowing for efficient matrix-vector multiplication operations, thereby enhancing computational throughput, power efficiency, and reducing costs.
This approach results in significantly higher memory bandwidth and power efficiency compared to conventional systems, with potential cost savings of up to 100x compared to LPDDR4 SDRAM dies, while maintaining high memory capacity.
Smart Images

Figure US2024034736_31072025_PF_FP_ABST
Abstract
Description
SYSTOLIC Al PROCESSOR COMPILERCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit under 35 U.S.C. §119 of U.S. Provisional Patent Application No. 63 / 591,513 filed on October 19, 2023, and entitled "Systolic Al Processor Compiler’, which is hereby incorporated by reference herein in its entirety.STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH
[0002] This invention was made with government support under FA8702-15-D-0001 awarded by The U.S. Air Force. The government has certain rights in the invention.BACKGROUND
[0003] Artificial intelligence (Al) is becoming ubiquitous as increasing performance is enabling new capabilities to improve peoples’ lives. Current Al processing, including generative Al processing, requires high computational throughput, and it is likely that next-generation Al systems will require even more processing, cost more, and consume much more power, contributing significantly to increasing levels of global warming. Therefore, it is critical to develop low-cost Al-processing technology7for future systems that will provide much higher performance while also consuming much less power.SUMMARY
[0004] Fig. 1 shows a simple example neural network 100, where the circles represent neurons. The neurons are arranged in multiple layers including an input layer 102, a hidden layer 104, and an output layer 106. The input layer 102 of neurons receives external input signals. These signals are then multiplied by unique w eights and sent to the next layer of neurons (hidden layer 104 in Fig. 1), as illustrated with arrows between layers 102 and 104. The neurons in hidden layer 104 sum the received signals, multiply the sums with unique weights, and then send the product signals to the next layer of neurons, shown as the output layer 106 in the simple example of Fig. 1. The neurons of output layer 106 sum the input signals and output the results. In practice, there can be a very large number of neurons in each layer, and there can also be a large number of hidden layers.
[0005] The computations in neural networks can be modeled as matrix computations. The sums of the input signals into a neuron layer can be thought of as an input vector, and the products between input elements and unique weights can be thought of as partial products. The sums of the partial products in the next neuron layer can be thought of as an output vector. The multiplication by weights and summing operations can be thought of as a matrix-vector multiplication operation. Therefore, processing systems designed for neural network implementations are generally optimized for matrix computations.
[0006] In recent years, efforts have been made to speed up processing for neural network implementations for Al applications. Many different architectures were typically used, including central processing units (CPUs), graphics processing units (GPUs), tensor processing units (TPUs), and other specialized architectures. Al processing tends to require high memory bandwidth because it generally does relatively few7processing tasks per memory access. For each neural network weight value accessed from memory, only one weight multiplication is performed, whereas other types of scientific and numerical processing perform many arithmetic computations for each memory access. Some TPUs and GPUs process large batches of data in parallel and perform many multiplications and accumulates per w eight accessed from memory, thus requiring few er memory' accesses per each computation performed.
[0007] Unlike some large data center Al applications, other Al applications such as embedded and edge computing applications generally do not process large batches of data in parallel, and thus need high memory bandwidth and also low memory power consumption per memory access. Many low-latency applications such as self-driving vehicles and military applications can also require high memory bandwidth and low' power consumption.
[0008] Al processors typically use some sort of dynamic memory (DRAM) since DRAMs provide high memory capacity' at reasonable costs. Many different forms of DRAMs are used, including synchronous DRAM (SDRAM), graphics double data rate DRAM (GDDR DRAM), and high bandwidth memory (HBM). Although these DRAMs provide high memory capacity at reasonable costs, they are limited in memory’ bandwidth because they reside in a separate package from the processor and the communication path to the memory' often becomes the principal performance bottleneck. The communicationpath also significantly increases the system power consumption because it consumes much more power per bit transmitted / received than the memory bit operations.
[0009] Some Al processors try to eliminate the memory access bottleneck by using static random-access memory (SRAM) built into the processor application-specific integrated circuit (ASIC) die. Normally, processors use SRAMs on the same die mostly for cache memory, and the main memory is DRAM that is on a separate die / package. However, some Al processors use SRAMs as main memory on the same die to achieve high memory bandwidth. In order to get the necessary higher memory capacity, a very large die may be required (e.g., a dinner plate-sized die). Although a larger processor die with high-capacity SRAMs on board can eliminate the memory bottleneck, it comes at a significant cost because larger dies cost significantly more to fabricate than a small die. Even with such large dies, it is difficult to match the capacity of DRAMs, so it is difficult for on-chip-SRAM-based Al processors to have high capacity main memory. This limits these processors’ work on large-scale Al problems and increases the processor cost.
[0010] Another attempt to alleviate the memory bottleneck is using the processing in memory (PiM) concept. SAMSUNG developed DRAM dies that use DRAM fabrication process devices within the DRAM die to implement logic devices to perform matrix multiplications and other arithmetic functions. However, because the logic devices implemented in DRAM fabrication process are much larger, slower, and power hungry than the logic devices in ASIC fabrication processes, the performance gain, although significant, was limited. In addition, up to half the die area in the DRAM die had to be dedicated to the logic devices, thus reducing the DRAM capacity.
[0011] This disclosure describes systems and methods for parallelized computation of a product of a matrix and an input vector using a plurality of processing cores and a plurality of memory devices. More specifically, this disclosure describes how a computational kernel of matrix-vector multiplication, as used for example in neural networks, can be implemented efficiently in architecture that have more than one processing core and more than one memory' device. Such an architecture may, for example, be a systolic or semi-systolic array. This results in higher computational throughput, higher power efficiency, and lower cost per throughput than conventional processing systems and methods (e g., as used in conventional Al processors), while stillproviding the high memory capacity necessary to be able to work on large problems. Other types of memories may be used with the disclosed concepts and structures. For example, non-volatile memories such as FLASH memory may be used. It is appreciated herein that such memories may have higher capacity compared to DRAM, but lower read / write speeds and lower endurance with limited read / write cycles.
[0012] The general concepts described herein may be used to provide processing systems and methods for parallelized computation of a product of a matrix and an input vector suitable for use in various applications, including Al applications.
[0013] With existing processing systems, input and output signals from multiple memory banks go through a common interface, such as double data rate (DDR) interfaces in case of SDRAMs. It is appreciated herein that the common interface may be eliminated by providing simple individual low-power interfaces for each memory bank or utilizing such interfaces provided by existing memory banks. Individual interfaces can be simpler than a common interface in that they do not have to coordinate inputs and outputs associated with many memory banks. The processing systems and methods described herein may take advantage of such individual interfaces, for example in distributing data across a plurality of processing core and a plurality of memory devices.
[0014] While the general concept described herein is well suited for neural network processing, it is possible to extend the concept for other applications. For example, processor cores that directly interface with individual memory banks can be optimized for different types of applications. The general concepts and structures disclosed herein can provide high memory bandwidth, low memory access power per bit access, high memory capacity, and low memory cost for various applications for which such merits are important. In some cases, processor cores for such applications can be designed to read or write to the memory one complete word at a time, further improving efficiency.
[0015] Potential applications of the disclosed technique include, but are not limited to: Al inference, Al training, various levels of self-driving, high-bandwidth memory' systems, matrix processing, sparse matrix processing, graph processing, real-time trading / finance, data security, database processing, database search, finite element methods, digital signal processing, image processing, video processing, sensor array processing, communications, sorting, medical imaging, entertainment, gaming, graphics processing, personal computing,parallel computing, supercomputing, CAD / CAM, mobile computing, and Internet of Things (loT).
[0016] For at least some of these applications, the performance improvement, power savings, and cost savings are projected to be very' significant. For example, compared to LPDDR4 (Low-Power Double Data Rate) SDRAM dies, executing the disclosed methods on the die stack described herein may provide up to lOOx memory- bandwidth and lOOx power efficiency with associated cost savings.
[0017] Disclosed parallelized processing systems and methods can be used to provide providing edge Al computing (e.g., in mobile and automotive applications) with low power consumption. Although high memory bandwidth provides the low latency needed, the duty cycle can be kept low because computing data input, such as voice input or visual input, w ill likely be provided sporadically (a few w ords per second or less, tens of frames per second or less) for critical applications such as natural language user interface and various levels of self-driving.
[0018] Disclosed concepts, structures, and techniques may be used in conjunction with concepts, structures, and techniques disclosed in International Pat. App. No.PCT / US2024 / 019512, which was filed on March 12, 2024, is entitled '‘3D Processor." and is hereby incorporated by reference herein in its entirety.
[0019] According to one aspect of the disclosure, a system for parallelized computation of a product of a matrix and an input vector includes a plurality of processing cores. Each processing core has at least one accumulator (and in some cases more than one accumulator). The system also includes a plurality- of memory devices, each memorydevice coupled to and associated with at least one of the plurality of processing cores. The system further includes a controller coupled to the plurality of processing cores. Each processing core of the plurality of processing cores is configured to receive from the controller to (a) store a plurality- of matrix elements of the matrix in at least one memory' device associated with the processing core. Each processing core is further configured to (b), during a first set of clock cycles, receive at least one vector element of the input vector from the controller. Each processing core is also configured to (c), during the first set of clock cycles, multiply the at least one vector element with at least one matrix element stored in the at least one memory device associated w ith the processing core. Eachprocessing core is configured to (d), during the first set of clock cycles, add the result of the multiplication to a value of the at least one accumulator and store the result of the addition in the at least one accumulator. Each processing core is also configured to (e) repeat steps (b) through (d) during at least a second set of clock cycles. Each processing core is configured to (f) transmit the value of the at least one accumulator as an output to the controller. The general concepts described herein may be used to provide system for parallelized computation suitable for use in various applications, including Al applications.
[0020] In some embodiments, each memory device is coupled to and associated with one processing core. In some embodiments, the plurality of processing cores forms one of a one-dimensional semi-systolic array and a two-dimensional semi-systolic array. In some embodiments, each processing core is configured to receive one input vector element per clock cycle from the controller. In other embodiments, each processing core is configured to receive more than one vector element per clock cycle from the controller.
[0021] In some embodiments, the controller is configured to store one of a row of the matrix and a column of the matrix in each processing core. In other words, the controller may be configured to store a row of the matrix in each processing core, or the controller may be configured to store a column of the matrix in each processing core. In some embodiments, the controller is configured to store a plurality of rows of the matrix in each processing core. In some embodiments, the controller is configured to store one of a portion of the columns of the matrix and a portion of the row s of the matrix in each processing core.
[0022] In some embodiments, the semi-systolic array is two-dimensional and arranged in rows and columns, and w herein each processing core in a row of the two-dimensional semi-systolic array is configured to receive half of the input vector elements from the controller. In some embodiments, the controller is configured to sum the outputs of the processing cores in a column of the two-dimensional semi-systolic array to result in at least one element of the product. In some embodiments, the matrix is a weight matrix associated with a neural network.
[0023] According to another aspect of the disclosure, a computer-implemented method for parallelized computation of a product of a matrix and an input vector is executed on aplurality of processing cores, each processing core having at least one accumulator, a plurality of memory devices, each memory device coupled to and associated with at least one of the plurality of processing cores, and a controller coupled to the plurality of processing cores. The method includes (a) receiving instructions from the controller and storing, by a processing core of the plurality of processing cores, a plurality7of matrix elements of the matrix in at least one memory device associated with the processing core. The method also includes (b). during a first set of clock cycles, receiving, by the processing core and from the controller, at least one vector element of the input vector. The method further includes (c), during the first set of clock cycles, multiplying, by the processing core, the at least one vector element with at least one matrix element stored in the at least one memory device associated with the processing core. The method includes (d), during the first set of clock cycles, adding, by the processing core, the result of the multiplication to a value of the at least one accumulator and storing the result of the addition in the at least one accumulator. The method also includes (e) repeating, by the processing core, steps (b) through (d) during at least a second set of clock cycles. The method further includes (f) transmitting, by the processing core and to the controller, the value of the at least one accumulator as an output.
[0024] In some embodiments, the processing core stores one of a row and a column of the matrix and receives one input vector element per clock cycle, and the output is an element of the product. In some embodiments, each memory^ device is coupled to and associated with one processing core.
[0025] In some embodiments, each processing core stores one of a portion of the columns of the matrix and a portion of the rows of the matrix, and the method further includes summing, by the controller, the outputs of more than one processing core of the plurality of processing cores to result in an element of the product. In some embodiments, the plurality of processing cores is arranged in one of a systolic and a semi-systolic array and wherein the array is one of one-dimensional and two-dimensional. In some embodiments, the matrix is a weight matrix associated with a neural network.
[0026] In some embodiments, the method further includes determining, by the controller, a distribution of matrix elements to the plurality of processing cores. In some embodiments, the distribution of matrix elements to the plurality of processing cores isdetermined by a compiler and provided to the controller. In some embodiments, the distribution of matrix elements to the plurality of processing cores is determined at run time by the controller. In some embodiments, the method further includes determining, by at least one of the controller and a compiler, a sequence of computation operations and data flow.100271 It should be appreciated that individual elements of different embodiments described herein may be combined to form other embodiments not specifically set forth above. Various elements, which are described in the context of a single embodiment, may also be provided separately or in any suitable sub-combination. It should also be appreciated that other embodiments not specifically described herein are also within the scope of the following claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The manner of making and using the disclosed subject matter may be appreciated by reference to the detailed description in connection with the drawings, in which like reference numerals identify like elements.
[0029] Fig. 1 is a schematic diagram showing a simplified neural network.
[0030] Fig. 2 is a perspective view of a 3D processor, according to some embodiments.
[0031] Fig. 3 is a cross-sectional view illustrating a die integration that may be used with a 3D processor, according to some embodiments.
[0032] Fig. 4A is schematic diagram illustrating a one-dimensional (I D) systolic array architecture, according to some embodiments.
[0033] Fig. 4B is schematic diagram illustrating a two-dimensional (2D) systolic array architecture, according to some embodiments.
[0034] Fig. 5 is a block diagram showing an implementation of a 3D processor die having a ID semi-systolic array architecture, according to some embodiments.
[0035] Fig. 6 is a block diagram illustrating a systolic processing core (SPC) architecture, according to some embodiments.
[0036] Fig. 7 A is a schematic diagram illustrating a ID toroidal array architecture, according to some embodiments.
[0037] Fig. 7B is a schematic diagram illustrating a 2D toroidal array architecture, according to some embodiments.
[0038] Fig. 8 is a flow diagram of a computer-implemented method for parallelized computation of a product of a matrix and an input vector, according to some embodiments.
[0039] Fig. 9 is a flow diagram of computer-implemented method for parallelized computation of a product of a matrix and an input vector, according to some embodiments.
[0040] The drawings are not necessarily to scale, or inclusive of all elements of a system, emphasis instead generally being placed upon illustrating the concepts, structures, and techniques sought to be protected herein.DETAILED DESCRIPTION
[0041] Fig. 2 shows a 3D-integrated processor 200, according to some embodiments. A processor die 202 having a plurality of processor cores 206a, 206b, etc. (206 generally) can be situated or disposed on top of a memory die 204 having a plurality of memory banks 208a, 208b, etc. (208 generally). Processor 200 having both memory die 204 and processor die 202 may be said to have a “3D die stack"’ or “3D die” for short.
[0042] Processor cores 206 and memory banks 208 may be connected using 3D signal paths, meaning signal paths that run not just alone the plane of the die (e.g., the x-y plane in Fig. 2), but also along the perpendicular axis (e.g.. the z axis in Fig. 2).
[0043] In the example shown, each of the processor cores 206a, 206b, etc. may be disposed over a corresponding one of the memon' banks 208a, 208b, etc. Also, in the example of Fig. 2, the processor cores 206 and memon' banks 208 can be arranged as 2D arrays on the respective dies 202 and 204. While Fig. 2 shows twelve (12) processor cores and twelve (12) corresponding memoiy7banks, arranged in a 4x3 array, the generally concepts and structures described herein can provide for 3D processors having othernumbers of processor cores and memory banks. As one example, a 3D processor may have 1024 processor cores and 1024 memory’ banks.
[0044] In some embodiments, each of the processor cores 206 may have similar, if not identical, surface areas (e.g., dimensions in the x-y plane) to each of the memory' banks 208. For example, as shown in Fig. 2, processor core 206a can have the same width and length as memory bank 208a, processor core 206b can have the same width and length as memory bank 208b, etc. (where width and length are measured along the x and y axes, respectively, in the figure). In some cases, the area of the process cores can be substantially equal to that of the memory banks. For example, the areas can be within ±1%, ±5%, ±10%, or ±20% of each other. This allows each processor core to be perfectly, or nearly perfectly aligned, with each corresponding memory bank in the x-y plane and allows them to be compactly stacked on top of each other. This can be done to increase (and ideally maximize) density on the dies, allowing for a greater number of 3D processor cores on a given die surface; to reduce (and ideally minimize) the total surface area of the dies for a given number of 3D processor cores; and / or to improve (ideally maximize) performance.
[0045] A given memory’ bank 208 may be configured to store many bits of data in a memory cells. Data can be read and written to a memory bank 208 in “words.” each word corresponding to many bits (e.g., thousands of bits). That is, all the bits belonging to a word can be tied together by a “w ord line” and all the bits tied to a w ord line are read or written together. In more detail, a w ord line in a DRAM bank may consist of a few thousand bits (e.g., 4K or 8K bits), may be read out through a DRAM interface in smaller portions (e.g., 64 bits at a time). In some embodiments, memory banks 208 may be provided as DRAMs. In some embodiments, memory’ banks 208 may be provided as nonvolatile memories such as FLASH memory. In other embodiments, memory banks 208 may be provided as non-volatile memory components. Individual memory banks 208 may have their own input / output (I / O) interfaces such that different memory banks 208 can concurrently perform different read / write operations. Processor cores 206 interface directly with memory’ banks 208 using these individual I / O interfaces. For example, to access a 64-bit portion of a word line from a memory bank, a processor core can specify (over the 3D signal paths) address bits of a particular word line and of a particular portion thereof (e.g., a beginning and an end of a 64-bit portion thereof). To access multipleportions of a particular word line, or of different word lines, the memory' bank can send read data 64-bits at a time to the processor core.
[0046] Processor cores 206 can be configured to read or write one complete word at a time to take full advantage of the available memory' bandwidth. Memory banks 208 may be organized so that a full word is read or written during each memory access cycle. Fortunately, for matrix-based Al computations, it is relatively straightforw ard to organize memory ■ access in this yvay, as a yvord can represent entire row' or column (or partial royv or column) of the neural netw ork w eight matrix.
[0047] A processor core 206 can include circuitry (e.g., an ASIC) configured to execute a set of instructions, such as arithmetic instructions, memory operations, and control operations. The particular circuitry and instruction set provided by a processor core 206 can be selected based on the ty pe of application the 3D processor 200 is intended to be used with. In some embodiments, all processor cores 206 on the die 202 may be identically configured. In other embodiments, different processor cores 206 on the die 202 may be differently configured. For example, different processor cores 206 may be provided as different types of ASICs to complete different sub-tasks of a larger Al problem / application. Additional details of processor cores 206 including how they can be interconnected and controlled are provided below and in subsequent figures.
[0048] The 3D signal paths between the two dies 202, 204 allow data to flow directly between memory banks 208 and processor cores 206 situated there above. This arrangement eliminates the common interface bottleneck found in conventional DRAM. For example, with processor 200, a w ord of data (yvhich may include thousands of bits) can be read from a memory' bank 208 and output through one or more 3D signal paths to a processor core 206 directly above it. Conversely, a word from a processor core 206 can be written into a memory’ bank 208 below it.
[0049] The memory configuration illustrated in Fig. 2 has many advantages over conventional memory (e.g., conventional DRAM modules). Since there are individual I / O interfaces to each memory bank rather than one common interface for all the memory banks, the memory bandwidth is much larger than conventional memory'. All the memory banks can be accessed in parallel by all the cores rather than accessing one bank at a time as in conventional memory. In addition, since the vertical signal paths between thememory banks and processor cores can be made relatively short, associated parasitic capacitances are minimal. This greatly reduces the power consumption associated with memory access in terms of pow er spent per bit accessed (written or read). In addition, the disclosed memory7configuration retains the high memory capacity and low cost per bit features of existing memory7(e.g., DRAM).
[0050] The 3D processor illustrated in Fig. 2 is merely one example of the general concept of a processing system for executing the methods disclosed herein. Other variations to this concept are possible as well. For example, for applications that require lower memory bandwidth, a 3D memory interface may be shared betw een a single processor core and multiple memory banks. As another example, a 3D memory interface may be shared between multiple processor cores and a single memory bank. The general concepts described herein can be extended to provide 3D dies having three or more layers. For example, a three-layer stack may include a processor die, a DRAM die, and a nonvolatile memory7die. The non-volatile memory can be used as a high bandwidth on- package disk drive for the die stack, for example. As another example, a die stack may include a processor die and two or more DRAM dies.
[0051] Fig. 3 shows an example of 3D die integration. In this example, multiple layers of metal interconnect layers, and insulator layers are used to make electrical connections between device layers.
[0052] As shown, a structure 300 can include a first device layer 302 and a second device layer 304 disposed thereover. The two device layers 302. 304 may be attached by a layer-to-layer bond 306. Each device layer may include signal paths formed along a 2D plane. For example, first device layer 302 can include a series of signal paths 308a (e.g., wires or other electrically conductive paths) oriented in the x-axis direction and another series of signal paths 308a oriented in the y-axis direction. Similarly, second device layer 304 can include a series of signal paths 310a oriented in the y-axis direction and another series of signal paths 310b oriented in the x-axis direction.
[0053] The different device layers 302. 304 may correspond to different wafer dies such as DRAM and ASIC dies. For example, device layers 302 and 304 may correspond to dies 204 and 202, respectively, of Fig. 1. As shown in the figure, bottom device layer 302 may be thicker than top device layer 304 to provide mechanical integrity.
[0054] The two device layers 302, 304 may be connected via one or more layer interconnects 312. In more detail. 2D signal paths of device layer 302 may be connected to 2D signal paths of device layer 304 via perpendicular layer interconnects 312 such that the overall structure 300 can be said to have 3D signal paths. The layer interconnects 312 can have a length of a few routing layer thicknesses on the dies compared to the die dimension which is many orders of magnitude larger. Using very short vertical signal paths between the memory banks and processor cores means that the parasitic capacitances associated with the signal paths are minimized.
[0055] Various commercially available 3D die stacking processes can be used to form structures that are the same as or similar to that shown in Fig. 3.
[0056] Turning to Figs. 4A and 4B, to handle large computations (e.g., large neural network computations using the methods disclosed herein), multiple processor cores may be used. Thus, as illustrated in Fig. 2. a single 3D die can have multiple interconnected processor cores. To further increase processing capacity, multiple 3D dies each having multiple processor cores may be interconnected. In such cases, it is critical that the processor cores within a die and across multiple dies be connected in such a way as to enable efficient implementation of such coordinated computations.
[0057] Fig. 4A shows an example if a ID systolic array architecture 400 having processing nodes 402a, 402b, 402c, 402d, etc. (402 generally). Fig. 4B show s an example of a 2D systolic array architecture 440 having processing nodes 442a, 442b, ..., 442p, etc. (442 generally). Both the ID and 2D architectures can effectively implement large matrix multiplications across multiple processing nodes.
[0058] Systolic arrays generally consist of ID or 2D arrays of replicated processing nodes that perform identical tasks. The signal communication paths generally include only nearest neighbor connections, keeping the paths short and efficient. For example, in ID architecture 400, node 402b is connected to neighboring node 402a and 402c. As another example, in 2D architecture 440, node 442f is connected to neighboring nodes 442b, 442j, 442e, and 442j (which may be referred to as the north, south, east, and west neighbors).
[0059] Systolic arrays are efficient in implementing matrix operations, including the neural network-weight multiplication operations disclosed herein. In addition, thecommunication bandwidth between the processor cores in matrix multiplication is much less than memory access bandwidth, and the communication paths can be implemented with relatively low-bandwidth links. The architectures also reduce the design effort because the processor cores are generally kept small, enabling a high degree of design optimization, and the entire computing structure is generated by replication. Therefore, such processor systems can be designed by relatively small teams.
[0060] The illustrated processing nodes 402, 442 may correspond to processing cores within a single 3D die, or to separate 3D dies w ithin a processor network. In other w ords, a systolic array configuration can be used to interconnect multiple cores within a 3D die and / or to interconnect multiple 3D dies within a processor network. Designing a processor network in this case can be helpful in producing efficient implementations that can work w ell with different problem sizes. Moreover, because disclosed 3D dies can achieve low- average pow er consumption, multiple 3D dies can be stacked on top of each other to achieve a small form factor.
[0061] It should be appreciated that a 2D array of 3D processor cores can be connected to form either a ID or 2D systolic array. For example, the 3x4 array illustrated in Fig. 2 can be connected as a 2D systolic array where each 3D processor core is connected to up to four (4) of its nearest neighbors. As another example, the 3x4 array illustrated in Fig. 2 can be connected as a ID systolic array where each 3D core is connected to one of its nearest neighbors, e.g., all cores within each row can be connected together serially, and the rows themselves can also be connected together serially, in a serpentine manner.
[0062] Next described are examples of systolic hardware implementations that can efficiently implement matrix-vector multiplications (e.g., for Al inference). With the described implementations. ID semi-systolic array architecture is used within the 3D processor die, and either a ID or 2D systolic array architecture can be used for connecting multiple 3D dies together. It will be understood from earlier discussions that a 3D die stack includes both a processor die and a memory' die. As previously discussed, is possible that more than one layer of memory’ is included in a 3D die stack.
[0063] Fig. 5 shows an implementation of a 3D processor die 500 having a ID semi- systolic array architecture, according to some embodiments. Processor die 500 includes aplurality of systolic processor nodes (SPNs) 502a, 502b, 502c, ..., 502n (502 generally) each connected to a control & I / O processor 504 via a broadcast bus 506 and via a control bus 508, as shown. The processor die 500 may, for example, be configured and utilized to execute the methods described in detail below with reference to Figs. 8 and 9.
[0064] Control & I / O processor 504 can include data SRAM 518 and instruction SRAM 520. SRAMs 518. 520 can be used by control & I / O processor 504 for cache memory, scratch pad, instruction cache, instruction memory, etc. In addition to SRAMs 518, 520, control & I / O processor 504 can include circuitry configured to send and / or receive data over buses 506, 508 and I / O ports 510. Moreover, control & I / O processor 504 can include circuitry configured to control all the processing, control, and timing of circuitry' within the SPNs 502, including 512, 514, and 516.
[0065] Control & I / O processor 504 has four I / O ports 510 that can be connected to up to four other 3D processor for sending and receiving data. If 3D processor die 500 is used within a processor network having a 2D systolic array architecture, the four I / O ports 510 may be connected to nearest neighbors to the north, south, east, and west. If 3D processor die 500 is used within a processor network having a ID systolic array architecture, only- two of the four I / O ports 510 may be used. Alternatively, the implementation of 3D processor die 500 may be modified to have only tw o I / O ports.
[0066] The SPNs 502 may be similar in structure and function. Representative SPN 502a includes a routing module 512 and a systolic processor core (SPC) 514. SPC 514 may be directly connected to a memory bank 516 via an individual low-power memory interface. Memory bank 516 can be provided on a separate die that is situated under 3D processor die 500, such as illustrated in Fig. 2. In other words, while the 3D processor die 500 of Fig. 5 is shown as having a plurality of memory banks, it should be understood that these memory- banks may be disposed on a separate die situated under the 3D processor die.
[0067] 3D processor die 500 may be used to perform matrix-vector multiplications associated with neural networks, for example. Control & I / O processor 504 can be configured to broadcast input vector elements to all the SPNs 502 simultaneously over the broadcast bus 506. Control & I / O processor 504 may broadcast one input vector element per clock cycle, or it may broadcast more than one input vector element per clock cycle.Control & I / O processor 504 can receive such input vectors via one or more of the I / O ports 510 and store (e g., on a temporary basis) input vectors in data SRAM 518. At each SPN 502, the routing module 512 can communicate the broadcasted input vector elements to the SPC 514. Due to such broadcasting, 3D processor die 500 can be said to have a “semi-systolic” array architecture. In some cases, control & I / O processor 504 can broadcast an input vector of size N one element at time over N successive clock cycles, as described below with reference to Figs. 8 and 9.
[0068] A top-level controller (not shown) can be connected to the systolic array that controls all the processing and handles the input / output that are going on in all the systolic array dies. Input vectors can be inputted to the 3D die stack through the ports 510 and stored in 518 before they are used or used as they come in. Output vectors can also be stored before they are outputted from the die stack or outputted as they are generated. I / O of the input and output vectors to / from the individual cores are handled through routing module 512 and is controlled by the control & I / O processor 504.
[0069] The SPC 514 can be configured to multiply the input vector elements by neural network weights read from the SPN’s individual memory' bank 516, and produce output vector elements. Each memory bank 516 may store a subset of weight matrix rows that belong to the output vector elements to be computed for the corresponding SPC 514. Various techniques may be employed to program / store weight matrix rows onto the various memory' banks. For example, each core / bank can store subset of the weight matrix rows. This can be done ahead of lime by the control & I / O processor through the control bus 508. All the programming, control, timing, etc. can happen over the control bus. Examples for techniques for distributing and storing weight matrix elements in memory banks 516 are given below with reference to Figs. 8 and 9.
[0070] The computed output vector elements can then be shifted out serially to control & I / O processor 504 through the routing module 512. In other words, over successively clock cycles, each SPN can send its output vector to its nearest neighbor (e.g., to its rightmost neighbor in Fig. 5), which in turn passes the output vector to its nearest neighbor, and so on until the output vector reaches control & I / O processor 504. While the operation of a single SPN is described, it should be appreciated that multiple SPNs 502 (or even all SPNs) can perform these steps in parallel. The output vector elements may beused as the input vector elements for the next layer in a neural network. In that case, the control & I / O processor 504 may transmit an output vector element over the broadcast bus 506 as soon as it is received without waiting for subsequent elements or the whole output vector. For example, control & I / O processor 504 may transmit the first output vector element over the broadcast bus immediately after receiving it without waiting for the second or subsequent elements. Transmitting the individual vector elements as they are received minimizes idle times for the SPNs 502 as they may perform the next neural network layer matrix-vector multiplication even if the previous layer multiplications are not completed yet.
[0071] When a single 3D die is used to implement a matrix-vector multiplication, each SPC 514 contributes to computation of specific output vector elements. Each memory bank 516 stores the corresponding weight matrix elements necessary’ for computation of the specific output vector elements.
[0072] While 3D processor die 500 can be used to perform matrix-vector multiplications associated with neural networks, the general concepts and structures described can be applied to various ty pes of compute applications. In general, an SPC 514 can include circuitry’ (e.g., an ASIC) configured to execute a set of instructions, including but not limited to various types of arithmetic instructions. The particular circuitry and instruction set / AP provided by a SPC 514 can be selected based on the type of application the 3D processor is intended to be used with. In some embodiments, all SPCs 514 on may be similarly configured to accomplish similar task. In other embodiments, different SPCs 514 may be differently configured to accomplish different tasks.
[0073] In some embodiments, one or more SPCs may implement one or more of the following matrix operations / instructions: matrix-vector multiplication, matrix-matrix multiplication, vector-vector multiplication, matrix addition, matrix subtraction, vector addition, vector subtraction, matrix element-wise product, vector element-wise product, matrix scaling, and vector scaling.
[0074] Control bus 508 is generally used for all I / O, control, timing, etc. of the SPNs 502, except for the input / output of the input vectors and output vectors which can be done through 512 and broadcast bus 506. As illustrated in Fig. 5, control bus 508 may be configured to be directly accessible to means external to 3D processor die 500. Forexample, a top-level controller (not shown) may use control bus 508 to program weight vectors into memory banks 516 via the respective SPCs 514. In the case where 3D processor die 500 is one of many processor dies in a processing network, control bus 508 may correspond to a global control bus connecting to each of the many 3D processor dies. In contrast, each of the 3D processor dies may have its own local broadcast bus 506.
[0075] Fig. 6 shows an example of an SPC architecture that may be provided within the 3D processor die of Fig. 5. For example, illustrative SPC 600 of Fig. 6 may correspond in whole or in part to any of the SPCs 514 show n in Fig. 5. For convenience, elements of Fig. 5 are referenced in the following discussion.
[0076] Illustrative SPC 600 includes one or more computer units 602a, 602b, 602c, ..., 602n (602 generally) each of which can include computational modules 604, accumulators 606, and register files 608. The computation models 604 may include, for example, a multiplier-adder, an algorithmic logic unit (ALU), a nonlinear function unit such as a Rectified Linear Unit (ReLU), and potentially other computational modules. Non-linear functions, such ReLU, can be used as activation functions in neural networks.
[0077] Input from the broadcast bus 506 is broadcasted to the one or more CUs 602, either directly as illustrated in Fig. 6 and / or indirectly via routing modules 512 as illustrated in Fig. 5. In the latter case, the routing modules 512 may simply forward the broadcast data to the CUs 602. All the CUs 602 can receive the same data or different data depending on whether they are computing different output vector elements or the same output vector element. Examples for routing the same data and examples for routing different data to the CUs 602 are given below with reference to Figs. 8 and 9.
[0078] The number of CUs 602 may be selected based on the access bandwidth of the memory bank 516 (e.g.. DRAM) associated with the SPC 600. For example, if the memory bank 516 can read sixty -four (64) bits of data per clock cycle and if a weight matrix element consists of eight (8) bits, then 8 CUs can be provided to perform eight (8) weight multiplications in parallel, provided that a CU can perform one w eight multiplication-accumulation per clock cycle. In this case, the broadcast bus 506 may be broadcasting eight (8) input vector elements per clock cycle.
[0079] The CUs 602 can be designed to support many weight matrix element and vector element formats including various integer, fixed point, floating point, and potentially other formats as needed. The design can be tailored to support a single data format or multiple data formats interchangeably.
[0080] In the embodiment of Fig. 6, SPC 600 can include an output add unit 610 configured to add outputs from multiple CUs 602 before outputting the results via the routing module 512. In other embodiments, output add unit 610 may be omitted and CU 602 outputs can pass directly to the routing module 612. The optional output add unit 610 can be provided when multiple CUs are used to compute a single output vector element, for example.
[0081] Reading, writing, and refreshing of memory bank 516 is done by the SPC 600. Since all SPCs within the same die (e.g. processor die 500) work synchronously with the same broadcasted input, the reading, writing, and refreshing of different memory banks may likewise be synchronous. In some embodiments, the computations, I / O and / or memory operations, and other functions of the SPC 600, CUs 602, and control & I / O processor 504 may be pipelined as known to the skilled person to increase throughput and minimize idle times.
[0082] Dividing up the weight matrix elements storage and output vector element computations for weight matrix times input vector can be done differently for different situations to optimize the memory’ usage, memory bandwidth, and computational resources. Several such situations are described next and also in detail below with reference to Figs. 8 and 9.
[0083] When one processor die is used and when the number of output vector elements are similar to but not greater than the number of SPNs. then each SPN may be used to compute one output vector. In this case, each SPN can store one row' of the weight matrix. Each CU can store a subset of the row elements, and computation results from multiple CUs within the same SPN can be added to compute one output vector element.
[0084] When the number of output vector elements are similar to but not greater than the total number of CUs in the die, then each CU may be used to compute one output vector element and CU output addition is not needed.
[0085] When the number of output vector elements are much greater than the number of CUs, then each CU may be used to compute multiple output vector elements, with each CU multiplying an input vector element with multiple weights belonging to multiple output vector elements utilizing multiple accumulators per CU.
[0086] When one processor die is used and when the number of output vector elements are much less than the number of SPNs, then multiple SPNs may be used to compute each output vector element, each SPN storing the subset of the weight matrix row elements and computing the corresponding sum of the partial products. In this case, the control & I / O processor 504 can be used to perform the final additions needed to compute the output vector elements, as the intermediate sums of the partial products are shifted out to the control & I / O processor 504.
[0087] When the weight matrix is large, multiple dies may be used to distribute the weight matrix storage and output vector computation. Multiple dies can be useful when the weight matrix is too large to fit within the DRAM on one die and / or when high-speed computation with low latency is required.
[0088] For example, the ID systolic array architecture shown in Fig. 4A may be used to connect multiple dies. In this case, if the number of output vector elements are similar to but not greater than the total number of SPNs across all the dies, then each SPN may be used to compute one output vector element. When the number of output vector elements are much greater than the total number of SPNs across all dies, but is similar to and not greater than the total number of CUs across all the dies, then each CU may be used to compute one output vector element, with each CU multiplying an input vector element with appropriate weights belonging to the output vector element. When the number of output vector elements are much greater than the total number of CUs across all dies, then each CU may be used to compute multiple output vector elements, with each CU multiplying an input vector element with multiple sets of weights belonging to the multiple output vector elements. In all these cases, the weight matrix may be distributed evenly (or as evenly as possible) across all the SPNs and SPCs to balance memory storage and computational throughput.
[0089] For very large weight matrices, it may be more efficient to use multiple dies connected in 2D systolic array architecture (Fig. 4B) to distribute the storage of weightmatrix elements and output vector computation. In this case, each weight matrix row can be distributed to multiple rows of the 2D die array within the same column. The 2D matrix weight distribution can result in the same number of weights per die as ID distribution, and the computational throughput requirement per die can also be the same. However, the data flow may possibly be faster or easier in some cases by enabling parallel input vector element distribution for each subset of weight matrix columns stored in different die rows. The sum of the partial products from different rows of the dies can be added to compute the output vector elements in this case when the weight matrix row elements belonging to one output vector element are distributed to multiple rows of the dies.
[0090] Turning to Figs. 7A and 7B, while ID and 2D systolic array architecture provides many design and performance advantages, embodiment of the present disclosure can achieve further improvement by utilizing the ID and 2D toroidal systolic array architectures. For example, the general 3D processor die concept illustrated in Fig. 5 may be adapted to use a ID toroidal array. Fig. 7A shows an example of a ID toroidal array architecture 700 having processing nodes 702a, 702b, 702c, 702d, etc. (702 generally). Fig. 7B shows an example of a 2D toroidal array architecture 740 having processing nodes 742a, 742b, ..., 742p, etc. (742 generally). The toroidal systolic array allows for more general data flow and makes input vector element and output vector element communications easier.
[0091] Fig. 8 is a flow diagram of a computer-implemented method 800 for parallelized computation of a product of a matrix and an input vector. Method 800 may. for example, be executed by a system that includes a plurality of processing cores, a plurality of memory’ devices, and a controller, as described above. This system may, for example, include one or more 3D processor dies as disclosed above with reference to Figs. 5 and 6. Each processing core has an accumulator, and each memory device is coupled to and associated with at least one of the plurality of processing cores. The controller is also coupled to the plurality of processing cores. In some examples, the system executing method 800 may be a systolic or semi-systolic processor array as shown and described in detail above w ith reference to Figs. 5 and 6, including a plurality’ of processing cores 514, a plurality of memory devices 516, and a controller 504. However, it is also expressly noted that method 800 may be executed on any other architecture that includes a plurality of processing cores, a plurality of memory devices, and a controller.
[0092] The method 800 allows the system to efficiently and advantageously distribute elements of the matrix to the processing cores 514 and associated memory devices 516 and therefore, as shown above, to improve processing speed and reduce power consumption of the system. To this end, the matrix may be a weight matrix associated with a neural network, and the matrix elements may be weights associated with an artificial neuron. A variety of examples for performing a multiplication of a matrix, such as a weight matrix, with a vector, such as an input vector for an artificial neuron, are now disclosed. While the terms ‘'weight matrix" and “input vector” are used throughout this disclosure, it is expressly contemplated that method 800 may be used to perform any multiplication of a matrix and a vector and is not limited to artificial neurons and / or neural networks. It is also noted that the matrix-vector multiplication systems and methods disclosed herein are not limited to the examples below but may be generalized by the skilled person based on these examples.
[0093] For a matrix-vector multiplication in a one-dimensional systolic array, the weight matrix is distributed evenly across all the cores. An example for a multiplication is y = Wx where y is the output vector, W is the weight matrix, and x is the input vector. In order to perform the matrix-vector multiplication, the system needs to compute y (k) =li=i w(k, I) x %(t) for k = 1. 2, 3, . . . , M, where N is the size of the input vector x, and M is the size of the output vector y . In some examples, the system includes 1024 processing cores 514 (P = 1024, where / 5is the number of processing cores), and there are 1024 elements in the input vector (N = 1024), and 1024 elements in the output vector (M = 1024). The weight matrix W thus has the dimension of M rows by N columns. Illustratively, one row of the weight matrix W may be stored in each one of the f systolic processing cores 514. The k-th core 514 may store the matrix elements w(k, z) for i = 1, 2, 3. ... . N. The controller 504 may instruct the processing cores 514 to store certain elements and transmit the respective matrix elements to the respective processing cores. The processing cores 514 may then store those elements received from the controller in a memory device 516 associated with the respective processing core. It is also assumed that the systolic array allows broadcasting ofdata, which means that the array is actually a semi-systolic array rather than a strictly systolic array. One element x(z) is then broadcast to all cores 514 per clock cycle, from i = 1 to i = N successively over N clock cycles. The broadcasting may, for example, be initiated and controlled by controller 504. During each clock cycle, each core 514 receives the input vector element x(i), multiplies the vector element by the corresponding weight value w(k. z), as received from the controller 504 and stored in an associated memory device 516, and accumulates the result in accumulator 606. Computation of one output vector element y(k) therefore requires N clock cycles. Since M = P and AT multiplications are performed in parallel on processing cores 514, the multiplication of the entire matrix W with input vector x is also performed in N clock cycles. The output vector y can be shifted out serially, for example controlled by controller 504, or used (e.g., broadcast by controller 504) as an input for a subsequent matrix-vector multiplication. In some embodiments, the computations and data flow shown for this example and for all other examples in this disclosure may be pipelined, as known to the skilled person, to increase computational throughput and / or clock speed.
[0094] In another example, each processing core 514 stores and processes multiple rows of the weight matrix. For example, if N= 1024 and M= 2048 and the system includes P = 1024 processing cores 514, then each processing core 514 may store and process two rows of weight matrix IF. The matrix elements w(2 / - 1. z) and w(2j, z) for z = 1,2,3... , N may be stored in the / -th processing core 514, where j ' = 1, 2 ,... , y. Each processing core 514 may multiply each broadcasted input vector element with two weights, one from each w eight matrix row, and may accumulate the results of the multiplication in two separate accumulators 606. While two rows per core are shown in this example, the concept can be generalized to any number of rows per core.
[0095] In another example, each processing core 514 may receive more than one weight matrix element per clock cycle. This allows the system to broadcast multiple input vector elements per clock cycle and therefore speed up the computation. For example, if each processing core 514 receives eight weight matrix elements per clock cycle, the controller 504 may broadcast eight input vector elements per clock cycle. The computation of the matrix-vector product is therefore sped up by a factor of eight.
[0096] In another example, computation time may also be reduced in a onedimensional systolic array that includes more processing cores 514 than the number of rows of the weight matrix W. For example, the dimensions of weight matrix W may be 512 x 512 (N= 512, M= 512), and the processing system may include 1024 processing cores (P = 1024 = 2M). The controller 504 may then instruct each processing core 514 to store N / 2 = 256 weights. The w(k, i) for z= 1.2,... 2 may be stored in the (2k -l)-th core, and the w . .. , N may be stored in the 2Zc-th core. The controller 504 maybroadcast two elements of input vector x per clock cycle. The elements x(z') +for 1 < i < would be broadcasted substantially simultaneously, one element to the first set of AT processing cores 514 that store the first half of the columns of the weight matrix, and the other vector element to the second set of Al processing cores 514 that store the second half of the columns of the weight matrix. At the end of computation, the accumulator values from the first set of cores 514 and the second set of cores 514 are summed, for example by the controller 504 or by the processing cores 514, to generate Al output elements. Even in a systolic array this is an easy process, because the two accumulator values to be added are found in adjacent processing cores.
[0097] In a similar example, a 2D systolic array may be used. Illustratively, the first set of Al processing cores 514 may be arranged in a first row of the 2D systolic array, and a second set of Al processing cores 514 may be arranged in a second row of the 2D systolic array. The controller 504 may broadcast the input vector elements can separately to each array row, with the first set x(i) for i = 1, 2,... , N / 2 being broadcast to the first row of the array one vector element at a time, and the second set of x(z) for z = y+1, y+2, . . . , N being broadcast to the second row of the array one vector element at a time. The resulting accumulator values can be communicated vertically in the array and summed, either by the controller 504 or by the processing cores 514, to produce the output vector elements y(k).
[0098] In a further example, the method 800 described herein for parallelizing matrixvector multiplications may be used to handle very large matrices across many processing units 514 in 2D systolic arrays of dies. Multiple dies are required to provide sufficient memory' capacity as well as to speed up computation. Illustratively, 2D arrays of processing unit dies in a systolic array configuration may be used. For example, a 4096 x4096 weight matrix (N = 4096, M = 4096) may be multiplied with a 4096-element input vector x. Each processor die may include a one-dimensional 1024-element systolic array of processing cores 514 (P = 1024). 16 processor dies may be arranged in 4 rows x 4 columns in a two-dimensional systolic array configuration (R = 4 and C = 4, where R is the number of rows and C is the number of columns of the die array). Each processing NM core 514 of each die may store — = 1024 weight matrix elements. More specifically, the following matrix elements may be stored in a processor die in the r-th row and c-th column:Each processing core 514 within the processor die stores w(k.i) for a single value of k that is in the appropriate range for the die, and i = — - - 1- j where j = 1, 2, ■■■, -■During each clock cycle, the controller 504 may broadcast one input vector element i =W(~r 1')+ j to the r-th row of the processor dies. The processing cores within each die then multiply the received vector element with the appropriate matrix element w(k, i). This multiplication is performed substantially simultaneously in all the rows of processor dies. After N / R clock cycles of computation, all R accumulator values (one in each row of processor dies) that belong to the same output vector element are summed together to yield an output vector element value y(k). This example may further be generalized to any number of rows and / or columns of the processor dies and any number of processing cores per processor die and allows the processing system to handle large matrix-vector multiplications computations efficiently.
[0099] In addition, the processing cores 514 within each processor die may be arranged in various ID or 2D systolic array architectures. If the output vector y(k) is required as an input vector to the next layer of neural netw ork computation, the controller 504 may broadcast its elements to the rows of the systolic arrays of the dies or process it in any other way known to the skilled person.
[0100] In step 810, a processing core 514 receives instructions from the controller 504 to store a plurality of matrix elements of the matrix in at least one memory device 516 associated with the processing core 514. The processing core 514 then also stores the plurality of matrix elements in the at least one memory device 516, as instructed by the controller 504. The controller 504 may select the matrix elements to be stored in the respective processing cores 514 as described in the various examples above.
[0101] Illustratively, each processing core 514 executes steps 820 through 840 during a single clock cycle. However, it is also contemplated that steps 820 through 840 may be executed during more than one clock cycle. It is further contemplated that steps 802 through 840 may be pipelined for higher efficiency, as known to the skilled person.
[0102] In step 820, the processing core 514 receives at least one vector element of the input vector from the controller 504. Again, the controller 504 may select the vector element or vector elements to transmit to the respective processing cores 514 as described in detail the various examples above and / or the generalized versions of those examples.
[0103] In step 830, the processing core 514 multiplies the at least one vector element with at least one matrix element stored in the at least one memory device 516 associated with the processing core 514. The multiplication may, for example, be performed in a CU 602 as shown above with reference to Fig. 6.
[0104] In step 840, the processing core 514 adds the result of the multiplication to a value of the accumulator 606 and stores the result of the addition in the accumulator 606. Illustratively, the value of the accumulator 606 may have been reset to zero before the start of the matrix-vector multiplication. In other examples, the value of the accumulator may not have bene reset or may have been set to a different value before the start of the matrixvector multiplication.
[0105] In step 850, the processing core 514 repeats steps 820 through 840 during at least a second clock cycle or a second plurality of clock cycles.
[0106] In step 860, the processing core 514 transmits the value of the accumulator 606 as an output. The processing core 514 may transmit the value of the accumulator 606 tothe controller 504, or it may transmit the value of the accumulator 606 to one or more different processing cores 514.
[0107] Fig. 9 is. is a flow diagram of a computer-implemented method 900 for parallelized computation of a product of a matrix and an input vector. Method 900 may, for example, be executed by a system that includes a plurality of processing cores, a plurality of memory devices, and a controller, as described above. Each processing core has an accumulator, and each memory device is coupled to and associated with at least one of the plurality' of processing cores. The controller is also coupled to the plurality of processing cores. In some examples, the system executing method 900 may be a systolic or semi-systolic processor array as shown and described in detail above with reference to Figs. 5 and 6, including a plurality of processing cores 514, a plurality of memory devices 516, and a controller 504. However, it is also expressly noted that method 900 may be executed on any other architecture that includes a plurality of processing cores, a plurality of memory devices, and a controller.
[0108] In step 910, the controller 504 determines a distribution of matrix elements to the plurality of processing cores 514. The controller 504 may select the matrix elements to be distributed to the respective processing cores 514 as described in detail in the various examples above and / or the generalized versions of those examples. The controller 504 may determine the distribution of the matrix elements at run time of the matrix-vector multiplication. In other words, the controller 504 may receive an instruction to multiply a matrix with a vector, determine the size of the matrix and the size of vector and then determine the distribution of matrix elements to the processing cores 514 based on the size of the matrix, the size of the vector, and the number and / or arrangement of available processing cores 514 as described above. In other embodiments, the controller 504 may receive a predetermined distribution of matrix elements to processing cores 514 from another component. For example, the controller 504 may receive a distribution of matrix elements to processing cores from the executable code that includes the matrix-vector multiplication. In this example, a compiler may have determined the distribution at compile time and may have embedded a representation of the distribution in the executable code for the multiplication. The controller 504 may be configured to accept the distribution that has been predetermined by the compiler and follow it. In other examples, the controller 504 may be configured to use the predetermined distribution to determine itsown, possibly improved, distribution. For example, if the dimensions of the matrix and vector and the number and arrangement of processing cores given to the compiler at compile time matches the dimensions of the matrix and vector and the available number and arrangement of processing cores at run time when the multiplication is executed, the controller 504 may determine the distribution of matrix elements to processing cores bycopying the predetermined distribution made by the compiler at compile time. If the matrix dimensions and / or number and arrangement of processing cores do not match between compile time and run time, the controller 504 may determine a distribution that is different from the one determined by the compiler. In yet other embodiments, the controller 504 may terminate execution of the multiplication if the matrix dimensions and / or number and arrangement of processing cores do not match, and / or raise an exception or other signal.
[0109] If a compiler is being used to determine the distribution of matrix elements to the processing cores 514, it is critical to have an appropriate application programming interface (API) in the programming language being used. The programming language may be any suitable language known to the skilled person. While a compiler is mentioned herein, it is expressly noted that the programming language may also be an interpreted language utilizing an interpreter for execution instead of a compiler.
[0110] For the systolic-array -based matrix-vector computations disclosed herein, a matrix operation-based format would likely yield the most natural and efficient API. Other matrix computations can also be supported by the matrix operation-based API. For example, a matrix-matrix multiplication can consist of multiple matrix-vector operations. Matrix or vector addition and subtraction operations can also be supported with systolic array architectures. In addition, matrix or vector scaling and specialized Al operations such as element-wise nonlinear operations can also be supported with a matrix-based API.
[0111] Table 1 shows some examples of matrix-based API operations. The variables A, B, and C are matrices, and x, y, and z are vectors. The variable s is a scalar. The dimensionalities of matrices and vectors have to match in order for each operation to be valid.Table 1Matrix-Based API Examples
[0112] Based on the matrix-based Al operations specified with this matrix-based API, the compiler maps the matrices, such as weight matrices, to the memory devices of the systolic array processing units and determines the sequence of computation operations and data flow. The matrix-to-memory device mapping is relatively straightforward, using the examples described in detail above, and can easily be automated for the compiler design. Combinations and variations of those methods may also be appropriate. The optimal memory' mapping and data flow may depend on the systolic array configuration and the problem size (i. e.. the weight matrix size). If the problem size is small, the compiler and / or controller may not utilize all the processing cores in the array, but when the problem is large, all processing cores should be utilized for high performance and efficiency. The memory' mapping may also take into account any necessary routing of vector elements and output elements to and from the processing units and / or memory devices to make that routing more efficient. Once the compiler and / or controller has determined the memory mapping for the weight matrices, planning of the broadcasting of input vector elements and routing of the output vector elements are relatively straightforward given the systolic array structures and memory' mapping of the weight matrices.
[0113] In step 920, the processing core 514 receives instructions from the controller 504 to store a plurality of matrix elements of the matrix in at least one memory device 16associated with the processing core 514, as described above with reference to step 810 of Fig. 8.
[0114] In step 930, the processing core 514 receives at least one vector element of the input vector from the controller 504 as described above with reference to step 820 of Fig. 8.
[0115] In step 940, the processing core 514 multiplies the at least one vector element with at least one matrix element stored in the at least one memory device 516 associated with the processing core 514, as described above with reference to step 830 of Fig. 8
[0116] In step 950, the processing core 514 adds the result of the multiplication to a value of the accumulator 606 and stores the result of the addition in the accumulator 606, as described above with reference to step 840 of Fig. 8.
[0117] In step 960, the processing core 514 repeats steps 930 through 950 during at least a second clock cycle or a second plurality of clock cycles.
[0118] In step 970, the processing core 514 transmits the value of the accumulator 606 as an output, as described above with reference to step 860 of Fig. 8.
[0119] As used herein and in the claims, ordinal identifiers for process or method steps such as “(a)”, “(b)”. “first,” “second,” etc., do not by itself connote any priority. precedence, or order of one claim element over another or the temporal order in which acts of a method are performed, but are used merely as labels to distinguish one claim element from another element.
[0120] As used herein, the terms “processor” and “controller” are used to describe electronic circuitty7that performs a function, an operation, or a sequence of operations. The function, operation, or sequence of operations can be hard coded into the electronic circuit or soft coded by way of instructions held in a memory- device. The function, operation, or sequence of operations can be performed using digital values or using analog signals. In some embodiments, the processor or controller can be embodied in an application specific integrated circuit (ASIC), which can be an analog ASIC or a digital ASIC, in a microprocessor with associated program memory, in a digital signal processor (DSP), and / or in a discrete electronic circuit, which can be analog or digital. A processoror controller can include internal processors or modules that perform portions of the function, operation, or sequence of operations. Similarly, a module can include internal processors or internal modules that perform portions of the function, operation, or sequence of operations of the module. A single processor or other unit may fulfill the functions of several means recited in the claims.
[0121] As used in the claims or elsewhere herein, the term '‘comprising’7does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality7. The term “set” includes at least one member.
[0122] As used herein, the term “predetermined.” when referring to a value or signal, is used to refer to a value or signal that is set, or fixed, in the factory at the time of manufacture, or by external means, e.g., programming, thereafter. As used herein, the term “determined,” when referring to a value or signal, is used to refer to a value or signal that is identified by a circuit during operation, after manufacture.
[0123] Various embodiments of the concepts systems and techniques are described herein with reference to the related drawings. Alternative embodiments can be devised without departing from the scope of the described concepts. It is noted that various connections and positional relationships (e.g., over, below, adjacent, etc.) are set forth between elements in the claims, detailed description, and drawings. These connections and / or positional relationships, unless specified otherwise, can be direct or indirect, and the claimed inventions are not intended to be limiting in this respect. Accordingly, a coupling / connection of entities can refer to either a direct or an indirect coupling / connection, and a positional relationship between entities can be a direct or indirect positional relationship. As an example of an indirect positional relationship, references in the present description to element or structure A coupled / connected to element or structure B include situations in which one or more intermediate elements or structures (e.g., element C) is provided between elements A and B regardless of whether the characteristics and functionalities of elements A and / or B are substantially changed by the intermediate element(s).
[0124] Furthermore, it should be appreciated that relative, directional or reference terms (e.g. such as “above,” “below,” “left,” '‘right,” “top,” “bottom,” “vertical,” “horizontal,” “front,” “back,” “rearward,” “forward,” etc.) and derivatives thereof are usedonly to promote clarity in the description of the figures. Such terms are not intended as, and should not be construed as. limiting. Such terms may simply be used to facilitate discussion of the drawings and may be used, where applicable, to promote clarity of description when dealing with relative relationships, particularly with respect to the illustrated embodiments. Such terms are not, however, intended to imply absolute relationships, positions, and / or orientations. For example, with respect to an object or structure, an “upper’ or “top” surface can become a “lower” or “bottom” surface simply by turning the object over. Nevertheless, it is still the same surface and the object remains the same.
[0125] The terms “disposed over.” “overlying,” “atop,” “on top.” “positioned on” or “positioned atop” mean that a first element, such as a first structure, is present on a second element, such as a second structure, where intervening elements or structures (such as an interface structure) may or may not be present between the first element and the second element. The term “direct contact” means that a first element, such as a first structure, and a second element, such as a second structure, are connected without any intermediary elements or structures between the interface of the two elements. The term “connection” can include an indirect connection and a direct connection.
[0126] In the foregoing detailed description, various features are grouped together in one or more individual embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that each claim requires more features than are expressly recited therein. Rather, inventive aspects may lie in less than all features of each disclosed embodiment.
[0127] References in the disclosure to “one embodiment,” “an embodiment,” “some embodiments,” or variants of such phrases indicate that the embodiment(s) described can include a particular feature, structure, or characteristic, but every embodiment can include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment(s). Further, when a particular feature, structure, or characteristic is described in connection knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0128] The disclosed subject mater is not limited in its application to the details of construction and to the arrangements of the components set forth in the detailed description or illustrated in the drawings. The disclosed subject mater is capable of other embodiments and of being practiced and carried out in various ways. As such, those skilled in the art will appreciate that the conception, upon which this disclosure is based, may readily be utilized as a basis for the designing of other structures, methods, and systems for carrying out the several purposes of the disclosed subject mater. Therefore, the claims should be regarded as including such equivalent constructions insofar as they do not depart from the spirit and scope of the disclosed subject mater.
[0129] Although the disclosed subject mater has been described and illustrated in the foregoing exemplary embodiments, it is understood that the present disclosure has been made only by way of example, and that numerous changes in the details of implementation of the disclosed subject mater may be made without departing from the spirit and scope of the disclosed subject mater.
[0130] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.
[0131] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to obtain an advantage.
[0132] Any reference signs in the claims should not be construed as limiting the scope.
[0133] All publications and references cited herein are expressly incorporated herein by reference in their entirety.
Claims
CLAIMS1. A system for parallelized computation of a product of a matrix and an input vector, the system comprising: a plurality of processing cores, each processing core having at least one accumulator: a plurality^ of memory7devices, each memory7device coupled to and associated with at least one of the plurality7of processing cores; and a controller coupled to the plurality of processing cores, wherein each processing core of the plurality of processing cores is configured to:(a) receive from the controller to store a plurality of matrix elements of the matrix in at least one memory7device associated with the processing core;(b) during a first set of clock cycles, receive at least one vector element of the input vector from the controller;(c) during the first set of clock cycles, multiply the at least one vector element with at least one matrix element stored in the at least one memory device associated with the processing core;(d) during the first set of clock cycles, add the result of the multiplication to a value of the at least one accumulator and store the result of the addition in the at least one accumulator;(e) repeat steps (b) through (d) during at least a second set of clock cycles; and(f) transmit the value of the at least one accumulator as an output to the controller.
2. The system of claim 1, wherein each memory device is coupled to and associated with one processing core.
3. The system of claim 1. wherein the plurality of processing cores forms one of a onedimensional semi-systolic array and a two-dimensional semi-systolic array.
4. The system of claim 1, yvherein each processing core is configured to receive one input vector element per clock cycle from the controller.
5. The system of claim 1, wherein the controller is configured to store one of a row of the matrix and a column of the matrix in each processing core.
6. The system of claim 1. wherein the controller is configured to store a plurality of rows of the matrix in each processing core.
7. The system of claim 1, wherein the controller is configured to store one of a portion of the columns of the matrix and a portion of the rows of the matrix in each processing core.
8. The system of claim 3, wherein the semi-systolic array is two-dimensional and arranged in rows and columns, and wherein each processing core in a row of the two- dimensional semi-systolic array is configured to receive half of the input vector elements from the controller.
9. The system of claim 8, wherein the controller is configured to sum the outputs of the processing cores in a column of the two-dimensional semi-systolic array to result in at least one element of the product.
10. The system of claim 1, wherein the matrix is a weight matrix associated with a neural network.
11. A computer-implemented method for parallelized computation of a product of a matrix and an input vector on a plurality of processing cores, each processing core having at least one accumulator, a plurality of memory devices, each memory device coupled to and associated with at least one of the plurality of processing cores, and a controller coupled to the plurality of processing cores, the method comprising:(a) receiving instructions from the controller and storing, by a processing core of the plurality of processing cores, a plurality of matrix elements of the matrix in at least one memory device associated with the processing core;(b) during a first set of clock cycles, receiving, by the processing core and from the controller, at least one vector element of the input vector;(c) during the first set of clock cycles, multiplying, by the processing core, the at least one vector element with at least one matrix element stored in the at least one memory device associated with the processing core;(d) during the first set of clock cycles, adding, by the processing core, the result of the multiplication to a value of the at least one accumulator and storing the result of the addition in the at least one accumulator;(e) repeating, by the processing core, steps (b) through (d) during at least a second set of clock cycles; and(f) transmitting, by the processing core and to the controller, the value of the at least one accumulator as an output.
12. The method of claim 11, wherein the processing core stores one of a row' and a column of the matrix and receives one input vector element per set of clock cycles, and wherein the output is an element of the product.
13. The method of claim 11, wherein each memory device is coupled to and associated with one processing core.
14. The method of claim 11, wherein each processing core stores one of a portion of the columns of the matrix and a portion of the rows of the matrix, and wherein the method further comprises summing, by the controller, the outputs of more than one processing core of the plurality of processing cores to result in an element of the product.
15. The method of claim 11, wherein the plurality of processing cores is arranged in one of a systolic and a semi-systolic array and wherein the array is one of onedimensional and two-dimensional.
16. The method of claim 11, further comprising determining, by the controller, a distribution of matrix elements to the plurality of processing cores.
17. The method of claim 17, wherein the distribution of matrix elements to the plurality of processing cores is determined by a compiler and provided to the controller.
18. The method of claim 17, wherein the distribution of matrix elements to the plurality of processing cores is determined at run time by the controller.
19. The method of claim 17, further comprising determining, by at least one of the controller and a compiler, a sequence of computation operations and data flow.
20. The method of claim 11, wherein the matrix is a weight matrix associated with a neural network.
Citation Information
Patent Citations
Normalization unit for signed operands
US20190339943A1