Processor architecture supporting in-memory matrix operation based on static memory and processor
By introducing SRAM-based in-memory computing modules into the processor architecture, it supports in-memory matrix computing, and solves the performance and energy efficiency shortcomings of existing accelerator hardware, achieving more efficient computing and reducing energy consumption.
Patent Information
- Application Number
- CN202510314035.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-27
AI Technical Summary
Existing accelerator hardware such as GPU and TPU have high power consumption and relatively limited computing energy efficiency in performance, which is difficult to meet the high-energy-efficiency needs of artificial intelligence generative models in large-scale inference deployment.
In-memory (CIM) module based on static random access memory (SRAM) is introduced into the processor architecture as a matrix computing unit, and a two-dimensional pulsating array is formed through the pulsating data path, and the static memory array is used to support in-memory matrix operations.
It improves the overall performance of the processor, can better meet the energy-efficient needs of generative model inference and training, significantly reduces energy consumption, and improves computing efficiency.
Smart Images

Figure CN120216453A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of chip technology, and in particular, to a processor architecture and a processor that support in-memory matrix operations based on static memory. Background Art
[0002] Generative models are a type of machine learning models that are widely used in various fields because of their excellent performance in generating various modal contents. Currently, graphics processing units (GPUs) and tensor processing units (TPUs) are the main hardware platforms for generative model inference and training, which have large-scale parallel computing components. However, these processors usually have problems such as high power consumption and relatively limited computing energy efficiency in terms of performance. Therefore, cross-level innovation is urgently needed to improve performance. Summary of the Invention
[0003] Embodiments of this application provide a processor architecture, a processor, a chip, and an electronic device that support in-memory matrix operations based on static memory, so as to solve or at least partially solve the problems mentioned in the above background art. Among them,
[0004] In the first embodiment, this application provides a processor architecture that supports in-memory matrix operations based on a static processor. The processor architecture includes:
[0005] A tensor core module, where the tensor core module includes a plurality of matrix operation units; among them,
[0006] The matrix operation unit has a plurality of in-memory computing macro units;
[0007] The plurality of in-memory computing macro units are connected by a systolic data path to form a two-dimensional systolic array, and each in-memory computing macro unit includes a plurality of parallel memory banks, and the memory banks include a static memory array that supports in-memory matrix operations.
[0008] In the second embodiment, this application provides a processor. The processor has the processor architecture provided in the first embodiment above.
[0009] In the third embodiment, this application provides a chip. The chip includes the processor provided in the second embodiment above.
[0010] In the fourth embodiment, this application provides an electronic device. The electronic device includes the chip provided in the third embodiment above.
[0011] In the technical solutions provided by the embodiments of the present application, the tensor core module on the processor includes multiple matrix operation units, and each matrix operation unit has multiple in-memory computing macro units. These multiple in-memory computing macro units are connected by a systolic data path to form a two-dimensional systolic array. Each in-memory computing macro unit includes multiple parallel memory banks, and each memory bank includes a static memory array that supports in-memory matrix operations. It can be seen that the processor provided by this solution introduces in-memory computing macro units based on static memory as matrix operation units to support in-memory matrix operations. Using such matrix operation units is beneficial to improving the overall performance of the processor and can better meet the high energy efficiency requirements for generative model inference and training. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0013] Figure 1 is a schematic structural diagram of a processor architecture provided by an exemplary embodiment of the present application;
[0014] Figure 2a is a schematic structural diagram of the operation matrix unit (CIM-MXU) in the processor architecture provided by an exemplary embodiment of the present application and the structure of the in-memory computing macro unit (CIM macro unit) in the operation matrix unit;
[0015] Figure 2b is a column storage block in a certain row of the static memory array provided by an exemplary embodiment of the present application;
[0016] Figure 3 is an example diagram of the mapping engine of the present application mapping the model data of the generative model to the corresponding cache for storage;
[0017] Figure 4 is a bar chart comparing the inference latency and energy consumption of the generative model using the processor architecture provided by the present application and the benchmark processor architecture in an embodiment of the present application;
[0018] Figure 5 is a bar chart of the inference latency and energy consumption of the processor architecture with different matrix operation units (CIM-MXU) in an embodiment of the present application;
[0019] Figure 6 is a bar chart comparing the inference throughput of the processor for the generative model under different numbers in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] Artificial intelligence generative models (i.e., generative models) are a type of machine learning models. For example, large language models (LLMs) and diffusion models (DMs) are generative models, which have excellent performance in content generation across various modalities. For instance, LLMs dominate in natural language processing (NLP) tasks, driving the development of popular products such as ChatGPT (a language model based on the Transformer architecture). DMs have also achieved leading performance in image and video generation, such as various generation tools like DALL-E 3, Stable Diffusion by Stability AI, and OpenSORA. Thus, the growing demand for generative models highlights the necessity of designing high-performance acceleration hardware to support model deployment. Currently, mainstream acceleration hardware such as GPUs and TPUs are the main hardware platforms for generative model inference and training, with large-scale parallel computing components (such as the matrix multiply unit (MXU) in TPUs). However, these mainstream acceleration hardware usually consume more than 350W of thermal design power (TDP), so cross-layer innovation is needed to improve computational efficiency.
[0021] Compute-in-Memory (CIM) based on static random access memory (SRAM) is an emerging computational circuit paradigm that exhibits excellent energy efficiency and computational density. During the development of CIM technology, the initial array-level CIM macro performance was relatively limited, usually not exceeding 500 GOPS. With technological progress and design evolution, multiple CIM macro modules have been integrated into larger-scale designs, achieving a computational capacity of over 1 TOPS. In recent years, CIM-based artificial intelligence chips have emerged one after another, further demonstrating the feasibility of improving CIM performance to meet the requirements of high-performance chips. However, there is still a significant performance gap between current CIM designs and GPUs or TPUs, etc. For example, some existing high-performance computing GPUs (such as the NVIDIA A100 GPU) achieve a performance of 312 TFLOPS at BF16 precision, and some high-performance computing TPUs designed specifically for deep learning tasks (such as TPUv4) have a performance of 275 TFLOPS at BF16 precision, far exceeding existing CIM-based GPU and TPU solutions.
[0022] Therefore, the existing accelerator hardware such as GPUs and TUPs mainly has the following two key problems:
[0023] 1) Currently, the mainstream accelerator hardware, such as NVIDIA GPUs and TPUv4, etc., has relatively limited computing energy efficiency and is difficult to meet the high energy efficiency requirements of artificial intelligence generative models in large-scale inference deployment.
[0024] 2) Currently, the scale of artificial intelligence chip designs based on CIM is small, and the computing power has not reached the level of current mainstream artificial intelligence acceleration chips. There is an urgent need to conduct further research on large computing power and large-scale CIM designs.
[0025] To solve the above problems, the embodiments of the present application provide a solution. The basic idea is: introducing a CIM computing module based on SRAM as a matrix operation unit in a processor architecture (such as the TPU architecture).
[0026] To make the purpose, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0027] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the drawings.
[0028] First, the vocabulary involved in the embodiments of the present application will be described. It can be understood that this description is for a clearer understanding of the embodiments of the present application and does not necessarily constitute a limitation on the embodiments of the present application.
[0029] Matrix operation unit: It is a hardware acceleration unit for processing matrix operations. Among them, the matrix operations processed in the present application mainly include matrix multiply-accumulate.
[0030] Compute-In-Memory (CIM): It is a technology aimed at solving the data processing bottleneck problem in the traditional von Neumann architecture. CIM mainly reduces the need for data movement by directly performing computing operations inside the storage unit, thereby significantly improving energy efficiency and speed. Thus, that is: CIM technology is computing in the storage unit, and the storage unit can be but is not limited to any one of the following types of memories: Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Magnetoresistive Random Access Memory (MRAM).
[0031] SRAM array: It refers to a storage array composed of multiple SRAMs arranged in a two-dimensional matrix form. Each SRAM memory cell is composed of several transistors and is used to store one bit of data.
[0032] Core: It is a module / unit that can implement specific functions, such as a hardware module / unit.
[0033] Tensor Core: It is a computing unit specific to processors such as TPU, mainly used to accelerate matrix operations, especially for tensor calculations in deep learning and machine learning.
[0034] Systolic Array: It is a hardware architecture specifically designed for efficient execution of matrix operations, especially suitable for Multiply-Accumulate (MAC) operations. A systolic array usually consists of multiple Processing Elements (PEs), and each PE can perform basic MAC operations. The MAC (Multiply-Accumulate) operation is an operation that multiplies two numerical values and accumulates the result into an accumulator. Thus, a systolic array can also be called a PE array.
[0035] Systolic data path: It improves computing efficiency through pipelining and parallel processing. It usually consists of multiple Processing Elements (PEs), which transfer data in a pipelined manner and perform local calculations in each clock cycle.
[0036] Common Memory (CMEM): It refers to the cache shared among processor cores, used to store recently accessed data and instructions to reduce the access frequency to the main memory. The common memory can be caches at levels such as L2 and L3, depending on the design.
[0037] Vector Memory (VMEM): It usually refers to the cache or storage area specifically for the Vector Processing Unit (VPU). The vector processing unit VPU is a hardware component designed for executing large-scale parallel computing tasks and is commonly found in devices such as graphics processing units (GPUs) and AI accelerators.
[0038] Figure 1Schematic diagram of a processor architecture that supports in-memory matrix operations based on static memory provided by this application. This processor can be implemented by introducing a CIM computing module based on static memory such as SRAM as a matrix operation unit in a processor architecture such as a TPU (such as TPUv4i, etc.) or a GPU, and a TensorCore module is adopted in the processor architecture for generative model inference calculations. Among them, the matrix operation unit MXU implemented by the above SRAM-based CIM computing module can be denoted as SRAM CIM-MXU or CIM-MXU. Among them, the CIM computing module includes multiple CIM macro units, and the CIM macro units are based on SRAM.
[0039] Specifically, as shown in Figure 1 The processor architecture includes: a TensorCore module 1 for inference calculations of machine learning models, and the machine learning model includes a generative model; this TensorCore module is a key component for accelerating matrix operations in generative models. Among them,
[0040] The TensorCore module includes multiple matrix operation units 11 (i.e., Figure 1 the CIM-MXU shown in ). The matrix operation unit 11 has multiple in-memory computing macro units (i.e., CIM macro units). The multiple in-memory computing macro units are connected by a systolic data path to form a two-dimensional systolic array, and each of the memory computing macro units includes a static memory array. The static memory is preferably SRAM.
[0041] The reason for this application to introduce the matrix operation unit 11 with the above structural form in the processor architecture is as follows:
[0042] In traditional processor architectures such as TPU, ordinary MXUs are often used, such as MXUs based on systolic arrays. The MXU based on systolic arrays uses a systolic array of 128×128 multiply-accumulate (MAC) units for large-scale general matrix multiplication (General Matrix-Matrix Multiplication, GEMM). However, the power consumption of such MXUs is relatively high and it is difficult to meet the high-performance requirements of generative models in large-scale inference deployments. Therefore, in this application's solution, the above SRAM CIM-based MXU (CIM-MXU) is used to replace the ordinary MXU (such as the MXU based on systolic arrays) to improve the computing efficiency, such as accelerating GEMM and GEMV (General Matrix Vector Multiply) operations.
[0043] As can be learned from the above content, the main computational effect that the above matrix operation unit 11 needs to achieve is to accelerate GEMM and GEMV operations. To achieve this effect, multiple CIM macro cells in the matrix operation unit 11 need to be organized into a CIM-MXU architecture that can handle the parallelism of GEMM and GEMV calculations. For this purpose, in this application, a systolic data path is used to connect multiple CIM macro cells to form a two-dimensional systolic matrix, and this two-dimensional systolic matrix is a CIM-MXU (matrix operation unit 11).
[0044] Figure 2a An example of a CIM-MXU (which is the matrix operation unit 11) given in this application is shown.
[0045] As shown in the example of FIG. 2, when implementing the CIM-MXU, its top layer is composed of multiple CIM macro cells such as X×Y (such as 16×8), forming a two-dimensional systolic array. Among them, each CIM macro cell performs 128 MAC (multiply-accumulate) operations per cycle. The computational form of the CIM macro cell is the same as that of a typical weight stationary (WS) CIM macro cell. In the computational form of weight stationary WS, the weight data remains stationary, while the activations and partial sums move during the calculation. Also, within each CIM macro cell, the weight matrix is loaded into the CIM macro cell through the weight input / output channels (weight I / O).
[0046] Further, continue to refer to Figure 2aIn the example shown, each CIM macro cell includes multiple banks, and the banks are completely parallel to each other and there is no dependency between them. It can be understood that each bank is independent, but they share control signals, so the start and end of computing executed by different banks are synchronized. For this reason, in this application, each of the above-mentioned banks corresponds to an output channel, and for the input channel, each bank shares the same input channel. Here, it is designed that each bank has its own exclusive output channel but they share the same input channel, which means that all banks share the input, that is, all banks obtain data from the same input channel, then process it according to their respective configurations or tasks, and then each bank outputs its processing result to its own dedicated output channel. Preferably, the design adopted in this application is that each bank corresponds to an output channel but shares the same input channel. Additionally, further, each bank has an SRAM array, and the SRAM array refers to a storage array composed of multiple SRAM cells. In each bank, the SRAM array is divided into R storage sub-arrays by rows (such as the shown sub-array 1, sub-array 2,...., sub-array 32), and each of the R storage sub-arrays corresponds to an input channel. The weight matrix is loaded into each storage sub-array in the form of N columns for participating in subsequent computing tasks. Specifically in implementation, each column storage block in each of the above-mentioned storage sub-arrays contains at least one storage bit cell, and the storage bit cell is the SRAM bit cell shown in Figure 2. This Figure 2a and Figure 2b shows an example where each column storage block (such as column 1, column 8) includes 4 SRAM bit cells, and a local local readout calculation circuit is equipped for the 4 SRAM bit cells. For example, the 4 SRAM cells are connected in parallel to this one local readout calculation circuit locally, so as to realize that the local readout calculation circuit can directly perform calculations at the storage location, reducing data transmission.
[0047] Based on the above content, that is, for each in-memory computing macro cell (CIM macro cell) in the matrix operation unit, a feasible technical solution is that: the in-memory computing macro cell includes multiple parallel banks. Each of the banks includes a static memory array (such as the aforementioned SRAM array), and the static memory array supports in-memory matrix operations.
[0048] Specifically, the implementation method of parallel operation of the above multiple memory banks is as follows: each memory bank corresponds to an output channel but shares the same input channel. Moreover, for the static memory array included in each memory bank, the design is that each row in the static memory array corresponds to an input channel, and the input channels corresponding to different rows are different. And each column storage block in each row includes at least one memory bit cell and one arithmetic circuit (i.e., the aforementioned local readout calculation circuit), and at least one memory bit cell is connected to the arithmetic circuit in parallel. The number of at least one memory bit cell can be but is not limited to 1, 2, 3, 4, etc., and can be arranged according to actual requirements, and no specific limitation is made here. Preferably, the number of at least one memory bit cell is 4.
[0049] Further, before performing calculations such as multiplication using the above arithmetic circuit, one of the at least one memory bit cells connected to the arithmetic circuit is selected to read the weight value of the selected one memory bit cell and store it in this arithmetic circuit locally for subsequent calculations. Among them, one bit storage cell is selected from at least one memory bit cell by means of address fetching.
[0050] It should be supplemented here that the weight value stored in one memory bit cell refers to one bit number of the weight (i.e., weight bit), because: in this application, each row in the static memory array performs bit-level multiply-accumulate calculation on the input data (1-bit data) and the corresponding weight through the arithmetic circuit thereon, so the weight corresponding to one row is represented by m bits (such as 8 bits), and different bit numbers of the weight are respectively stored in different column storage blocks corresponding to the row (specifically, in one memory bit cell among the at least one memory bit cell included in the column storage block).
[0051] For example, continue to refer to Figure 2aAs shown, taking column 1 in the shown sub-array 1 as an example, before the local readout calculation circuit (operation circuit) in column 1 performs calculations, it will select one of the corresponding 4 SRAM bit cells through the address fetching function of the SRAM, so as to read out the weight value stored in the selected SRAM bit cell, and store the weight value locally to participate in subsequent calculations (such as bit-level multiply-accumulate calculations of input data and weights). Among them, the reasons for selecting one of the 4 SRAM bit cells and reading and locally storing the weight value of the selected SRAM bit cell are as follows: First, as a "bit cell", the SRAM bit cell only has a storage function. If you want to perform calculations, you need to know that the data value information (such as the weight value) stored in the SRAM bit cell is then used to participate in subsequent calculations; Second, because only one operation circuit is equipped for 4 SRAM bit cells, if the operation circuit wants to perform calculation processing, it needs to determine which data value stored in which SRAM bit cell is needed for calculation processing currently.
[0052] Furthermore, continue to refer to Figure 2a As shown, in addition to including multiple parallel memory banks, the in-memory computing macro cell (CIM macro cell) also includes: a shift accumulator and a first memory unit (the partial sum buffer unit shown in Figure 2). The shift accumulator is used to perform shift and accumulation operations, which is particularly suitable for fixed-point numbers. Its core idea is to use shift and accumulation operations to achieve multiplication. Among them, the purpose of performing shift accumulation through the shift accumulator here is: to calculate the multiplication result of the activation values (activation) input in multiple cycles and the weights, because only 1-bit activation value is input in each cycle of the in-memory computing macro cell. And, each memory bank included in the in-memory computing macro cell also includes an adder tree unit. The adder tree is a hardware architecture for fast summation, and is often used to add multiple input data. In the memory bank, each row in its static memory array shares this adder tree unit. By adopting this design method of sharing the same adder tree by different rows, hardware resources can be saved and the calculation efficiency can be improved. During the calculation process, the calculation result output by the operation circuit is specifically processed in parallel by the adder tree unit and the shift accumulator, and the processing result will be stored in the first memory unit (partial sum buffer unit). Among them, during the calculation process, the input vector is broadcast to all input channels in a bit-serial manner.
[0053] Specifically, in this application, the above-mentioned addition tree is used to accumulate the calculation results corresponding to different rows in the static memory array, and the obtained accumulated result is passed to the shift accumulator. The shift accumulator will perform a shift accumulation operation on the accumulated result. Specifically, in implementation, the shift accumulator will perform a shift accumulation on the results output by the addition tree in different cycles. Among them, when performing the shift operation, for example, certain preset rules can be used to determine whether to shift the accumulated result output by the addition tree left or right by a certain number of bits to adjust its value. The result output by the shift accumulator will be stored in the first memory unit.
[0054] Among them, the calculation result corresponding to each row output in the above-mentioned static memory array is obtained by performing a bit-level multiply-accumulate calculation on the input activation value (activation, which is 1-bit data) and the corresponding weight of this row through the arithmetic circuit on the row. Exemplarily, referring to Figure 2a the shown memory bank 1, the size of the static memory array included in this memory bank 1 is 32×8 (with 32 rows and 8 columns). Then, the weight corresponding to each row in the static memory array is represented by 8 bits, and each column storage block in each row stores 1 bit of the weight (such as a certain bit (bit) in the binary representation of the weight). Assume that the 8-bit weight W = 10010110 (decimal is 150) corresponding to the first row (sub-array 1), and the corresponding input data A[0] = 1. Then, through the corresponding arithmetic circuit on the first row, a bit-level multiply-accumulate calculation is performed on [A0] and W, and the output calculation result is 150.
[0055] Here, it needs to be supplemented and explained that: in this application, each row in the static memory array in each memory bank corresponds to a word line (WL). When a certain word line is activated, it will select all the memory bit cells corresponding to the row and allow read or write operations on them. That is, the word line is the horizontal control line in the memory array (such as a static memory array) used to select a specific memory row (Row). In addition, each column in the static memory array corresponds to a bit line (BL) and a bit line inverse (BLB) for data transmission. The local bit lines (LBL) and local bit line inverses (LBLB) shown for each column in each row in Figure 2 are used to transmit data to the corresponding arithmetic circuit.
[0056] Furthermore, continuing to refer to Figure 2a the shown example, the in-memory computing macro cell may also include other units, such as a driving unit (i.e., the word line and input drive shown in Figure 2a ) and a control unit.
[0057] The above-mentioned driving unit has the functions of word line driving and input driving. The word line driving is used to activate the word line to select a row corresponding to the word line in the static memory array; among them, the read / write state of each column storage block in the selected row is allowed to read and write. And, the input driving, which can be understood as belonging to the bit line driving, is used to load data onto the corresponding bit line and write it into the corresponding storage bit cell (such as an SRAM bit cell) during a write operation.
[0058] For example, after the address signal of an input data (such as an input vector) is transmitted to the driving unit, the driving unit will select the corresponding word line according to the row address in the address signal and activate the word line. In addition, according to the column address in the address signal, the input data will be loaded onto the corresponding bit line and written into the corresponding storage bit cell.
[0059] The functions of the above-mentioned control unit include but are not limited to: controlling the behavior of the driving unit according to the received read / write command, determining whether to activate the word line and how to process data transmission (such as the input vector is broadcast to the input channels in a bit-serial manner, that is, the input vector is broadcast in a bit-serial manner, that is, transmitted bit by bit in sequence).
[0060] In addition, the in-memory computing macro cell (CIM macro cell) also includes a dedicated weight input / output (weight I / O) channel. Through the dedicated weight I / O channel on it, the CIM macro cell can perform corresponding read / write operations in the column dimension, such as writing the weight matrix (such as the aforementioned weight matrix is loaded into the CIM macro cell through the weight I / O).
[0061] In summary, for Figure 2aThe CIM-MXU (Matrix Arithmetic Unit) example given in [reference] uses a systolic array with an output stationary (OS) data stream. Among them, after the input and weights are calculated in each CIM macro cell in the CIM-MXU, they will propagate in the corresponding directions in turn. In the row dimension, 32-bit input vectors (i.e., the input vectors such as IF1 and IF2 shown in Figure 2, which are 32 bits (i.e., 32b)) propagate in a systolic manner, moving from one column of CIM macro cells to the next column in turn. In the column dimension, each CIM macro cell performs read and write operations of CIM through its dedicated weight I / O. Before calculation, weight initialization requires 32 cycles. During the calculation process, the weight array (which is the weight matrix, the 256b shown in Figure 2 represents the data bandwidth of the weight array, and in this case, the weight matrix can be, for example) is transferred to the CIM macro cells in adjacent rows through interleaved SRAM read and write operations. To meet the need for frequent weight updates, the CIM macro cell supports performing calculations and reading and writing weights simultaneously through the weight I / O. Using this systolic data path maximizes the reuse opportunities of weights and inputs, saves I / O bandwidth, and enables the system to scale to a larger array of CIM macro cells.
[0062] The above-mentioned CIM macro cell supports performing calculations and reading and writing the kernel weights simultaneously through the weight I / O. It can be understood that: because the arithmetic circuit in the CIM macro cell needs to latch the data in the corresponding SRAM cell into the arithmetic circuit before performing the calculation, that is, the arithmetic circuit can store data. Specifically, it can store the value read from the corresponding SRAM bit cell within multiple cycles. At this time, the SRAM bit cell in the CIM macro cell can perform normal read and write operations. Thus, it can be realized that within multiple cycles, the arithmetic circuit can calculate data while reading data from the corresponding SRAM bit cell.
[0063] Return to Figure 1 , in this application, the tensor core module 1 may further include, in addition to multiple matrix arithmetic units 11: a second memory unit 12 (i.e., a vector cache unit (Vector Memory, VMEM)) and a vector processing unit 13 (Vector Processing Unit, VPU).
[0064] The above vector processing unit VPU is responsible for other general parallel computations in the generative model. For example, non-linear operations such as normalization and activation functions (such as ReLU, Sigmoid, Tanh, etc.). For example, the matrix multiplication result output by the matrix arithmetic unit CIM-MXU can be passed to the vector processing unit VPU so that the vector processing unit VPU can perform the corresponding activation function (such as ReLU).
[0065] The above vector cache unit (vector memory) VMEM is used to store input vectors, which are obtained by slicing through corresponding operator operations in the generative model, and the operator operations include matrix multiplication. For example, assuming the algorithm operation is matrix A multiplied by matrix B, then A and B can be sliced respectively, and the obtained input vectors are stored in the vector cache unit VMEM. In addition, the vector cache unit VMEM can also store other data such as weight matrices. The aforementioned matrix operation unit CIM-MXU can obtain each required input from the vector cache unit.
[0066] The foregoing combination Figure 1 and Figure 2a mainly introduced the tensor core module 1 in the processor architecture. Continue to refer to Figure 1 , in addition to including the tensor core module 1, the processor architecture also includes other modules, such as the third memory unit 3 (shared cache unit (Common Memory, CMEM)), the inter-chip interconnect module 4 (Inter-chip Interconnect, ICI), and the main memory controller 2.
[0067] The above shared cache unit CMEM (i.e., the third memory unit 3) is a cache memory used to share data among multiple units and reduce the access frequency to the main memory. In design and implementation, the storage capacity of the shared cache unit can be but is not limited to 128M. In addition, both the shared cache unit CMEM and the aforementioned vector cache unit VMEM are implemented as on-chip SRAMs, that is: the memory types of the shared cache unit CMEM and the aforementioned vector cache unit VMEM are both SRAMs. Also, in the processor architecture adopted in this application, the data transfer between the vector cache unit VMEM and the shared cache unit CMEM is modeled by simulating the available on-chip bandwidth; among them, the on-chip bandwidth simulation can be performed according to some data metrics (such as the amount of data that can be transferred per second) reported by some existing processors (such as GPUs, TPUs).
[0068] It can be seen that in the processor (such as TPU) architecture provided in this application, there is a two-level on-chip memory hierarchy, and this two-level on-chip memory hierarchy includes: two memory hierarchies, namely the shared cache unit and the vector cache unit.
[0069] Also, the above inter-chip interconnect module 4 (ICI) is a high-speed interconnect technology used to efficiently transfer data among multiple units.
[0070] The above main memory controller 2 is responsible for managing the interaction with the external memory (such as the main memory) and controlling data read and write operations on the external memory. For example, it can manage the address mapping of the external memory.
[0071] Furthermore, in the processor architecture, in addition to having the main memory controller 2, a top-level control circuit can be set on top of it. The main function of the top-level control circuit is to manage and coordinate various operations inside the processor. For example, data transfer (such as controlling the transfer of data from an external memory to the CMEM) and cache coherence (such as coordinating data coherence between the CMEM and the external memory), etc. There is a mapping engine (the model mapping engine) in this top-level control circuit. This mapping engine can be hardware, software, or a combination of hardware and software, and no specific limitation is made here. The main function of the mapping engine is to read out and map the model data of the generative model, etc., so as to write the model data into the corresponding cache (such as the shared cache model and the vector cache unit), and allocate it to the corresponding units in the processor architecture (such as the vector processing unit VPU, the matrix operation unit CIM-MXU, etc.) for calculation processing.
[0072] Among them, in this application, considering that the search space for effective mapping is extremely large, a heuristic method is adopted to prune the search space for mapping. It can be understood here that: if all possible mapping strategies are considered, there are a very large number of possibilities. If every mapping method is to be enumerated, the latency and computational overhead for data transfer, data calculation, etc. will be very large. Therefore, this application does not enumerate every possible mapping strategy, but adopts a heuristic algorithm to predict and consider some possible better mapping strategies for use. The mapping engine explores different mapping strategies to find the optimal mapping scheme to better utilize hardware resources.
[0073] Figure 3 Shows an example of mapping a single operator in the generative model by the mapping engine. As shown in Figure 3 As shown, the model data (model configuration) of the generative model is stored in the external main memory. Generally, the generative model includes multiple hierarchical blocks, such as a normalization layer, a QKV generation layer (a key component in the Transformer model for processing the self-attention mechanism), a feed-forward neural network layer, etc. Each layer has some operator operations (such as matrix multiplication), etc. Assume that there is a matrix multiplication on a certain hierarchical block of the generative model, and this matrix multiplication is [L, D] × [D, D] = [L, D], which means a matrix of size L×D is multiplied by another matrix of size D×D to obtain a matrix of size L×D. For the matrix multiplication in the above generative model, through the mapping engine, this matrix multiplication operation can be divided into multiple blocks of size [L tileM , D tileK × [D tileK , D tileN to adapt to the storage limit of the on-chip shared cache unit. In addition, [L tileM , DtileK , [D tileK , D tileN Matrices such as ] will be further divided to adapt to the storage space limitation of the vector cache unit, and finally processed by the matrix operation unit CIM-MXU or the vector processing unit VPU. It can be understood from the above example that the mapping engine reads the [D, D] matrix in the main memory row by row according to the pre-configured mapping strategy and writes the read row block (which is a row vector, such as the row vector of the second row, etc.) into the shared cache unit. After writing into the shared cache unit, the next step is to write it into the vector cache unit at the upper level. Since the storage capacity of the vector cache unit is relatively small, the read row block will be further divided into multiple sub-blocks and then written into the vector cache. Similarly, the mapping engine can perform mapping processing on the [L, D] matrix in the main memory to map it into the corresponding shared cache unit and vector cache unit for storage. Among them, for the [L, D] matrix, when a data block in a certain row cannot be written into, for example, the vector cache unit due to its size, the data block can be further divided into multiple sub-blocks to be written into the vector cache unit.
[0074] It should be added here that the mapping strategy in the mapping engine is pre-configured by software and then implemented by hardware.
[0075] In addition, in order to overlap the computing and memory access cycles, the present application also adopts double buffering and memory access coalescing as scheduling options in each memory hierarchy (such as the shared cache unit and the vector cache unit) to improve the overall performance.
[0076] The above "double buffering" can be understood as follows: for example, the vector cache unit has a certain storage space. Usually, when performing calculations, the actual access to this vector cache unit is occupied because the calculation needs to continuously read data from the vector cache unit. At this time, it is actually impossible to transfer new data from the shared cache unit into the vector cache. By adopting double buffering, a part of the space in the vector cache unit will be used for the read and write access of the calculation, and another part of the space will be used for data sharing with the shared cache unit for the next calculation. This is equivalent to sacrificing a part of the storage space, but it can achieve the parallelism of calculation and data access.
[0077] That is: "double buffering" is to use two sets of buffer areas, one for the current calculation and the other for preparing data for the next stage. In this way, data prefetching and write-back operations can be performed without interrupting the calculation, thereby improving the throughput and efficiency of the system. Among them, the working process of "double buffering" can be as follows:
[0078] 1) Initialization:
[0079] Create two buffer areas in the vector cache unit: one buffer area is used for the current calculations of vector processing units VPU and / or matrix operation units CIM-MXU (referred to as the working buffer area), and the other is used to prepare the input data required for the next calculations of VPU and / or CIM-MXU (referred to as the preparatory buffer area).
[0080] 2) Calculation stage: Use the data in the working buffer area for calculations. Meanwhile, prepare the input data required for the next calculation in the preparatory buffer area.
[0081] 3) Switching stage: When the current calculation is completed, swap the roles of the working buffer area and the preparatory buffer area. The working buffer area becomes the new preparatory buffer area, and the preparatory buffer area becomes the new working buffer area.
[0082] 4) Repeat steps 2) and 3): Continuously loop the above process to maintain the continuity of calculations and data preparation.
[0083] Moreover, the above-mentioned "memory access merging" can be understood as follows: For example, there are four arithmetic circuits, and all four arithmetic circuits simultaneously require a piece of data. At this time, actually only the transfer time of one piece of data is considered during data transfer because after this one piece of data is transferred up, the four arithmetic circuits can use it simultaneously.
[0084] Table 1 below shows an example comparison of the processor architecture parameters provided in this application (taking the processor type as TPU as an example) with the parameters of existing processor architectures (such as the TPUv4i architecture).
[0085] Table 1
[0086]
[0087] Note: The above CIM macro module is the aforementioned matrix operation unit. The CIM macro module dimension of 16×8 indicates that the matrix operation unit contains an array of CIM macro cells (in-memory computing macro cells) with a size of 16×8.
[0088] To verify the processor architecture provided by this application, the performance of this processor architecture was also analyzed and evaluated. Taking the processor type of TPU as an example, in this case, the TPU architecture based on SRAM CIM provided by this application, namely the CIM-TPU architecture for short, supports the evaluation of multiple core operators in generative models, including matrix multiplication GEMM and other non-linear functions. The GEMM calculation is performed on three-dimensional data, where the input matrix is sliced into smaller matrices to fit the size of the SRAM cache. As a comparison benchmark for the CIM-MXU, the matrix operation unit in the architecture, this application uses SCALE-Sim to evaluate the GEMM calculation performance of the systolic array for a given array dimension. In addition, this application also uses a similar method to model the calculations of Softmax (a non-linear function), LayerNorm (layer normalization), and GeLU (a non-linear activation function). These operators are mainly processed by the vector processing unit VPU in the architecture, rather than the CIM macro unit in the matrix operation unit. Specifically, this application uses the Online Softmax algorithm and uses the tanh approximation for GeLU calculation. In the performance evaluation, tensor parallelism and pipeline parallelism are used to expand the computing power of the CIM-TPU to improve the overall performance.
[0089] To verify the advantages of CIM compared to the systolic array, this application also presents a comparison between the MXU based on the systolic array and the 16×8 CIM-MXU (the TPU-related parameters are shown in Table 1 above). Among them, this application uses Gemmini (an open-source hardware accelerator generator) to generate a 128×128 systolic array as the reference MXU, and performs digital synthesis through Cadence Genus (a synthesis tool mainly used for logic synthesis in integrated circuit (IC) design), while the layout of the CIM macro unit is manually drawn. Both the reference MXU and the CIM-MXU are implemented under the same TSMC 22nm process. As shown in Table 2, the implementation of the CIM-MXU has an energy consumption and area efficiency of 7.26 TOPS / W and 1.31 TOPS / mm2, respectively, which are 9.43 times and 2.02 times higher than those of the reference MXU, while maintaining the same per-cycle MAC throughput. These results verify the significant advantages of the CIM architecture in terms of performance and efficiency.
[0090] Table 1 Parameter Comparison between Reference MXU and CIM-MXU
[0091] Evaluation parameter Baseline MXU CIM-MXU Speedup MAC per cycle 16384 16384 1× Energy consumption efficiency 0.77 TOPS / W 7.26 TOPS / W 9.43× Area efficiency <![CDATA[0.648 TOPS / mm 2 > <![CDATA[1.31TOPS / mm 2 > 2.02×
[0092] To further evaluate the effectiveness of generative model inference on TPU, this application also selected two representative generative models for evaluation: GPT-3 and DiT-XL / 2. During the evaluation, the existing benchmark (such as TPUv4i) architecture parameters were used as the benchmark, while the CIM-MXU was used to replace the systolic array-based MXU architecture, while keeping other hardware specifications unchanged, such as memory capacity and bandwidth. The benchmark TPU and CIM-TPU were both scaled to the same process node and clock frequency to ensure a fair comparison of performance and energy consumption. Taking the generative model as a large language model (LLM) based on Transformer as an example, the model configuration of the Transformer layer in the generative model is shown in Table 3. During the simulation process, the batch size was set to 8, and the inference process of a single Transformer layer of GPT-3 and a DiT layer in DiT-XL / 2 (a picture generation model based on the Transformer architecture) was simulated, with an image resolution of 512×512 and using INT8 data precision. During the LLM prefill stage, the input length was set to 1024. During the LLM decoding stage, the 256th decoding process was simulated. Figure 4 Shows the performance of the evaluated generative models in terms of inference latency and energy consumption.
[0093] Table 2 Model configuration of the generative model
[0094] Generative model Number of model layers Number of attention heads Model dimension GPT3-30B 48 56 7168 DiT-XL / 2 28 16 1152
[0095] LLM prefill stage: During the prefill stage, QKV generation, output projection, and the feed forward neural network layer (FFN) consist of large-dimension GEMM calculations. These layers account for 84.9% of the TPU inference latency and are the main computational bottlenecks. In contrast, the attention layer, including Query-Key dot product, Score-Value dot product, and Softmax operations, only accounts for 13.1% of the latency. Since the systolic array in the benchmark MXU has been optimized for large-scale GEMM operations, the CIM-MXU does not show a significant improvement in inference latency. However, the CIM-MXU shows an obvious advantage in energy efficiency, and its energy consumption during the prefill stage is 9.21 times lower than that of the benchmark MXU.
[0096] LLM Decoding Stage: In the LLM decoding stage, the input sequence length is 1, significantly reducing the input dimension of GEMM, resulting in a lower operator arithmetic intensity and memory bandwidth bottleneck. In the baseline TPU design, the attention layer accounts for 33.7% of the inference latency, mainly composed of Query-Key dot product and Score-Value dot product. Compared with the baseline design, CIM-TPU achieves a 72.7% acceleration for these GEMV operations, significantly reducing the inference latency by 29.9%. This is attributed to the fact that the input activation vectors in the CIM macrocells are broadcast to all output channels in a bit-serial manner, which is more efficient than the way the input activation values need to traverse all the previous MAC units in the traditional systolic array. Thanks to the improvement in latency and efficiency, the energy consumption of CIM-MXU is 13.4 times lower than that of the baseline MXU, significantly improving the LLM decoding efficiency.
[0097] DiT Layer: In a single DiT layer, the GEMM operations of QKV generation, output projection, and FFN account for 35.65% of the inference latency. In these GEMM computations, the performance of the baseline MXU and CIM-MXU is similar. However, the Softmax operation in the attention layer accounts for up to 36.9% of the inference latency, becoming the computational bottleneck in DiT inference. Notably, CIM-MXU achieves a 30.3% performance improvement for the Query-Key dot product and Score-Value dot product in the attention layer. Overall, compared with the baseline design, CIM-TPU reduces the inference latency by 6.67% and the energy consumption by 10.4 times.
[0098] Based on the above model inference evaluation results, this application summarizes the design observations of using CIM in the TPU architecture. First, CIM significantly improves the area and energy efficiency. The CIM-MXU proposed in this application contains 128 CIM macrocells, providing the same peak performance as the baseline MXU but occupying only 50% of its area. At the same time, CIM-MXU can reduce the energy consumption by about an order of magnitude, significantly improving the energy efficiency of matrix calculations. Second, CIM has different performance improvements for different generative models. As can be seen from the above model evaluation, CIM achieves the largest performance improvement in the layers dominated by GEMV (such as the LLM decoding stage). Since the LLM has a large output length, the decoding process consumes the most latency in LLM inference. Therefore, CIM-TPU can effectively improve the performance and efficiency of LLM inference. During DiT inference, the Softmax and GEMM operations constitute the main inference latency. Therefore, CIM mainly contributes to the improvement of area and energy consumption efficiency in DiT inference, rather than performance improvement.
[0099] Based on the aforementioned proposed CIM-TPU architecture, this application further optimizes the design to find the best design points. Table 4 shows several architecture design options. Figure 5 Shows the inference latency and energy consumption evaluations of GPT-3-30B and DiT-XL / 2 under different CIM-MXU architecture settings and compares them with the baseline design.
[0100] Table 3 Architecture design options for CIM-MXU
[0101]
[0102] LLM inference includes a prefill stage and a decoding stage. In this application, the input and output sequence lengths are set to 1024 and 512. Since LLM decoding is limited by memory bandwidth, when the number of MXUs and the array dimension continue to increase, the performance improvement of CIM-MXU tends to saturate. For example, although 8 CIM-MXUs configured with 16×16 CIM macrocells have twice the peak performance compared to the same number of CIM-MXUs with 16×8 CIM macrocells, the performance improvement for LLM inference is only 2.5%, and the energy consumption increases by 95%. To effectively utilize the efficiency advantage of CIM technology, using smaller-sized CIM-MXUs can achieve significant energy efficiency improvement with the least reduction in performance. For example, even if only 2 CIM-MXUs are configured in the TPU, a smaller 8×8 CIM macrocell array will result in a 38% increase in latency, but the single energy consumption is reduced by 27.3 times. Considering the balance between latency and energy consumption, this application adopts 4 CIM-MXUs with 8×8 arrays as the CIM-TPU architecture optimized for LLM inference, which is called Design A.
[0103] For computationally intensive DiT inference, this application observes that CIM-TPUs equipped with more or larger CIM-MXU arrays can significantly reduce inference latency by leveraging higher peak performance. Specifically, CIM-TPUs equipped with 4 and 8 CIM-MXUs (16×16 CIM macrocells) achieve 25.3% and 33.8% reduction in inference latency respectively. However, such performance improvement is also accompanied by an increase in energy consumption. Thanks to the high energy efficiency of CIM, even when configuring the highest-performance CIM-MXU, that is, 8 16×16 CIM macrocells, its power consumption is still 3.56 times lower than that of the baseline MXU. Considering the trade-off between the latency, energy, and area of the MXU, this application selects CIM-MXUs with 8 16×8 CIM macrocells for the CIM-TPU architecture optimized for DiT inference, which is called Design B.
[0104] In addition to the architecture exploration of a single CIM-TPU, this application also extends the evaluation to multi-TPU inference scenarios to meet the large-scale deployment requirements of generative models. To accommodate large-batch model inference, this application increases the number of TPUs and implements 4-way pipeline parallelism. The 4 TPUs are interconnected in a ring topology to make full use of the two ICI links on each TPU chip, which is consistent with the default configuration in TPUv4i. This application evaluates the throughput of the benchmark TPU and CIM-TPU (i.e., Design A and Design B) in generative model inference, and selects output sequences per second (token / s) and images per second (image / s) as the inference evaluation metrics for LLM and DiT.
[0105] Figure 6 Shows the inference throughput of GPT-3-30B and DiT-XL / 2 when using 1, 2, and 4 TPUs. When comparing the LLM inference performance, Design A achieved an average 28% speedup compared to the benchmark MXU, while reducing the energy consumption by 24.2 times compared to the benchmark MXU. Thanks to the higher peak performance, Design B achieved a 33% throughput increase compared to the benchmark, and at the same time, CIM-MXU also reduced the energy consumption by 6.34 times compared to the benchmark MXU.
[0106] This application also provides a processor. The processor includes the processor architecture provided by this application above. The type of this processor can be but is not limited to TPU, GPU, etc.
[0107] This application also provides a chip. The chip includes the processor provided by this application above.
[0108] This application also provides an electronic device. The chip provided in other embodiments of this application is installed and deployed on the electronic device. The electronic device can be various terminal devices such as smartphones, personal computers (such as desktop computers, laptop computers), tablets, etc.; or it can also be a server device such as a single server, a server cluster, a virtual server, etc.
[0109] Furthermore, the electronic device also includes other components, such as an interface module, a control panel, and other components.
[0110] It should be noted that the term "including", "comprising" or any other variant thereof in this application is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.
[0111] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A processor architecture based on static memory supporting in-memory matrix operations, characterized in that: include: A tensor core module, wherein the tensor core module includes a plurality of matrix operation units; wherein, The matrix operation unit has a plurality of in-memory computing macro units; The multiple in-memory computing macro units are connected by systolic data paths to form a two-dimensional systolic array, and each of the in-memory computing macro units includes multiple parallel storage bodies, each of which includes a static memory array that supports in-memory matrix operations.
2. The processor architecture according to claim 1, characterized in that: Each row in the static memory array corresponds to an input channel, and each column storage block in each row includes at least one storage bit unit and an operation circuit, and the at least one storage bit unit is connected to the operation circuit in parallel; The operation circuit is used to perform a corresponding calculation operation based on the numerical information stored in the target storage bit unit; wherein the target storage bit unit is one of the at least one storage bit unit, the numerical information includes a bit number of the weight, and the calculation operation includes a multiplication operation.
3. The processor architecture according to claim 2, characterized in that: The in-memory computing macro unit also includes: a shift accumulator; and the memory body also includes an addition tree unit, each row in the static memory array in the memory body shares the addition tree unit, and each row performs bit-level multiplication and accumulation calculations on input data and weights through the corresponding operation circuit and outputs the calculation results to the addition tree unit; The addition tree unit is used to accumulate the calculation results outputted from different rows in the static memory array, and transmit the accumulated results to the shift accumulator; The shift accumulator is used to perform a shift accumulation operation on the accumulation result.
4. The processor architecture according to claim 3, characterized in that: The in-memory computing macro unit also includes: a first memory unit; The first memory unit is used to store the output result of the shift accumulator.
5. The processor architecture according to claim 2, characterized in that: Each column storage block includes four storage bit units; and The operation circuit is also used to: select a target storage bit unit from the corresponding four storage bit units through an address access method; read the weight value stored in the target storage bit unit, and store the weight value locally.
6. The processor architecture according to any one of claims 1 to 5, characterized in that: Each memory bank of the plurality of memory banks corresponds to an output channel but shares the same input channel.
7. The processor architecture according to any one of claims 1 to 5, characterized in that: The tensor core module further includes: a vector processing unit and a second memory unit; The vector processing unit is used to process nonlinear operation; The second memory unit has two cache areas, one of which is used for read and write access of the current calculation of the matrix operation unit and / or vector processing unit, and the other cache area is used to store input data required for the next calculation of the matrix operation unit and / or vector processing unit.
8. A processor, characterized in that: A processor architecture having any one of claims 1 to 7.
9. A chip, characterized in that: The processor comprises the processor as claimed in claim 8 above.
10. An electronic device, characterized in that: The chip comprises the chip as claimed in claim 9.
Citation Information
Cited By
Large-model-oriented adaptive hybrid computing architecture
CN121029686A
Expert model reasoning system and reasoning method based on RRAM memory internal calculation
CN121352033A