Near memory calculation acceleration circuit supporting weight prefetching in controller

By integrating near-memory processing unit and matrix vector multiplication processing block in the memory controller, prefetching and parallel processing of the weight matrix are achieved, which solves the shortcomings of existing acceleration circuits in terms of data transmission delay, energy consumption and cost, and improves system performance and cost-effectiveness.

CN120126525APending Publication Date: 2025-06-10FUDAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510151061.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing acceleration circuits have shortcomings in reducing data transmission delay and energy consumption, especially in PIM and PNM solutions, where cost and heat dissipation problems remain.

Method used

A near-memory computing acceleration circuit supporting weight prefetching in the controller is designed. This circuit integrates a near-memory processing unit in the memory controller, and realizes prefetching and parallel processing of the weight matrix through technologies such as asynchronous request and response queues, matrix vector multiplication processing blocks, etc.

Benefits of technology

By prefetching the weight matrix, the read latency of DRAM is reduced, the interface throughput and system performance are improved, while the energy consumption and cost are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126525A_ABST
    Figure CN120126525A_ABST
Patent Text Reader

Abstract

The invention provides a near memory calculation acceleration circuit supporting weight prefetching in a controller, which is oriented to a three-dimensional stacked DRAM (Dynamic Random Access Memory) architecture, adopts a PIC (Peripheral Interface Controller) technology of integrating a near memory processing module in the controller, and supports packaging and interconnection between the PIC and the three-dimensional stacked DRAM through a TSV (Through Silicon Via) or a hybrid bonding mode, so that the weight prefetching in the controller can be realized at a lower cost and a shorter storage-calculation path. Remarkable data migration efficiency and energy consumption income are realized; in the near memory processing module, the waiting time delay of reading the weight matrix from the DRAM is eliminated through the operation of prefetching the weight matrix of the next operation from the DRAM to a calculation circuit in the module in parallel in the matrix vector multiplication operation process, so that the higher interface throughput rate is realized; moreover, read-write command / data transmission between the near memory processing module and the controller processing engine is completed through a group of asynchronous FIFO, flexible clock frequency configuration is supported, low-power-consumption data migration of an interface and a system is realized through frequency reduction processing in near memory processing, and the scale of the GEMV can be flexibly configured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of acceleration circuits, and particularly to a near-memory computing acceleration circuit supporting weight prefetching in a controller. Background Art

[0002] With the development of deep learning and large language models, the sharp growth of model parameters and data volume has put forward higher requirements for memory bandwidth. However, the unbalanced development between memory bandwidth and processor performance has led to the "memory wall" problem, causing a storage bottleneck. HBM and 3D stacked DRAM memories have greatly expanded the storage bandwidth through 3D stacking advanced technologies, effectively alleviating the performance bottleneck brought by the memory wall. However, the high energy consumption and latency of data transmission still have a significant impact on the system.

[0003] To further break through the memory wall bottleneck and reduce the latency and energy consumption caused by the frequent transmission of data between memory and processor, in recent years, Processing-in-Memory (PIM) integrating computing processing units inside DRAM and Processing-near-Memory (PNM) close to the CPU / GPU have become the main solutions.

[0004] Currently, the main types of PIM are as follows: Samsung's PIM, SK Hynix's AiM (Accelerator-in-Memory), and Upmem's PDU (Processing-in DRAM Unit), etc. Among them, Samsung integrates the processing unit into HBM, SK Hynix integrates the acceleration unit into GDDR6 memory, and Upmem integrates the computing unit into a standard DRAM chip. Figure 10 The structural schematic diagram of Upmem's PIM solution in the prior art is shown.

[0005] The current PNM solution is to place the computing processing unit in a dedicated acceleration engine close to the CPU / GPU. Figure 11 The structural schematic diagram of the PNM solution in the prior art is shown.

[0006] Although the above existing solutions have reduced the latency and energy consumption caused by the frequent transmission of data between memory and processor to a certain extent, there are still some deficiencies.

[0007] Specifically, the PIM solution requires embedding computing units inside DRAM chips. The high-concurrency operation of PIM effectively reduces data transmission and has great advantages in terms of energy efficiency. However, PIM faces complex co-development of memory architecture, controller, and processor, as well as high manufacturing costs for the DRAM embedding process. At the same time, embedding computing units in the 3D stacked DRAM storage layer also introduces local heat accumulation, leading to heat dissipation and power consumption problems.

[0008] The PNM solution adopts the same compromise approach, integrating computing and processing units near the CPU / GPU side at a lower process cost. Based on the 3D stacked DRAM load, through the multi-channel parallel method, it achieves greater cost performance in terms of energy consumption and cost. However, the current PNM is mainly integrated in dedicated acceleration engines, and a large number of data access operations between the dedicated acceleration engine and memory still lead to excessive energy consumption and data latency, which will limit the overall system performance. Summary of the Invention

[0009] The present invention is made to solve the above problems, and aims to provide an acceleration circuit that can further improve data transfer efficiency and energy consumption benefits and can be implemented at a relatively lower cost. The present invention adopts the following technical solutions:

[0010] The present invention provides a near-memory computing acceleration circuit supporting weight prefetching in a controller, which is arranged in a storage controller interconnected with a three-dimensional stacked dynamic random access memory and connected to a processor. The acceleration circuit has the following technical features: an interface processing and parsing unit for receiving operation instructions from the processor and sending corresponding operation results to the processor; a processor in the controller, including a near-memory processing unit for processing general matrix-vector multiplication, an asynchronous request first-in-first-out queue, and an asynchronous response first-in-first-out queue; a controller command processing engine for sending a corresponding command sequence to the dynamic random access memory according to the request commands in the asynchronous request first-in-first-out queue and writing the returned read data into the asynchronous response first-in-first-out queue. Among them, the near-memory processing unit includes a plurality of matrix-vector multiplication processing blocks, each of the matrix-vector multiplication processing blocks having a prefetch weight cache. While performing the current matrix-vector multiplication operation, the matrix-vector multiplication processing block prefetches the weight matrix corresponding to the next matrix-vector multiplication from the dynamic random access memory and stores it in the prefetch weight cache.

[0011] The in-memory computing acceleration circuit with weight prefetch support in the controller provided by the present invention may further have the following technical features. The in-memory computing acceleration circuit is oriented to a three-dimensional stacked dynamic random access memory architecture. The in-memory computing acceleration circuit is interconnected with the three-dimensional stacked dynamic random access memory through through-silicon vias or hybrid bonding to the memory controller. The in-memory computing acceleration circuit can achieve bank-level parallel processing of the three-dimensional stacked dynamic random access memory.

[0012] The in-memory computing acceleration circuit with weight prefetch support in the controller provided by the present invention may further have the following technical features. The matrix-vector multiplication processing block includes: a storage array for serving as the prefetch weight cache; a write driver for writing the prefetch weight matrix to the storage array; an address decoder for decoding the read address; and a multiplier and accumulator tree for performing the matrix-vector multiplication operation. The processor in the controller further includes: a top-level state machine control module for controlling the reception and transmission of the request command and the read data; a decoder for decoding the instructions from the processor; a command generation module for generating corresponding request commands according to the decoded instructions and storing them in the asynchronous request first-in-first-out queue; a configuration module for storing configuration information; and an output cache. The in-memory processing unit obtains the returned read data from the asynchronous response first-in-first-out queue, performs matrix-vector multiplication operations based on the read data and the configuration information, and stores the operation results in the output cache.

[0013] The in-memory computing acceleration circuit with weight prefetch support in the controller provided by the present invention may further have the following technical features. Each matrix-vector multiplication processing block is an in-memory computing macro. The matrix-vector multiplication processing block traverses the in-memory computing macro address from low address to high address to read a complete row of weight vectors from the storage array. During the reading process, each element in the weight vector row is multiplied and accumulated with each element in the activation column vector to obtain the dot product result of the two vectors. During the multiplication and accumulation operation, the processor in the controller prefetches the weight matrix for the next operation from the dynamic random access memory. The matrix-vector multiplication processing block traverses the in-memory computing macro address from low address to high address and writes the prefetched weight matrix to the storage array, thereby updating the prefetched weight matrix in the cache.

[0014] The in - controller near - memory computing acceleration circuit supporting weight pre - fetch provided by the present invention may further have the following technical features. Among them, the bit - width of the asynchronous response first - in - first - out queue is set to the data length that the storage controller can respond to in one read request command. The pre - fetched weight cache includes multiple cache units, and the width of each cache unit is the bit - width of the asynchronous response first - in - first - out queue. Multiple cache units read out the weight sub - vectors of the pre - fetched weight matrix from low address to high address in parallel. The weight sub - vectors are used to complete dot - product operations with the corresponding activation sub - vectors in the activation vector. At the same time, the weight sub - vectors for the next operation pre - fetched from the dynamic random access memory are updated to the corresponding read addresses from low address to high address in sequence.

[0015] The in - controller near - memory computing acceleration circuit supporting weight pre - fetch provided by the present invention may further have the following technical features. Among them, the read - write control signals of the processor in the controller include a read enable signal, a read address signal, a read data signal, a partial - sum output signal, an updated write enable signal, a write address signal, and a pre - fetched weight data signal. Starting from when the first - pen weight data returns to the asynchronous response first - in - first - out queue, the read enable signal is pulled high, controlling the read address signal to change from low to high in sequence. After at least one cycle after the read enable signal is pulled high, the updated write enable signal is pulled high, so that the pre - fetched new weight data sequentially updates and overwrites the original weight data in the pre - fetched weight cache.

[0016] The in - controller near - memory computing acceleration circuit supporting weight pre - fetch provided by the present invention may further have the following technical features. Among them, for each request command, the near - memory processing unit checks whether the update of the last - round weight sub - vectors has been completed before sending the last - beat data, and sends the last - beat data when it has been completed.

[0017] The in - controller near - memory computing acceleration circuit supporting weight pre - fetch provided by the present invention may further have the following technical features. Among them, when the number of matrix row data>the number of parallel banks, the matrix - vector multiplication processing block completes the matrix - vector multiplication operation in blocks. The processing time for each block is n time - periods, and the number of blocks is the total number of rows h of the weight matrix / the number of parallel processing rows m of the weight matrix. The scale of the general matrix - vector multiplication is:

[0018] W∈R H×W×op_width

[0019] V∈R H×1×op_width

[0020] In the formula, H is the height of the matrix, W is the width of the matrix, and op_width is the width of the operand.

[0021] The in - controller near - memory computing acceleration circuit supporting weight pre - fetching provided by the present invention may also have the following technical feature: after the general matrix - vector multiplication starts to calculate and after a predetermined delay time, the cached pre - fetched weight matrix is updated. When n = 1, the delay time is 1T; when n≥2, the delay time is selected from the range of [1T, (n - 1)T], where T is the clock cycle.

[0022] The in - controller near - memory computing acceleration circuit supporting weight pre - fetching provided by the present invention may also have the following technical feature: the interface processing and parsing unit and the in - controller processor operate in a low - frequency clock domain, and the controller command processing engine operates in a high - frequency clock domain.

[0023] Functions and effects of the invention

[0024] Compared with the acceleration circuits in the prior art, the in - controller near - memory computing acceleration circuit supporting weight pre - fetching provided by the present invention has the following advantages in many aspects:

[0025] First, the acceleration circuit of the present invention adopts the technology of integrating a near - memory processing unit in the controller (PIC technology). Compared with the solutions of integrating PIM in DRAM or integrating a dedicated acceleration engine near the CPU / GPU side, the acceleration circuit of the present invention can achieve significant data transfer efficiency and energy consumption benefits with lower cost and a shorter storage - computing path.

[0026] Second, in the near - memory processing unit, by parallelly pre - fetching the weight matrix for the next operation from DRAM to the in - module computing circuit during the matrix - vector multiplication operation, the waiting delay for reading the weight matrix from DRAM is eliminated, thereby achieving a higher interface throughput rate; further combined with the ping - pong pipelining processing of the calculation, low - delay and high - bus throughput rate can be achieved.

[0027] Third, the near - memory processing unit and the controller processing engine complete the transmission of read - write commands / data through a group of asynchronous FIFOs. It supports flexible clock frequency configuration. The near - memory processing achieves low - power data transfer between the interface and the system through frequency - reduction processing.

[0028] Fourth, the acceleration circuit of this embodiment supports GEMV calculations with a flexible and configurable matrix scale, and supports any matrix scale and operand type. It can achieve DRAM bank - level parallel transfer and energy - efficient GEMV calculation processing. Therefore, it can be adapted to various current AI applications. For example, when the size is small, it supports applications such as CNN and DNN image processing; when the size is very large, it supports applications such as transformers. Description of the drawings

[0029] Figure 1It is a schematic diagram of the interaction between the storage controller and the memory in the embodiment of the present invention;

[0030] Figure 2 It is a schematic diagram of the structure of the storage controller in the embodiment of the present invention;

[0031] Figure 3 It is a schematic diagram of the structure of the processor in the controller in this embodiment;

[0032] Figure 4 It is a schematic diagram of the working principle of the near-memory processing unit in this embodiment;

[0033] Figure 5 It is a schematic diagram of the principle of the prefetch weight matrix in this embodiment;

[0034] Figure 6 It is a schematic diagram of time-domain partitioning in the prefetch weight cache in this embodiment;

[0035] Figure 7 It is a spatial schematic diagram of the prefetch weight cache in this embodiment;

[0036] Figure 8 It is a timing schematic diagram of parallel read and write for in-memory computing in this embodiment;

[0037] Figure 9 It is a timing constraint schematic diagram of one GEMV processing in this embodiment;

[0038] Figure 10 It is a schematic diagram of the structure of the PIM scheme in the prior art;

[0039] Figure 11 It is a schematic diagram of the structure of the PNM scheme in the prior art.

[0040] Reference numerals:

[0041] Storage controller 100; Interface processing and parsing unit 11; Read / write arbiter 111; Command scheduler 112; Data scheduler 113; Processor 12 in the controller; General matrix-vector multiplication near-memory processing unit 121; Matrix-vector multiplication processing block 1210; Storage array 1211; Address decoder 1212; Write driver 1213; Multiplier 1214; Accumulation tree 1215; Asynchronous request FIFO queue 122; Asynchronous response FIFO queue 123; Top-level state machine 124; Decoder 125; Configuration module 126; Command generator 127; Output buffer 128; Processor 20; 3D stacked dynamic random access memory 200. Detailed implementation manners

[0042] In order to make the technical means, creative features, achieved objectives and effects realized by the present invention easy to understand, the in-memory computing acceleration circuit supporting weight prefetch in the controller of the present invention will be specifically described below in conjunction with embodiments and the accompanying drawings.

[0043] Figure 1 is an interaction schematic diagram between the memory controller and the memory in this embodiment, Figure 2 is a structural schematic diagram of the memory controller in this embodiment.

[0044] As Figure 1 and Figure 2 shown, the system-on-chip (SoC) in this embodiment includes a memory controller 100 and a processor 20. The processor 20 can be a CPU or a GPU. The system-on-chip is interconnected with a three-dimensional stacked dynamic random access memory (3D-stacked DRAM) 200. In this embodiment, the 3D-stacked DRAM is interconnected with the memory controller 100 by means of through-silicon vias (TSV) or hybrid bonding.

[0045] The memory controller 100 includes an in-memory computing acceleration circuit supporting weight prefetch in the controller (hereinafter simply referred to as the acceleration circuit). The outside of the acceleration circuit is connected to the CPU or the CPU through an interconnect bus.

[0046] In this embodiment, the memory controller 100 internally integrates an in-memory processing unit (processor in the controller), and the inventor names it the Processing-in-Controller (PIC) scheme. The moved data is processed by the internal in-memory processing unit, thereby reducing the output load on the interface of the memory controller 100 and significantly reducing the energy consumption.

[0047] Specifically, the acceleration circuit includes an interface processing and parsing unit 11 (referred to as the "parsing unit" in the figure), a processor in the controller 12 (i.e., the processor in the controller), and a controller command processing engine 13 (referred to as the "command processing engine" in the figure).

[0048] Among them, the interface processing and parsing unit 11 includes a read / write arbiter 111, a command scheduler 112, and a data scheduler 113. The interface processing and parsing unit 11 is used to interface with the processor 20, receive arithmetic instructions from the processor 20, and send the corresponding arithmetic results to the processor 20.

[0049] The processor 12 in the controller mainly includes a general matrix vector multiplication near-memory processing unit 121 for processing the General Matrix Vector Multiplication (GEMV) algorithm and a group of asynchronous FIFO queues with the number of DRAM banks as the parallel granularity, including an asynchronous request first-in first-out (FIFO) queue 122 and an asynchronous response first-in first-out queue 123.

[0050] The controller command processing engine 13 interacts with the asynchronous request FIFO queue 122 and the asynchronous response FIFO queue 123. According to the request commands in the asynchronous request FIFO queue 122, it sends out the corresponding command sequence to the 3D stacked DRAM according to the DDR protocol requirements, and writes the returned read data into the asynchronous response FIFO queue 123.

[0051] As Figure 2 shown, in this embodiment, the interface processing and parsing unit 11 and the processor 12 in the controller operate in a low-frequency clock domain, and the controller command processing engine 13 operates in a high-frequency clock domain. Among them, low frequency usually refers to 100 MHz to 400 MHz, and high frequency usually refers to above 400 MHz. For example, DDR4 is 800 MHz to 1600 MHz, DDR5 is 1600 MHz to 3200 MHz, and DDR6 is 6400 MHz or even higher.

[0052] Figure 3 is the structural schematic diagram of the processor in the controller in this embodiment.

[0053] As Figure 3 shown, in addition to the general matrix vector multiplication near-memory processing unit 121, the asynchronous request FIFO queue 122, and the asynchronous response FIFO queue 123, the processor 12 in the controller further includes a top-level state machine 124, a decoder 125, a configuration module 126, a command generator 127, and an output buffer 128.

[0054] The top-level state machine 124 is used to control the reception and transmission of instructions and data. The instructions from the CPU or GPU are decoded by the decoder 125, and then the command generator 127 generates the corresponding request commands and stores them in the asynchronous request FIFO queue 122. The general matrix vector multiplication near-memory processing unit 121 obtains the returned read data from the asynchronous response FIFO queue 123, performs GEMV calculations based on the read data and the configuration information stored in the configuration module 126, and then stores the calculation results in the output buffer 128.

[0055] Figure 4 is the schematic diagram of the working principle of the near-memory processing unit in this embodiment.

[0056] As Figure 4As shown, the general matrix-vector multiplication near-memory processing unit 121 supports two working modes: direct passing of DRAM read data and GEMV calculation processing of DRAM read data. Among them, the GEMV block instantiates a group of matrix-vector multiplication processing blocks 1210, and each matrix-vector multiplication processing block 1210 is a computing-in-memory macro (CIM Macro). The CIM Macro consists of a storage array 1211 composed of 8T cells, an address decoder 1212, a write driver 1213, a multiplier 1214, and an accumulation tree 1215.

[0057] Among them, the storage array 1211 is used to store the weight matrix prefetched from the DRAM, that is, as a prefetched weight cache. During the GEMV calculation process of the matrix-vector multiplication processing block 1210, it traverses the CIM Macro address from low to high to read out the complete weight vector row from the storage array 1211, and during the process of reading the weight vector, it completes the multiply-accumulation operation for each element of the read weight vector row and each element of the activation column vector, and finally obtains the dot product result of the two vectors. And during the above GEMV calculation process, it reads the weight matrix for the next operation from the DRAM, traverses the CIM Macro address from low to high, and stores the read weight matrix into the storage array 1211, that is, updates the weight matrix of the next address prefetched from the DRAM to the prefetched weight cache. This processing method that supports parallel calculation and DRAM read operations avoids the waiting delay introduced by reading the weight matrix and significantly improves the interface throughput.

[0058] Figure 5 It is a schematic diagram of the principle of prefetched weight matrix in this embodiment. The upper row in the figure shows the schematic diagram of the data flow processing process without support for prefetching, and the lower row in the figure shows the schematic diagram of the data flow processing process with support for prefetching.

[0059] As Figure 5 shown in the upper row of the figure, in the case of not supporting the prefetch operation, the data flow processing process is as follows: when receiving a request command, first obtain the activation vector from the activation buffer, then read back the weight matrix from the DRAM, then perform the GEMV operation, and finally output the operation result data in burst form. That is, in the case of no prefetch operation, before each GEMV calculation, the weight needs to be moved to the GEMV module first, and then the GEMV calculation is performed after completion, which will cause a large waiting delay.

[0060] As Figure 5As shown in the following figure, in the case of supporting the prefetch operation, the GEMV calculation is processed in parallel with the operation of reading weights from the DRAM, and the read-back weights are updated into the prefetch weight cache of the GEMV module. To ensure that the weight matrix has been updated into the prefetch weight cache before the current request command processing ends, it is necessary to determine whether the weight reading status has ended. After the end, the last beat rlast of the last burst output is pulled high.

[0061] Figure 6 This is a schematic diagram of the time-domain partitioning in the prefetch weight cache in this embodiment. This partitioning scheme is set according to the requirements of the general matrix-vector multiplication operation in the DRAM data transfer scenario.

[0062] As Figure 6 shown, the scales of the weight matrix and the activation vector are: W ∈ R H×W×op_width , V ∈ R H×1×op_width , where H is the height, W is the width, and op_width is the width of the operand (operator). Set the bit width fifo_width of the asynchronous FIFO queue to the data length that the storage controller 10 can respond to in one read request command. Then, any row vector of the weight matrix can be divided into n sub-vectors w ij , where m in the figure is the number of banks. The result of the matrix multiplication operation:

[0063]

[0064] Starting from when the storage controller 10 starts to read weights until the first weight sub-vector is updated to the prefetch weight buffer, it takes nT to complete the update of all the weight vectors in one row, where T is one clock cycle. The number of parallel banks is set to m, so there are m weight vectors participating in the update and the multiplication operation with the activation column vector.

[0065] Figure 7 This is a schematic diagram of the space of the prefetch weight cache in this embodiment.

[0066] As Figure 7 shown, in the prefetch weight cache (i.e., the above-mentioned storage array 111), the width of each cache unit (Buffer) is fifo_width, and the depth is n; there are m Buffers for parallel processing. The Buffer reads out the prefetch weight sub-vectors in sequence from the low address to the high address, and completes the dot product operation with the activation sub-vectors of the same bit width as the activation vector. Almost simultaneously, the weight sub-vectors prefetched from the DRAM are updated to the corresponding read addresses in sequence from the low address to the high address.

[0067] Figure 8 This is a schematic diagram of the timing of parallel read and write for in-memory computing in this embodiment.

[0068] AsFigure 8 As shown, the relevant read and write control signals of the processor 12 in the controller include read enable rd_en, read address rd_addr, read data rd_data, dot product partial sum result psum_data, update write enable update_en, write address wr_addr, and prefetch weight data prefetch_data.

[0069] Starting from when the first weight data returns to the asynchronous response FIFO queue 123, the CIM read enable rd_en is pulled high, and the read address rd_addr is controlled to change sequentially from low to high. In the figure, psum_data is the dot product partial sum result of each sub-vector. Starting at least one cycle after the read enable rd_en is pulled high, the update write enable update_en is pulled high, so that the new weight data sequentially updates and overwrites the original weight data in the prefetch weight buffer. This operation ensures the parallel processing of GEMV calculation and weight reading, avoiding the inefficient processing of first retrieving the weights and then calculating GEMV.

[0070] Figure 9 It is a schematic diagram of the timing constraints for one GEMV processing in this embodiment.

[0071] As Figure 9 shown, when the matrix-vector multiplication operation W·V can be processed in one command request (one Read_dram_with_gemv CMD), T decode +T prefetch_initial +T W·V +T burst_output is the time for the near-memory processing unit 121 to process one request command. Among them, T decode is the time for one address decoding, T prefetch_initial is the time from issuing the prefetch request for weights to the asynchronous response FIFO queue 123 receiving the first weight sub-vector, T delay is the predetermined delay time, T update is the time to update the weight matrix at the next pre-fetched address to the prefetch weight cache, T W·V is the time for one GEMV operation, T burst_output is the time to write the GEMV operation result to the output cache module 128. That is, after the asynchronous response FIFO queue 123 receives the first weight sub-vector, it enters the stage where GEMV calculation and weight update proceed in parallel.

[0072] When the matrix-vector multiplication operation W·V cannot be processed in one command request, that is, W·V in =V out of V outIf it cannot be fully output in one burst operation, multiple read output buffer instructions requests (multiple Read_obuf CMDs) need to be sent subsequently to read the remaining data from the obuffer. The complete matrix-vector multiplication result is the concatenation of all output results.

[0073] Among them, if the scale of the weight matrix is very large, the GEMV calculation and weight update operations need to be completed in blocks. Specifically, when the number of matrix row data > the number of parallel banks, buffer resources need to be multiplexed, and block operations are required. The processing time for each block is n time cycles, and the number of blocks is: the total number of rows h of the weight matrix / the number of rows m processed in parallel of the weight matrix.

[0074] For example, for an operand of 16×16×8b, there are only 8 banks, so the 16-row matrix data is divided into 2 blocks for processing (16 rows / 8 banks). Each bank occupies 1 CIM block and processes one row of matrix data. The bit width of one row of data is 16x8bit = 128bit. If the width of the FIFO is 64bit, 2 cycles need to be processed.

[0075] Regarding the setting of the delay time, it is required that the start time of weight update is at least one cycle later than the start time of GEMV calculation, but it is also necessary to ensure that the weight matrix of the next address is updated before the start of the next GEMV calculation. Therefore, the delay requirements for GEMV calculation and weight update are: when n = 1, the delay time is 1T; when n ≥ 2, the delay time can be selected between [1T, (n - 1)T].

[0076] Since the weight update in the weight prefetch cache also multiplexes the transmission delay of burst output, it is necessary to ensure that the weight update has ended before the burst output is about to end, that is, it is required that T update <T W·V +T burst_output 。

[0077] In addition, to ensure that the near-memory processing unit 121 for general matrix-vector multiplication can correctly execute the next request command, each request command needs to check whether the last round of weight update (update of weight sub-vectors) has been completed before sending the last beat of data. Only when it has been completed can the last beat of data be sent. If the scale of the weight matrix to be processed is very large and the GEMV result cannot be output within one request command with the existing maximum bus bit width, multiple read output Buffer request commands need to be continuously sent until all data is read out.

[0078] Functions and effects of the embodiments

[0079] Compared with the acceleration circuits in the prior art, the in-controller near-memory computing acceleration circuit supporting weight prefetch provided by this embodiment has the following advantages in multiple aspects:

[0080] First, the acceleration circuit of this embodiment adopts the technology of integrating a near-memory processing unit in the controller (PIC technology). Compared with the solutions of integrating PIM in DRAM or integrating a dedicated acceleration engine near the CPU / GPU side, the acceleration circuit of the present invention can achieve significant data transfer efficiency and energy consumption benefits with lower cost and a shorter storage-computation path.

[0081] Second, in the near-memory processing unit, by prefetching the weight matrix for the next operation from DRAM in parallel during the matrix-vector multiplication operation to the computing circuit in the module, the waiting delay for reading the weight matrix from DRAM is eliminated, thereby achieving a higher interface throughput rate; further combined with the ping-pong pipelining processing of computing, low latency and high bus throughput rate can be achieved.

[0082] Third, the near-memory processing unit and the controller processing engine complete the transmission of read / write commands / data through a group of asynchronous FIFOs. It supports flexible clock frequency configuration, and the near-memory processing achieves low-power data transfer between the interface and the system through frequency reduction processing.

[0083] Fourth, the acceleration circuit of this embodiment supports GEMV calculations with a flexible and configurable matrix size, and supports any matrix size and operand type. It can achieve DRAM bank-level parallel transfer and energy-efficient GEMV calculation processing. Therefore, it can be adapted to various current AI applications. For example, when the size is small, it supports applications such as CNN and DNN image processing; when the size is very large, it supports applications such as transformers.

[0084] In the embodiment, the parallel processing of GEMV calculation and weight reading is ensured through the setting of relevant read / write control signals, and it is checked whether the update of the last round of weight sub-vectors has been completed before the last beat of data of each command is sent. Therefore, the correct execution of the next command can be ensured, thereby improving the data transfer efficiency and avoiding problems in data transmission or operation caused by this special mechanism, and ensuring the accuracy of the operation.

[0085] Furthermore, for a very large weight matrix, the GEMV calculation and weight update operations can be completed in blocks, and the corresponding delay time can be set to ensure that the weight update operations are parallel and completed in a timely manner. Therefore, the acceleration circuit of the embodiment can also be used for GEMV operations of very large weight matrices.

[0086] The above embodiments are only used to illustrate the specific embodiments of the present invention, and the present invention is not limited to the scope of description of the above embodiments. Those skilled in the art should understand that the present invention is not restricted by the above embodiments. What is described in the above embodiments and the specification only explains the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A near memory computing acceleration circuit supporting weight pre-fetching in a controller, arranged in a storage controller interconnected with a three-dimensionally stacked dynamic random access memory and connected to a processor, characterized in that: include: An interface processing and parsing unit, used for receiving operation instructions from the processor and sending corresponding operation results to the processor; A processor in the controller, including a near memory processing unit for processing general matrix-vector multiplication, an asynchronous request first-in-first-out queue, and an asynchronous response first-in-first-out queue; The controller command processing engine sends a corresponding command sequence to the dynamic random access memory according to the request command in the asynchronous request first-in-first-out queue, and writes the returned read data into the asynchronous response first-in-first-out queue. The near memory processing unit includes a plurality of matrix-vector multiplication processing blocks. Each of the matrix-vector multiplication processing blocks has a pre-fetch weight cache, While performing the current matrix-vector multiplication operation, the matrix-vector multiplication processing block pre-fetches the weight matrix corresponding to the next matrix-vector multiplication from the dynamic random access memory and stores it in the pre-fetch weight cache.

2. The near memory computing acceleration circuit supporting weight pre-fetching in the controller according to claim 1, characterized in that: in, The near memory computing acceleration circuit is oriented to a three-dimensional stacked dynamic random access memory architecture, and the near memory computing acceleration circuit and the three-dimensional stacked dynamic random access memory are interconnected with the storage controller through silicon vias or hybrid bonding. The near-memory computing acceleration circuit can realize bank-level parallel processing of the three-dimensional stacked dynamic random access memory.

3. The near memory computing acceleration circuit supporting weight pre-fetching in the controller according to claim 1, Features: Wherein, the matrix-vector multiplication processing block includes: A storage array, used as the pre-fetch weight cache; A write driver, used for writing the pre-fetched weight matrix into the storage array; An address decoder, for decoding the read address; and A multiplier and an accumulator tree for performing the matrix-vector multiplication operation, The processor in the controller also includes: A top-level state machine, used for controlling the reception and transmission of the request command and the read data; A decoder, for decoding instructions from the processor; A command generator, used for generating a corresponding request command according to the decoded instruction, and storing it in the asynchronous request first-in-first-out queue; A configuration module, used to store configuration information; and Output caching, The near memory processing unit obtains the returned read data from the asynchronous response first-in-first-out queue, performs a matrix-vector multiplication operation based on the read data and the configuration information, and stores the operation result in the output cache.

4. The near memory computing acceleration circuit supporting weight pre-fetching in the controller according to claim 3, characterized in that: in, Each of the matrix-vector multiplication processing blocks is an in-memory calculation macro. The matrix vector multiplication processing block traverses the in-memory calculation macro address from the low address to the high address to read out the complete weight vector row from the storage array. During the reading process, each element in the weight vector row is multiplied and accumulated with each element in the laser column vector to obtain the dot product result of the two vectors. During the multiplication and accumulation operation, the processor in the controller pre-fetches the weight matrix for the next operation from the dynamic random access memory, and the matrix-vector multiplication processing block traverses the in-memory calculation macro address from low address to high address, and writes the pre-fetched weight matrix into the storage array, thereby updating the cached pre-fetched weight matrix.

5. The near memory computing acceleration circuit supporting weight pre-fetching in the controller according to claim 4, characterized in that: in, The bit width of the asynchronous response FIFO queue is set to the data length that the storage controller can respond to in one read request command. The pre-fetch weight cache comprises a plurality of cache units, The width of each cache unit is the bit width of the asynchronous response first-in-first-out queue. The plurality of cache units read out the pre-fetched weight sub-vectors of the weight matrix in parallel from low address to high address in sequence, and the weight sub-vectors are used to complete the dot multiplication operation with the corresponding activation sub-vector in the activation vector. At the same time, the weight sub-vectors of the next operation pre-fetched from the dynamic random access memory are updated in sequence from low address to high address to the corresponding read address.

6. The near memory computing acceleration circuit supporting weight pre-fetching in the controller according to claim 4, characterized in that: in, The read and write control signals of the processor in the controller include a read enable signal, a read address signal, a read data signal, an update write enable signal, a partial sum output signal, a write address signal, and a pre-fetch weight data signal. Starting from the first weight data being returned to the asynchronous response FIFO queue, the read enable signal is pulled high to control the read address signal to change from low to high in sequence. Starting at least one cycle after the read enable signal is pulled high, the update write enable signal is pulled high, so that the pre-fetched new weight data sequentially updates and covers the original weight data in the pre-fetch weight cache.

7. The near memory computing acceleration circuit supporting weight pre-fetching in a controller according to claim 4, characterized in that: in, For each of the request commands, the near memory processing unit checks whether the last round of updating of the weight sub-vector has been completed before sending the last beat of data, and sends the last beat of data when it has been completed.

8. The near memory computing acceleration circuit supporting weight pre-fetching in a controller according to claim 4, characterized in that: in, When the matrix row data > the number of parallel banks, the vector multiplication unit completes the matrix-vector multiplication operation in blocks. The processing time of each block is n time cycles, and the number of blocks is the total number of rows of the weight matrix h / the number of parallel processing rows of the weight matrix m. The size of the general matrix-vector multiplication is: W∈R H×W×op_width V∈R H×1×op_width Where H is the height of the matrix, W is the width of the matrix, and op_width is the width of the operand.

9. The near memory computing acceleration circuit supporting weight pre-fetching in a controller according to claim 8, characterized in that: in, After the general matrix-vector multiplication starts to be calculated and a predetermined delay time has passed, the cached pre-fetch weight matrix is ​​updated, When n=1, the delay time is 1T; When n≥2, the delay time is selected from the range of [1T, (n-1)T], Where T is the clock period.

10. The near memory computing acceleration circuit supporting weight pre-fetching in a controller according to claim 1, characterized in that: in, The interface processing and parsing unit and the processor in the controller operate in a low-frequency clock domain, and the controller command processing engine operates in a high-frequency clock domain.