Apparatus and Method for Matrix Multiplication Using In-Memory Processing
By dividing the PIM block array into memory mode and computing mode groups, and using MUX redirected data streams and VVM engine computing, the problems of low energy efficiency and data redundancy in the existing ReRAM architecture are solved, and high energy efficiency and flexible matrix multiplication operations are achieved.
Patent Information
- Application Number
- CN202080001761.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-07-07
AI Technical Summary
The existing ReRAM-based NVPIM architecture has problems such as low energy efficiency, excessive data copying and excessive writes when performing matrix multiplication, especially in deep convolutional neural networks, which affects system energy efficiency.
Using a reconfigurable PIM architecture, the PIM block array is divided into groups of memory mode and computing mode, redirecting data streams through MUX, reducing the use of ADC/DAC, using the VVM engine for dot product calculations, avoiding high-cost data exchange, and optimizing data mapping and layout to reduce redundant writes.
It improves the energy efficiency and flexibility of matrix multiplication, reduces data copying and writing, reduces system energy consumption, and adapts to the needs of different computing tasks.
Smart Images

Figure CN114158284B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to processing-in-memory (PIM). Background Art
[0002] Ultra-low power machine learning processors are crucial for performing cognitive tasks in embedded systems because the power budget is limited, for example, in the case of battery or energy harvesting sources. However, the data generated by deep convolutional neural networks (DCNNs) results in heavy traffic between the memory and the computing units in a conventional von Neumann architecture and adversely affects the energy efficiency of these systems. As a promising solution to accelerate DCNN execution, resistive random access memory (ReRAM)-based non-volatile PIM has emerged. The high cell density of ReRAM allows large on-chip ReRAM arrays to be implemented on a chip to store the parameters of DCNNs, while suitable functions, such as vector matrix multiplication (VMM), can be directly executed in the ReRAM array and its peripheral circuits. Summary of the Invention
[0003] Embodiments of apparatuses and methods for matrix multiplication using PIM are disclosed herein.
[0004] In one example, an apparatus for matrix multiplication includes an array of PIM blocks in row and column form, a controller, and an accumulator. Each PIM block is configured to be in a compute mode or a memory mode. The controller is configured to divide the array of PIM blocks into: a first group of PIM blocks, each PIM block being configured to be in the memory mode; and a second group of PIM blocks, each PIM block being configured to be in the compute mode. The first group of PIM blocks is configured to store a first matrix, and the second group of PIM blocks is configured to store a second matrix and calculate partial sums of a third matrix based on the first and second matrices. The accumulator is configured to output the third matrix based on the partial sums of the third matrix.
[0005] In some embodiments, the first group of PIM blocks includes the rows of the array of PIM blocks, and the second group of PIM blocks includes the remainder of the array of PIM blocks.
[0006] In some embodiments, the size of the first group of PIM blocks is smaller than the size of the first matrix. In some embodiments, the controller is configured to map the first matrix to the first group of PIM blocks such that the first group of PIM blocks stores a portion of the first matrix simultaneously based on the size of the first group of PIM blocks.
[0007] In some embodiments, the apparatus further includes a multiplexer (MUX) between the first and second sets of PIM blocks. In some embodiments, the controller is further configured to control the MUX to direct data from each PIM block of the first set of PIMs to a PIM block in a corresponding column of the second set of PIM blocks.
[0008] In some embodiments, at least one PIM block of the first set of PIM blocks is not in the same corresponding column.
[0009] In some embodiments, the size of the second set of PIM blocks matches the size of the second matrix.
[0010] In some embodiments, the second matrix has three or more dimensions. In some embodiments, the controller is configured to map the second matrix to the second set of PIM blocks such that each PIM block of the second set of PIM blocks stores a corresponding part of the second matrix.
[0011] In some embodiments, the controller is configured to control the second set of PIM blocks to perform a convolution of the first and second matrices when calculating a partial sum of a third matrix.
[0012] In some embodiments, the first matrix includes a feature map in a convolutional neural network (CNN), and the second matrix includes a kernel of the CNN.
[0013] In some embodiments, each PIM block in the PIM block array includes a memory array and a vector-vector multiplication (VVM) engine configured to be deactivated in memory mode.
[0014] In another example, a PIM device includes a memory array, a vector-vector multiplication (VVM) engine, and control circuitry. The memory array is configured to store a first vector. The control circuitry is configured to enable the VVM engine in compute mode and control the VVM engine to perform a dot product between the first vector and a second vector to generate a partial sum. The control circuitry is further configured to deactivate the VVM engine in memory mode and control the memory array to write or read the first vector.
[0015] In some embodiments, the VVM engine includes a bit counter, a shift accumulator, and a plurality of AND gates.
[0016] In some embodiments, the PIM device further includes a first buffer configured to receive and buffer the second vector from another PIM device.
[0017] In some embodiments, the PIM device further includes a second buffer configured to buffer the partial sum and send the partial sum to another PIM device.
[0018] In some embodiments, the memory array includes a ReRAM array.
[0019] In yet another example, a method for matrix multiplication implemented by an array of PIM blocks in the form of rows and columns is disclosed. Each of a first set of PIM blocks of the PIM block array is configured by a controller to be in a memory mode, and each of a second set of PIM blocks of the PIM block array is configured by the controller to be in a compute mode. A first matrix is mapped by the controller to the first set of PIM blocks, and a second matrix is mapped by the controller to the second set of PIM blocks. Partial sums of a third matrix are calculated by the second set of PIM blocks based on the first and second matrices. The third matrix is generated based on the partial sums of the third matrix.
[0020] In some embodiments, the first set of PIM blocks includes the rows of the PIM block array, and the second set of PIM blocks includes the remainder of the PIM block array.
[0021] In some embodiments, the size of the first set of PIM blocks is smaller than the size of the first matrix. In some embodiments, a portion of the first matrix is stored simultaneously by the first set of PIM blocks based on the size of the first set of PIM blocks.
[0022] In some embodiments, data from each PIM block of the first set of PIM blocks is directed by a MUX between the first set of PIM blocks and the second set of PIM blocks to a PIM block in a corresponding column of the second set of PIM blocks.
[0023] In some embodiments, at least one PIM block of the first set of PIM blocks is not in the same corresponding column.
[0024] In some embodiments, the size of the second set of PIM blocks matches the size of the second matrix.
[0025] In some embodiments, the second matrix has three or more dimensions. In some embodiments, a corresponding portion of the second matrix is stored by each PIM block of the second set of PIM blocks.
[0026] In some embodiments, to calculate the partial sums of the third matrix, a convolution of the first and second matrices is performed.
[0027] In some embodiments, the first matrix includes a feature map in a CNN, and the second matrix includes a kernel of the CNN. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings incorporated herein and forming a part of the specification illustrate embodiments of the present disclosure and, together with the specification, further serve to explain the principles of the present disclosure and enable one of ordinary skill in the art to use the present disclosure.
[0029] Figure 1 A block diagram of an apparatus including an array of PIM blocks is shown.
[0030] Figure 2 A block diagram of an exemplary apparatus including an array of reconfigurable PIM blocks according to some embodiments of the present disclosure is shown.
[0031] Figure 3 A block diagram of an exemplary apparatus including an array of PIM blocks for matrix multiplication according to some embodiments of the present disclosure is shown.
[0032] Figure 4 An exemplary PIM block in the apparatus according to some embodiments of the present disclosure is shown. Figure 3 in the apparatus according to some embodiments of the present disclosure is shown.
[0033] Figure 5A An exemplary PIM block in the apparatus according to some embodiments of the present disclosure is shown. Figure 4 A detailed block diagram of the PIM block shown in memory mode according to some embodiments of the present disclosure is shown.
[0034] Figure 5B An exemplary PIM block in the apparatus according to some embodiments of the present disclosure is shown. Figure 4 A detailed block diagram of the PIM block shown in compute mode according to some embodiments of the present disclosure is shown.
[0035] Figure 6A An exemplary mapping scheme for PIM blocks in compute mode in matrix multiplication according to some embodiments of the present disclosure is shown.
[0036] Figure 6B An exemplary mapping scheme for PIM blocks in memory mode in matrix multiplication according to some embodiments of the present disclosure is shown.
[0037] Figure 7 An exemplary data flow between different PIM blocks in matrix multiplication according to some embodiments of the present disclosure is shown.
[0038] Figure 8 An exemplary compute flow between different PIM blocks in matrix multiplication according to some embodiments of the present disclosure is shown.
[0039] Figure 9 A flowchart of an exemplary method for matrix multiplication implemented by an array of PIM blocks according to some embodiments of the present disclosure is shown.
[0040] Embodiments of the present disclosure will be described with reference to the accompanying drawings. DETAILED DESCRIPTION
[0041] Although the configurations and arrangements of the present invention have been discussed, it should be understood that this discussion is for illustrative purposes only. Those skilled in the art can understand that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. It is obvious to those skilled in the art that the present invention can also be used in many other applications.
[0042] It should be noted that the "one embodiment", "one implementation", "exemplary embodiment", "some embodiments", etc. mentioned in the specification of the present invention mean that the described embodiments may include specific features, structures or characteristics, but not every embodiment necessarily includes such specific features, structures or characteristics. In addition, such expressions do not necessarily refer to the same embodiment. In addition, when a specific feature, structure or characteristic is described in connection with a certain embodiment, it is within the knowledge scope of those skilled in the art to implement such specific features, structures or characteristics in combination with other embodiments, regardless of whether it is explicitly stated herein.
[0043] Generally speaking, terms can be understood at least in part according to their use in the context. For example, the term "one or more" used herein can be used, at least in part according to the context, to describe any feature, structure or characteristic in the singular form or to describe a combination of features, structures or characteristics in the plural form. Similarly, terms such as "a", "an", or "the" can also be understood at least in part according to the context as expressing singular usage or expressing plural usage.
[0044] Existing ReRAM-based NVPIM architectures are mainly in a mixed-signal manner. For example, some NVPIM devices include a large number of digital / analog converters and analog / digital converters (ADC / DAC) to respectively transform digital inputs into analog signals for PIM operations and then transform the calculation results back into digital format. However, the ADC / DAC occupies most of the area and power consumption of these PIM designs. Digital PIM architectures have been proposed to improve energy efficiency by eliminating A / D conversion and to improve the overall design flexibility in dealing with such randomness. Some NVPIM devices, for example, attempt to use in-memory "NOR" logic operations to implement VMM. However, in-memory "NOR" logic operations require initializing memory cells to a low-resistance state (LRS) at the start of each "NOR" operation. When combining "NOR" logics to calculate multiplication or accumulation, additional memory space is also required to store each intermediate "NOR" result. In addition, the high performance of existing NVPIM architectures is achieved at the cost of excessive data replication or writing.
[0045] At a higher level, when performing matrix multiplication in a CNN, it is known that NVPIM architectures typically use NVPIM blocks as computing devices in combination with random access memory (RAM) to store input and output data as well as intermediate results from the calculations. For example, Figure 1A block diagram of an apparatus 100 including an array of PIM blocks is shown. The apparatus 100 includes a RAM 102, an array of PIM blocks 104, and a bus 106. Each PIM block 104 includes an ADC and a DAC for converting a digital input into an analog signal for PIM operations and then converting the calculation result back into a digital format, respectively. Each PIM block 104 further includes a ReRAM array, each of which is configured to switch between two or more levels by applying electrical excitations with different amplitudes and durations. Each ReRAM element represents a vector by the voltage at the input, and then the bit-line current collected at the output forms the VMM result. The input and output data are transmitted through the bus 106, such as the main / system bus of the apparatus 100, and stored in the RAM 102, such as the main / system memory of the apparatus 100, rather than in the PIM blocks 104. During matrix multiplication, intermediate data also needs to be frequently exchanged between each PIM block 104 and the RAM 102.
[0046] According to various embodiments of the present disclosure, a reconfigurable PIM architecture is provided, which has higher energy efficiency and flexibility in various matrix multiplication applications, such as convolution in CNN. At the PIM block level, each PIM block can be reconfigured to be in either a computing mode or a memory mode. In the computing mode, partial sums of the dot product between vectors can be calculated in a digital circuit, i.e., a VVM engine, to eliminate the high-cost ADC / DAC. At a higher level, an array of the same PIM blocks and multiple MUXs can be reconfigured into an optimized arrangement according to the specific task to be performed. In some embodiments, the PIM block array is configured with a first group of PIM blocks in the memory mode and a second group of PIM blocks in the computing mode for performing matrix multiplication, such as convolution in CNN. For example, the data of the feature map and the weights of the kernel can be aligned based on the configuration of the PIM array using an efficient data mapping scheme, thereby avoiding excessive data replication or writing in the previous NVPIM architecture.
[0047] Figure 2A block diagram of an exemplary apparatus 200 including an array of reconfigurable PIM blocks in accordance with some embodiments of the present disclosure is shown. The apparatus 200 may include an array of PIM blocks 202, a plurality of MUXs 204, a controller 206, an accumulator 208, a global functional unit 210, and a bus 212. Each PIM block 202 may be the same and configured to be either in a memory mode for storing data such as vectors or matrices in two or more dimensions or in a computing mode for storing data and performing vector / matrix computations such as VMM or VVM. As the specific task to be performed changes, such as the computations in the convolutional layer or the fully connected (FC) layer in a CNN, each PIM block 202 may be reconfigured between the computing mode and the memory mode based on the computational scheme of the specific task. In some embodiments, even though the layout of the PIM block 202 array is preset, such as in the form of orthogonal rows and columns, the configuration of the MUX 204 can still be flexibly changed according to the specific task to be performed. For example, by enabling and disabling certain MUXs 204 between different rows of PIM blocks 202, the arrangement of the PIM block 202 array can be configured to adapt to the computational scheme and data flow corresponding to the specific task. According to some embodiments, the enabled MUXs 204 divide the PIM block 202 array into two or more groups, each group being configured to be in the same computing or memory mode. Additionally, although the default data flow between PIM blocks 202 is in the same row and / or column, the enabled MUXs 204 can further redirect the data flow between different rows and / or columns as needed for the specific task.
[0048] The bus 212 may be the main / system bus of the apparatus 200, which is used to transfer input data such as matrices to the PIM block 202 array. Different from the apparatus 100 in Figure 1 —where the apparatus 100 includes a centralized RAM 100 for storing input and output as well as intermediate results—, a group of PIM blocks 202 in the apparatus 200 may be configured to be in the memory mode to replace the RAM 102. As a result, according to some embodiments, the data flow is no longer between each PIM block 104 and the centralized RAM 102, but follows a specific path based on the arrangement of the PIM block 202 array, such as the layout of the PIM block 202 array and / or the configuration of the MUX 204. The output of the PIM block 202 array, such as a partial sum, may be sent to the accumulator 208, which may be further configured to generate an output matrix based on the partial sums.
[0049] In some embodiments, the global functional unit 210 is configured to perform any suitable global miscellaneous functions, such as pooling, activation, and encoding schemes. For example, the global functional unit 210 can perform a zero flag encoding scheme to save unnecessary writes of zero values into the memory array. Additional columns can be added to the memory array by the global functional unit 210 to store zero flags, which correspond to high-precision data stored in multiple columns of the memory array. The default value of the zero flag can be "0", which indicates that the data is non-zero. If the data is zero, the zero flag can be set to "1", which indicates that no write to the storage unit is performed to save write energy. During the preparation phase, which will be described in detail below, the loaded bits of the data marked as zero can all be set to "0" if the zero flag of the data is "1". The controller 206 is configured to control the operation of other components of the device 100, such as data mapping and data flow, and the computing schemes of the PIM block 202 array and the MUX 204. In some embodiments, the controller 206 is also configured to reconfigure the data precision of each VVM operation.
[0050] Figure 3 A block diagram of an exemplary device 300 including an array of PIM blocks for matrix multiplication according to some embodiments of the present disclosure is shown. The device 300 can include: a plurality of PIM banks 302, each of which is for matrix multiplication; and an input / output (I / O) interface 304, which is configured to exchange data with other devices such as a host processor and / or system memory. In some embodiments, the device 300 is configured to perform various operations in a CNN. For example, some groups 306 of the PIM banks can receive images from the I / O interface 304, perform operations in the convolutional layer of the CNN, and generate intermediate feature maps. Another group 308 of the PIM banks can receive the intermediate feature maps, perform operations in the fully connected (FC) layer of the CNN, and generate labels of the images, which are sent back to the I / O interface 304.
[0051] Figure 3 An exemplary PIM bank 302 for performing matrix multiplication in the convolutional layer and the FC layer of a CNN is also shown. Figure 3 The PIM bank 302 in Figure 2An example of the apparatus 200, as described in detail below, the apparatus 200 is configured based on matrix multiplication tasks in the convolutional layer and the FC layer in the CNN. In some embodiments, the controller 206 is configured to divide the PIM block 202 array into: a first group of PIM blocks 312, each PIM block being configured to be in memory mode; and a second group of PIM blocks 314, each PIM block being configured to be in compute mode. The first group of PIM blocks 312 can be configured to store a first matrix, such as an input feature map in the CNN. The second group of PIM blocks 314 can be configured to store a second matrix, such as a kernel of the CNN, and calculate a partial sum of a third matrix, such as an output feature map, based on the first and second matrices. In some embodiments, the accumulator 208 is configured to receive the partial sums from the second group of PIM blocks 314 and output the third matrix based on the partial sums of the third matrix. For example, the accumulator 208 can generate each element of the third matrix based on the corresponding partial sums.
[0052] As Figure 3 shown, according to some embodiments, the first group of PIM blocks 312 includes rows of the PIM block 202 array, such as the first row of the PIM block 202 array. According to some embodiments, the second group of PIM blocks 312 then includes the remaining portion of the PIM block 202 array. In some embodiments, only one MUX 204 between the first and second groups of PIM blocks 314 and 312 in the MUX 204 is enabled. That is, the controller 206 can be configured to enable the MUX 204 between the first and second rows of the PIM block 202 array and disable the remaining MUX 204 (not shown in Figure 3 to divide the PIM block 202 array into a first group of PIM blocks 312 in memory mode and a second group of PIM blocks 314 in compute mode. As will be described in detail below, the controller 206 is also configured to control the MUX 204 to direct data from each PIM block 202 of the first group of PIM blocks 312 to the PIM block 202 in the corresponding column of the second group of PIM blocks 314. The data flow path between the first and second groups of PIM blocks 312 and 314 can be changed under the control of the controller 206 based on the corresponding matrix multiplication scheme at different stages in the matrix multiplication. For example, at least one PIM block 202 of the first group of PIM blocks 312 is not in the same corresponding column. In one example, the data loaded in the first column of the first group of PIM blocks 312 can be redirected by the MUX 204 to the second column, the third column, or any column other than the first column of the second group of PIM blocks 314.
[0053] As described above, each PIM block 202 can be the same PIM device that can be configured to be in either memory mode or compute mode. Refer to Figure 4, each PIM block 202 may include a memory array 402 and a VVM engine 404, and the VVM engine 404 is configured to be deactivated in the memory mode. In some embodiments, the memory array 402 includes a ReRAM array. It can be understood that in other examples, the memory array 402 may include any other suitable memory. By way of example, the memory includes, but is not limited to: phase change random access memory (PRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM). The memory array 402 may store a first vector. The PIM block 202 may further include a control circuit 406, and the control circuit 406 is configured to enable the VVM engine 404 in the computing mode and control the VVM engine 404 to perform a dot product between the first vector and the second vector to generate a partial sum. The control circuit 406 may further be configured to deactivate the VVM engine 404 in the memory mode and control the memory array 402 to write or read the first vector.
[0054] In some embodiments, the PIM block 202 further includes peripheral circuits, including a row decoder 416 and a column decoder 414 for facilitating the operation of the memory array 402. The peripheral circuits may include any suitable digital, analog, and / or mixed-signal peripheral circuits for facilitating the operation of the memory array 402. For example, the peripheral circuits may include one or more of the following: page buffers, decoders (such as row decoder 416 or column decoder 414), sense amplifiers, drivers, charge pumps, current or voltage references, or any active or passive components of the circuit (such as transistors, diodes, resistors, or capacitors). The PIM block 202 may further include a memory I / O interface 412, and the memory I / O interface 412 is operatively coupled between the memory array 402 and the memory bus to write and read the first vector between the memory bus and the memory array 402. The PIM block 202 may further include various buffers for intermediate data storage, including: a column buffer 408, which is configured to receive and buffer, for example, a second vector from other PIM devices through the memory bus; and a partial sum buffer 410, which is configured to buffer the partial sum and send the partial sum to another PIM device through the partial sum bus.
[0055] Figure 5A A detailed block diagram of the PIM block 202 in the memory mode according to some embodiments of the present disclosure is shown. Figure 5B A detailed block diagram of the PIM block 202 in the computing mode according to some embodiments of the present disclosure is shown. Figure 4 The VVM engine 404 may include a bit counter 502, a shift accumulator 508, and a plurality of AND gates 506. As Figure 5AAs shown, the control circuit 406 can deactivate the VVM engine 404 and the partial sum buffer 410 (shown in dashed lines) in the memory mode, such that the PIM block 202 acts as a memory element for storing the first vector in the memory array 402. As Figure 5B shown, the control circuit 406 can activate the VVM engine 404 and the partial sum buffer 410 (shown in solid lines) in the compute mode, such that the first vector stored in the memory array 402 and the second vector buffered in the column buffer 408 can be sent to the VVM engine 404 to compute the dot product of the first and second vectors, and the dot product can be buffered as partial sums in the partial sum buffer 410.
[0056] In an example of performing the convolution of the feature map and the kernel in the CNN, when the PIM block 202 is in the memory mode, the PIM block 202 acts as a memory that supports read and write accesses; when the PIM block 202 is in the compute mode, the PIM block 202 can execute the VVM to compute partial sums of the output feature map during the CNN execution with reconfigurable data precision. In the memory mode, when accessing the feature map in the PIM block 202, only the column buffer 408 can be activated (the memory array 402 and the control circuit 406 are always on during any operation). The column buffer 408 can be used to buffer the input data being accessed on each row of the memory column. The memory array 402 may need to support both row and column accesses. In the compute mode, the partial sum buffer 410 and the VVM engine 404 are activated, and the computation of the executed VVM can be expressed as:
[0057]
[0058] where a and b are respectively the columns of the feature map originally stored in the first set of PIM blocks 312 (PIM blocks 202 in the memory mode) and the kernels stored in the second set of PIM blocks 314 (PIM blocks 202 in the compute mode). It should be noted that a may have been loaded into the column buffer 408 of the PIM module 202 during the preparation phase in the compute mode. N and M are the bit widths of a and b respectively. After b is read out from the column of the memory array 402 of the PIM block 202 in the compute mode, it can be sent to the VVM engine 404 to perform the dot product between a and b using the AND gate 506, or a·b = [a0b0,...,a row- 1b row-1The product a·b can be sent to the bit counter 502 to calculate the sum of the elements of a·b. If the bit widths of the feature map and the kernel are not 1 (i.e., binary precision), the above process can be repeated N×M times to calculate the dot product between the higher-order bits of a and b. Shift-and-add operations may be required to calculate the sum of dot products with different weights, and the partial sum buffer 410 can be used to store intermediate partial sums. To handle negative values, the VVM engine 404 can be replicated to process both positive and negative accumulations in parallel. The sign bit of each element in the vector can first be read from the memory array 402, and it can be determined which VVM engine 404 can perform the accumulation of positive or negative data. Then, the negative partial sums can be subtracted from the positive partial sums during partial sum accumulation.
[0059] Referring back Figure 3 , the controller 206 can be configured to perform a data mapping process to map the first matrix to the first set of PIM blocks 312 and the second matrix to the second set of PIM blocks 314. In some embodiments, the size of the first set of PIM blocks 312 is smaller than the size of the first matrix. For example, the first matrix can be a feature map in a CNN. The controller 206 is configured to map the first matrix to the first set of PIM blocks 312 such that the first set of PIM 312 stores a portion of the first matrix simultaneously based on the size of the first set of PIM blocks 312. That is, the first matrix can be "folded" to become multiple blocks, each block adapted to the size of the first set of PIM blocks 312. In some embodiments, the size of the second set of PIM blocks 314 matches the size of the second matrix. For example, the second matrix such as a kernel in a CNN can have three or four dimensions. The controller 206 can be configured to map the second matrix to the second set of PIM blocks 314 such that each PIM block 202 of the second set of PIM blocks 314 stores a corresponding portion of the second matrix. That is, the number of PIM blocks 202 and the layout of the second set of PIM blocks 314 can be determined based on the number of elements and the layout of the second matrix to form a one-to-one mapping between the PIM blocks 202 in the second set of PIM blocks 314 and the elements in the second matrix.
[0060] Continuing with the above example of performing convolution of a feature map and a kernel in a CNN, it can be inferred that the weights of the network layer (i.e., the kernel) can be stored in the second set of PIM blocks 314, and the feature map can be stored in the first set of PIM blocks 312. Since the feature map and the kernel can be high-dimensional tensors, they can first be unfolded into multiple vectors and then stored column by column in the memory array 402 of the PIM blocks 202. As described in detail below, such a columnar arrangement can conveniently support reconfigurable data precision.
[0061] Assume the configuration of the kernel is [In ch ,Out ch,K,K], where In ch is the number of input channels, Out ch is the number of output channels, and K is the size of the kernel. The configuration of the feature map is represented by [In ch ,H,W], where H and W are the height and width of the feature map, respectively. The size of the memory array 402 of each PIM block 202 is [Row,Col], where Row and Col are the number of rows and columns, respectively. The bit widths of the kernel weights and the feature map parameters are N and M, respectively. Each memory cell can only reliably store binary values, i.e., "0" and "1". For a kernel of size K×K corresponding to the same output channel In ch , the elements at the same position of the In ch kernel can be grouped into a vector, and this vector can be mapped to a column of a PIM block 202 in the second set of PIM blocks 314. The elements at different positions of the kernel can be mapped to different PIM blocks 202 in the second set of PIM blocks 314. Therefore, the kernel can be mapped to a total of K×K PIM blocks 202 in the second set of PIM blocks 314, as shown in Figure 6A . If the bit width N of the kernel weight > 1, then N memory cells in different columns in the memory array 402 may be required to represent a kernel element, assuming In ch ≤Row and Out ch ×N≤Col. If In ch >Row or Out ch ×N>Col, then the mapping may need to be extended to more PIM blocks 202 in the second set of PIM blocks 314.
[0062] Figure 6B shows the mapping of the feature map to the first set of PIM blocks 312. The In chElements at the same position in the same row are grouped into vectors, and the vectors can be mapped to columns of a PIM block 202 in the first group of PIM blocks 312. Then, the element positions can slide horizontally, and the grouped vectors can be mapped to different columns in the same PIM block 202 in the first group of PIM blocks 312. There may be a total of W such vectors that cover the elements on the same row of all feature maps and are mapped to the PIM blocks 202 in the first group of PIM blocks 312. If the bit width N of the feature map > 1, then N storage units at different columns of the memory array 402 may be required to represent a feature map parameter. After that, move to the next row of the feature map, and all the above operations can be repeated to map this new row of the feature map to another PIM block 202 in the first group of PIM blocks 312. It should be noted that only K rows of the feature map need to be stored at a time to facilitate the calculation. When K + 1 rows are needed, it can overwrite the position of the first row in the first group of PIM blocks 312. If In ch >Row or W×N>Col, the mapping may need to be extended to more PIM blocks 202 in the first group of PIM blocks 312.
[0063] For example, to calculate the CONV3×3 layer in a VGG network with a kernel [256, 32, 3, 3] and a feature map [256, 32, 32] using PIM blocks 202 with a memory array size of 256×256, the feature map and the kernel may need to be mapped to the first group of PIM blocks 312 (hereinafter referred to as "Comp.block") with 3 PIM blocks 202 and the second group of PIM blocks 314 (hereinafter referred to as "Mem.block") with 9 PIM blocks 202, as Figure 7 shown. The Comp.block block (i, j) (i, j = 0, 1, 2) represents the element (i, j) in the In ch ×ch kernel, and the Mem.block block (3, j) (j = 0, 1, 2) represents the j-th row of the feature map participating in the calculation.
[0064] To map the kernel of the FC layer to the PIIM blocks 202 in the second group of PIM blocks 314, the columns of the weight matrix are directly mapped to the columns of the PIM blocks 202, and multiple columns of the PIM blocks 202 can be combined to support the high precision of the weights. Furthermore, if the size of the weight matrix is greater than the size of the second group of PIM blocks 314, the weight matrix can be mapped to multiple second groups of PIM blocks 314.
[0065] According to some embodiments, after data mapping, the controller 206 is configured to control the second set of PIM blocks 314 to perform the convolution of the first matrix and the second matrix when calculating the partial sum of the third matrix. Continuing the example of performing the convolution of the feature map and the kernel in the CNN above, the calculation of the convolution can be divided into three stages - preparation, calculation, and translation, as Figure 8 shown.
[0066] In the preparation stage, the feature map can be first read column by column from the Mem.block and then sent to the input of each row of the Comp.block. The data transfer direction can be determined by the row index R out of the output feature map to be calculated. For example, R out = 0 at the start of the convolution. The column i (i = 0, 1, 2) of the block (3, j) (j = 0, 1, 2) can be read out and sent to the input of the block (i, j) (i, j = 0, 1, 2) to calculate the row R out = 0. After obtaining the row R out = 0, R out moves to 1. The position of the 0th row of the input feature map stored in the block (3, 0) may be overwritten by the 3rd row. The column i (i = 0, 1, 2) of the block (3, j) (j = 0, 1, 2) can be read out and sent to the input of the block (i, (j + 2) % 3) (i, j = 0, 1, 2) to calculate the row R out = 1, as Figure 8 shown. Generally, the column i (i = 0, 1, 2) of the block (3, j) (j = 0, 1, 2) can be read out and sent to the block (i, f(j)) (i, j = 0, 1, 2) to calculate the row R out . The function f(j, R out ) can be expressed as:
[0067] f(j, R out ) = (j + K - R out % K) % K (2)
[0068] where K is 3 in this example.
[0069] In the calculation stage, the Comp.block can calculate the VVM result between the feature map and the weight according to the following formula:
[0070]
[0071] where i, j are the indices of the PIM blocks 202 in the array, ic, oc are the indices of the input and output channels respectively, X i,j,ic is the ic-th input channel of the feature map stored in the block (i, j), W oc,ic,i,jis the ic-th input channel of the oc output channels of the kernel stored in block (i, j). In this example, a total of 3×3 = 9 partial sums can be generated simultaneously in blocks (0,0)–(2,2). All these 9 partial sums can be accumulated by accumulator 208 to produce the corresponding element of the output feature map. Note that it can be Out ch Repeat the above calculation Out times to generate the elements at the same positions of all Out ch feature maps. In this example, Out ch is 32.
[0072] During the calculation phase, R out +K rows of the input feature map can be stored into Mem.block and overwrite the positions of R out rows, as described above. During the shift phase, the columns C out +K–1 of block (3, j) (j = 0, 1, 2) can be read, where C out is the column index of the output feature map to be calculated. Then, the input of block (i, j) can be shifted to block (i–1, j) (i, j = 0, 1, 2), and the column C out +K–1 of block (3, j) can be sent to block (2, f(j)) as described above. Note that such a delicate shift design avoids duplicating feature map elements multiple times and thus reduces the energy consumption associated with memory writes.
[0073] Figure 9 is a flowchart of an exemplary method 900 for matrix multiplication implemented by a PIM block array according to some embodiments of the present disclosure. Figure 9 An example of the PIM block array shown in Figure 3 includes the PIM block 202 array shown in Figure 9 It can be understood that the operations shown in method 900 are not exhaustive, but other operations can also be performed before, after, or between the shown operations. Additionally, some of the operations can be performed simultaneously or in an order different from that shown in Figure 9 shown.
[0074] Referring to Figure 9 , method 900 begins at operation 902, in which a first group of PIM blocks of the PIM block array are each configured by a controller to be in a memory mode, and a second group of PIM blocks of the PIM block array are each configured by the controller to be in a calculation mode. In some embodiments, the first group of PIM blocks includes the rows of the PIM block array, and the second group of PIM blocks includes the remaining part of the PIM block array. As Figure 3 shown, controller 206 can configure each PIM block 202 in the first group of PIM blocks 312 to be in a memory mode and each PIM block 202 in the second group of PIM blocks 314 to be in a calculation mode.
[0075] Method 900 proceeds to operation 904, which, as Figure 9 shown, in operation 904, the first matrix is mapped by the controller to a first set of PIM blocks, and the second matrix is mapped by the controller to a second set of PIM blocks. In some embodiments, the first matrix includes a feature map in a CNN, and the second matrix includes a kernel of the CNN. In some embodiments, the size of the first set of PIM blocks is smaller than the size of the first matrix, and a portion of the first matrix is stored simultaneously by the first set of PIM blocks based on the size of the first set of PIM blocks. In some embodiments, the size of the second set of PIM blocks matches the size of the second matrix. For example, the second matrix may have three or more dimensions, and corresponding portions of the second matrix may be stored by each PIM block of the second set of PIM blocks. As Figure 3 shown, the controller 206 may map the first matrix, such as the feature map, to the first set of PIM blocks 312, and map the second matrix, such as the kernel, to the second set of PIM blocks 314.
[0076] Method 900 proceeds to operation 906, which, as Figure 9 shown, in operation 906, a partial sum of a third matrix is calculated by the second set of PIM blocks based on the first and second matrices. In some embodiments, data from each PIM block of the first set of PIM blocks is directed by a MUX between the first set of PIM blocks and the second set of PIM blocks to a PIM block in a corresponding column of the second set of PIM blocks. In some embodiments, at least one PIM block of the first set of PIM blocks is not in the same corresponding column. In some embodiments, calculating the partial sum of the third matrix includes performing a convolution of the first and second matrices. As Figure 3 shown, the MUX 204 between the first and second sets of PIM blocks 312 and 314 may direct data between them, and the second set of PIM blocks 314 may calculate the partial sum of the third matrix based on the first and second matrices.
[0077] Method 900 proceeds to operation 908, which, as Figure 9 shown, the third matrix is generated based on the partial sum of the third matrix. As Figure 3 shown, the third matrix may be generated by the accumulator 208 based on the partial sum of the third matrix calculated by the second set of PIM blocks 312.
[0078] The foregoing detailed description of various specific embodiments is intended to fully disclose the general nature of the present invention, so that others can easily modify / adjust these specific embodiments to suit various applications without undue experimentation and without departing from the basic concepts of the present invention by applying common general knowledge within the field of application. Therefore, the above adjustments and modifications are based on the teachings and guidance of the present invention and are intended to keep these modifications and adjustments within the meaning and scope of the equivalents of the embodiments described in the present invention. It is understood that the vocabulary or terms used herein are for descriptive purposes only, so that those with professional knowledge can understand these vocabulary and terms under the inspiration and guidance of the present invention, and should not be used to limit the content of the present invention.
[0079] The present invention describes the implementation cases in the present invention by explaining specific functions and specific relationships with the aid of functional modules. For the convenience of narration, the definition of the above functional modules is arbitrary. As long as the required specific functions and specific relationships can be achieved, other alternative definitions can also be adopted.
[0080] The description of the invention and the abstract may set forth one or more embodiments of the present invention, but do not include all exemplary embodiments conceived by the inventor. Therefore, it is not intended to limit the scope of the present invention and the claims in any way.
[0081] The scope of the present invention is not limited to any of the above embodiments, but should be defined in accordance with the claims and their equivalents.
Claims
1. An apparatus for matrix multiplication, comprising: An in-memory processing (PIM) block array in row and column form, each in-memory processing PIM block being configured to be in a compute mode or a memory mode, wherein each PIM block in the PIM block array includes a memory array, a buffer, and a vector-vector multiplication (VVM) engine, wherein the memory array of the PIM block is configured to store a first vector and the buffer of the PIM block is configured to load a second vector from another PIM block; A controller configured to partition the PIM block array into: a first group of PIM blocks, each PIM block being configured to be in a memory mode such that the VVM engine in each PIM block in the first group of PIM blocks is deactivated; And a second group of PIM blocks, each PIM block being configured to be in a compute mode such that the VVM engine in each PIM block in the second group of PIM blocks is enabled, wherein the first group of PIM blocks is configured to store a first matrix, and the second group of PIM blocks is configured to store a second matrix and calculate partial sums of a third matrix based on the first and second matrices; And An accumulator configured to output the third matrix based on the partial sums of the third matrix.
2. The apparatus according to claim 1, wherein the first group of PIM blocks includes rows of the PIM block array, and the second group of PIM blocks includes the remaining part of the PIM block array.
3. The apparatus according to claim 2, wherein the size of the first group of PIM blocks is smaller than the size of the first matrix, and the controller is configured to map the first matrix to the first group of PIM blocks such that the first group of PIM blocks stores a part of the first matrix simultaneously based on the size of the first group of PIM blocks.
4. The apparatus according to claim 3, further comprising a multiplexer (MUX) between the first and second groups of PIM blocks, wherein the controller is further configured to control the MUX to direct data from each PIM block of the first group of PIMs to the PIM blocks in the corresponding columns of the second group of PIM blocks.
5. The apparatus according to claim 4, wherein at least one PIM block of the first group of PIM blocks is not in the same corresponding column.
6. The apparatus according to claim 1, wherein the size of the second group of PIM blocks matches the size of the second matrix.
7. The apparatus according to claim 6, wherein the second matrix has three or more dimensions, and the controller is configured to map the second matrix to the second group of PIM blocks such that each PIM block of the second group of PIM blocks stores a corresponding part of the second matrix.
8. The apparatus according to claim 1, wherein the controller is configured to control the second group of PIM blocks to perform a convolution of the first and second matrices when calculating the partial sums of the third matrix.
9. The apparatus according to claim 8, wherein the first matrix includes a feature map in a convolutional neural network (CNN), and the second matrix includes a kernel of the CNN.
10. An in-memory processing (PIM) device, comprising: A memory array configured to store a first vector; A first buffer configured to receive and buffer a second vector from another PIM device; A vector-vector multiplication (VVM) engine; and A controller, which is configured to: Enable the VVM engine in a computing mode and control the VVM engine to perform a dot product between a first vector and a second vector to generate a partial sum; and Disable the VVM engine in a memory mode and control the memory array to write or read the first vector.
11. The PIM device according to claim 10, wherein the VVM engine includes a bit counter, a shift accumulator, and a plurality of AND gates.
12. The PIM device according to claim 10, further comprising a second buffer, which is configured to buffer the partial sum and send the partial sum to another PIM device.
13. The PIM device according to claim 10, wherein the memory array includes a resistive random access memory (ReRAM) array.
14. A method for matrix multiplication performed by the apparatus for matrix multiplication according to claim 1, the method comprising: Configuring, by the controller, each of a first group of PIM blocks of a PIM block array to be in a memory mode, such that the vector-vector multiplication (VVM) engine in each PIM block of the first group of PIM blocks is disabled, and configuring each of a second group of PIM blocks of the PIM block array to be in a computing mode, such that the VVM engine in each PIM block of the second group of PIM blocks is enabled; Mapping, by the controller, a first matrix to the first group of PIM blocks and mapping a second matrix to the second group of PIM blocks; Calculating, by the second group of PIM blocks, partial sums of a third matrix based on the first and second matrices; And Generating the third matrix based on the partial sums of the third matrix.
15. The method according to claim 14, wherein the first group of PIM blocks includes rows of the PIM block array, and the second group of PIM blocks includes the remaining part of the PIM block array.
16. The method according to claim 15, wherein the size of the first set of PIM blocks is smaller than the size of the first matrix, and the method further comprises: Simultaneously storing, by the first group of PIM blocks, a part of the first matrix based on the size of the first group of PIM blocks.
17. The method according to claim 16, further comprising: Directing, by a multiplexer (MUX) between the first and second groups of PIM blocks, data from each PIM block of the first group of PIMs to a PIM block in a corresponding column of the second group of PIM blocks.
18. The method according to claim 17, wherein at least one PIM block of the first group of PIM blocks is not in the same corresponding column.
19. The method according to claim 14, wherein the size of the second group of PIM blocks matches the size of the second matrix.
20. The method according to claim 19, wherein the second matrix has three or more dimensions, and the method further comprises: Storing, by each PIM block of the second group of PIM blocks, a corresponding part of the second matrix.
21. The method according to claim 14, wherein calculating the partial sums of the third matrix includes performing a convolution of the first and second matrices.
22. The method according to claim 14, wherein the first matrix includes a feature map in a convolutional neural network (CNN), and the second matrix includes a kernel of the CNN.
Citation Information
Patent Citations
Matrix multiplier
CN109992743A
Apparatuses and methods for in-memory operations
CN110476210A