In-Memory Computing Method and Device
By preprocessing and multiplication operations in memory, filtering and post-processing the calculation results, the problems of reduced performance and power consumption in data-intensive applications are solved, and an efficient computing system is realized.
Patent Information
- Application Number
- CN202111010149.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-25
- Filing Date
- 2021-08-31
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-08-31
AI Technical Summary
When performing data-intensive applications, traditional computing systems need to perform a large amount of computing and frequent data movement, resulting in reduced system performance and high power consumption.
A memory calculation method and device are provided, by preprocessing the input data and weight data, distinguishing it into main parts and secondary parts, and multiplication and addition operations are performed in the memory, filtering out the calculation results with smaller values, and post-processing to obtain output data.
It improves the efficiency of the computing system, reduces data movement and power consumption, and supports complex computing.
Smart Images

Figure CN114153420B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an operation method and apparatus, and more particularly to a method and apparatus for in-memory computing. Background Art
[0002] In traditional computing systems, when executing data-intensive applications, a large amount of computation is required, and data needs to be frequently moved between the processor and the memory. Among them, performing a large amount of computation will reduce the system performance, and a large amount of data movement will cause high power consumption.
[0003] To solve the above problems of performance limitation and power consumption, new algorithms and / or memory architectures have been proposed in recent years, including Nearest Neighbor Search, Decision tree learning, distributed systems, In-memory computing, etc. However, Decision tree learning still requires a large amount of data movement, distributed systems have problems of high cost and communication between devices, and In-memory computing cannot support complex operations.
[0004] Disclosure
[0005] In view of the above, the present disclosure provides an in-memory computing method and an in-memory computing apparatus that can improve the performance of a computing system.
[0006] The present disclosure provides an in-memory computing method suitable for a processor to perform multiply-accumulate (MAC) operations on a memory. The memory includes a plurality of input lines and a plurality of output lines that cross each other, a plurality of memory cells disposed at the intersection points of the input lines and the output lines, and a plurality of sense amplifiers respectively connected to the output lines. The method includes the following steps: preprocessing the input data and weight data to be written into the input lines and the memory cells respectively to distinguish them into a main part and a secondary part; writing the input data and weight data distinguished into the main part and the secondary part into the input lines and the memory cells in batches to perform multiply-accumulate operations to obtain a plurality of operation results; filtering the operation results according to the numerical magnitudes of the respective operation results; and post-processing the filtered operation results according to the parts corresponding to the operation results to obtain output data.
[0007] In an embodiment of the present disclosure, the step of filtering the operation results according to the numerical magnitudes of the respective operation results includes filtering out the operation results whose numerical magnitudes are not greater than a preset threshold value, sorting the filtered operation results, and selecting at least one of the operation results ranked ahead for post-processing.
[0008] In one embodiment of the present disclosure, the method further includes encoding the input data and the weight data when preprocessing the input data and the weight data, and performing a weighted operation corresponding to the encoding on the operation result when postprocessing the filtered operation result.
[0009] In one embodiment of the present disclosure, the step of performing a weighted operation corresponding to the encoding on the operation result includes multiplying the operation result by a first weight to obtain a first product in response to the operation result corresponding to the main part of the input data and the main part of the weight data, multiplying the operation result by a second weight to obtain a second product in response to the operation result corresponding to the main part of the input data and the minor part of the weight data, multiplying the operation result by a third weight to obtain a third product in response to the operation result corresponding to the minor part of the input data and the main part of the weight data, multiplying the operation result by a fourth weight to obtain a fourth product in response to the operation result corresponding to the minor part of the input data and the minor part of the weight data, and accumulating the first product, the second product, the third product, and the fourth product obtained by performing the weighted operation on the operation result, and outputting the accumulated result as output data.
[0010] The present disclosure provides an in-memory computing device, which includes a memory and a processor. The memory includes a plurality of input lines and a plurality of output lines that cross each other, a plurality of cells respectively disposed at the intersection points of the input lines and the output lines, and a plurality of sense amplifiers respectively connected to the output lines. The processor is coupled to the memory and is configured to preprocess the input data and the weight data to be written to the input lines and the memory cells respectively to distinguish them into a main part and a minor part, batch-write the input data and the weight data that are distinguished into a main part and a minor part to the input lines and the memory cells for multiplication and addition operations, accumulate the sensed values of the sense amplifiers to obtain multiple operation results, filter the operation results according to the numerical magnitudes of the operation results, and perform postprocessing on the filtered operation results according to the corresponding parts of the operation results to obtain output data.
[0011] In one embodiment of the present disclosure, the main part is the most significant bit (MSB) of multiple bits of the data to be processed, and the minor part is the least significant bit (LSB) of multiple bits of the data to be processed.
[0012] In one embodiment of the present disclosure, the in-memory computing device further includes a filter for filtering operation results whose numerical magnitudes are not greater than a preset threshold value, wherein the processor includes sorting the filtered operation results and selecting at least one operation result ranked ahead for postprocessing.
[0013] In one embodiment of the present disclosure, when preprocessing the input data and the weight data, the processor encodes the input data and the weight data, and when postprocessing the filtered operation result, the processor performs a weighted operation corresponding to the encoding on the operation result.
[0014] In one embodiment of the present disclosure, the processor includes multiplying the operation result by a first weight to obtain a first product in response to the operation result corresponding to the main part of the input data and the main part of the weight data, multiplying the operation result by a second weight to obtain a second product in response to the operation result corresponding to the main part of the input data and the secondary part of the weight data, multiplying the operation result by a third weight to obtain a third product in response to the operation result corresponding to the secondary part of the input data and the main part of the weight data, multiplying the operation result by a fourth weight to obtain a fourth product in response to the operation result corresponding to the secondary part of the input data and the secondary part of the weight data, and accumulating the first product, the second product, the third product, and the fourth product obtained by performing the weighted operation on the operation result, and outputting the accumulated result as output data.
[0015] To make the foregoing features and advantages of the present disclosure more understandable, embodiments of the drawings are described in detail below. Description of the Drawings
[0016] Figure 1 Schematic diagram of an in-memory computing device according to an embodiment of the present disclosure.
[0017] Figure 2 Flowchart of an in-memory computing method according to an embodiment of the present disclosure.
[0018] Figure 3 Schematic diagram of data encoding according to an embodiment of the present disclosure.
[0019] Figure 4 Schematic diagram of data postprocessing according to an embodiment of the present disclosure.
[0020] Description of Reference Numerals
[0021] 10: Computing device
[0022] 12: Memory
[0023] 14: Processor
[0024] B7~B0, B70~B77, B60~B63: Bits
[0025] IL i : Input line
[0026] OL j : Output line
[0027] Rij : Resistance
[0028] S202~S208: Steps
[0029] SA: Sense amplifier Detailed implementation manners
[0030] Figure 1 is a schematic diagram of an in-memory computing device according to an embodiment of the present disclosure. Please refer to Figure 1 , the in-memory computing device 10 of this embodiment is, for example, a memristor, and the memristor is configured to implement process-in-memory (PIM), and is applicable to data-intensive applications such as face search. The computing device 10 includes a memory 12 and a processor 14, and their functions are described as follows:
[0031] The memory 12 is, for example, a NAND flash memory, a NOR flash memory, a phase change memory (PCM), a spin-transfer torque random-access memory (STT-RAM), or a resistive random-access memory (ReRAM) with a 2D or 3D structure, which is not limited herein. In some embodiments, various volatile memories, such as static random access memory (SRAM), dynamic random access memory (DRAM), and various non-volatile memories, such as ReRAM, PCM, flash, magnetoresistive RAM, ferroelectric RAM, can be integrated for in-memory computing, which is not limited herein.
[0032] The memory 12 includes a plurality of input lines IL intersecting with each other i and a plurality of output lines OL j , and a plurality of memory cells (represented by the resistor R i respectively disposed at the intersection points of the input line IL j and the output line OL ij ), and a plurality of sense amplifiers SA respectively connected to the output line OL j for sensing the current I j output from the output line OL j . In some embodiments, the input line IL i is a word line and the output line OL j is a bit line, and in some embodiments, the input line IL i is a bit line and the output line OL jis a word line, which is not limited herein.
[0033] The processor 14 is, for example, a central processing unit (CPU) or other programmable general or special microprocessor, microcontroller (MCU), programmable controller, application specific integrated circuit (ASIC), programmable logic device (PLD), or other similar devices, or a combination of such devices, and the present embodiment does not limit it. In the present embodiment, the processor 14 is configured to execute instructions for performing in-memory operations. The in-memory operations can be implemented in various artificial intelligent (AI) applications, such as fully connected layers, convolution layers, multi-layer perceptrons, support vector machines, or other applications implemented using memristors, which are not limited herein.
[0034] Figure 2 is a flowchart of an in-memory operation method according to an embodiment of the present disclosure. Please refer to Figure 1 and Figure 2 , the method of the present embodiment is suitable for the above in-memory operation device 10, and the detailed steps of the in-memory operation method of the present embodiment will be described below with reference to various devices and components of the in-memory operation device 10.
[0035] First, in step S202, the processor 14 preprocesses the input data and weight data to be written to the input line and the storage unit, respectively, to distinguish them into a main part and a secondary part. In one embodiment, the processor 14 divides the input data into the most significant bit (MSB) of multiple bits and the least significant bit (LSB) of multiple bits, and divides the weight data into the MSB of multiple bits and the LSB of multiple bits. When the input data is 8 bits, the processor 14, for example, divides the input data into the MSB of 4 bits and the LSB of 4 bits, and divides the weight data into the MSB of 4 bits and the LSB of 4 bits. In other cases, the processor 14 can divide the input data and the weight data into one or more MSBs and one or more LSBs of the same number or different numbers according to the implementation requirements, and the present embodiment does not limit this here. In other embodiments, the processor 14 can also mask or filter one or more unimportant bits (i.e., the secondary part) in the input data and only retain the more important bits (i.e., the main part) for subsequent operations, and the present embodiment also does not limit this.
[0036] In other embodiments, the processor 14 may further encode the input data and the weight data. For example, the processor 14 may convert the most significant bits (MSB) and the least significant bits (LSB) of multiple bits of the input data or the weight data from a binary format to a unary code value format. The processor 14 may then copy the converted unary code to unfold it into a dot product format.
[0037] For example, Figure 3 is a schematic diagram of data encoding according to an embodiment of the present disclosure. Refer to Figure 3 , in this embodiment, it is assumed that there are N-dimensional input data and weight data to be written, where N is a positive integer, and each piece of data has 8 bits B0 to B7 represented in binary. For the N-dimensional input data <1> to <n>For example, in this embodiment, each input data <1> ~ <n>It is divided into an MSB vector and an LSB vector, where the MSB vector includes 4-bit MSB B7 to B4, and the LSB vector includes 4-bit LSB B3 to B0. Then, each bit of the MSB vector and the LSB vector is converted into a one-hot encoding according to the value. For example, bit B7 is converted into bits B70 to B77, bit B6 is converted into B60 to B63, bit B5 is converted into B50 to B51, and bit B4 remains unchanged. Then, the converted one-hot encoding is copied to expand into a dot product format. For example, the (2 4 -1) one-hot encodings after conversion of the MSB vector of each input data are copied (2 4 -1) times to expand into 225 bits, and data in the unfolding dot product (unFDP) format shown in Figure 3 is generated. Similarly, the weight data can also be preprocessed in the same encoding manner as the above input data, which will not be elaborated here.
[0038] Returning to Figure 2 the process, in step S204, the processor 14 writes the input data and the weight data, which are divided into a main part and a secondary part, into the input line and the storage unit in batches for multiplication and addition operations to obtain multiple operation results. Specifically, for example, the processor 14 writes the weight data divided into the main part into the corresponding storage unit in the memory 12, and inputs the input data divided into the main part into the corresponding input line IL i in the memory 12, so that the sense amplifier SA connecting each output line OL j senses the current I j output from the output line OL j , and then accumulates the sensed value of the sense amplifier SA via a counter or an accumulator to obtain the operation result of the multiplication and addition operation of the input data and the weight data. Similarly, for example, the processor 14 writes the weight data divided into the main part into the corresponding storage unit in the memory 12, and inputs the input data divided into the secondary part into the corresponding input line IL i in the memory 12 to obtain the operation result of the multiplication and addition operation; writes the weight data divided into the secondary part into the corresponding storage unit in the memory 12, and inputs the input data divided into the main part into the corresponding input line IL i in the memory 12 to obtain the operation result of the multiplication and addition operation; and writes the weight data divided into the secondary part into the corresponding storage unit in the memory 12, and inputs the input data divided into the secondary part into the corresponding input line IL i in the memory 12 to obtain the operation result of the multiplication and addition operation.
[0039] In some embodiments, the memory 12 may also support operations such as Inverse, AND, OR, Exclusive OR (XOR), Exclusive NOR (XNOR), etc., not limited to multiply-add operations. In addition, the memory 12 is not limited to being implemented using digital circuits, but may be implemented using analog circuits, and the implementation manner thereof is not limited in this embodiment.
[0040] For example, in a digital circuit, the processor 14 may divide the input data into a multi-bit MSB and a multi-bit LSB (the number of bits is not limited), and after being processed by different encoding (i.e., preprocessing) methods, it is sent to the memory 12 to perform Inverse, AND, OR, Exclusive OR, Exclusive NOR, multiply-add operations or a combination of the above operations. Finally, after being filtered by corresponding post-processing, the final operation result can be obtained. In an analog circuit, the processor 14 may mask or filter out some bits of the input data (i.e., preprocessing), and then send it to the memory 12 to perform Inverse, AND, OR, Exclusive OR, Exclusive NOR, multiply-add operations or a combination of the above operations. Finally, after being filtered by corresponding post-processing, the final operation result can be obtained. The above is only an example, and the processor 14 may perform any type of preprocessing and post-processing on the input data to obtain a dedicated operation result.
[0041] In step S206, the processor 14 filters the operation results according to the numerical magnitudes of the operation results. In one embodiment, the in-memory arithmetic device 10 includes, for example, a filter (not shown), and is used to filter out operation results whose numerical magnitudes are not greater than a preset threshold value. The processor 14 will then sort the filtered operation results and select the top N operation results for post-processing, where N is, for example, 3, 5, 10, 20 or any positive integer, and is not limited herein.
[0042] In step S208, the processor 14 performs post-processing on the filtered operation results according to the corresponding parts of the operation results to obtain output data. In one embodiment, when preprocessing the input data and the weight data, the processor 14, for example, encodes the input data and the weight data, and when performing post-processing on the filtered operation results, it performs a weighted operation corresponding to the encoding on the operation results.
[0043] Specifically, in response to the operation result corresponding to the main part of the input data and the main part of the weight data, the processor 14 multiplies the operation result by a first weight to obtain a first product; in response to the operation result corresponding to the main part of the input data and the secondary part of the weight data, the processor 14 multiplies the operation result by a second weight to obtain a second product; in response to the operation result corresponding to the secondary part of the input data and the main part of the weight data, the processor 14 multiplies the operation result by a third weight to obtain a third product; in response to the operation result corresponding to the secondary part of the input data and the secondary part of the weight data, the processor 14 multiplies the operation result by a fourth weight to obtain a fourth product. Finally, the processor 14 accumulates the first product, the second product, the third product, and the fourth product obtained by performing the weighted operation on the above operation results, and outputs the accumulated result as the output data.
[0044] For example, Figure 4 is a schematic diagram of data post-processing according to an embodiment of the present disclosure. Please refer to Figure 4 This embodiment illustrates the post-processing corresponding to the Figure 3 encoding method. Among them, in response to the operation result corresponding to the main part of the input data (i.e., MSB) and the main part of the weight data, the corresponding weight value is 16×16; in response to the operation result corresponding to the main part of the input data and the secondary part of the weight data (i.e., LSB), the corresponding weight value is 16×1; in response to the operation result corresponding to the secondary part of the input data and the main part of the weight data, the corresponding weight value is 1×16; in response to the operation result corresponding to the secondary part of the input data and the secondary part of the weight data, the corresponding weight value is 1×1. By multiplying the operation result obtained by batch-writing the input data and the weight data into the memory 12 by the corresponding weight value, the operation result of the multiplication and addition operation of the original input data and the weight data can be restored.
[0045] After completing the multiplication and addition operation of each piece of input data and weight data and obtaining the operation result, the processor 14 will return to step S204, continue to write the next piece of input data and weight data into the memory 12 for multiplication and addition operation until all the operation results of the input data and weight data are completed, and the operation in the memory is completed.
[0046] In summary, the in-memory computing method and apparatus according to the embodiments of the present disclosure combine in-memory computing and a hierarchical filtering scheme. By preprocessing the input data and weight data to be written into the memory, operations on bits with a lower proportion in the data value (i.e., LSB) are selectively deleted, and operations are preferentially performed on bits with a higher proportion (i.e., MSB). And by filtering the operation results, operation results with higher values are selected for corresponding post-processing of the data, and finally output data is obtained. Therefore, the efficiency of the computing system can be improved without overly affecting the numerical value of the operation results.
[0047] Although the present disclosure has been disclosed through the above embodiments, the embodiments are not intended to limit the present disclosure. It will be obvious to those skilled in the art that various modifications and changes can be made to the structure of the present disclosure without departing from the scope or spirit of the present disclosure. Therefore, the protection scope of the present disclosure falls within the scope of the appended claims.< / n> < / n>
Claims
1. A method for in-memory computing, suitable for a processor to perform multiply-accumulate operations using a memory, wherein the memory includes a plurality of input lines and a plurality of output lines that cross each other, a plurality of memory cells respectively disposed at the intersection points of the input lines and the output lines, and a plurality of sense amplifiers respectively connected to the output lines, and the method includes: The input data and the weight data are respectively divided into the most significant bits of multiple bits and the least significant bits of multiple bits. Each bit of the most significant bits of multiple bits and the least significant bits of multiple bits of the input data and the weight data is converted into a one-hot encoding according to the value, and then the converted one-hot encoding is copied to generate the input data and the weight data in an expanded dot product format; The input data and the weight data in the expanded dot product format are written into the input line and the storage unit in batches. The sense amplifier senses the current output from the output line to generate a sensed value, and the sensed values of the sense amplifier are accumulated to obtain multiple operation results; Filter out the operation results whose numerical magnitudes are not greater than a preset threshold value; and Perform post-processing on the filtered operation results according to the corresponding parts of the operation results to obtain output data.
2. The method for in-memory computing according to claim 1, wherein the step of filtering out operation results whose numerical magnitudes are not greater than a preset threshold further includes: Sort the filtered operation results and select at least one operation result ranked at the front for the post-processing.
3. The method for in-memory computing according to claim 1, further includes: When generating the input data and the weight data in the expanded dot product format, encode the input data and the weight data; and When performing the post-processing on the filtered operation results, perform a weighted operation corresponding to the encoding on the operation results.
4. The method for in-memory computing according to claim 3, wherein the step of performing the weighted operation corresponding to the encoding on the operation result includes: In response to the operation result corresponding to the most significant bits of multiple bits of the input data and the most significant bits of multiple bits of the weight data, multiply the operation result by a first weight to obtain a first product; In response to the operation result corresponding to the most significant bits of multiple bits of the input data and the least significant bits of multiple bits of the weight data, multiply the operation result by a second weight to obtain a second product; In response to the operation result corresponding to the least significant bits of multiple bits of the input data and the most significant bits of multiple bits of the weight data, multiply the operation result by a third weight to obtain a third product; In response to the operation result corresponding to the least significant bits of multiple bits of the input data and the least significant bits of multiple bits of the weight data, multiply the operation result by a fourth weight to obtain a fourth product; and Accumulate the first product, the second product, the third product, and the fourth product obtained by performing the weighted operation on the operation result, and output the accumulated result as the output data.
5. An in-memory computing device, suitable for a processor to perform multiply-accumulate operations using a memory, includes: A memory, comprising: Multiple input lines and multiple output lines that cross each other; Multiple storage units, respectively disposed at the intersection points of the input lines and the output lines; and Multiple sense amplifiers, respectively connected to the output lines and configured to sense the current output from the output lines to generate sensed values; A processor, coupled to the memory and configured to: The input data and the weight data are respectively divided into the most significant bits of multiple bits and the least significant bits of multiple bits. Each bit of the most significant bits of multiple bits and the least significant bits of multiple bits of the input data and the weight data is converted into a one-hot encoding according to the value, and then the converted one-hot encoding is copied to generate the input data and the weight data in an expanded dot product format; Unfold the input data and the weight data in the form of dot product, write them in batches to the input line and the storage unit for the multiplication and addition operation, and accumulate the sensed values of the sense amplifier to obtain multiple operation results; Filter out the operation results whose numerical magnitudes are not greater than a preset threshold value; and Perform post-processing on the filtered operation results according to the corresponding parts of the operation results to obtain output data.
6. The in-memory computing device according to claim 5, wherein the processor includes sorting the operation results after filtering, and selecting at least one operation result ranked at the front for the post-processing.
7. The in-memory computing device according to claim 5, wherein the processor includes encoding the input data and the weight data when generating the input data and the weight data in the expanded dot product format, and performing a weighted operation corresponding to the encoding on the filtered operation result when performing the post-processing on the operation result.
8. The in-memory computing device according to claim 7, wherein the processor includes: In response to the operation result corresponding to the most significant bits of multiple bits of the input data and the most significant bits of multiple bits of the weight data, multiply the operation result by a first weight to obtain a first product; In response to the operation result corresponding to the most significant bits of multiple bits of the input data and the least significant bits of multiple bits of the weight data, multiply the operation result by a second weight to obtain a second product; In response to the operation result corresponding to the least significant bits of multiple bits of the input data and the most significant bits of multiple bits of the weight data, multiply the operation result by a third weight to obtain a third product; In response to the operation result corresponding to the least significant bits of multiple bits of the input data and the least significant bits of multiple bits of the weight data, multiply the operation result by a fourth weight to obtain a fourth product; and Accumulate the first product, the second product, the third product and the fourth product obtained by performing the weighted operation on the operation result, and output the accumulated result as the output data.
Citation Information
Patent Citations
High-cardinal-number approximate Booth encoding method and mixed-cardinal-number Booth encoding approximate multiplier
CN111488133A
Convolution memory
US5014235A