A memory-computing integrated chip adder
By designing a high-performance matrix adder on the memory and computing integrated chip, using the MCU module for data preprocessing and encoding, and the memory and computing integrated chip module for addition operations, the problem of speed bottleneck in the memory and computing integrated chip during addition operations is solved, and efficient addition operations are realized.
Patent Information
- Application Number
- CN202111332543.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-11
AI Technical Summary
When the memory and computing integrated chip implements addition operations, it is difficult to increase the computing speed while ensuring accuracy, which becomes the bottleneck of its speed.
A memory-computer integrated chip adder is designed, which is preprocessed and encoded through the MCU module, and the memory-computer integrated chip module is used for addition operations to realize a high-performance matrix adder. The adder decomposes the to-process data into single integer addition, single floating-point number addition, vector integer addition and vector floating-point number addition, respectively, and calls related adders to improve operational efficiency.
On the premise of ensuring the calculation accuracy, the calculation speed of matrix addition is significantly improved and the addition performance of the memory and computing integrated chip is improved.
Smart Images

Figure HDA0003349442310000011 
Figure HDA0003349442310000021 
Figure HDA0003349442310000031
Abstract
Description
Technical Field
[0001] The present invention relates to the field of high performance computing technology, and in particular to a storage and computing integrated chip adder. Background Art
[0002] Adders are basic units in computing devices. They have specific implementations in CPUs, GPUs, MCUs, FPGAs, etc. The main implementation is to split the addition of different types of data into binary logic operations. Adders are commonly used in computer arithmetic logic units, performing logical operations, shifting and instruction calls, etc. In electronics, adders are digital circuits that can perform digital addition operations.
[0003] The rapid development of integrated storage and computing chip technology in recent years has created opportunities for high-speed real-time operation of matrix operations. With the rise of application fields such as the Internet of Things and artificial intelligence, the technology has been widely studied and applied by academic and industrial circles at home and abroad. In 2016, Professor Xie Yuan's team at the University of California, Santa Barbara (UCSB) proposed using RRAM to build a deep learning neural network based on an integrated storage and computing architecture (PRIME). This neural network based on an integrated storage and computing architecture has attracted widespread attention in the industry. Experimental results show that compared with the traditional architecture based on the von Neumann computing architecture, PRIME can achieve a power consumption reduction of about 1 / 20 and a speed increase of about 50 times. This integrated storage and computing solution can efficiently implement vector-matrix multiplication operations and has broad application prospects in the field of deep learning neural network accelerators. In addition, many internationally renowned universities and companies such as Duke University, Purdue University, Stanford University, University of Massachusetts, Nanyang Technological University, Singapore, HP, Intel, and Micron have carried out related research and released test chip prototypes.
[0004] At present, the integrated storage and computing chip has a large number of product series. The main manufacturers include Mythic and Syntiant in the United States, and China's Flash Easy Semiconductor, Zhicun Technology, Xinyi Technology, Hengshuo Semiconductor, etc., and have received industrial investment from mainstream domestic and foreign semiconductor companies and capital including Intel, ARM, Bosch, Amazon, Microsoft, Softbank, SMIC, etc. Although the integrated storage and computing chips designed by these manufacturers are of various types and different, their principles and internal basic structures are similar. They all adopt a homogeneous multi-core architecture. Each storage computing core (MPU) contains: computing engine (PE, processing engine), cache (cache), control (CTRL) and input and output (I / O). The cache can be SRAM, MRAM or similar high-speed random access memory. Each MPU is connected through an on-chip network (NoC, network on chip). Each MPU accesses its own cache, which can achieve high-performance parallel computing. From the perspective of the combination of storage and computing, storage and computing integrated chips can be divided into two categories: one is to implant logical operation units in DRAM, which can be called in-memory processing. This method is suitable for big data and neural network training in the cloud; the other is to fully combine storage and computing. The storage device is also the computing unit. AI chips based on NOR flash memory architecture have such characteristics. The main characteristics are low energy consumption, high computing efficiency, fast speed and low cost, which are suitable for applications in neural network reasoning and other aspects.
[0005] With the booming development of artificial intelligence today, storage and computing chips have unique conditions for processing algorithms such as neural networks. Their fast multipliers based on analog signals can quickly implement various artificial intelligence algorithms. However, they only have simple functions when implementing fast addition operations and cannot achieve large-scale calculations. Large-scale addition operations have become the bottleneck of the speed of storage and computing chips. The main problem is that it is impossible to ensure accuracy while increasing the calculation speed. Summary of the invention
[0006] In view of the shortcomings of the background technology, the present invention relates to a storage-computation integrated chip adder. According to the above problems, a high-performance matrix adder used on a storage-computation integrated chip is designed, which can greatly improve the calculation speed of matrix addition while ensuring the calculation accuracy. The storage-computation integrated chip adder has two parts: an MCU module and a storage-computation integrated chip module, and the two parts complement each other. The MCU module preprocesses data, and the storage-computation integrated chip performs adder operations, and the two perform their respective duties with a certain degree of independence.
[0007] The present invention relates to a storage-computing integrated chip adder, comprising an MCU module and a storage-computing integrated chip module, which are implemented on the MCU and the storage-computing integrated chip; the adder function of the storage-computing integrated chip is the same as the adder function on the MCU, and the operation steps are as follows: S1: writing data to be processed into the MCU module; S2: the MCU module establishes communication with the storage-computing integrated chip module; S3: the MCU module executes an encoding instruction; S4: the MCU module executes a transmission instruction, and transmits the encoded data to the SRAM of the storage-computing integrated chip; S5: waiting for the transmission to be completed, the MCU executes the instruction, and the control chip performs the operation; S6: the storage-computing integrated chip module receives the operation end information, and the MCU module reads the SRAM; S7: the MCU module executes a decoding instruction; S8: the MCU outputs the addition result.
[0008] By adopting the above scheme, the performance of the adder is tested with a 30-dimensional floating point array, and the data is randomly generated. The calculation accuracy of the adder using the integrated storage and computing chip can reach 10 -4 , compensating for the limited multiplication precision of the integrated storage and computing chip.
[0009] Furthermore, the MCU module is used to receive data to be processed, encode and decode data, and communicate with the storage and computing integrated chip. The storage and computing integrated chip module decomposes the adder process into several types, including single integer addition, single floating-point addition, vector integer addition and vector floating-point addition, so that the relevant adder can be called each time to process data.
[0010] Furthermore, the storage-computing integrated chip module includes FLASH, SRAM, and DATABUS located inside the storage-computing integrated chip, which are used to implement the adder body, transmit data with the MCU module, and write the results into the SRAM.
[0011] Furthermore, the storage-computing integrated chip module uses the data interaction interface DATABUS to interact with the MCU, stores temporary data on the SRAM, and uses FLASH to implement the main structure of the adder inside the chip to complete basic operations such as addition, judgment, and looping.
[0012] Furthermore, the MCU module is implemented by the memory, processor, basic input and output, and counter inside the MCU. The main implementation method is that the memory stores the input data, the input data is encoded on the MCU using a designed encoder, the basic input and output is responsible for data interaction with the storage and computing integrated chip, the data sent back by the chip is decoded on the MCU using a designed decoder, and the counter completes the overall timing logic.
[0013] Furthermore, the main contents of the MCU module include a pre-calculation part and an interactive communication part. The MCU module is mainly implemented by an STM single-chip microcomputer, and the control program is written and used using STM32. The specific implementation method is: first, the data to be processed is imported into the MCU through the serial port, and the data is preprocessed using a preprocessing program. Then, the processed data is written into the SRAM of the chip through the serial port, and then the addition program in the chip is called out using the serial port. Given the relevant SRAM address and operation data dimension, the addition operation is performed, and finally the result is read out from the SRAM, and the decoding result is used as the output of the addition program in the MCU through data decoding.
[0014] Furthermore, the MCU module is implemented in C language, and the storage-computing integrated chip module is implemented in assembly language.
[0015] Furthermore, the STM single chip microcomputer is STM32H743IITxLQFP176.
[0016] The beneficial effects of the present invention are as follows:
[0017] 1. The various operation sub-modules in the operation module are independent of each other and there is data interaction between them, but the core module is a storage and computing integrated chip module, which can independently implement the adder function on the chip.
[0018] 2. The use of high-performance matrix adders can greatly improve the speed of matrix addition while ensuring calculation accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The present invention is further described below in conjunction with the accompanying drawings and embodiments.
[0020] Figure 1 It is the overall internal structure of the storage and computing integrated chip adder and the input and output signal flow diagram of each module in the embodiment of the present invention.
[0021] Figure 2 It is the internal structure and signal flow diagram of the adder in the storage-computation-in-one module of an embodiment of the present invention.
[0022] Figure 3 It is an operation flow chart of the storage-computation integrated chip adder according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] The following will clearly and completely describe and discuss the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings of the present invention. Obviously, what is described here is only a part of the examples of the present invention, not all the examples. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0024] In order to facilitate the understanding of the embodiments of the present invention, the following will be further explained by taking specific embodiments as examples in conjunction with the drawings, and each embodiment does not constitute a limitation on the embodiments of the present invention.
[0025] Embodiment 1 of the present invention refers to Figure 1 , Figure 2 As shown, the adder based on the storage-computing integrated chip adopts the STM32H743IITxLQFP176 of STMicroelectronics as the MCU, and the adder function is implemented on the storage-computing integrated chip provided by FlashEasy Semiconductor.
[0026] The MCU module is implemented by the memory, processor, basic input and output, and counter inside the MCU. The main implementation method is that the memory stores the input data, the designed encoder is used to encode the input data on the MCU, the basic input and output is responsible for data interaction with the storage and computing chip, the designed decoder is used to decode the data sent back by the chip on the MCU, and the counter completes the overall timing logic.
[0027] The main contents of the MCU module include the pre-calculation part and the interactive communication part. The MCU module is mainly implemented by the STM microcontroller. The control program is written and used using STM32CubeIDE. The specific implementation method is: first import the data to be processed into the MCU through the serial port, use the preprocessor to preprocess the data, and then write the processed data into the SRAM of the chip through the serial port, and then use the serial port to call out the addition program in the chip, give the relevant SRAM address and operation data dimension, perform the addition operation, and finally read the result from the SRAM, and use the data decoding to use the decoding result as the output of the addition program in the MCU. This is the implementation of the overall MCU module, which is mainly divided into the basic encoding part, the chip interaction part, and the decoding part.
[0028] The preprocessing part mainly deals with the data encoding operation, encoding a single int type data into 4 8-bit data, encoding a single floating point number into 8 parts (float type) or 12 parts (double type), and encoding the vector into a vector similar to the above, that is, a vector of 4 times, 8 times or 12 times the length.
[0029] The chip interaction part has the function of transmitting data and controlling chip behavior. By writing into SRAM, the encoded data, as well as the designed data input and output addresses and lengths are stored in SRAM. By calling the corresponding addition module that has been implanted in the chip, the chip is controlled to perform addition operations.
[0030] The decoding part is responsible for restoring the read SRAM data to the addition result, and the specific operation is the inverse operation of the encoding operation.
[0031] The storage and computing integrated chip module is implemented by the chip's internal SRAM, FLASH, DATABUS, etc. Its main implementation method is to use the data interaction interface DATABUS to interact with the MCU, store temporary data on the SRAM, and use FLASH to implement the main structure of the chip's internal adder to complete basic operations such as addition, judgment, and looping.
[0032] The storage and computing chip module refers to the part of the storage and computing chip that performs the addition operation by calling the relevant program after the MCU module sends the instruction to execute the addition program. The main implementation process is to first implement the main adder program through machine code at the bottom level, then use the serial port of the MCU to communicate with the chip, copy the written machine code into the chip, and by giving the program identifier, you can use the MCU to write the corresponding interface function to achieve the purpose of interacting with the chip.
[0033] The main content of the storage-computing integrated chip module is the mechanical code (assembly) implementation of the adder. Considering that the storage-computing integrated chip has its own set of instruction sets, its corresponding application and implementation functions are different from those of traditional chips. The mechanical code module of the adder mainly utilizes the addition, shift, truncation, jump, data reading and writing, loop judgment and other instructions in the instruction set. These instructions are relatively common, so the program is highly reusable.
[0034] The various adder programs of the module are burned into the chip, and a single int type data adder, a single floating point type data adder, an int vector adder and a floating point vector adder are realized through assembly. And they are made into functions for easy calling. When using this module, the MCU module is used to communicate and call the relevant programs and data in SRAM. The storage and calculation integrated chip stores the adder results in SRAM, and then uses SRAM to read the addition results.
[0035] In this embodiment, the various operation sub-modules in the operation module are independent of each other and there is data interaction between them, but the core module is a storage and computing integrated chip module and can independently implement the adder function on the chip.
[0036] The MCU module is implemented in C language, and the storage-computing chip module is implemented in assembly language with a specific instruction set. Figure 3 As shown:
[0037] (1) Load the data to be processed onto the MCU;
[0038] (2) The MCU module establishes communication with the storage and computing integrated chip module;
[0039] (3) The MCU module encodes the data to be processed;
[0040] (4) The MCU module executes instructions for transmitting data;
[0041] (5) Wait for the data transmission to end;
[0042] (6) Execute instructions to control the storage and computing chip to call the adder program;
[0043] (7) Wait for the storage and computing chip program to finish running and receive relevant information about the MCU module;
[0044] (8) The MCU module executes the program to read the data of the storage and computing integrated chip;
[0045] (9) The MCU module decodes the data;
[0046] (10) The decoded data is output as the adder result.
[0047] Among them, the working process of the storage and computing integrated chip module is as follows:
[0048] 1) The MCU module initiates a data storage instruction to store the data into SRAM;
[0049] 2) The MCU module initiates an instruction to call the corresponding adder program;
[0050] 3) The chip calls the corresponding adder program and executes the assembly program;
[0051] 4) When the program execution is completed, the chip returns the stop information to the MCU module.
[0052] The performance of the adder is tested with a 30-dimensional floating point array, and the data is randomly generated. The calculation accuracy of the adder using a storage and computing integrated chip can reach 10 -4 , compensating for the limited multiplication precision of the integrated storage and computing chip.
[0053] The adder algorithm is split and implemented as follows:
[0054] The adder algorithm is divided into several parts by data type, namely single int data addition, single floating point number addition, int vector addition, and floating point number vector addition. In terms of hierarchy, the basic part is single int data addition, and the rest can be achieved by adding corresponding structural transformations.
[0055] The integrated storage and computing chip contains basic addition instructions that can handle 16-bit addition (no carry). In order to implement single int data addition, it is necessary to use basic addition instructions for further design. Since the built-in addition instruction does not contain the carry part, the design uses a 16-bit adder to perform 8-bit addition operations and saves the carry result in the 9th bit. In practice, the data bit number of an int type data is 32 bits. It is converted into 4 8-bit data and imported into the integrated storage and computing chip. Then, the data with the same digits are added, and the carry data is added in the next level for addition. The result is read out through SRAM and reversely restored to int type data.
[0056] Mathematical expression: Consider int type data a and b, and divide them into a=[a1,a2,a3,a4], b=[b1,b2,b3,b4]. Among them, a1 corresponds to the lower 8 bits of data, and so on, a4 corresponds to the higher 8 bits of data. The carry is expressed as [t1,t2,t3,t4], that is, t1=carry(a1,b1), t2=carry(a2,b2), t3=carry(a3,b3), t4=carry(a4,b4). The carry represents the 9th bit of data after adding the two operands, that is, the carry data. Then the result c of a+b can be expressed as c=[c1,c2,c3,c4], where c1=sum(a1,b1), c2=sum(a2,b2,t1), c3=sum(a3,b3,t2), c4=sum(a4,b4,t3), where sum means adding the first 8 bits of the operands. The above operation can get the output result c. This is the mathematical representation of a single int type adder.
[0057] The implementation of a single floating-point adder (I) is to add a converter to the implementation of a single int type data adder. Considering that the storage and computing integrated chip does not have floating-point storage, a conversion operation is required to convert the floating-point data into int type data by pre-multiplying it by 10 to the fifth power, then rounding it to the integer, and then using a single int type data adder to complete the addition operation, and finally dividing it by 10 to the fifth power to restore it to double type data. This operation can realize a single double type data adder, and its adder accuracy can reach five decimal places, but there are requirements for the input data, and the data size is between -10000 and 10000 (approximately 2 to the 30th power divided by 100,000).
[0058] The implementation of a single floating-point adder (II) takes into account the established format of floating-point numbers, where the float data type is 32 bits, including 1 bit as a sign bit, 8 bits as an exponent bit, and a 23-bit decimal place. Similarly, the double data type is 64 bits, including 1 bit as a sign bit, 11 bits as an exponent bit, and 52 bits as decimal places. The following is the encoding of floating-point data and the step-by-step calculation after encoding. For float data, a single data is split into 6 parts, each with 8 bits, which are the sign bit, the exponent bit, and 4 parts of decimal places. The sign bit is operated with an AND operation, and the difference of the exponent bit is used as the shift of the corresponding decimal place (right shift and zero padding are preferred), and then the decimal places of the 4 parts are added using the addition of the int type. The last 6 parts are decoded back into float type data, thus completing the addition of the entire float type data. For the double data type, a single data will be split into 12 parts, including the sign bit, 2 exponent bits, and 9 parts of decimal places. The operation process is the same as that of float type data.
[0059] The implementation of the Int vector adder is based on a single int data type adder. The int vector is pre-split into 8-bit vectors that are four times the length of the previous vector, and stored in the SRAM of the storage and computing integrated chip through serial port communication. The int vector adder function address, SRAM address of input and output data, data length and other data are then given to the chip through DataBus. The chip calls the int vector adder function and stores the operation result in SRAM. After that, the output vector result is exported using SRAM to complete the implementation process of the int vector adder.
[0060] The entire int vector adder function is implemented through assembly. The main framework is equivalent to a single int adder, but a loop needs to be added on this basis to complete the entire vector addition process.
[0061] Method (I) of floating-point vector addition is to convert the floating-point vector into an int vector, use the int vector adder to get the addition result, and then convert it into a floating-point type vector to complete the adder operation. Method (II) considers the data composition principle of floating-point data, and through operations similar to floating-point adders, encodes each floating-point number into 8 or 16 parts (corresponding to float data and double data, respectively), and uses DataBus to batch transfer the above data to SRAM, and transfer the address of the floating-point vector adder, input and output SRAM address, data length and other information. After calculation, read the data in SRAM, export the output result, and restore it to floating-point type through decoding to complete the entire vector addition.
[0062] The working principle and steps of the present invention are as follows:
[0063] In the description of the present invention, it should be noted that the terms “first”, “second” and “third” are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0064] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A storage and computing integrated chip adder, characterized in that: It includes an MCU module and a storage-computing integrated chip module, wherein the MCU module and the storage-computing integrated chip module cooperate to implement matrix addition operations to solve the problem of precision loss of the storage-computing integrated chip in floating-point simulation calculations; The MCU module performs preprocessing of input data, including: Matrix addition splitting: split the input matrix into multiple vectors and perform calculations separately; Data format conversion: Due to the inherent precision limitation of analog computing of the storage and computing integrated chip, the MCU module converts the floating-point data into integer data by multiplying it by a predetermined factor, performs addition operation, and then decodes and restores it into floating-point data to improve the calculation accuracy; Data interaction optimization: SRAM is used for data caching, and the data bus is used to interact with the storage and computing chip module to improve data transmission efficiency; The storage-computation integrated chip module includes an addition operation program stored in FLASH, which performs single integer addition, single floating point addition, vector integer addition and vector floating point addition. This enables matrix addition operations, and under the test conditions of a 30-dimensional floating-point array, the calculation accuracy reaches 10⁻ 4 , effectively improving the accuracy of storage and computing integrated chips in floating-point calculations.
2. The storage-computation integrated chip adder according to claim 1, characterized in that: The MCU module is used to receive data to be processed, encode and decode data, and communicate with the storage and computing integrated chip. The storage and computing integrated chip module decomposes the adder process into several types, including single integer addition, single floating-point addition, vector integer addition and vector floating-point addition, so that the relevant adder can be called each time to process data.
3. The storage-computation integrated chip adder according to claim 2, characterized in that: The storage-computing integrated chip module includes FLASH, SRAM, and DATABUS located inside the storage-computing integrated chip, which are used to implement the adder body, transmit data with the MCU module, and write the result into the SRAM.
4. The storage-computation integrated chip adder according to claim 3, characterized in that: The storage-computing integrated chip module uses the data interaction interface DATABUS to interact with the MCU, stores temporary data on the SRAM, and uses FLASH to implement the main structure of the adder inside the chip to complete basic operations such as addition, judgment, and looping.
5. The storage-computation integrated chip adder according to claim 2, characterized in that: The MCU module is implemented by the memory, processor, basic input and output, and counter inside the MCU. The main implementation method is that the memory stores the input data, and the input data is encoded on the MCU using a designed encoder. The basic input and output are responsible for data interaction with the storage and computing integrated chip, and the data sent back by the chip is decoded on the MCU using a designed decoder, and the counter completes the overall timing logic.
6. The storage-computation integrated chip adder according to claim 5, characterized in that: The main contents of the MCU module include a pre-calculation part and an interactive communication part. The MCU module is mainly implemented by an STM single-chip microcomputer, and the control program is written and used using STM32. The specific implementation method is: first, the data to be processed is imported into the MCU through the serial port, and the data is preprocessed using a preprocessing program. Then, the processed data is written into the SRAM of the chip through the serial port, and then the addition program in the chip is called out using the serial port. Given the relevant SRAM address and operation data dimension, the addition operation is performed, and finally the result is read out from the SRAM, and the decoding result is used as the output of the addition program in the MCU through data decoding.
7. The storage-computation integrated chip adder according to claim 1, characterized in that: The MCU module is implemented in C language, and the storage-computing integrated chip module is implemented in assembly language.
8. The storage-computation integrated chip adder according to claim 6, characterized in that: The STM single chip microcomputer is STM32H743IITxLQFP176.
9. The storage-computation integrated chip adder according to claim 1, characterized in that: During the preprocessing process, the MCU module converts the format of the floating-point data, specifically including: converting the floating-point number into integer data using a predetermined multiplier for addition operation; restoring the integer data to floating-point data through decoding operation to compensate for the precision loss of the storage and computing integrated chip in floating-point operation and improve the calculation accuracy.
Citation Information
Patent Citations
Chip system for carrying out AI calculation based on NVM and operation method thereof
CN112988082A
Reduced instruction set storage and calculation integrated neural network coprocessor based on resistive memristor
CN113010213A