A hybrid-precision memory and computing control circuit based on Sense-Switch type pFLASH
By designing a hybrid precision memory control circuit based on Sense-Switch type pFLASH, the traditional memory architecture has solved the lack of computing accuracy adaptability, and realized a memory control circuit with configurable computing accuracy, which improves the flexibility and computing efficiency of the system.
Patent Information
- Application Number
- CN202510355931.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The traditional von Neumann architecture leads to storage wall and power consumption wall problems, while in-memory computing architectures have advantages in energy efficiency, but the existing FLASH memory computing chips are difficult to adapt to different computing accuracy requirements.
Design a hybrid accuracy memory control circuit based on Sense-Switch type pFLASH, including NOR FLASH type SSpF memory array, hybrid accuracy array controller, hybrid accuracy shift adder and AD module, which supports configurable calculation accuracy and realize 4, 8 or 16 bit convolutional calculations.
It realizes a memory control circuit with configurable computing accuracy, and can perform 4, 8 or 16 bit convolutional calculations in parallel, improving the flexibility and computing efficiency of the system.
Smart Images

Figure CN119862921B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of very large scale integrated circuit technology, and particularly to a mixed-precision memory-computation control circuit based on a Sense-Switch type pFLASH. Background Art
[0002] With the surge in the computing power required by edge AI chips, how to solve the computing energy efficiency problem and enhance the local processing ability has become the key. Due to the separation of memory and computing in the traditional von Neumann architecture, problems such as the "memory wall" and "power consumption wall" occur. While the in-memory computing architecture has great advantages in terms of energy efficiency due to the integration of computing and storage.
[0003] This concept was first proposed in the 1990s and began to receive more extensive attention and research in the 2010s. Compared with near-memory computing, in-memory computing directly executes computing operations in memory, which can significantly reduce the data transmission time. Since computing and storage are performed at the same location, the data movement between memory and the processor is reduced, and the performance in processing large data and complex computations is greatly improved. In addition, in-memory computing can be compatible with existing memory architectures and can be used in combination with various processors to enhance the flexibility of the system.
[0004] According to the signal type, the memory-computation architecture can be divided into three categories: analog memory-computation architecture, digital memory-computation architecture, and hybrid analog-digital memory-computation architecture. Among them, the digital memory-computation architecture is characterized by performing computing operations in digital circuits, with higher digital computing precision and calculation result accuracy. Therefore, this advantage can be utilized to implement complex computing operations. However, the disadvantages are higher hardware design complexity, larger power consumption, and the need for additional power supply and heat dissipation designs, etc.; the analog memory-computation architecture uses analog circuits to perform computing operations. Its advantages are lower hardware design complexity and lower energy consumption, so the system performance is greatly improved. However, due to the poor stability of analog devices and the sensitivity of analog circuits to changes in environmental conditions, the computing precision and calculation result accuracy of analog memory-computation are relatively low. To maintain computing accuracy, the circuit stability under various environmental conditions needs to be considered; the hybrid analog-digital memory-computation architecture combines the characteristics of analog and digital circuits, combines analog and digital computing units together to achieve hybrid computing functions. This type of architecture takes into account both the precision of digital computing and the efficiency of analog computing, and can flexibly select analog or digital computing units to perform computing operations according to specific application requirements. Its design complexity is higher, and issues such as the interface and signal conversion between analog and digital computing units need to be considered, and there is less existing research.
[0005] Since the analog memory-computation architecture has low power consumption and combines the advantages of large-scale integration of SSpF devices, it can significantly improve performance, reduce energy consumption and costs. However, there are still problems with the design flexibility of FLASH memory-computation chips: it is difficult to adapt to different computing precision requirements. Summary of the Invention
[0006] The object of the present invention is to provide a mixed-precision memory and computing control circuit based on Sense-Switch type pFLASH to solve the problems in the background technology.
[0007] To solve the above technical problems, the present invention provides a mixed-precision memory and computing control circuit based on Sense-Switch type pFLASH, including:
[0008] The NOR FLASH type SSpF memory and computing array is used to implement convolution operations, which is composed of reconfigurable differential Sense-Switch type pFLASH memory and computing units. The reconfigurable differential Sense-Switch type pFLASH memory and computing unit supports 4-bit signed weight storage and realizes 4-bit signed multiplication operations;
[0009] The mixed-precision array controller is composed of a state machine, configurable registers, counters, WL decoders and BL decoders. It divides the NOR FLASH type SSpF memory and computing array according to the configured computing precision, and completes the input control and output sampling of 4, 8 or 16-bit convolution operations, as well as the 4-bit weight programming and reading;
[0010] The mixed-precision shift adder is composed of a serial shift adder, a channel shift adder, a state machine and configuration registers. It adds the outputs of the mixed-precision array controller by shifting serially and by channels according to the configured computing precision to obtain 4, 8 or 16-bit convolution calculation results;
[0011] The AD module is composed of a current-voltage conversion circuit and an analog-to-digital conversion circuit, which realizes the conversion of the current output of the NOR FLASH type SSpF memory and computing array into a voltage output and converts the analog signal into a 4-bit digital signal.
[0012] In one embodiment, the reconfigurable differential Sense-Switch type pFLASH memory and computing unit is a differential structure composed of two Sense-Switch type pFLASH devices. A single Sense-Switch type pFLASH device is a three-terminal FLASH device containing a bit line BL, a source line SL, and a word line WL. The BL end is the input end, the WL end is the strobe end, and the SL end is the output end;
[0013] The i output current of the
[0014]
[0015] row is expressed as: wherei Row j The equivalent conductance of the column Sense-Switch type pFLASH device, is the j input voltage of the column. The multiply-accumulate operation is implemented in parallel for each row; Utilize the I-V characteristic of the device operating in the linear region:
[0016]
[0017] is the carrier mobility, is the oxide capacitance, is the aspect ratio of the device, is the gate-source voltage, is the threshold voltage, is the drain-source voltage.
[0018] Based on this formula, the output current at the SL end in the differential structure eliminates the second-order term in the original formula:
[0019]
[0020] is the positive weighted current of the i row and the j column, is the negative weighted current of the i row and the j column, is the input voltage of the positive weight of the j column, is the input voltage of the negative weight of the j column. At this time, is regarded as the input voltage, and the input is changed by changing the voltage value at the BL end; is regarded as the signed conductance G ij , and the conductance is changed by changing the threshold voltage: G ij .
[0021] In one implementation, the hybrid-precision array controller configures the calculation precision and divides the NOR FLASH type SSpF memory-computation array: When configuring 4-bit calculation precision, no division is performed; When configuring 8-bit calculation precision, the word line WL of the array is divided into two parts to calculate the high 4 bits and the low 4 bits respectively; When configuring 16-bit calculation precision, the word line WL of the array is divided into four parts, and each part calculates 4 bits among them. The erasure, weight writing, and reading of the NOR FLASH type SSpF memory-computation array are completed through the APB protocol.
[0022] In one embodiment, the mixed-precision shift adder configures the calculation precision and performs shift and addition operations on the output result of the mixed-precision array controller: the sequence shift adder adds the stored and calculated results input each time by shifting them in cycles, and the shift number and addition times are controlled by a counter; the channel shift adder combines the partitioning method of the NOR FLASH type SSpF storage and calculation matrix by the mixed-precision array controller, performs shift addition on the results of the same output channel according to the configured precision, and outputs the results.
[0023] A mixed-precision storage and calculation control circuit based on the Sense-Switch type pFLASH provided by the present invention, through the design of a mixed-precision array controller and a shift adder, supports configurable calculation precision, realizes 4, 8 or 16-bit convolution calculations, can be applied to a mixed-precision in-memory calculation core architecture based on SSpF, parallelly realizes 4, 8 or 16-bit convolution calculations, and the calculation precision is configurable. Description of the Drawings
[0024] Figure 1 It is the architecture diagram of the SSpF mixed-precision storage and calculation control circuit provided by the present invention.
[0025] Figure 2 It is the structure diagram of the NOR FLASH type SSpF storage and calculation array provided by the present invention.
[0026] Figure 3 It is the timing diagram of weight writing and erasing of the mixed-precision array controller provided by the present invention.
[0027] Figure 4 It is the timing diagram of weight reading of the mixed-precision array controller provided by the present invention.
[0028] Figure 5 It is the timing diagram of weight storage and calculation of the mixed-precision array controller provided by the present invention.
[0029] Figure 6 It is a schematic diagram of a mixed-precision storage and calculation array partitioning method provided by the present invention.
[0030] Figure 7 It is the flowchart of the mixed-precision algorithm provided by the present invention.
[0031] Figure 8 It is the structure diagram of the mixed-precision shift adder provided by the present invention. Detailed Embodiments
[0032] The following further elaborates on a mixed-precision memory and computing control circuit based on the Sense-Switch type pFLASH in conjunction with the accompanying drawings and specific embodiments. According to the following description, the advantages and features of the present invention will be clearer. It should be noted that the accompanying drawings are all in a very simplified form and use non-precise scales, only for the purpose of conveniently and clearly assisting in explaining the purpose of the embodiments of the present invention.
[0033] The present invention provides a mixed-precision memory and computing control circuit based on the Sense-Switch type pFLASH, and its overall architecture is as Figure 1 shown, including a NOR FLASH type SSpF memory and computing array, a mixed-precision array controller, a mixed-precision shift adder, and an AD module. Among them, the NOR FLASH type SSpF memory and computing array is composed of reconfigurable differential Sense-Switch type pFLASH (SSpF) memory and computing units, which are used to implement convolution operations; the mixed-precision array controller inputs the input data in the form of a bit stream, one bit at a time in ascending order every cycle, into the NOR FLASH type SSpF memory and computing array. The NOR FLASH type SSpF memory and computing array outputs the current obtained by multiplication and accumulation to the AD module. The AD module converts the current to a voltage and then converts the analog signal to a digital signal, and transmits the digital signal to the mixed-precision array controller. The mixed-precision array controller samples and calculates the results according to the output timing of the ADC in the AD module and outputs them to the mixed-precision shift adder. In addition, the mixed-precision array controller can implement the programming and reading of weights and the erasure of the array through the APB protocol. Each module will be described in detail below.
[0034] The weights in the convolutional neural network can be replaced by the number of charges written into the SSpF. When it works as a synapse, information is transmitted from the drain to the source. In addition, multiple units connected in parallel to the same gate control line can be regarded as a neuron to complete the multiplication and accumulation operation. An SSpF can be regarded as a three-terminal FLASH device containing BL, SL, and WL. BL is connected to the drain of the FLASH, WL is connected to the control gate terminal of the FLASH, and SL is connected to the source of the FLASH. When inputting data, the FLASH has two working regions corresponding to different input methods: inputting data from the BL terminal and inputting data from the WL terminal.
[0035] When the FLASH works in the linear region, its current is:
[0036]
[0037] Among them, is the carrier mobility, is the oxide capacitance, is the aspect ratio of the device, is the gate-source voltage, is the drain-source voltage, is the threshold voltage; based on this formula, the read current at the SL end can be obtained as the product of the two minus a second-order term, and at this time, the second-order term cannot be ignored. When two SSpFs form a differential structure and the two currents are subtracted, the second-order term in the original formula is eliminated, ensuring linearity. In addition, the weight can be represented by the difference in the threshold voltages of a group of SSpFs and can be expressed as a positive or negative number.
[0038] The following table shows the I-V characteristic relationship and the read current range in different operating intervals of SSpF
[0039]
[0040] As can be seen from the table, when the device operates in the subthreshold region, the read current range is small, so the power consumption is smaller, but the derivation formula contains a logarithmic operation and the calculation complexity is higher. When the device operates in the deep linear region, the voltage range that meets the conditions is small, and the drain-source current I DS and the drain-source voltage V DS are approximately linearly related, resulting in errors in the calculation results and requiring additional circuit design for correction. When the device operates in the linear region, the second-order term in the I-V formula can be cancelled out through a differential design to obtain a linear relationship. Compared with the previous two schemes, the calculation error is smaller and the calculation complexity is lower.
[0041] The present invention uses a NOR FLASH type SSpF memory-computation array as the memory-computation simulation array, and its structure is as Figure 2 shown. Each BL is respectively connected to the drain of each row of FLASH, each WL is respectively connected to the control gate of each column, and each SL is respectively connected to the source of each column. The following is the specific derivation for implementing vector-matrix multiplication. Given the vector-matrix multiplication formula:
[0042]
[0043] The voltage amplitude input method is used for vector-matrix multiplication, and the current is converted into a voltage through the conductance G L (IVC circuit). The input vector X and the output vector Y are both voltage signals, and the implemented vector-matrix multiplication is:
[0044]
[0045] The conductance G ij stores by programming and erasing to change the threshold voltage of the FLASH. The output current of a single device is accumulated based on the NOR FLASH type array and Kirchhoff's current law. For example, the relationship between the total current I i of the oi row and the input voltage V j can be expressed as:
[0046]
[0047] When operating in the linear region, let W ij = W i+j W i-j , W i+j represent the positive weight value of the i th row and j th column; W i - j represent the negative weight value of the i th row and j th column, and there is the following formula:
[0048]
[0049]
[0050] where is the positive-weight current of the i th row and j th column, is the negative-weight current of the i th row and j th column, is the input voltage of the positive weight of the j th column, is the input voltage of the negative weight of the j th column, is the carrier mobility, is the oxide capacitance, is the width-to-length ratio of the device, is the gate-source voltage of the i th row, is the drain-source voltage of the positive weight of the i th row and j th column, is the drain-source voltage of the negative weight of the i th row and j th column;
[0051] And it satisfies the condition:
[0052]
[0053] is the drain-source voltage of the i th row and j th column, is the gate-source voltage of the negative weight of the i th row, For the i Row positive weight gate-source voltage,
[0054] Then the output current of the differential SSpF storage unit can be obtained as:
[0055]
[0056] at this time can be regarded as the input voltage, that is , the input can be changed by changing the voltage value of the BL terminal. Can be regarded as the conductivity G ij , the conductance can be changed by changing the threshold voltage.
[0057] The main function of the mixed precision array controller is to output the digital signal to the NORFLASH type SSpF storage array according to the timing required by the analog circuit, and at the same time receive the digital signal output by the ADC and perform corresponding processing. In order to save the DAC circuit to reduce the design complexity, the present invention adopts a bit stream input method. For example, if you want to achieve 4-bit input accuracy, you need to input the data in 4 clock cycles, and input 1 bit in each clock cycle. At this time, the output result of the NOR FLASH type SSpF storage array cannot be directly transmitted to the operator module, but a shift accumulation operation is performed. The output of the 0th cycle remains unchanged, and the output result of the 1st cycle corresponds to the 1st input data. It needs to be shifted left by 1 bit and accumulated with the result of the 0th cycle, and so on. After 4 cycles, the final accumulation result is obtained.
[0058] In order to realize the configurability of information such as calculation accuracy, storage delay, write and erase pulse width, the designed digital circuit is divided into three modules, namely, configuration register module, read-write-erase control module and storage control module. Correspondingly, the control circuit of the present invention has three states, namely, configuration register state, write, erase and read weight state, and storage state.
[0059] The configuration register is written by APB when configuring parameters. After completing parameter configuration, write, erase and read weight instructions can be sent through APB. The addresses of erase and write are different. The write operation needs to enter the address of the corresponding FLASH unit, and only a fixed address is needed for erase. According to the requirements of the analog circuit, the BL and WL ends need to design corresponding decoding circuits. In addition, multiple buffer signals are required before starting the analog array. The write and erase timing design is as follows: Figure 3 As shown, Figure 3A counter configuration scheme for signal delay and hold is given, where WRITE represents the programming weight state, ERASE represents the erase state, and IDLE represents the initial state. The read operation needs to cooperate with the output timing of the AD module. By increasing the counter to control the sampling time, data is transmitted to the prdata port of the APB when the ADC output in the AD module is valid. The read timing design of the mixed-precision array controller is as Figure 4 shown, where READ represents the read weight state.
[0060] In the compute-in-memory state, the input data is continuously input in multiple cycles. For example, when the computing precision is 8 bits, it needs to be input in 8 cycles. The mixed-precision array controller internally shifts the upstream data. During this period, the handshake signal ready_o to the upstream is pulled low and raised until the end of a compute-in-memory operation. Its timing design is as Figure 5 shown. Among them, BLEN is input to the analog array, and is the BL strobe signal, and CIM represents the compute-in-memory state. The control circuit of the present invention transmits the array output result to the downstream shift and add circuit.
[0061] The method for realizing the mixed precision of weights is to divide the NOR FLASH type SSpF compute-in-memory array. For example, if the scale of the analog array is 2j*i, that is, 2j WL terminals and i BL terminals, and each NOR FLASH type SSpF compute-in-memory unit stores 4-bit precision weights. When the computing precision configuration is 4, the array does not need to be divided. When the computing precision configuration is 8, the array is divided into two blocks. The first j rows are responsible for the low 4-bit calculation, and the last j rows are responsible for the high 4-bit calculation. And so on, when the computing precision is 16, the array is divided into 4 blocks, each block has 64 rows and is responsible for 4-bit calculations in order of high and low. Figure 6 Shows the matrix division method when the precision configuration is 8.
[0062] Figure 7 Taking the implementation of 8*8-bit operation with a 4-bit weight storage as an example, it demonstrates the mixed-precision operation process through bit-stream input, serial shift adder, and channel shift adder. Therefore, different precision shift accumulation operations can be realized by changing the value of the configuration register. To ensure the stability of the calculation process, only the state is allowed to be changed during the configuration parameter stage, and the state remains unchanged in the compute-in-memory state. The shift adder is internally divided into two major parts: the serial shift adder and the channel shift adder. The internal structure is shown in Figure 8 . Among them, the number of serial shift adders is the same as the number of WLs of the compute-in-memory array, and the compute-in-memory results of each bit are calculated in parallel. The shift number and the number of addition operations are controlled by a counter. Therefore, the output of the serial shift adder is the complete compute-in-memory result of each WL terminal. If mixed precision is to be realized, it is also necessary to combine the matrix division method to perform shift addition on the results of the same output channel according to the precision configuration.
[0063] IVC (Integrate-and-Voltage Converter) is a circuit that converts an input current signal into a corresponding voltage signal. The IVC module has two functions: (1) clamping the voltages of the source lines SLP and SLN of the FLASH memory and computing array and providing the required memory and computing current for the memory and computing array; (2) converting the computing current of the FLASH memory and computing array into a voltage and sending it to the ADC for conversion into a digital code. In this structure, the source lines SLP and SLN of the FLASH memory and computing array are clamped through the negative input terminal of the operational amplifier, and the current flowing through the source lines is converted through a resistor to convert the current into a corresponding voltage signal. Subsequently, the converted voltage signal is sent to the backend ADC for quantization, thereby achieving accurate sampling and processing of data. The differential expression of the signal output by the IVC is:
[0064]
[0065] The ADC circuit (Analog to Digital Converter) is an important component for converting analog signals into digital signals, and its working principle is as follows: (1) Sampling and quantization: The ADC circuit first samples the input continuous analog signal and quantizes it into discrete digital values. This process usually includes two steps: time sampling and amplitude quantization to ensure that each instantaneous value of the analog signal can be accurately captured. (2) Encoding: After completing sampling and quantization, the ADC converts the quantized value into a binary code and outputs a digital signal for subsequent digital circuit processing.
[0066] The above description is only a description of the preferred embodiments of the present invention and does not limit the scope of the present invention in any way. Any changes and modifications made by those of ordinary skill in the field of the present invention based on the above disclosure shall fall within the scope of protection of the claims.
Claims
1. A mixed precision storage and calculation control circuit based on Sense-Switch type pFLASH, characterized in that: It includes a NORFLASH type SSpF storage and calculation array, a mixed precision array controller, a mixed precision shift adder and an AD module; The NOR FLASH type SSpF storage and calculation array is used to implement convolution operations, and is composed of reconfigurable differential Sense-Switch type pFLASH storage and calculation units. The reconfigurable differential Sense-Switch type pFLASH storage and calculation units support 4-bit signed weight storage and implement 4-bit signed multiplication operations. The mixed precision array controller is composed of a state machine, a configurable register, a counter, a WL decoder and a BL decoder. It divides the NOR FLASH type SSpF storage array according to the configured calculation precision, completes the input control and output sampling of 4, 8 or 16-bit convolution operations, and 4-bit weight burning and reading; The mixed precision shift adder is composed of a sequence shift adder, a channel shift adder, a state machine and a configuration register. The output of the mixed precision array controller is added in sequence and channel shift according to the configured calculation accuracy to obtain a 4, 8 or 16-bit convolution calculation result; The AD module is composed of a current-voltage conversion circuit and an analog-to-digital conversion circuit, which can convert the current output of the NOR FLASH type SSpF storage array into a voltage output and convert the analog signal into a 4-bit digital signal.
2. The mixed precision storage and calculation control circuit based on Sense-Switch type pFLASH according to claim 1, characterized in that: The reconfigurable differential Sense-Switch pFLASH storage unit is a differential structure composed of two Sense-Switch pFLASH devices. A single Sense-Switch pFLASH device is a three-terminal FLASH device including a bit line BL, a source line SL, and a word line WL. The BL terminal is an input terminal, the WL terminal is a selection terminal, and the SL terminal is an output terminal. No. i The output current of the line is expressed as: (1) in, For the i OK j The equivalent conductance of the Sense-Switch pFLASH device is shown in Figure 2. For the j The input voltage of the column is used to realize multiplication and accumulation operation in parallel for each row; the IV characteristics of the device working in the linear region are used: (2) is the carrier mobility, is the oxide layer capacitance, is the aspect ratio of the device, is the gate-source voltage, is the threshold voltage, is the drain-source voltage; Based on formula (2), the output current of SL terminal under the differential structure eliminates the second-order term in formula (2): (3) For the i Line j Column positive weight current, For the i Line j Column negative weight current, For the j The positive weighted input voltage, For the j The input voltage of the column is negative weighted. Considered as input voltage, the input is changed by changing the voltage value of BL terminal; As a signed conductivity G ij , the conductance is changed by changing the threshold voltage: G ij .
3. The mixed precision storage and calculation control circuit based on Sense-Switch type pFLASH as claimed in claim 2, characterized in that: The mixed precision array controller configures the calculation precision and divides the NOR FLASH type SSpF storage array: when the 4-bit calculation precision is configured, no division is performed; when the 8-bit calculation precision is configured, the word line WL of the array is divided into two parts, and the upper 4 bits and the lower 4 bits are calculated respectively; when the 16-bit calculation precision is configured, the word line WL of the array is divided into four parts, and each part calculates 4 bits. The erasing, weight burning and reading of the NOR FLASH type SSpF storage array are completed through the APB protocol.
4. The mixed precision storage and calculation control circuit based on Sense-Switch type pFLASH as claimed in claim 3, characterized in that: The mixed precision shift adder configures the calculation precision and performs shift and addition operations on the output results of the mixed precision array controller: the sequence shift adder shifts and adds the storage calculation results of each input in a cycle, and controls the shift bit number and the number of additions through a counter; the channel shift adder combines the mixed precision array controller to divide the NOR FLASH type SSpF storage calculation matrix, and performs shift addition on the results of the same output channel according to the configured precision and outputs the results.
Citation Information
Patent Citations
Precision-configurable convolution hardware structure of deep learning hardware accelerator
CN110458277A
Convolutional neural network accelerator based on mixed low-precision quantization and design method thereof
CN118211621A