Sparse digital memory calculation method and circuit for high-density RRAM (Resistive Random Access Memory)
By designing a high-density RRAM sparse digital in-memory computing circuit and logic gate instead of adder, the problems of complexity of high-density RRAM array read circuit and low utilization efficiency of sparse network hardware resources are solved, and efficient data read, write and calculate performance are achieved.
Patent Information
- Application Number
- CN202510740590.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing high-density RRAM array read circuits are complex and have high power consumption, and the hardware resource utilization efficiency of sparse neural networks is low. Traditional computing architectures cannot fully utilize the hardware resource efficiency of sparse computing.
A high-density RRAM sparse digital in-memory computing circuit is designed, and the RRAM in-memory computing unit is combined with a digital logic unit. The peripheral circuit unit realizes parallel reading and writing functions, and the hardware architecture is optimized through logic gates instead of adders, combining with the weight mask fine-tuning training method.
It significantly improves data reading and writing efficiency, reduces energy consumption and delay, and improves the computing efficiency and hardware resource utilization of sparse neural networks.
Smart Images

Figure CN120255849A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of integrated circuit design, and particularly relates to a high-density RRAM sparse digital in-memory computing method and circuit. Background Art
[0002] With the development of satellite technology, the complexity of payload data processing has been continuously increasing, and on-board intelligent computing systems are facing real-time and energy-efficient computing requirements. When dealing with these requirements, the traditional von Neumann architecture has significant bottlenecks, mainly reflected in the frequent transmission of computing data between the processor and the memory, resulting in low computing energy efficiency and high latency, especially in data-intensive applications such as satellite image processing, neural network inference, and large-scale data analysis. To solve this problem, in-memory computing technology has emerged, which embeds the computing process inside the memory, reducing the data transmission between the memory and the processor, and greatly improving the computing efficiency and energy efficiency. However, existing in-memory computing architectures, especially those based on resistive random-access memory (RRAM), still face some technical challenges.
[0003] First of all, the complexity of the read circuit for high-density RRAM arrays is an urgent problem to be solved. RRAM has a high array density and non-volatility, but its reading process requires converting the stored resistance into a digital signal, which requires a complex and power-consuming read circuit. Especially in high-density arrays, existing read circuits often cannot efficiently and accurately read the stored digital information, thus affecting the performance and energy efficiency of the overall system. Secondly, the utilization efficiency of hardware resources for existing sparse neural network computing is relatively low. Traditional hardware architectures, especially when dealing with neural network computing, rely on traditional adders and multipliers for operations. Although these components are effective in general, when dealing with sparse neural networks, many multipliers and adders are not fully utilized in the calculation.
[0004] Therefore, how to improve the computing efficiency of sparse networks by optimizing the hardware architecture has become a key technical problem. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present invention provides a high-density RRAM sparse digital in-memory computing method and circuit. The technical problems to be solved by the present invention are achieved through the following technical solutions: In a first aspect, the present invention provides a high-density RRAM sparse digital in-memory computing circuit, including: A plurality of RRAM in-memory computing units arranged in an array, and the RRAM in-memory computing units are configured to store data through resistance states, perform digital calculations, and obtain output results; A plurality of digital logic units arranged in an array, the digital logic units being electrically connected to the RRAM in-memory computing units, and the digital logic units being configured to perform multiply-accumulation on the output results to obtain a multiply-accumulation result; wherein, for multiple RRAM in-memory computing units in the same column of RRAM in-memory computing units, one digital logic unit is correspondingly provided; A peripheral circuit unit, the peripheral circuit unit being electrically connected to the RRAM in-memory computing units, the peripheral circuit unit including a read path and a write path, and being configured to simultaneously read and write the resistance states of the RRAM in-memory computing units; wherein, for one column of RRAM in-memory computing units, one peripheral circuit unit is correspondingly provided.
[0006] In a second aspect, the present invention further provides a high-density RRAM sparse digital in-memory computing method, including: Writing the resistance states of the RRAM in-memory computing units through the write path of the peripheral circuit unit; wherein, using the preset binarized weight matrix as the basis for writing the resistance states of the RRAM in-memory computing units, the preset binarized weight matrix is obtained according to the trained neural network, and there is at most one non-zero weight in each row of the preset binarized weight matrix; Setting the word line WL voltage and the bit line BL voltage input to the RRAM in-memory computing units; wherein, using the preset eigenvalue as the basis for the bit line BL voltage input to the RRAM in-memory computing units, the preset eigenvalue is obtained according to the trained neural network; The RRAM in-memory computing units perform digital calculations according to the received resistance states, word line WL voltage, and bit line BL voltage to obtain output results; The digital logic units perform multiply-accumulation calculations on the output results to obtain multiply-accumulation results.
[0007] Advantages of the present invention: A high-density RRAM sparse digital in-memory computing method and circuit provided by the present invention include RRAM in-memory computing units, digital logic units, and peripheral circuit units; wherein, the RRAM in-memory computing units are composed of RRAM and NMOS transistors and have in-memory computing functions, the peripheral circuit units include read paths and write paths, realize the function of simultaneously reading and writing a row of RRAM in-memory computing units, and have storage functions; the digital logic units are sparse logic gates and adder trees, perform multiply-accumulation on the output results, and output the final results; thus, aiming at the array density problem caused by the contradiction between the high-density RRAM memory-compute array and the complex read circuit, the peripheral circuit unit provided by the present invention has read paths and write paths, realizes the function of simultaneously reading and writing the same row of RRAM, can significantly improve the data reading and writing efficiency, and reduce the energy consumption and delay in the traditional RRAM computing architecture.
[0008] The following will further describe the present invention in detail with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 FIG. is a schematic diagram of a high-density RRAM sparse digital in-memory computing circuit provided by an embodiment of the present invention; Figure 2 FIG. is a flowchart of a high-density RRAM sparse digital in-memory computing method provided by an embodiment of the present invention; Figure 3 FIG. is a schematic diagram of digital logic unit calculation provided by an embodiment of the present invention; Figure 4 FIG. is a schematic diagram of weight mask fine-tuning training provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0010] The present invention will be further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0011] In the existing RRAM digital in-memory computing architecture, the most common solution is to perform calculations by reading the resistance values of RRAM cells. For example, a high-density 1T1R cell array is used to store data, and multiply-accumulate operations are performed by reading the resistance values. However, due to the poor direct compatibility between the resistance state of RRAM and digital calculations, existing calculation methods usually require complex analog-to-digital conversion circuits, which consume a large amount of power and are easily affected by accuracy. Especially when the array scale increases, the complexity and power consumption problems of the reading circuit become particularly prominent. In addition, the non-ideality of resistance changes, such as errors and resistance drift in read and write operations, is also a key factor limiting the wide application of RRAM as an in-memory computing unit.
[0012] In addition, the existing sparse neural networks are increasingly used in neural network inference and training. In the existing sparse neural networks, many connection weights are zero, and traditional hardware architectures cannot fully utilize this feature. Most current sparse computing architectures reduce storage requirements and computational burdens through hardware-level sparse representations (such as filtering out zero weights). However, the existing sparse optimization methods optimize sparsity in the training and inference stages of the network, while at the hardware level, many architectures do not truly achieve hardware acceleration for sparse calculations. This limitation makes most sparse neural network architectures unable to fully utilize the maximum efficiency of their hardware resources.
[0013] In summary, the existing methods have the following defects: (1) The high-density RRAM array requires a complex reading circuit, resulting in limited system speed and energy efficiency.
[0014] (2) The hardware adders and multipliers in traditional sparse network calculations occupy a large amount of hardware resources, resulting in low computational efficiency.
[0015] In view of this, the present invention provides a high-density RRAM sparse digital in-memory computing method and circuit, and a high-density RRAM array providing parallel read and write functions, so as to solve the contradiction between the traditional RRAM computing array and the complex read circuit.
[0016] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of a high-density RRAM sparse digital in-memory computing circuit provided by an embodiment of the present invention. A high-density RRAM sparse digital in-memory computing circuit provided by the present invention includes: A plurality of RRAM in-memory computing units 10 arranged in an array. The RRAM in-memory computing unit 10 is configured to store data through a resistance state, perform digital calculations, and obtain an output result; A plurality of digital logic units 20 arranged in an array. The digital logic unit 20 is electrically connected to the RRAM in-memory computing unit 10. The digital logic unit 20 is configured to perform multiply-accumulation on the output result to obtain a multiply-accumulation result. Among them, one digital logic unit 20 is correspondingly provided for multiple RRAM in-memory computing units 10 in the same column of RRAM in-memory computing units; A peripheral circuit unit 30. The peripheral circuit unit 30 is electrically connected to the RRAM in-memory computing unit 10. The peripheral circuit unit 30 includes a read path and a write path, and is configured to simultaneously read and write the resistance state of the RRAM in-memory computing unit 10. Among them, one peripheral circuit unit 30 is correspondingly provided for one column of RRAM in-memory computing units.
[0017] Specifically, please continue to refer to Figure 1 . The high-density RRAM sparse digital in-memory computing circuit provided in this embodiment includes an RRAM in-memory computing unit 10, a digital logic unit 20, and a peripheral circuit unit 30. Among them, the RRAM in-memory computing unit 10 is composed of an RRAM and an NMOS transistor 11 and has an in-memory computing function. The peripheral circuit unit 30 includes a read path and a write path, realizes the function of simultaneously reading and writing a row of RRAM in-memory computing units, and has a storage function. The digital logic unit 20 is a sparse logic gate and an adder tree, performs multiply-accumulation on the output result, and outputs the final result. In this way, aiming at the array density problem caused by the contradiction between the high-density RRAM memory computing array and the complex read circuit, the peripheral circuit unit 30 provided by the present invention has a read path and a write path, realizes the function of simultaneously reading and writing the same row of RRAMs, can significantly improve the data read and write efficiency, and reduce the energy consumption and delay in the traditional RRAM computing architecture.
[0018] In an optional embodiment of the present invention, please continue to refer to Figure 1 . The RRAM in-memory computing unit 10 includes an NMOS transistor 11, a first RRAM 12, and a buffer 13. Among them, The gate of the NMOS transistor 11 is electrically connected to the word line WL, the drain of the NMOS transistor 11 is electrically connected to the peripheral circuit unit 30, the source of the NMOS transistor 11 is electrically connected to the first node N1, the first end of the first RRAM 12 is electrically connected to the first node N1, the second end of the first RRAM 12 is electrically connected to the bit line BL, the input end of the buffer 13 is electrically connected to the first node N1, and the output end of the buffer 13 is electrically connected to the digital logic unit 20.
[0019] Specifically, please continue to refer to Figure 1 , in this embodiment, the RRAM in-memory computing unit 10 is composed of an RRAM and an NMOS transistor 11, has an in-memory computing function, and obtains the digital result stored in the RRAM by the voltage division of the RRAM and the gated NMOS transistor 11 to reduce the area. For multiple rows of RRAMs belonging to the same column, the multiplication results are transmitted to the adder tree for multiply-accumulate calculation to obtain the multiply-accumulate result, and this result is transmitted to the shift-accumulate circuit for shift-accumulation to obtain the final output. It can be understood that the digital result stored in the RRAM is obtained by the voltage division of the RRAM and the gated NMOS transistor and increasing the swing, realizing the compatible design of the RRAM and digital in-memory computing, and solving the challenge of large area overhead.
[0020] In an optional solution, for the RRAM in-memory computing unit 10, the resistance value of the RRAM needs to be read out before performing digital in-memory computing. Through the voltage division of the RRAM and the gated NMOS transistor 11, a 1-bit value of the RRAM is obtained. When the RRAM is in the low-resistance state, that is, storing 1, the voltage division point is at a high voltage; when the RRAM is in the high-resistance state, that is, storing 0, the voltage division point is at a low voltage; this high voltage or low voltage is input to the logic gate and the adder tree unit for digital calculation after the swing is increased by the buffer 13, and the multiply-accumulate result is obtained. The buffer 13 selects the circuit structure of cascaded inverters to simplify the circuit design; the buffer 13 adopting this unit circuit can effectively reduce the power consumption and area overhead of the in-memory computing core, and realize the compatibility of the RRAM and digital in-memory computing.
[0021] In an optional embodiment of the present invention, please continue to refer to Figure 1 , the peripheral circuit unit 30 includes a single-pole double-throw switch 31, a second RRAM 32, and an operational amplifier 33; wherein, When the common terminal of the single-pole double-throw switch 31 is conducted with the first terminal, a write operation is performed on the resistance state of the RRAM in-memory computing unit 10; When the common terminal of the single-pole double-throw switch 31 is electrically connected to the second terminal, the second terminal of the single-pole double-throw switch 31 is electrically connected to the second node N2, the first terminal of the second RRAM 32 is electrically connected to the second node N2, the second terminal of the second RRAM 32 is electrically connected to the output terminal of the operational amplifier 33, one input terminal of the operational amplifier 33 is connected to the second node N2, and the other input terminal of the operational amplifier 33 receives the Vclp voltage signal to perform a read operation on the resistance state of the RRAM in-memory computing unit 10.
[0022] Specifically, please continue to refer to Figure 1 , in this embodiment, the peripheral circuit unit 30 includes a read path and a write path, which realizes the simultaneous read and write functions for a row of RRAM in-memory computing units 10 and has a storage function. In order to accurately read the RRAM resistance, a low-power and high-precision negative feedback read circuit is designed. The resistors used in the negative feedback are RRAMs with the same process to reduce the read circuit error, realize the readout method of converting the RRAM resistance to voltage, improve the read accuracy, and effectively reduce the power consumption. Especially, it has great advantages in high-density storage arrays. Based on this, the high-speed and low-power RRAM parallel read and write function is realized, and a row of RRAMs can be read and written simultaneously.
[0023] In an optional solution, read paths and write paths are provided for each column of RRAM in-memory computing units in the SL drive. The write path realizes changing the resistance of a row of first RRAMs simultaneously. At this time, the BL voltage is set to VDD / 2, and the SL voltage of this path is 0. For the write process of the low-resistance state (SET) of the first RRAM 12 (i.e., editing the first RRAM 12 into the low-resistance state and storing the weight 1); for the write process of the high-resistance state (RESET) of the RRAM (i.e., editing the first RRAM 12 into the high-resistance state and storing the weight 0), by setting the BL voltage to VDD / 2 and the SL voltage of this path to VDD, and sequentially editing each row of the first RRAMs 12 through the controller. The read path uses a negative feedback circuit to achieve high-precision readout. By setting the BL voltage and the Vclp voltage, the resistance of the second RRAM 32 is converted into a voltage output form. The resistance of the second RRAM 32 in the negative feedback circuit and the first RRAM 12 in the RRAM in-memory computing unit 10 are RRAMs with the same process to reduce the read error problem caused by the negative feedback resistor error.
[0024] In an optional embodiment of the present invention, the structures of the first RRAM 12 and the second RRAM 32 are the same.
[0025] Based on the same inventive concept, please refer to Figure 2 , Figure 2It is a flowchart of a high-density RRAM sparse digital in-memory computing method provided by an embodiment of the present invention. The present invention also provides a high-density RRAM sparse digital in-memory computing method, which is applied to the circuit provided by the above embodiment of the present invention. For the circuit embodiment, please refer to the above and will not be elaborated here. The method includes: S101. Write the resistance state of the RRAM in-memory computing unit 10 through the write path of the peripheral circuit unit 30; wherein, the preset binarized weight matrix is used as the basis for writing the resistance state of the RRAM in-memory computing unit 10. The preset binarized weight matrix is obtained according to the trained neural network, and there is at most one non-zero weight in each row of the preset binarized weight matrix; it can be understood that there is at most one non-zero weight in each row, or there is no non-zero weight in each row; S102. Set the word line WL voltage and bit line BL voltage input to the RRAM in-memory computing unit 10; wherein, the preset eigenvalue is used as the basis for the bit line BL voltage input to the RRAM in-memory computing unit 10, and the preset eigenvalue is obtained according to the trained neural network; S103. The RRAM in-memory computing unit 10 performs digital calculations according to the received resistance state, word line WL voltage, and bit line BL voltage to obtain an output result; S104. The digital logic unit 20 performs multiply-accumulate calculations on the output result to obtain a multiply-accumulate result.
[0026] In an alternative embodiment of the present invention, writing the resistance state of the RRAM in-memory computing unit 10 through the write path of the peripheral circuit unit 30 includes: Drive the peripheral circuit unit 30 to turn on the write path and write the resistance states of a column of RRAM in-memory computing units; wherein, When the bit line BL voltage received by the RRAM in-memory computing unit 10 is VDD / 2 and the source line SL voltage is 0, the resistance state of the RRAM in-memory computing unit 10 is written as a low resistance state, and the storage logic is 1; when the bit line BL voltage received by the RRAM in-memory computing unit 10 is VDD / 2 and the source line SL voltage is VDD, the resistance state of the RRAM in-memory computing unit 10 is written as a high resistance state, and the storage logic is 0.
[0027] In an alternative embodiment of the present invention, when the resistance state of the RRAM in-memory computing unit 10 is written as a low resistance state, the first node N1 is at a high voltage, and when the resistance state of the RRAM in-memory computing unit 10 is written as a high resistance state, the first node N1 is at a low voltage.
[0028] In an alternative embodiment of the present invention, it further includes: Drive the peripheral circuit unit 30, turn on the read path, and read the resistance state of one row of the RRAM in-memory computing unit 10; among them, Through the read path of the peripheral circuit unit 30, set the bit line BL voltage received by the RRAM in-memory computing unit 10 and the Vclp voltage input to one input terminal of the operational amplifier 33, and read the resistance state of the RRAM in-memory computing unit 10.
[0029] In an alternative embodiment of the present invention, the digital logic unit 20 performs multiply-accumulate calculation on the output result to obtain a multiply-accumulate result, including: Multiply the preset binary weight matrix by the preset eigenvalue to obtain an output result in matrix form; Input the output voltages of the same row of the output result in matrix form into the same-level OR gate to obtain the calculation result of one row, and input the calculation results of all rows into the full adder for addition to obtain a calculation result in matrix form; Input multiple calculation results in matrix form into the 2b adder for addition to obtain multiple calculation results in matrix form, and accumulate them into the 3b adder tree for accumulation, and so on, to obtain the multiply-accumulate result.
[0030] Specifically, please refer to Figure 3 , Figure 3 is a schematic diagram of the calculation of the digital logic unit provided by the embodiment of the present invention. Based on the sparse in-memory computing architecture replaced by logic gates, the full adder in the conventional adder tree is replaced by an OR gate. In the prior art, as Figure 3 the left part in, without considering the sparsity of the neural network, for the first row 1-bit value of the 3 3 convolution results, such as Figure 3 X0, X1, X2 in, all need a full adder (FA) for addition and output a 2-bit result, and at the same time transmit this result to the subsequent adder for calculation, while the standard one-bit full adder circuit requires 28 transistors. In the sparse digital in-memory computing method proposed in this embodiment, by training the neural network convolution kernel to have at most one non-zero weight in each row, as Figure 3 shown on the right in, replacing the full adder with a three-input OR gate will not cause loss of calculation accuracy, while reducing the number of MOS transistors from 28T to 8T, and further reducing the number of adders in the subsequent calculation, effectively reducing the hardware overhead. This part only takes the 3 3 convolution kernel and setting at most one non-zero weight in each row as an example for illustration. In practical applications, the sparsity of the convolution kernel and the number of inputs of the OR gate can be determined according to requirements.
[0031] An alternative solution, as Figure 3 shown on the right in, is as follows: 1. The nine-bit weights of the convolution kernel are "1, 0, 0", "0, 1, 0", and "0, 0, 0" respectively, and the nine-bit inputs are "1, 1, 1", "0, 0, 0", and "1, 1, 1" respectively. Then the product results X0 to X8 are "1, 0, 0, 0, 0, 0, 0, 0, 0". 2. X0 to X2 are input to the first OR gate, and the output is 1. X3 to X5 are input to the second OR gate, and the output is 0. X6 to X8 are input to the third OR gate, and the output is 0. 3. The three outputs in the previous step are input to a full adder, and the output is two-bit data 0 and 1. Subsequently, the result is accumulated into the adder tree.
[0032] In this embodiment, innovatively, the full adder in the traditional adder tree is replaced by a logic gate (such as an "OR gate"). By training the neural network convolution kernel to have at most one non-zero weight in each row, the use of multiple full adders is avoided, thereby reducing the number of MOS transistors from 28T to 8T, effectively reducing the hardware overhead, lowering the power consumption, and improving the utilization rate of hardware resources without affecting the calculation accuracy.
[0033] In addition, this embodiment optimizes the calculation based on sparsity. Through the logic gate and adder tree module, the redundant calculation in the sparse neural network calculation is significantly reduced, and the calculation efficiency is improved. Without considering the sparsity of the neural network, the structure of using a logic gate instead of an adder can reduce the hardware resources in the calculation, and at the same time effectively reduce the energy consumption of the system.
[0034] In an alternative embodiment of the present invention, please refer to Figure 4 , Figure 4 which is a schematic diagram of the weight mask fine-tuning training provided by the embodiment of the present invention. The training process of the trained neural network includes: Adopting a preset floating-point training to obtain a floating-point model; according to the floating-point model, adopting a preset pruning training to obtain a sparse model; according to the sparse model, adopting a preset quantization training to obtain a quantization model, and performing weight mask fine-tuning training on the quantization model to obtain a fine-tuning model; Performing weight mask fine-tuning training on the quantization model to obtain a fine-tuning model, including: Recording the position of the maximum weight in each row of the neural network convolution kernel and generating a weight mask , , , and respectively represent the weights in each row; among them, the mask value at the position of the maximum weight in each row of the neural network convolution kernel is set to 1, and the mask values at the remaining positions in each row are set to 0; Applying the weight mask to the weights to obtain updated weights , , denotes the weight; the updated weight retains the maximum value of each row of the neural network convolution kernel, and the rest are all 0. At the same time, the updated weight is used for convolution calculation, denoted as , denotes the current convolution calculation result, denotes the previous convolution calculation result; Through backpropagation, the error is calculated and the weight is updated according to the learning rate , , denotes the learning rate, denotes the error. By iterating in this way, a fine-tuned model is obtained.
[0035] Specifically, in this embodiment, by recording the position of the maximum weight of each row of the convolution kernel and generating a weight mask, while maintaining the accuracy, the hardware acceleration efficiency of the sparse neural network is further improved through fine-tuning training. This method combines the advantages of quantization training and pruning training, providing an effective solution for the implementation of sparse neural networks on hardware. It can be understood that aiming at the computational energy efficiency problem caused by the contradiction between the lightweight sparse network and the utilization rate of digital in-memory computing hardware resources, a sparse digital in-memory computing architecture based on logic gate replacement and a sparse network training method based on weight mask fine-tuning training are proposed to realize a digital in-memory computing architecture with software and hardware co-optimization, solve the challenges of large power consumption and area overhead, and meet the accuracy requirements of neural networks.
[0036] In summary, a high-density RRAM sparse digital in-memory computing circuit provided by the present invention enables the RRAM array to realize parallel read and write functions by designing a high-density RRAM digital in-memory computing array, solving the contradiction between the high-density RRAM array and the complex read circuit, and significantly improving the read and write efficiency. Secondly, the present invention adopts algorithm-hardware co-design, effectively reducing the hardware overhead and optimizing the computational efficiency of the sparse neural network by using logic gates to replace traditional adders. On this basis, the present invention also proposes a sparse neural network training method, which further improves the utilization rate of hardware resources, reduces power consumption and area overhead through the fine-tuning of weight masks, so as to meet the requirements of high energy efficiency and high performance.
[0037] Through the above innovations, the present invention not only solves the problems of read circuit complexity and resource waste in the prior art, but also provides efficient hardware support for the calculation of sparse neural networks, with broad application prospects.
[0038] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variants are intended to cover non-exclusive inclusion, so that an article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the article or device comprising the element. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The orientation or positional relationship indicated by "above", "below", "left", "right", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on the present invention.
[0039] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0040] The above content is a further detailed description of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A high-density RRAM sparse digital in-memory computing circuit, characterized in that, Including: A plurality of RRAM in-memory computing units arranged in an array, where the RRAM in-memory computing units are configured to store data through resistance states, perform digital calculations, and obtain output results; A plurality of digital logic units arranged in an array, where the digital logic units are electrically connected to the RRAM in-memory computing units, and the digital logic units are configured to perform multiply-accumulation on the output results to obtain a multiply-accumulation result; among them, one of the digital logic units is correspondingly provided for a plurality of the RRAM in-memory computing units in the same column of RRAM in-memory computing units; A peripheral circuit unit, where the peripheral circuit unit is electrically connected to the RRAM in-memory computing units, and the peripheral circuit unit includes a read path and a write path, and is configured to read and write the resistance states of the RRAM in-memory computing units simultaneously; among them, one of the peripheral circuit units is correspondingly provided for one column of RRAM in-memory computing units.
2. The high-density RRAM sparse digital in-memory computing circuit according to claim 1, wherein The RRAM in-memory computing unit includes an NMOS transistor, a first RRAM, and a buffer; where The gate of the NMOS transistor is electrically connected to the word line WL, the drain of the NMOS transistor is electrically connected to the peripheral circuit unit, the source of the NMOS transistor is electrically connected to a first node, the first end of the first RRAM is electrically connected to the first node, the second end of the first RRAM is electrically connected to the bit line BL, the input end of the buffer is electrically connected to the first node, and the output end of the buffer is electrically connected to the digital logic unit.
3. The high-density RRAM sparse digital in-memory computing circuit according to claim 2, wherein The peripheral circuit unit includes a single-pole double-throw switch, a second RRAM, and an operational amplifier; where When the common terminal of the single-pole double-throw switch is turned on with the first end, a write operation is performed on the resistance state of the RRAM in-memory computing unit; When the common terminal of the single-pole double-throw switch is turned on with the second end, the second end of the single-pole double-throw switch is electrically connected to a second node, the first end of the second RRAM is electrically connected to the second node, the second end of the second RRAM is electrically connected to the output end of the operational amplifier, one input end of the operational amplifier is connected to the second node, and the other input end of the operational amplifier receives a Vclp voltage signal to perform a read operation on the resistance state of the RRAM in-memory computing unit.
4. The high-density RRAM sparse digital in-memory computing circuit according to claim 3, wherein The first RRAM and the second RRAM have the same structure.
5. A high-density RRAM sparse digital in-memory computing method, characterized in that, Including: Writing the resistance state of the RRAM in-memory computing unit through the write path of the peripheral circuit unit; where a preset binarized weight matrix is used as the basis for writing the resistance state of the RRAM in-memory computing unit, the preset binarized weight matrix is obtained according to a trained neural network, and there is at most one non-zero weight in each row of the preset binarized weight matrix; Setting the word line WL voltage and the bit line BL voltage input to the RRAM in-memory computing unit; where a preset eigenvalue is used as the basis for the bit line BL voltage input to the RRAM in-memory computing unit, and the preset eigenvalue is obtained according to a trained neural network; The RRAM in-memory computing unit performs digital calculations based on the received resistance state, word line WL voltage, and bit line BL voltage to obtain an output result; The digital logic unit performs multiply-accumulate calculations on the output result to obtain a multiply-accumulate result.
6. The high-density RRAM sparse digital in-memory computing method according to claim 5, wherein, The resistance state written into the RRAM in-memory computing unit through the write path of the peripheral circuit unit includes: Driving the peripheral circuit unit to turn on the write path and write the resistance states of a column of the RRAM in-memory computing units; where When the bit line BL voltage received by the RRAM in-memory computing unit is VDD / 2 and the source line SL voltage is 0, the resistance state of the RRAM in-memory computing unit is written as a low resistance state, and the stored logic is 1; when the bit line BL voltage received by the RRAM in-memory computing unit is VDD / 2 and the source line SL voltage is VDD, the resistance state of the RRAM in-memory computing unit is written as a high resistance state, and the stored logic is 0.
7. The high-density RRAM sparse digital in-memory computing method according to claim 6, characterized in that When the resistance state of the RRAM in-memory computing unit is written as a low resistance state, the first node is at a high voltage, and when the resistance state of the RRAM in-memory computing unit is written as a high resistance state, the first node is at a low voltage.
8. The high-density RRAM sparse digital in-memory computing method according to claim 5, characterized in that, It also includes: Driving the peripheral circuit unit to turn on the read path and read the resistance states of a row of the RRAM in-memory computing units; where Through the read path of the peripheral circuit unit, the bit line BL voltage received by the RRAM in-memory computing unit and the Vclp voltage input to one input terminal of the operational amplifier are set to read the resistance state of the RRAM in-memory computing unit.
9. The high-density RRAM sparse digital in-memory computing method according to claim 5, wherein The digital logic unit performs multiply-accumulate calculations on the output result to obtain a multiply-accumulate result, including: Multiplying the preset binary weight matrix by the preset eigenvalue to obtain an output result in matrix form; Inputting the output voltages of the same row of the output result in matrix form to the same-level OR gate to obtain the calculation result of one row, and inputting the calculation results of all rows to the full adder for addition to obtain a calculation result in matrix form; Inputting multiple calculation results in matrix form to the 2b adder for addition to obtain multiple calculation results in matrix form, and accumulating them to the 3b adder tree for accumulation, and so on, to obtain the multiply-accumulate result.
10. The high-density RRAM sparse digital in-memory computing method according to claim 5, wherein, The training process of the trained neural network includes: Adopting preset floating-point training to obtain a floating-point model; according to the floating-point model, adopting preset pruning training to obtain a sparse model; according to the sparse model, adopting preset quantization training to obtain a quantization model, and performing weight mask fine-tuning training on the quantization model to obtain a fine-tuning model; The performing weight mask fine-tuning training on the quantization model to obtain a fine-tuning model includes: Recording the positions of the maximum weights in each row of the neural network convolution kernel and generating a weight mask; where the mask value at the position of the maximum weight in each row of the neural network convolution kernel is set to 1, and the mask values at the remaining positions in each row are set to 0; Applying the weight mask to the weights to obtain updated weights; the updated weights retain the maximum value in each row of the neural network convolution kernel, and the rest are 0. At the same time, the updated weights are used for convolution calculations; Through backpropagation, the error is calculated and the weights are updated according to the learning rate. Through such iteration, a fine-tuned model is obtained.
Citation Information
Patent Citations
Circuit for parallel multiply-accumulate operation in binary neural network formed based on RRAM array
CN114254743A
In-memory computing circuit based on reusable Booth multiplication unit
CN116959517A
Boolean logic in-memory operational circuit based on 2T-2C ferroelectric storage unit
CN117894350A
Storage and calculation integrated calculation system and method supporting deep convolution channel full parallel calculation and storage and calculation integrated chip
CN118364883A
Two-Bit Memory Cell and Circuit Structure Calculated in Memory Thereof
US20220005525A1
Cited By
Memory, memory calculation method and electronic equipment
CN120932695A