Circuit device and operation method for implementing vector multiplication operation based on 1t1r

By using a 1T1R circuit architecture and multi-row parallel operation design, multi-bit vector multiplication is achieved using a single-bit memristor unit, which solves the problems of high production cost of memristor units and high demand for ADC/DAC, and realizes efficient multi-bit vector multiplication operation.

CN116127257BActive Publication Date: 2026-02-10INNOSTAR SEMICON (SHANGHAI) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210311144.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-28
Publication Date
2026-02-10
Estimated Expiration
2042-03-28

AI Technical Summary

Technical Problem

Existing memristors require precise control of memristor cells when performing vector matrix multiplication of multi-bit data, resulting in high production costs and high demands on ADC/DAC, which increases the overhead of the computing architecture.

Method used

The circuit architecture based on 1T1R is adopted. Through the design of 1T1R computing array and input circuit, multi-bit vector multiplication operation is realized by using single-bit memristor unit, which reduces the polymorphic requirements of memristor unit. Furthermore, by multi-row parallel operation and bit-by-bit splitting of input signal, the accuracy requirements of ADC circuit and the demand of DAC circuit are reduced.

Benefits of technology

It realizes multi-bit vector multiplication, reduces the production cost of memristor units and the overhead of ADC circuits, and improves the operation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116127257B_ABST
    Figure CN116127257B_ABST
Patent Text Reader

Abstract

The application provides a circuit framework for realizing vector multiplication operation based on 1T1R, comprising an input circuit, a 1T1R calculation array and an output circuit, wherein the 1T1R calculation array comprises at least two rows of operation units, the input ends of each row of operation units are connected to the input ends of the 1T1R calculation array and the input circuit after being connected to the same node, and the output ends of each row of operation units are connected to the output circuit; and each row of operation units comprises at least two parallel 1T1R memistor units, the input circuit is used for inputting input signals to each operation unit; the input signals and the 1T1R memistor units in each row of operation units are used for realizing vector multiplication operation and generating output signals; and the output signals are stored in the output circuit. The vector multiplication operation based on 1T1R provided by the application is used to solve the problem that a single memistor cannot realize multi-bit operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of memristor data storage design technology, and more specifically, to a circuit device based on 1T1R to realize vector multiplication operation. Background Technology

[0002] A memristor (ReRAM) is a non-linear resistor with memory function. Its resistance can be changed by controlling the current. If a high resistance value is defined as "1" and a low resistance value as "0", then this resistor can store data. Specifically, in practical applications, convolution operations are typically performed on the memristor to store different types of data. (See attached image.) Figure 1 As shown, taking the conductance G (the reciprocal of resistance R, R = 1 / G) of the memristor as the convolution weight and the voltage as the input data, then the current I = V * G for each parallel memristor circuit. If a series of memristors are connected in parallel, then the corresponding current output is the sum of all the weights and the data: I = I1 + I2 + I3 = V1 * G1 + V2 * G2 + V3 * G3.

[0003] Since memristors themselves store data using resistance, therefore, according to the appendix... Figure 1 As can be seen from the principle, the memristor is an excellent device for performing convolution operations using resistance. Existing operational architectures using memristors as convolution kernels are shown in the attached figure. Figure 2 As shown, by appendix Figure 2 It is evident that this in-memory computing architecture can improve computational performance and reduce power consumption. However, in practical applications, to meet storage requirements, this computing architecture places demands on the memristor cells themselves. For example, to achieve multi-bit multiplication, a single memristor cell needs to support multiple bits. To enable a single memristor cell to support multiple bits, the conductance value of each memristor needs to be precisely controlled individually, thereby achieving the goal of setting arbitrary weights for each memristor. Furthermore, even after quantization, the weights of the memristors typically require 6 to 8 bits. Therefore, precise control of the memristors during mass production consumes a significant amount of time, severely increasing the overall chip manufacturing cost.

[0004] It should be noted that implementing 1-bit (0, 1) data using memristors is the simplest method, but how to quickly reduce costs and complete vector matrix multiplication of multi-bit data using only 1-bit memristors is a problem that needs to be solved.

[0005] In addition, by appendix Figure 2It can be known that in the existing operation architecture using the memristor as the convolution kernel, the demand for ADC / DAC (Analog Digital Converter / Digital Analog Convert) is very high, and theoretically, each row needs a DAC and each column needs an ADC, which is a great cost for the operation architecture.

[0006] Based on the above technical problems, there is an urgent need for a method that can efficiently complete multi-bit operation and reduce the cost of ADC / DAC used by the operation unit. SUMMARY

[0007] In view of the above problems, the purpose of the present application is to provide a circuit architecture and operation method for implementing vector multiplication operation based on 1T1R, to solve the problem that a single memristor cannot realize multi-bit operation.

[0008] The circuit architecture for implementing vector multiplication operation based on 1T1R provided by the present application comprises an input circuit, a 1T1R calculation array and an output circuit, wherein,

[0009] The 1T1R calculation array comprises at least two rows of operation units, the input ends of each row of operation units are connected to the input ends of the 1T1R calculation array after being connected to the input circuit, and the output ends of each row of operation units are connected to the output circuit; and,

[0010] Each row of operation units comprises at least two 1T1R memristor units connected in parallel, the input circuit is used to input an input signal to each operation unit; the input signal and the 1T1R memristor units in each row of operation units interact to realize vector multiplication operation and generate an output signal; and the output signal is stored in the output circuit.

[0011] In addition, preferably, the vector multiplication operation to be implemented is the multiplication operation of two w-bit vectors, the 1T1R calculation array comprises k rows of operation units, each operation unit comprises m 1T1R memristor units connected in parallel; and,

[0012] Each parameter in the 1T1R calculation array satisfies the following constraint condition:

[0013] k*log2m=w.

[0014] In addition, preferably, the input circuit comprises a first shift register and a first control unit connected to the first shift register; wherein,

[0015] The input circuit splits the w-bit input signal into w analog input signals by the first shift register and the first control unit, and inputs the w analog input signals into the 1T1R computing array through w clock cycles; and

[0016] The input circuit inputs the analog input signals into the 1T1R computing array from low bits to high bits, one analog input signal per clock cycle; the analog input signals interact with the 1T1R memristor units in each row of operation units to realize vector multiplication operation and generate analog output signals; the analog output signals are converted into digital output signals by the output circuit and stored.

[0017] In addition, preferably, the output circuit comprises k rows of output units, the input end of each row of output units is connected to the output end of the corresponding row of operation units, and the digital output signals output by the output end of each row of output units are accumulated by a preset result adder to form a final digital output signal and stored in a preset final output register.

[0018] In addition, preferably, each row of output units comprises an ADC circuit and a circular shift adder, wherein,

[0019] The ADC circuit is configured to convert the analog output signals output by the corresponding row of operation units into the digital output signals;

[0020] The circular shift adder is configured to circularly shift and accumulate the digital output signals of w clock cycles to form a single-row accumulated digital output signal;

[0021] The single-row accumulated digital output signals of each row are accumulated by the result adder to form the final digital output signal and stored in the final output register.

[0022] In addition, preferably, the 1T1R memristor unit comprises a 1-bit memristor and a transistor connected in series, and the transistor is connected to a preset second control unit; wherein,

[0023] The second control unit controls the conduction and cutoff of the 1-bit memristor through the transistor.

[0024] In another aspect, the present application also provides a circuit architecture for implementing vector matrix multiplication operation based on 1T1R, comprising the circuit architecture for implementing vector multiplication operation based on 1T1R as described above; wherein,

[0025] The 1T1R computing array is provided with L, the input end of each 1T1R computing array is individually connected to the corresponding input circuit, and the output ends of the corresponding row of operation units of each 1T1R computing array are connected to the output circuit after being connected to a common node; and,

[0026] Suppose the vector matrix multiplication operation to be implemented is the multiplication operation of two n*n vector matrices; then, L = n*n.

[0027] Correspondingly, the application also provides a w bit vector multiplication operation method, which utilizes the circuit framework for implementing vector multiplication operation based on 1T1R to implement w bit vector multiplication operation, and the w bit vector multiplication operation method comprises the following steps of:

[0028] The input circuit splits the w bit input signal into w analog input signals by bit and inputs them into the 1T1R calculation array through w clock cycles; wherein the input circuit inputs the analog input signals into the 1T1R calculation array from low bit to high bit, and inputs one analog input signal in each clock cycle;

[0029] The analog input signal of each bit and the 1T1R memistor unit in each row operation unit of the 1T1R calculation array interact to implement vector multiplication operation of each bit, and generate analog output signals of each bit;

[0030] The analog output signals of each bit are converted into digital output signals of each bit by the ADC circuit of the output circuit;

[0031] After the digital output signals of each bit are cyclically shifted and accumulated by the cyclic shift adder, the single-row accumulated digital output signals are formed;

[0032] After the single-row accumulated digital output signals of each row are shifted and accumulated by the result adder, the final digital output signals are formed to implement w bit vector multiplication operation.

[0033] Correspondingly, the application also provides a w bit vector matrix multiplication operation method, which utilizes the circuit framework for implementing vector matrix multiplication operation based on 1T1R to implement w bit vector matrix multiplication operation, and the w bit vector matrix multiplication operation method comprises the following steps of:

[0034] Each input circuit splits the corresponding input signal in the input vector matrix into w analog input signals and inputs them into the corresponding 1T1R calculation array through w clock cycles; wherein each input circuit inputs the analog input signals into the corresponding 1T1R calculation array from low bit to high bit, and inputs one analog input signal in each clock cycle;

[0035] In the same 1T1R computing array, the analog input signal of each bit interacts with the 1T1R memristor unit in each row operation unit in the 1T1R computing array to realize the vector multiplication operation of each bit, and generates the analog output signal of each bit in each 1T1R computing array; and the analog output signals of each bit of the corresponding row operation units in all 1T1R computing arrays are added bit by bit to form the accumulated analog output signal of each bit;

[0036] The accumulated analog output signal of each bit is converted into a digital output signal of each bit by the ADC circuit of the output circuit;

[0037] After the digital output signal of each bit is cyclically shifted and accumulated by the cyclic shift adder, the single-row accumulated digital output signal is formed;

[0038] After the single-row accumulated digital output signal of each row is shifted and accumulated by the result adder, the final digital output signal is formed to realize the w-bit vector multiplication operation.

[0039] Compared with the prior art, the circuit architecture for realizing vector multiplication operation based on 1T1R according to the present application has the following beneficial effects:

[0040] The circuit architecture for realizing vector multiplication operation based on 1T1R and the operation method provided by the present application realize the multiplication and addition operation of multi-bit vector by using a single-bit memristor unit, which can reduce the multi-state requirement for the memristor unit; in addition, by using the multi-row parallel operation of the 1T1R computing array and the bit-by-bit splitting method for the input signal, the requirement for the precision of the ADC circuit can be reduced, and the requirement for the DAC circuit can be reduced to 0, thereby significantly reducing the overhead of the ADC circuit and the DAC circuit.

[0041] To achieve the above and related purposes, one or more aspects of the present application include features that will be explained in detail later and particularly pointed out in the claims. The following description and drawings detail certain illustrative aspects of the present application. However, these aspects indicate only some of the various ways in which the principles of the present application can be employed. In addition, the present application is intended to include all such aspects and their equivalents. BRIEF DESCRIPTION OF DRAWINGS

[0042] Other objects and results of the present application will become more apparent and easily understood by referring to the following description and claims in conjunction with the accompanying drawings, and with the more complete understanding of the present application. In the drawings:

[0043] Figure 1 The schematic diagram for multiplication and addition operation using existing memristor according to the embodiment of the present application;

[0044] Figure 2 A schematic diagram of vector matrix multiplication using existing memristors according to an embodiment of the present application;

[0045] Figure 3 A block diagram of an input circuit according to an embodiment of the present application;

[0046] Figure 4 A block diagram of implementing a first vector matrix multiplication operation based on 1T1R according to an embodiment of the present application;

[0047] Figure 5 A schematic diagram of an output circuit for implementing a first vector matrix multiplication operation based on 1T1R according to an embodiment of the present application;

[0048] Figure 6 A block diagram of implementing a second vector matrix multiplication operation based on 1T1R according to an embodiment of the present application;

[0049] Figure 7 A schematic diagram of an output circuit for implementing a second vector matrix multiplication operation based on 1T1R according to an embodiment of the present application;

[0050] The same reference numbers in all the drawings indicate similar or corresponding features or functions. DETAILED DESCRIPTION

[0051] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of one or more embodiments. It can be evident, however, that embodiments can be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form in order to facilitate describing one or more embodiments.

[0052] It needs to be explained in advance that since the circuit architecture for implementing vector multiplication operation based on 1T1R provided by the present application performs matrix convolution operation according to the circuit architecture formed by the memristor, before introducing the specific structure of the circuit architecture for implementing vector multiplication operation based on 1T1R provided by the present application, the principle of matrix convolution operation needs to be introduced first, and the principle of matrix convolution operation is as follows (taking 3*3 matrix and 2*2 matrix convolution as an example):

[0053]

[0054] It can be seen from the above formula that the convolution operation is mainly multiplication and addition operation. If the multiplication matrix size is inconsistent, there will be a shift operation, but the most basic operation is multiplication and addition, and each element of the matrix after convolution is the multiplication and addition result. Therefore, the multiplier is the core in the convolution operation, and therefore, the improvement and optimization of the multiplier directly determines the performance and overhead of the overall circuit architecture. Therefore, the most important problem actually solved by the present application is how to use a 1-bit memristor unit to realize a multi-bit multiplier.

[0055] The circuit architecture for implementing vector multiplication operation based on 1T1R and the operation method thereof provided by the present application will be described in detail below in combination with the accompanying drawings.

[0056] Figure 3 a framework of an input circuit according to an embodiment of the present application is shown, Figure 4 a framework for implementing first vector matrix multiplication operation based on 1T1R according to an embodiment of the present application is shown, Figure 5 a schematic diagram of an output circuit for implementing first vector matrix multiplication operation based on 1T1R according to an embodiment of the present application, Figure 6 a framework for implementing second vector matrix multiplication operation based on 1T1R according to an embodiment of the present application is shown, Figure 7 a schematic diagram of an output circuit for implementing second vector matrix multiplication operation based on 1T1R according to an embodiment of the present application.

[0057] In combination Figures 3 to 7 It is shown that the circuit architecture for implementing vector multiplication operation based on 1T1R provided by the present application includes an input circuit, a 1T1R calculation array and an output circuit, wherein the input circuit is used to provide a first vector data (usually voltage type data), the 1T1R calculation array (usually conductance type data) is used to represent a second vector data, when the first vector data provided by the input circuit is input into the 1T1R calculation array, the multiplication operation of the first vector data and the second vector data (the product of voltage type data and conductance type data) can be realized, and the third vector data (multiplication operation result, usually current type data) is generated and output and saved through the output circuit.

[0058] Specifically, the 1T1R calculation array includes a plurality of rows (at least two rows) of operation units, the input ends of the operation units in each row are connected to the input end of the 1T1R calculation array and the input circuit after being connected in common node, and the output ends of the operation units in each row are connected to the output circuit; and each row of operation units includes a plurality of rows (at least two rows) of 1T1R memristor units connected in parallel, the input circuit is used to input an input signal (i.e. the first vector data) to each operation unit; the input signal and the 1T1R memristor units in each row of operation units interact to realize vector multiplication operation and generate an output signal (i.e. the third vector data); and the output signal is stored in the output circuit.

[0059] It should be noted that the circuit architecture for vector multiplication based on 1T1R provided by this invention can realize the multiplication of any two vector data with the same number of bits (binary, represented by bits). Let the vector multiplication operation to be implemented be the multiplication of two w-bit vectors. The 1T1R computing array includes k rows of operational units, and each operational unit includes m 1T1R memristor units connected in parallel.

[0060] The parameters in the 1T1R computation array satisfy the following constraint: k*log2m=w.

[0061] Specifically, by Figure 3 It is known that the input circuit includes a first shift register and a first control unit connected to the first shift register; wherein, the input circuit, through the first shift register and the first control unit, splits the w-bit input signal into w analog input signals bit by bit, and inputs them to the 1T1R computing array through w clock cycles; furthermore, the input circuit inputs analog input signals to the 1T1R computing array from the least significant bit to the most significant bit, inputting one analog input signal per clock cycle; the analog input signals interact with the 1T1R memristor units in each row of the arithmetic unit to perform vector multiplication operations and generate analog output signals; the analog output signals are converted into digital output signals by the output circuit and stored.

[0062] By using a first shift register and a first control unit connected to the first shift register, a digital input signal can be split bit by bit into w analog input signals and input to the 1T1R computing array through w clock cycles; thus realizing the conversion between digital and analog signals, eliminating the need for a DAC module and significantly reducing overhead.

[0063] Specifically, the output circuit includes k rows of output units. The input terminal of each row of output units is connected to the output terminal of the corresponding row's arithmetic unit. The digital output signals output by the output terminals of each row of output units are accumulated by a preset result adder to form the final digital output signal and stored in a preset final output register.

[0064] More specifically, each row output unit includes an ADC circuit and a cyclic shift adder, wherein,

[0065] The ADC circuit is used to convert the analog output signal output by the corresponding row's arithmetic unit into a digital output signal; the cyclic shift adder is used to cyclically shift and accumulate the digital output signal for w clock cycles to form a single-row accumulated digital output signal; the single-row accumulated digital output signals of each row are accumulated by the result adder to form the final digital output signal and stored in the final output register.

[0066] By employing multi-row parallel operation in the 1T1R computing array, only one ADC circuit is used per row, which effectively reduces the lifespan of the ADC and the requirements for ADC accuracy, further reducing overhead.

[0067] Specifically, the 1T1R memristor unit includes a 1-bit memristor and a transistor connected in series, with the transistor connected to a preset second control unit. The second control unit controls the on and off states of the 1-bit memristor through the transistor, thereby achieving the switching between the high-resistance state and the low-resistance state of the 1-bit memristor. Furthermore, the conductance values ​​of all the 1-bit memristors in the 1T1R memristor unit together constitute the representation of the second vector data.

[0068] Obviously, the circuit architecture for vector multiplication based on 1T1R provided by this invention can realize the operation logic and method of multi-bit vector operation by using a 1-bit memristor (single-bit memristor unit), which reduces the polymorphic requirements of the memristor.

[0069] To illustrate the working principle of the circuit architecture for vector multiplication based on 1T1R provided by this invention, this invention also provides a method for multiplying w-bit vectors. This method utilizes the aforementioned circuit architecture for vector multiplication based on 1T1R to perform w-bit vector multiplication. The w-bit vector multiplication method includes:

[0070] The input circuit splits the w-bit input signal into w analog input signals, and inputs them to the 1T1R computing array over w clock cycles. The input circuit inputs analog input signals to the 1T1R computing array from the least significant bit to the most significant bit, with one analog input signal input per clock cycle.

[0071] The analog input signal of each bit interacts with the 1T1R memristor unit in each row of the 1T1R computing array to realize the vector multiplication operation of each bit and generate the analog output signal of each bit.

[0072] The analog output signal of each bit is converted into a digital output signal of each bit by the ADC circuit of the output circuit;

[0073] Each bit of the digital output signal is cyclically shifted and accumulated by a cyclic shift adder to form a single-line accumulated digital output signal;

[0074] The single-row accumulated digital output signals of each row are shifted and accumulated by the result adder to form the final digital output signal, so as to realize the vector multiplication operation of w bits.

[0075] It should be noted that the circuit architecture for vector multiplication based on 1T1R provided by the present invention can only implement the multiplication of two vectors with the same number of bits. It cannot implement the multiplication of two vector matrices with the same number of bits.

[0076] To this end, the present invention also provides a circuit architecture for implementing vector matrix multiplication based on 1T1R, including the aforementioned circuit architecture for implementing vector multiplication based on 1T1R; wherein, L 1T1R computing arrays are provided, the input terminal of each 1T1R computing array is individually connected to the corresponding input circuit, and the output terminals of the corresponding rows of the computing units of each 1T1R computing array are connected to the output circuit after sharing a common node; and, assuming that the vector matrix multiplication operation to be implemented is the multiplication operation of two n*n vector matrices; then, L = n*n.

[0077] By increasing to n*n 1T1R computation arrays, and connecting the input terminals of each 1T1R computation array to their respective input circuits, and connecting the output terminals of the corresponding rows of the computation units of each 1T1R computation array to the output circuit after sharing a common node, the product of each 1T1R computation array and its corresponding input circuit can represent a multiplication operation. Since the result of the multiplication of an n*n vector matrix is ​​the sum of n*n multiplication operations, it is clear that by setting up n*n 1T1R computation arrays and controlling the input signals and the resistance state of the 1-bit memristor in each 1T1R computation array, the representation of vector matrix multiplication can be achieved.

[0078] Accordingly, this invention also provides a corresponding method for multiplying a w-bit vector matrix. This method utilizes the circuit architecture described above, which implements vector matrix multiplication based on 1T1R, to perform the multiplication of a w-bit vector matrix. The method for multiplying a w-bit vector matrix includes:

[0079] Each input circuit decomposes the corresponding input signal in the input vector matrix into w analog input signals and inputs them to the corresponding 1T1R computing array through w clock cycles; wherein, each input circuit inputs analog input signals to the corresponding 1T1R computing array from the least significant bit to the most significant bit, and inputs one analog input signal per clock cycle;

[0080] Within the same 1T1R computing array, the analog input signal of each bit interacts with the 1T1R memristor unit in each row of the 1T1R computing array to realize the vector multiplication operation of each bit and generate the analog output signal of each bit in each 1T1R computing array; and, the analog output signals of each bit of the corresponding row of the computing units in all the 1T1R computing arrays are added bit by bit to form the accumulated analog output signal of each bit.

[0081] The accumulated analog output signal of each bit is converted into a digital output signal of each bit by the ADC circuit of the output circuit;

[0082] Each bit of the digital output signal is cyclically shifted and accumulated by a cyclic shift adder to form a single-line accumulated digital output signal;

[0083] The single-row accumulated digital output signals of each row are shifted and accumulated by the result adder to form the final digital output signal, so as to realize the vector multiplication operation of w bits.

[0084] To further explain the working principle of the circuit architecture and operation method for vector multiplication based on 1T1R provided by this invention, the following example of the multiplication of two 8-bit vector matrices (matrix size n*n) will be used for further explanation.

[0085] like Figure 4 As shown, the entire circuit architecture can be divided into three parts: the input circuit, the 1T1R computing array, and the output circuit.

[0086] First, we explain the operation principle of the 1T1R computing array:

[0087] exist Figure 4 In the diagram, the central box represents a 1T1R calculation array for 8-bit multiplication, signifying the multiplication of two 8-bit data (one 8-bit data is input by the input circuit, and the 1T1R calculation array itself represents one 8-bit data). Here, the invention uses 16 memristors to represent 4-bit numbers (as one row, i.e., the arithmetic unit, for a total of two rows). The number of low-resistance memristors represents, for example, 7 low-resistance units, which represents a value of 7. The upper and lower rows represent the high 4 bits and low 4 bits of the 8-bit data (the 8-bit data represented by the 1T1R calculation array itself), respectively. The calculation principle is as follows:

[0088] X = A * (B H *2 4 +B L )

[0089] Here, A represents the input, which is fed into the 1T1R computing array through the input circuit. Then, the first row calculates the operation of the high 4 bits, the second row calculates the operation of the low 4 bits, and finally the operation results of the high and low 4 bits are sent to the output circuit. The output circuit then performs shift addition to obtain the final multiplication result.

[0090] exist Figure 4In this context, the entire computation array (the collective term for all 1T1R computation arrays) has two rows, each with multiple 4-bit operation units, the number of which is equal to the number of units in a vector-matrix multiplication. For an n*n matrix, there are n^2 operation units. For example, for a common 3*3 matrix, then... Figure 4 The first row has nine 4-bit operation units, each corresponding to an input. In this way, two rows of operation units can complete vector-matrix multiplication.

[0091] Next, we describe the input circuit. Returning to Equation 1, the input value is A. To reduce reliance on the input DAC module, A is bitwise split, thus Equation 1 can be simplified to:

[0092]

[0093] Thus, there is only one bit at each clock input, as shown in the following equation:

[0094]

[0095] The intermediate 1T1R computing array calculates one bit of data per clock cycle, and then sends the next bit in the next cycle. The i-th clock cycle performs the operation as shown in Equation 2 (where i starts from 0). Thus, the specific structure of this input circuit is as follows: Figure 2 As shown, by Figure 2 It can be seen that by using a shift register (first shift register), 1 bit of data is shifted out to the 1T1R computing array every clock cycle, and then the control (first control unit) will continuously send data out to the first shift register.

[0096] As can be seen from the formulas above, the output circuit requires shifting and addition operations on the output result, and also needs to combine the results from the two paths (two rows). The specific circuit architecture is as follows: Figure 5 As shown, the first line of input represents the calculation result of the high 4 bits in the circuit, and the second line of input represents the calculation result of the low 4 bits in the circuit. Then, the results of each line are calculated. The calculation process is as follows:

[0097] 1) The result enters the ADC circuit, where the analog calculation result is converted into a digital signal. The required accuracy of the ADC is 4 bits + log2n. 2 ((The first part of this equation represents the original accuracy requirement, and the second part is log2n) 2 The accuracy requirement is formed by the accumulation of vector matrix elements.

[0098] 2) After the calculation result is output to the 8-bit register (by default n<=4; if n increases, the register will increase), the calculation result will be processed according to the formula: Perform a shift operation, shifting i bits in the i-th clock cycle.

[0099] 3) The result of the shift will be fed into the adder, and calculated according to the formula: Perform an addition operation on the previous calculation results.

[0100] 4) After 8 clock cycles, all the accumulation calculations for the first and second rows have been completed, that is, the operations of the high 4 bits and the low 4 bits have been completed.

[0101] 5) In the last clock cycle, the calculation results of the high 4 bits are shifted and then fed together with the calculation results of the low 4 bits into the adder (according to the formula). The final output result is obtained.

[0102] This completes all operations for 8-bit vector-matrix multiplication.

[0103] It should be noted that in the above description, this invention uses 16 memristors to implement a 4-bit resistance ladder, and employs two rows to complete 8-bit circuit operations. The row number k, the number of memristors m in a single 1T1R calculation array, and the bit width w of the data to be calculated have the following relationship:

[0104] k * log2m = w

[0105] The number of memristors M in each row is related to the size of the matrix. Assuming the matrix size is n*n, then the number of memristors in each row is:

[0106] M = m * n * n

[0107] Of course, the two companies mentioned above can also be used to calculate vector matrix multiplication of arbitrary bit width and the required number of memristors. For example, if it is a 7*7 32-bit vector matrix multiplication, then n=7 and w=32.

[0108] Specifically, for example, assuming 16 memristors are used to replace 4 bits of value, then m = 16. According to the above formula, k = 8 (a total of 8 rows are needed), and the number of memristors in each row is M = 16 * 7 * 7 = 784. Therefore, the total number of memristors needed is 784 * 8.

[0109] For example, if 4 memristors are used instead of 2 bits, then m = 4. According to the above formula, k = 16 (a total of 16 rows are needed), the number of memristors in each row is M = 2 * 7 * 7 = 98, and the total number of memristors needed is 98 * 16.

[0110] Similarly, when the rows and columns of the memristors are adjusted according to the above formula, the entire circuit structure will change. To further illustrate its structure, we will now implement an 8-bit n*n matrix in another way, replacing the 2-bit values ​​with 4 memristors, so m = 4. According to the above formula, k = 4 (a total of 4 rows are needed), and the number of memristors in each row is M = 4*n*n. The corresponding 1T1R calculation array circuit structure will change as follows: Figure 6 The structure shown remains the same for the input circuit.

[0111] exist Figure 6 In this version, the number of memristors in each small operational unit changes to 4, and the total number of rows increases to 4. Similarly, the output operational circuit also changes, and the changed structure is as follows: Figure 7 As shown, by Figure 7 It can be observed that besides the output becoming 4 lines, the shift of the output numbers has also changed, from... Figure 5 The 4 is transformed into a shifted 6, a shifted 4, a shifted 2, and then the results of the four numbers are added together.

[0112] Therefore, the entire circuit can be adjusted by modifying the number of memristors, and any circuits and methods that adopt this method are within the scope of patent protection of this invention.

[0113] As per the above reference Figures 1 to 7 The circuit architecture for vector multiplication based on 1T1R according to the present invention is described by way of example. However, those skilled in the art should understand that various modifications can be made to the circuit architecture for vector multiplication based on 1T1R proposed in the present invention without departing from the scope of the invention. Therefore, the scope of protection of the present invention should be determined by the contents of the appended claims.

Claims

1. A circuit device for implementing vector multiplication based on 1T1R, characterized in that, It includes input circuitry, a 1T1R computing array, and output circuitry, among which, The 1T1R computing array includes at least two rows of arithmetic units. The input terminals of each row of arithmetic units share a common node and are connected to the input circuit as the input terminal of the 1T1R computing array. The output terminals of each row of arithmetic units are connected to the output circuit. Each row arithmetic unit includes at least two parallel 1T1R memristor units. Assuming the vector multiplication operation to be implemented is the multiplication of two w-bit vectors, the input circuit splits the w-bit input signal bit-by-bit into w analog input signals and inputs them to the 1T1R computing array over w clock cycles. The input circuit inputs the analog input signals to the 1T1R computing array from least significant bit to most significant bit, one analog input signal per clock cycle. The analog input signals interact with the 1T1R memristor units in each row arithmetic unit to perform the vector multiplication operation and generate an analog output signal. The analog output signal is converted into a digital output signal by the output circuit and stored.

2. The circuit device for implementing vector multiplication based on 1T1R as described in claim 1, characterized in that, The 1T1R computing array comprises k rows of arithmetic units, and each arithmetic unit includes m 1T1R memristor units connected in parallel; therefore, The parameters in the 1T1R computing array satisfy the following constraints: k*log2m=w.

3. The circuit device for implementing vector multiplication based on 1T1R as described in claim 2, characterized in that, The input circuit includes a first shift register and a first control unit connected to the first shift register; wherein, The input circuit, in conjunction with the first control unit, splits the w-bit input signal into w analog input signals bit by bit.

4. The circuit device for implementing vector multiplication based on 1T1R as described in claim 3, characterized in that, The output circuit includes k rows of output units. The input terminal of each row of output units is connected to the output terminal of the corresponding row's arithmetic unit. The digital output signals output by the output terminals of each row of output units are accumulated by a preset result adder to form the final digital output signal and stored in a preset final output register.

5. The circuit device for implementing vector multiplication based on 1T1R as described in claim 4, characterized in that, Each row output unit includes an ADC circuit and a cyclic shift adder, wherein, The ADC circuit is used to convert the analog output signal output by the arithmetic unit of the corresponding row into the digital output signal; The cyclic shift adder is used to cyclically shift and accumulate the digital output signal for w clock cycles to form a single-row accumulated digital output signal; The single-row accumulated digital output signals of each row are accumulated by the result adder to form the final digital output signal and stored in the final output register.

6. The circuit device for implementing vector multiplication based on 1T1R as described in claim 5, characterized in that, The 1T1R memristor unit includes a 1-bit memristor and a transistor connected in series, the transistor being connected to a preset second control unit; wherein, The second control unit controls the on and off of the 1-bit memristor through the transistor.

7. A circuit device for implementing vector matrix multiplication based on 1T1R, characterized in that, Includes the circuit device for implementing vector multiplication based on 1T1R as described in claim 5 or 6; wherein, The 1T1R computing array is configured with L units. The input terminal of each 1T1R computing array is individually connected to the corresponding input circuit. The output terminals of the corresponding rows of the arithmetic units of each 1T1R computing array are connected to the output circuit after sharing a common node. Suppose the vector-matrix multiplication operation to be implemented is the multiplication of two n*n vector matrices; then, L=n*n.

8. A method for multiplication of w bit vectors, characterized in that, The multiplication of a w-bit vector is implemented using the circuit device based on 1T1R for vector multiplication as described in claim 5 or 6, wherein the multiplication method of the w-bit vector includes: The input circuit splits the w-bit input signal into w analog input signals bit by bit, and inputs them to the 1T1R computing array over w clock cycles; wherein, the input circuit inputs the analog input signals to the 1T1R computing array from the least significant bit to the most significant bit, and inputs one analog input signal per clock cycle; The analog input signal of each bit interacts with the 1T1R memristor unit in each row of the 1T1R computing array to realize the vector multiplication operation of each bit and generate the analog output signal of each bit. The analog output signal of each bit is converted into a digital output signal of each bit by the ADC circuit of the output circuit; The digital output signal of each bit is cyclically shifted and accumulated by the cyclic shift adder to form the single-row accumulated digital output signal; The single-row accumulated digital output signals of each row are shifted and accumulated by the result adder to form the final digital output signal, so as to realize the vector multiplication operation of w bits.

9. A method for multiplication of a w-bit vector matrix, characterized in that, The circuit device for vector matrix multiplication based on 1T1R as described in claim 7 is used to implement the multiplication operation of a w-bit vector matrix, wherein the multiplication operation method of the w-bit vector matrix includes: Each input circuit decomposes the corresponding input signal in the input vector matrix into w analog input signals and inputs them to the corresponding 1T1R computing array through w clock cycles; wherein, each input circuit inputs the analog input signal to the corresponding 1T1R computing array from the least significant bit to the most significant bit, and inputs one analog input signal per clock cycle; Within the same 1T1R computing array, the analog input signal of each bit interacts with the 1T1R memristor unit in each row of the 1T1R computing array to realize the vector multiplication operation of each bit and generate the analog output signal of each bit in each 1T1R computing array; and, the analog output signals of each bit of the corresponding row of the computing units in all the 1T1R computing arrays are added bit by bit to form the accumulated analog output signal of each bit. The accumulated analog output signal of each bit is converted into a digital output signal of each bit by the ADC circuit of the output circuit. The digital output signal of each bit is cyclically shifted and accumulated by the cyclic shift adder to form the single-row accumulated digital output signal; The single-row accumulated digital output signals of each row are shifted and accumulated by the result adder to form the final digital output signal, so as to realize the vector multiplication operation of w bits.

Citation Information

Patent Citations

  • Matrix vector multiplication circuit and calculation method

    CN110597487A