Vector multiplication circuit for look ahead optimization
By using a vector multiplication circuit optimized for prediction and employing hierarchical weighting and input splitting techniques, the problems of long computation delay and large weighting circuit area in existing technologies are solved, resulting in a significant reduction in computation delay and circuit resources, and improved computational efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2023-03-31
- Publication Date
- 2026-04-24
AI Technical Summary
Existing vector multiplication circuits based on novel non-volatile memories suffer from long computation delays and excessively large weighting circuit areas, especially as the number of bits of weight data increases, leading to frequent memory accesses and increased power consumption.
A vector multiplication circuit for prediction optimization is adopted. By using hierarchical weighting and input splitting, the input data is divided into two parts using a digital time converter. The weighting is performed using a 1T1R memory array and a weighting circuit. Combined with an analog-to-digital converter, the current is hierarchically weighted and the predicted output is performed, which reduces the computation delay and circuit resource consumption.
It effectively reduces computational latency by 80% and weighted circuit area by 85%, improving computational efficiency and circuit resource utilization.
Smart Images

Figure CN116401505B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a multiply-accumulate predictor circuit, specifically to a vector multiplication circuit for predictor optimization. Background Technology
[0002] With the rapid development of artificial intelligence, there is a corresponding need for higher-performance hardware. In AI algorithms represented by neural networks, vector multiplication accounts for more than 95% of the computation. Due to the computationally intensive nature of neural networks, the traditional vector multiplication architecture requires frequent memory accesses, and memory access latency has become the main factor limiting computation speed. In addition, due to the large amount of computation required by neural networks, a large amount of memory is needed, and memory read power consumption also dominates.
[0003] To overcome this bottleneck, researchers have proposed a novel computing architecture that can perform some computations within memory, significantly reducing the number of memory accesses and the latency of memory access during computation. Meanwhile, new non-volatile memories (such as RRAM) proposed by industry offer high integration density, low power consumption, and low latency, making them particularly suitable for high-density reads. Therefore, research on vector multiplication circuits based on novel non-volatile memories is crucial for improving the performance of neural networks.
[0004] The computational data of the circuit is divided into two parts: input and weights. The weights are stored in memory, and the input serves as the control signal for the memory. Weighting between weights of different bit lengths requires a weighting circuit. Currently, several vector multiplication circuits based on novel non-volatile memories still have some problems. The circuit requires digital-to-analog conversion of the input during computation, which causes accuracy and delay issues when processing the input data. Furthermore, as the number of bit lengths of the weight data increases, the area of the weighting circuit increases exponentially, and area consumption of the weighting circuit has become dominant. Summary of the Invention
[0005] This invention discloses a vector multiplication circuit for prediction optimization, which reduces the computation delay of vector multiplication; reduces the area consumption of the weighting circuit through hierarchical weighting; and reduces the response time between the output and input by predicting the output, so that the output can be applied to subsequent calculations more quickly.
[0006] The present invention provides a vector multiplication circuit for prediction optimization, comprising a digital time converter (DTC), a 1T1R memory array, a weighting circuit, and an analog-to-digital converter.
[0007] The DTC is used to convert the input data, which is symmetrically split into two parts, into a fully encoded time signal; the 1T1R memory array is used to receive the time information obtained by the DTC, and to activate the memory according to the received time information to obtain the accumulated current IBL on the nodes of a column of memory cells. j The 1T1R memory array includes multiple memory (1T1R) cells, each of which includes a select transistor and an RRAM. The select transistor controls whether the RRAM generates current, and the select transistor is controlled by an input signal generated by the DTC. The RRAM is a memory whose resistance changes.
[0008] The weighting circuit is used to weight the current generated by the memory. The weighting circuit includes a current mirror, a switching circuit, and a regulator. The current mirror includes a mirrored part and a mirrored part, which are connected by the switching circuit. The two parts of the current mirror complete the first level of weighting between different bit values. The switching circuit realizes the different correspondences between the two parts of the current mirror, and works with the analog-to-digital converter to realize the weighting between the two inputs. The regulator's function is to improve the linearity of the current generated by the 1T1R memory.
[0009] The analog-to-digital converter circuit is used to perform a second-level weighting of the current and generate an analog voltage, and then perform analog-to-digital conversion on the generated analog voltage to obtain the resulting digital output.
[0010] As a preferred embodiment of the present invention, since the input is split into two parts, DTC needs to work twice in the same calculation.
[0011] As a preferred embodiment of the present invention, the RRAM has two states: a high-resistance state and a low-resistance state. The high-resistance state represents a stored value of 0, and the low-resistance state represents a stored value of 1. When the stored value is 0, the RRAM will not generate current. The RRAM will contribute current to the node only when the stored value is 1 and the selector is turned on.
[0012] IBL j =IN0·W 0,j +IN j ·W 1,j +…+IN m ·W m,j ;
[0013] Among them, IBL j (j = 0, 1, 2, 3, 4, 5, 6, 7) represents the sum of the currents generated by the j-th column of memory cells, W i,j (i = 0, 1, 2…m; j = 0, 1, 2, 3, 4, 5, 6, 7) represents the j-th bit of the i-th weight in the weight vector, INi (i = 0, 1, 2, ..., m) represents the i-th data in the input vector.
[0014] As a preferred embodiment of the present invention, the weighting circuit performs a first-stage weighting on the two input components: the high 4 bits and low 4 bits of the weights are weighted separately to generate two partial sum currents: ISUMH = IBL7·1+
[0015] IBL7·1 / 2+IBL7·1 / 4+IBL7·1 / 8;ISUML=IBL3·1+IBL2·1 / 2+IBL1·1 / 4+IBL0·1 / 8 where ISUMH is the current in the high 4 bits, ISUML is the current in the low 4 bits, and IBL7 is the current formed by accumulating the current in the 8th weight bit to the same node.
[0016] In a preferred embodiment of the present invention, when the current mirror is input in the first part (i.e., the high 4 bits), SH is turned on. The connection relationship between the current mirrors is G7-M7, G6-M6, ... G0-M0. The currents ISUMH and ISUNL generated by the two sets of current mirrors converge into two capacitors and are converted into voltage. The LSB of the high 4 bits of the input is 16 times that of the low 4 bits. Therefore, when the low 4 bits of the input are input, SL is turned on, and the correspondence of the current mirrors is changed to G7-M3, G6-M2, G5-M1, G4-M0, G3-COMP. SH is the switching signal for sampling the high-order input. i (i = 0, 1, 2, 3, 4, 5, 6, 7) represents the i-th current mirror at the input terminal of the current mirror; the M mentioned above... i (i = 0, 1, 2, 3, 4, 5, 6, 7) represents the i-th current mirror at the output of the current mirror; SL is the low-order input sampling switch signal; COMP is the compensation circuit at the output of the current mirror.
[0017] As a preferred embodiment of the present invention, the analog-to-digital conversion circuit includes a capacitor array (ADC) for performing second-level weighting of the weights. SH and SL control the sampling of the IPH and IPL capacitor arrays, respectively. IPH is a capacitor array that samples the high-weight current ISUMH; IPL is a capacitor array that samples the low-weight current ISUMH. A 16:1 charge redistribution is performed between IPL and IPH by a CRD signal. CRD is a charge redistribution control signal that is closed before ADC quantization.
[0018] In a preferred embodiment of the present invention, the second-stage weighting is achieved internally by capacitor charge redistribution. The capacitor connected to ISUMH undergoes charge redistribution with a 16C0 capacitance, and the capacitor connected to ISUML undergoes charge redistribution with a C0 capacitance. This achieves a weighting ratio of 16:1 between ISUMH and ISUML. The charge redistribution formula is as follows:
[0019] so
[0020] Where: V ADC The analog voltage after the second stage of weighting within the ADC
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0022] (1) Because the present invention adopts hierarchical weighting, it overcomes the problem that the weighting circuit in the prior art has too large an area due to the need to complete a high weighting ratio, thereby reducing the circuit resource consumption of the weighting circuit. After hierarchical weighting, the area of the weighting circuit can be reduced by 85%.
[0023] (2) Because the present invention adopts input splitting and prediction mechanism, it overcomes the problem of excessive computation delay in the prior art, thereby reducing computation delay by 80%. Attached Figure Description
[0024] Figure 1 Overall architecture diagram of the vector multiplication circuit;
[0025] Figure 2 A diagram illustrating digital time conversion;
[0026] Figure 3 Schematic diagram of a 1T1R memory array;
[0027] Figure 4 Weighted circuit diagram;
[0028] Figure 5 A schematic diagram of the switching of the weighting circuit;
[0029] Figure 6 Input a calculation example diagram after decomposition;
[0030] Figure 7 Example diagram of prediction calculation;
[0031] Figure 8 Circuit diagram of an analog-to-digital converter. Detailed Implementation
[0032] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.
[0033] To address the problems existing in the background technology, this invention proposes a vector multiplication circuit oriented towards prediction optimization. A multiply-accumulate prediction circuit is proposed, which utilizes the principle of successive approximation of SAR ADCs to incorporate prediction into the vector multiplication circuit, reducing the circuit's output response time. The input is symmetrically split and post-weighted, and the weights are applied hierarchically, achieving a vector multiplication circuit with lower input delay and lower circuit resource consumption.
[0034] The multiply-accumulate operation uses 8 bits of data. The overall circuit architecture is as follows: Figure 1 As shown. The DTC (Digital Time Converter) converts the input data into a fully coded time signal, such as... Figure 2 As shown in the diagram, 'm' represents the size of the vector, i.e., the number of inputs. An input of 1 corresponds to the pulse duration of one clock cycle, an input of 2 corresponds to the pulse duration of two clock cycles, and so on. This time control mechanism activates the selector transistor in the 1T1R memory unit. The proposed circuit performs operations on 8-bit data. Using a fully encoded time signal would require 255 clock cycles, which is too long. Therefore, we propose symmetrically decomposing the input into two 4-bit data segments before serially inputting them to the DTC for digital-to-time conversion. This reduces the input time from 255 clock cycles to 30 clock cycles. After symmetrical input decomposition, there is a 16:1 ratio between the two input segments. This ratio is achieved through a weighting circuit, implemented using analog circuitry to reduce quantization errors. The vector multiplication circuit mainly consists of three parts: the 1T1R memory array, the weighting circuit, and the analog-to-digital converter.
[0035] (1) 1T1R memory array
[0036] The timing information obtained after input conversion by the DTC is split into two parts, IN[7:4] and IN[3:0]. Therefore, the DTC runs twice, obtaining two timing signals. These two timing signals are input sequentially to the 1T1R memory array to control the memory's enable time and also control the charging time of the ADC's internal capacitor by the subsequent current mirror current. IN[7:4] and IN[3:0] represent the high four bits and low four bits of IN, respectively. Figure 3As shown, the 1T1R memory array includes multiple memory (1T1R) cells; each 1T1R cell includes a selector transistor and an RRAM. The selector transistor controls whether the RRAM generates current, and the selector transistor is controlled by the input signal generated by the DTC. The column number m in the figure represents the number of weights contained in the weight vector. The RRAM stores the weight information bit by bit, for example, W in the figure. 0,0 W 0,1 …W 0,7 , representing the 8-bit weight of W0; the W i,j (i = 0, 1, 2…m; j = 0, 1, 2, 3, 4, 5, 6, 7) represents the j-th bit of the i-th weight in the weight vector. The RRAM has two states: high impedance and low impedance, representing the stored values 0 and 1 respectively. When the stored value is 0, the RRAM is in a high impedance state; even if the selector is on, no current will be generated. Current will only be generated when the stored value is 1 and the selector is on. The current of the j-th weight is accumulated at the same node to form the IBL. j The IBL j (j = 0, 1, 2, 3, 4, 5, 6, 7) represents the sum of the currents generated by the j-th column of memory cells; the IN... i (i = 0, 1, 2, ..., m) represents the i-th data in the input vector.
[0037] IBL j =IN0·W 0,j +IN j ·W 1,j +…+IN m ·W m,j ;
[0038] (2) Weighting circuit
[0039] The weighted circuit accepts IBL generated by the memory array. j These currents are weighted, and their on-time is controlled by the input time. Therefore, the output current generated by the current mirror also charges the ADC capacitor under the control of the input time. The weighting circuit weights the currents generated by the memory. In addition, due to the symmetrical splitting of the inputs, the weighting circuit also needs to perform a 16:1 weighting between the two input parts. The weighting circuit is as follows: Figure 4 As shown, the weighting circuit consists of a current mirror, a switching circuit, and a regulator. The current mirror performs the weighting of the inputs, the switching circuit performs the 16:1 weighting between the two split inputs, and the regulator improves the linearity of the current generated by the 1T1R memory.
[0040] The weighting circuit weights the weights. The proposed hierarchical weighting circuit performs a first-stage weighting, which weights the high 4 bits and low 4 bits of the weights separately, generating two current components: ISUMH = IBL7·1 + IBL7·1 / 2 + IBL7·1 / 4 + IBL7·1 / 8; ISUML = IBL3·1 + IBL2·1 / 2 + IBL1·1 / 4 + IBL0·1 / 8. ISUMH represents the current generated by the high 4 bits of weighting, and ISUML represents the current generated by the low 4 bits of weighting. Hierarchical weighting effectively reduces the area of the weighting circuit, which only needs to perform two 2x2 operations. 3 -2 0 The weighted average, compared to 2 7 -2 0 The current mirror area is reduced by more than 80%, which can also reduce the effects of parasitic capacitance and channel length modulation.
[0041] A weighting circuit weights the inputs from two parts. For example... Figure 5 As shown, when the first part of the input (high 4 bits) is active, SH is enabled. The connection relationship between the current mirrors is G7-M7, G6-M6, ... G0-M0. The currents ISUMH and ISUNL generated by the two sets of current mirrors converge into two capacitors and are converted into voltages. SH is the switching signal for sampling the high-order input; G... i (i = 0, 1, 2, 3, 4, 5, 6, 7) represents the i-th current mirror at the input terminal of the current mirror; the M mentioned above... i (i = 0, 1, 2, 3, 4, 5, 6, 7) represents the i-th current mirror at the current mirror output. The LSB of the high 4 bits of the input is 16 times that of the low 4 bits. Therefore, when the low 4 bits are input, SL is enabled, changing the correspondence of the current mirrors to G7-M3, G6-M2, G5-M1, G4-M0, G3-COMP. This is similar to shifting the weight bits to the right by 4 bits, which is equivalent to reducing it to 1 / 16 of the original. SL is the switching signal for low-order input sampling; COMP is the compensation circuit at the current mirror output. Because M7-M4 are not active at this time, their gates are temporarily turned off by connecting to VDD to reduce power consumption. The compensation circuit COMP is used to compensate for the carry effect that may be generated by the low-order weights when inputting low-order bits.
[0042] (3) Analog-to-digital conversion circuit
[0043] The ISUMH and ISUML inputs from the weighting circuit need to undergo a second stage of weighting within the ADC. This weighting is achieved through charge redistribution of capacitors. The capacitors connected to ISUMH and ISUML undergo charge redistribution in a 16:1 ratio to achieve weighting between ISUMH and ISUML. Furthermore, the ADC needs to wait for the inputs from both parts to complete their time before performing analog-to-digital conversion to obtain the digital output. Since the response time between the ADC's digital output and the inputs is relatively long, we analyzed the relationship between the two inputs and their respective contributions to the final result, incorporating prediction within the ADC. After analysis, as shown... Figure 6 As shown, the high four bits of the result are only related to the carry that may be generated by the high four bits and low four bits of the input. Therefore, a prediction is added between the high four bits and low four bits of the input. After the high four bits of the input and the weight are calculated, if the remaining value after subtracting OUT[7:4] is greater than 11000000, then the carry is determined to be 1; if it is less than 11000000, then the carry is determined to be 0. OUT[7:4] is the high four bits of the final calculation result. After subtracting OUT[7:4] and the carry, the remaining value will be accumulated with the low four bits of the input and quantized to produce OUT[3:0]; OUT[3:0] represents the low four bits of the final calculation result. Figure 7 (a) and (b) represent two different calculation scenarios. The specific implementation process is as follows:
[0044] ADC circuit diagram as follows Figure 8 As shown, the modification of the capacitor array can be seen quite intuitively. Generally speaking, the capacitor array of an 8-bit SARADC has 2... 0 -2 7 C. Because the ADC needs to perform a 16:1 weighting, we divide the original capacitor array into two IPL sampling capacitor arrays. 0 -2 3 and IPH sampling capacitor array 2 4 -2 7When sampling the first part of the input IN[7:4], SH and SL are closed. IPH receives current from ISUMH, and IPL receives current from ISUMH. After receiving the first part of the input, SH and SL are turned off, CRD is turned on, and IPL and IPH redistribute charge with a capacitance ratio of 1:16 to complete the second stage of weighting. After the charge redistribution of IPH and IPL, a partial voltage V1 is obtained. The SAR ADC starts to perform analog-to-digital conversion on V1 and controls the 128C-16C capacitor to quantize the 4b output OUT[7:4]. According to the quantization process of the SAR ADC, the voltage inside the ADC at this time (the voltage difference between IP and IN) is also reduced by the corresponding value of OUT[7:4]. After quantizing OUT[7:4], a prediction comparison is performed to check if the remaining value (the voltage difference between IP and IN) is greater than 3 / 4 VLSB (the VLSB at this time is the voltage added or subtracted from 16C). At this point, it is necessary to control 8C and 4C in the INL branch to connect to Vref1, that is, to add 3 / 4 VLSB to the IN terminal. This is equivalent to subtracting 3 / 4 VLSB from the difference between IP and IN. Then compare; if it is greater than 0, it means that the remaining value is greater than 3 / 4 VLSB and a carry is needed; otherwise, no carry is needed. If a carry is needed, after generating the carry output, control the 16C capacitor of INL to subtract one VLSB of voltage; if no carry is needed, no subtraction is required. Afterward, disconnect CRD, connect the 8C and 4C capacitors back to GND, and wait for the input of the second part. When sampling the second part of the input, SL is closed, SH and CRD are open, and only IPL receives current from ISUML. After the input is completed, SH and SL are open, CRD is closed, and IPL:IPH is redistributed in a 1:16 ratio. The result of the second part is added to the high-order residual voltage in a 1:16 ratio, which is equivalent to amplifying the residual value by 16 times and then adding the input result of the second part. Then, the low-order output OUT[3:0] is quantized.
[0045] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A vector multiplication circuit for prediction optimization, characterized in that, This includes a digital time converter (DTC), a 1T1R memory array, a weighting circuit, and an analog-to-digital converter; The DTC is used to convert the symmetrically split input data into two parts into a fully encoded time signal. The DTC operates twice in the same calculation. The 1T1R memory array is used to receive the time information obtained by the DTC and to activate the memory according to the received time information to obtain the accumulated current IBL on the nodes of a column of memory cells. j The 1T1R memory array includes multiple 1T1R memory cells, each including a selector transistor and an RRAM. The selector transistor controls whether the RRAM generates current, and is controlled by an input signal generated by the DTC. The RRAM is a memory whose resistance changes. The RRAM has two states: high resistance and low resistance. The high resistance state represents a stored value of 0, and the low resistance state represents a stored value of 1. When the stored value is 0, the RRAM will not generate current. The RRAM will contribute current to the node only when the stored value is 1 and the selector is turned on. IBL j =IN0·W 0,j +IN1·W 1,j +…+IN m ·W m,j ; Among them, IBL j The sum of the currents generated by the memory cells in the j-th column, where j = 0, 1, 2, 3, 4, 5, 6, 7; The weighting circuit is used to weight the current generated by the memory. The weighting circuit includes a current mirror, a switching circuit, and a regulator. The current mirror includes a mirrored part and a mirrored part, which are connected by the switching circuit. The two parts of the current mirror complete the first level of weighting between different bit values. The switching circuit realizes different correspondences between the two parts of the current mirror and, together with the analog-to-digital conversion circuit, realizes the weighting between the two inputs. The regulator improves the linearity of the current generated by the 1T1R memory. The weighting circuit performs the first stage of weighting on the two inputs: the high 4 bits and low 4 bits of the weights are weighted separately to generate two parts and currents: ISUMH = IBL7·1 + IBL6·1 / 2 + IBL5·1 / 4 + IBL4·1 / 8; ISUML = IBL3·1 + IBL2·1 / 2 + IBL1·1 / 4 + IBL0·1 / 8; where ISUMH is the current of the high 4 bits and ISUML is the current of the low 4 bits; The analog-to-digital converter circuit is used to perform a second-stage weighting of the current to generate an analog voltage, and then performs analog-to-digital conversion on the generated analog voltage to obtain the resulting digital output. The analog-to-digital converter (ADC) circuit includes a capacitor array for performing the second-stage weighting of the inputs. The high-order input sampling switch signal SH and the low-order input sampling switch signal SL control the sampling of the IPH and IPL capacitor arrays, respectively. IPH is a capacitor array that samples the high-order weight current ISUMH, and IPL is a capacitor array that samples the low-order weight current ISUMH. A 16:1 charge redistribution is performed between IPL and IPH through a CRD signal. CRD is a charge redistribution control signal that is closed before ADC quantization.
2. The vector multiplication circuit according to claim 1, characterized in that, When the current mirror is input in the first part (i.e., the high 4 bits), the high-bit input sampling switch signal SH is turned on. The connection relationship between the current mirrors is G7-M7, G6-M6, ..., G0-M0. The currents ISUMH and ISUML generated by the two sets of current mirrors converge into two capacitors and are converted into voltages. The LSB of the high 4 bits is 16 times that of the low 4 bits. When the low 4 bits are input, the low-bit input sampling switch signal SL is turned on, and the correspondence of the current mirrors changes to G7-M3, G6-M2, G5-M1, G4-M0, G3-COMP. COMP is the compensation circuit at the output of the current mirror.
3. The vector multiplication circuit according to claim 2, characterized in that, The second-stage weighting is achieved internally by capacitor charge redistribution. The capacitor connected to ISUMH undergoes charge redistribution with a 16C0 capacitance, while the capacitor connected to ISUML undergoes charge redistribution with a C0 capacitance. This achieves a weighting ratio of 16:1 between ISUMH and ISUML. The charge redistribution formula is as follows: ; so ; Where: V ADC This is the analog voltage after the second-stage weighting within the ADC.
Citation Information
Patent Citations
Convolution operation and full connection operation circuit used for convolutional neural network
CN108764467A
Convolution calculation accelerator based on 1T1R memory array and operation method of convolution calculation accelerator
CN110569962A