Method and circuit for performing multiplication and accumulation operations in memory
By performing MAC operations directly in memory using current sources and capacitive lines with time or current weighting, the method addresses the power efficiency limitations of the Von Neumann architecture, achieving high throughput and low power consumption for machine learning on embedded systems.
Patent Information
- Application Number
- FR2022012440
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-11-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-11-28
AI Technical Summary
The existing Von Neumann architecture incurs high power consumption due to data movement between memory and processor, limiting the efficiency of machine learning on embedded systems, particularly in performing multiplication and accumulation (MAC) operations.
The method and circuit for performing MAC operations directly in memory, utilizing current sources and capacitive lines to accumulate charges, with time or current weighting techniques to enhance calculation speed and reduce power consumption.
This approach significantly reduces power consumption by minimizing data movement and achieves high throughput, suitable for applications like keyword spotting in audio processing, with a throughput of 50 GOPS for a 5-bit MAC operation.
Smart Images

Figure 00000027_0000 
Figure 00000027_0001 
Figure 00000028_0000
Abstract
Description
Title of the invention: Method and circuit for carrying out multiplication and accumulation operations in memory Technical field
[0001] The present invention relates to the field of in-memory computing, and more particularly relates to a method for performing multiplication and accumulation in memory and an associated circuit, in particular for a neural network where the basic operation for a perceptron is multiplication and accumulation (MAC). Prior art
[0002] The emergence of the Internet of Things is creating a growing demand for ultra-low-power computing solutions to bring artificial intelligence to such devices. The additional power cost preventing machine learning on embedded systems comes from the power consumption due to data movements.
[0003] In a classic “Von Neumann” architecture, stored data must be moved from memory to the processor to be calculated step by step. In the example given in Boris Murmann’s “Mixed-Signal Processing Opportunities for AI” publication, to perform a MAC operation (a multiplication and an addition), the data must be retrieved from memory four times. Taking the cost of accessing the data alone and estimating this access at 50fJ / byte, the overall consumption is 200fJ / MAC, capping the efficiency at 10 TOPS / W. One way to alleviate this constraint, which is called the memory wall or the “Von Neumann bottleneck”, is to bring the processing elements inside the memory to avoid wasting energy accessing the data.
[0004] In-memory computing is a technique of performing computer calculations entirely in the computer's memory. In-memory processing is a method for solving the disadvantages, particularly in terms of performance and energy cost, caused by moving data between the processor and memory.
[0005] [Fig.l] shows the principle of in-memory calculation of MAC operations in the prior art. A digital vector [XB ..., Xm] is distributed over several lines. The multiplication of this vector by elements of a weight matrix W is carried out inside the memory where the weight matrix W is stored. The accumulation, in the current or voltage domain, is carried out at each column in a capacitive line, also called an accumulation line AL. This signal method mixed outputs the result of MAC operations in analog, an analog-to-digital converter A / D at the end of each accumulation line AL converts the result to digital, represented by the vector [Ob ..., On], so that this result can be reused. Each accumulation line AL is the equivalent of a perceptron without the activation function.
[0006] There are different approaches in the prior art to performing such a calculation in memory.
[0007] In a resistive approach illustrated in [Fig. 2] and the subject of the article Q. Liu et al. "A Fully Integrated Analog ReRAM Based 78.4 TOPS / W Compute-In-Memory Chip with Fully Parallel MAC Computing", memristors are used to perform the multiplication. The vector X is sent as an analog voltage on all lines. The weights W are stored as conductance values of the memristors, modulating the current through each accumulation line AL. However, this approach requires digital-to-analog converters with large output currents and a current comparator. It is possible to perform multi-bit operations with this approach using up to 16 conductance states, according to the aforementioned article. Process variations prevent achieving a higher number of bits.Additionally, in terms of power dissipation, writing requires large current spikes (of the order of 6qA in the publication Nguyen Cong Dao et al. “Memristor-based Reconfigurable Circuits: Challenges in Implementation”). It should be noted that a trade-off exists between power consumption and variability.
[0008] In a capacitive approach illustrated in [Fig.3] and the subject of the publication D. Bankman et al. “An 8-bit, 16 input, 3.2 pj / op switched-capacitor dotproduct circuit in 28-nm FDSO1 CMOS”, the multiplication is based on sharing charges or currents on a capacitive line, typically using XNOR gates for binary multiplication. The unit capacitors are charged according to each binary multiplication and the charges are redistributed across all capacitors at an accumulation line AL, resulting in a voltage to be converted by the analog-to-digital converter. To perform a multi-bit operation with switched capacitors, the article Boris Murmann “Mixed-Signal Computing for Deep Neural Network Inference” shows a topology similar to a digital multiplier using an accumulation line for each bit.However, this method requires additional circuitry to combine all row results, which is sufficient for a small number of rows (< 5) but increases power consumption for a higher number of bits. Statement of the invention
[0009] There is therefore a need to improve memory multiplication and accumulation (MAC) techniques, particularly in terms of energy consumption.
[0010] Method of performing multiplication and accumulation (MAC) operations in memory
[0011] The invention aims to meet this objective and has as its object, according to one of its aspects, a method of performing multiplication and accumulation (MAC) operations in memory, in particular for a neural network, in which we determine the result of the matrix product O of a vector X — [ X । Xi... Xm ] composed of m elements each represented on (nx - 1) bits and of a matrix W = Fli composed of mxn elements each represented on (nw - 1) bits:
[0012] each element of the matrix product O being the sum of scalar products O,=='' 1 *' "" ' 1 *' with the bit vector X j = [X; [ 0 ], ..., Xj [ n Y - 2 ] ] and the bit vector of weight Wji = [Wjf[0], ..., 2]} the multiplication operations and accumulation leading to the scalar product Rjj by applying to a set of logic gates the values of the vectors Xj and Wjj and those of time masks associated, the time masks PX = [PX ( 0 ), ..., PX(nx - 2)] being associated with the bits of the elements of the vector Aj and the time masks PW — [PW( 0 ), ..., PW(nw - 2)] being associated with the bits of the elements of the vector IF yf, so as to generate at least one control signal SC from at least one source of current, the control signal SC activating said at least one current source so as to accumulate capacitive charges in an accumulation line, in particular in a manner proportional to the scalar product Rj,i.
[0013] By "performing multiplication and accumulation (MAC) operations in memory" is meant the execution of these operations in a memory of the electronic chip where the operands Xj and VFjj are stored.
[0014] By "activating the current source" is meant allowing said source to deliver a current circulating in the accumulation line.
[0015] The invention provides a new method for calculating MAC operations in memory while reducing the power consumed.
[0016] Preferably, at least one of the two vectors is signed, X(sign] being the sign bit of the vector Xj and / or being the sign bit of the vector of significant bits Wÿ, the set of logic gates generating a signal of polarity SP = X loading or unloading the accumulation line.
[0017] The time masks are preferably generated in the form of pulses.
[0018] The PX and PW time masks can be generated by a time generator block. time masks located outside memory, that is, outside the memory of the electronic chip where the operands Xj and W ÿ are stored.
[0019] In one embodiment, the time masks PW(d) are generated so as to activate successively for a duration T d substantially equal to A / j- being a unit time corresponding to the activation duration of the time mask associated with the bit of the vector X, of the lowest weight, the time masks PX(c) being generated so as to activate successively for each Td for a duration substantially equal to 2c+(l* T.
[0020] This scenario corresponds to the presence of a single control signal controlling a single current intensity value in absolute value, by charging or discharging the accumulation line, such as: [°02n sc = sc^=^'^pwid) C) with 161 '■ - ">•j G {1, ..., m}.
[0022] Here there is a weighting by the activation time of the current source. In this case, the voltage Vq across the accumulation line for a multiplication and accumulation operation can be equal to: / *7*0, V o. -
[0023] C being the value of the capacitance of the accumulation line and I the intensity of the current of the current source. Note that the capacitance of the accumulation line can be a parasitic and / or distributed capacitance.
[0024] In the case where at least two current sources are present, each being controlled by a respective control signal SC, the time masks PW(d) can be generated so as to activate successively for a duration Td substantially equal to Td = with {0, -, max(^ - aj)}, the - 1) bits — I of PW being distributed in k intervals, b^], k being the number of current sources or control signals, with 1 e {0,..., k -1}, b} to { a,. b]} E [0, (niy-2)], aietbi being chosen such that the set of intervals [ aj, h,] represents a partition of the interval [0 , nw - 2], the longest interval being [0 , max(fy-r^)] , T being a unit time corresponding to the activation duration of the time mask associated with the least significant bit of the vector Xj, the time masks PX(c) being generated so as to activate successively for each Td for a duration substantially equal to 2c+d* T.
[0025] This scenario corresponds to the presence of several control signals SCqj controlling k current sources, k > 2, by charging or discharging the accumulation line, such as: 100261 SCi^^XpwÇd-a,) \ J. |r 1 " {0 -
[0027] Here there is a weighting by the current, by having several current sources, weighted by a power of two and controlled by control signals representative of the values of the bits associated with this weighting. This advantageously increases the calculation speed by reducing the duration of the control signals. In this case, the voltage Vq across the accumulation line for a multiplication and accumulation operation can be equal to: v o =------------c-------------
[0028] _ _}!^Oi _ _ _ -
[0029] C being the value of the capacitance of the accumulation line and I the intensity of the current of the least significant current source (i.e. for ^=0).
[0030] For example, in the case where - 1) — 4 and k — 2, it is possible to distribute the 4 bits equally between the two sources: 100311 SC 0JJ = ( d ) * W(EX2 / ' M - > V j< |: SC y J=(e' =2 pw (d - 2 ) * (EX 2 ™ ( c ) ™ kl)
[0032] Or to distribute the bits unequally: 100331 sc 0JJ =(EX™' ( d ) (EX 2 ™ ( e 100341 sc w =pw ( o ) *^3] (EX 2 ™ ( c ™ PD
[0035] Circuit for performing multiplication and accumulation (MAC) operations in memory
[0036] The invention also relates, according to another of its aspects, to a circuit for carrying out multiplication and accumulation operations in memory, in particular for a neural network, configured to determine the result of the matrix product O of a vector X = [X( X2 ... Xm] composed of m elements each represented on (nx - 1) bits and a matrix W = composed of mxn elements W . ■ ■■ w, each represented on (nw - 1) bits: ■Wu ■ ■ IV. 1 b. 0= [XjXo... Xm] x ■ ■ ■ = [OYO2... on .^1 ■ ■ W^.
[0037] each element of the matrix product O being the sum of scalar products o,= Ri ,= 1 e ' 1 .....">• j e < with the bit vector X z = [Xy [ 0 ], ..., Xj [ nx - 2 ] ] and the weight bit vector Wjj = [Wÿ[ 0], ..., [ ilw - 2]pe circuit comprising:
[0038] - a logic block comprising logic gates to which the values of the vectors Xj and W jj and those of associated time masks, the time masks PX = [PX( 0 ), .... PX(nx - 2)] being associated with the bits of the elements of the vector Xj and the time masks PW = [PW ( 0 ), ..., PW(nw - 2)] being associated with the bits of the elements of the vector Wjj, the logic block being configured to generate in output at least one control signal SC intended to activate at least one source of current, and
[0039] - said at least one current source configured to be activated by the signal of SC control so as to accumulate capacitive charges in an accumulation line.
[0040] Preferably, at least one of the two vectors is signed, X^s^gzt] being the sign bit of the vector Xj and / or VVypgn] being the sign bit of the vector of bits of weight Wjj, said accumulation line being configured to be charged or discharged according to a polarity signal SP = Xpfl] x Wjp'gn] this signal SP being in particular obtained at the output of an XOR or XNOR logic gate receiving at input the sign bit Xpgn] and the sign bit Wjpgn]-
[0041] In one embodiment, the logic block comprises NAND and / or NOR logic gates.
[0042] The logic block may comprise a first logic sub-block for the time masking of the bits Xj and a second logic sub-block for the time masking of the bits of weight Wjj, these two sub-blocks being cascaded so that the outputs of the first sub-block are at the input of the second sub-block.
[0043] The first sub-block, receiving as input the time masks PX and PW as well as the bits Xj(): nx - 2], can comprise two stages of logic gates: - a first stage comprising several levels of logic gates, the levels being connected in cascade so that the signal or signals at the output of one level are the signal or signals at the input of the following level until they reach a single output signal, first stage in which each bit Xy[0 : nx - 2] and the associated time mask PX are at the input of a first level logic gate, in particular a NAND or NOR gate; and - a second stage comprising at least one level of logic gates, in which the signal at the output of the first stage and each time mask PW are at the input of a logic gate of said level, in particular a NAND or NOR gate.
[0044] The second sub-block, receiving as input the output signals of the first sub-block and the weight bits WylO: nw - 2], can comprise several levels of logic gates, the levels being connected in cascade so that the signal or signals at the output of one level are the signal or signals at the input of the following level until reaching a single output signal corresponding to the control signal SC, each first-level logic gate receiving as input an output signal of the first sub-block and the corresponding weight bit VF^O: llw - 2]. The second sub-block may include an XOR or XNOR gate receiving as input the sign bit of the bit vector Aj and the sign bit yFyJsz^ / l] of the bit vector W. If the positive sign is interpreted as a binary '0' and the negative sign as a binary '1', an XOR gate may be used for the sign bits. If the positive sign is interpreted as a binary '1' and the negative sign as a binary '0', an XNOR gate may be used for the sign bits.
[0045] In one embodiment, said at least one current source is created by current mirrors using CMOS transistors.
[0046] The current mirrors preferably comprise PMOS transistors for charging the accumulation line and NMOS transistors for discharging said line.
[0047] Preferably, the current mirrors have a cascoded architecture. This makes it possible to increase the output impedance of the current mirrors and to have a more stable current value despite the variation of the voltage of the accumulation line.
[0048] In one embodiment, the circuit comprises at least one electronic switch, being in particular a transmission gate, controlled by the control signal SC for the activation of said at least one current source.
[0049] The circuit may comprise at least one capacitor at the output of said at least one in- electronic switch to increase the capacity of the accumulation line.
[0050] The circuit preferably comprises a secondary accumulation line connected to the accumulation line via a follower amplifier. Such a secondary accumulation line is intended to control the current.
[0051] Preferably, the circuit comprises a follower amplifier connected between the accumulation line and the secondary accumulation line. This has the advantage of allowing each of the two accumulation lines to remain at the same voltage level in order to counteract the charge sharing effect when the switches open and close.
[0052] In one embodiment, the circuit comprises an analog-to-digital converter connected either directly to the accumulation line in the absence of a secondary accumulation line, or to the secondary accumulation line in particular by means of a follower amplifier.
[0053] The analog-digital converter may comprise: - two analog comparators, one comparing to +LSB and the other to -LSB the value of the charge of the accumulation line to which the analog-digital converter is connected; - a counter configured to increment or decrement depending on the result of the comparison of the comparators, and whose output is that of the converter; - a charge injector block configured to return the charge of the accumulation line connected to the converter to its initial value in the event that the meter indicates that this line has been discharged; and - a discharger block configured to discharge the accumulation line connected to the converter and return it to its initial value in the event that the meter indicates that it has been charged.
[0054] The analog-digital converter may comprise a look-up table (LUT: “Look Up Table” in English) receiving as input the output of the counter and being configured to apply as output a substitution value chosen during calibration and / or corresponding to an activation function.
[0055] In the case in particular where the values of the LUT table are chosen during calibration, the output of the LUT acts via a feedback loop on the values of the reference inputs +LSB and -LSB of the analog comparators and / or on the unit time and / or on the unit current.
[0056] The correspondence table can be used to correct non-linearities in the circuit or to respect any other form of response.
[0057] The unit time can be acted upon in the case where there is a single control signal controlling a current source.
[0058] We can act on the unit current in the case where there are several signals of control controlling several current sources generating a multiple of this current.
[0059] The values of the reference inputs +LSB and -LSB can be acted upon in either of the aforementioned cases. This calibration can be carried out at least once to measure the non-linearities of the circuit, calibrate it and reduce the output error.
[0060] It is also possible, by choosing appropriate values for the LUT, to imitate an activation function in artificial intelligence, by adjusting the value of a parameter of the generation of the control signal.
[0061] An activation function is used to modify the data in a non-linear manner. This can be a sigmoid, or Tanh, or ReLU, etc. It represents in particular the predefined non-linear response that we are seeking to obtain.
[0062] In one embodiment, the circuit comprises a time mask generator block located outside memory.
[0063] In one embodiment, the generator block is configured to generate the time masks so that the time masks PW(d) are activated successively for a duration Td substantially equal to and that the I *^^0 “ J time masks PX(c) are activated successively for each T d for a duration substantially equal to 2c+d* T, T being a unit time corresponding to the activation duration of the time mask associated with the least significant bit of the vector Xj.
[0064] In this case, the circuit according to the invention preferably comprises a single current intensity value in absolute value, controlled by a single control signal SC.
[0065] In another embodiment, the circuit according to the invention comprises at least two current sources, each being controlled by a respective control signal SC, the generator block being configured to generate the time masks so that the time masks PW(d) are activated successively for a duration Td substantially equal to 2 with [maxf^-aj], the (nw - 1) bits of PW being distributed in k intervals [ a / , b J, k being the number of current sources or control signals, with 1 e {0,..., k- 1], bf > { ab b)} G [0, (nw-2) ], aietb; being chosen such that the set of intervals [r / / , b]] represents a partition of the interval [0 , nw - 2], the longest interval being [0 , max(bi-ü] )] , T being a unit time corresponding to the activation duration of the time mask associated with the least significant bit of the vector Xj, and that the time masks PX(c) are activated successively for each Td for a duration substantially equal to 2c+d* T.
[0066] The invention also relates, according to another of its aspects, to a set of circuits for carrying out MAC multiplication and accumulation operations in memory, especially for a neural network, configured to determine the result of the matrix product of a vector X — [ Xj X2 ... Xm ] by a weight matrix W= comprising a plurality of circuits according to the invention implemented / 77.1 W vv mjt parallel to each other and sharing a single time mask generator block located out of memory. Brief description of the drawings
[0067] The invention may be better understood by reading the detailed description which follows, of non-limiting examples of its implementation, and by examining the attached drawing, in which:
[0068] [Fig-1] [Fig. 1] schematically illustrates the principle of in-memory calculation of MAC operations in the prior art;
[0069] [Fig.2] [Fig.2] schematically represents an example of in-memory calculation of MAC operations according to a resistive approach in the prior art;
[0070] [Fig.3] [Fig.3] schematically illustrates an example of in-memory calculation of MAC operations according to a capacitive approach in the prior art;
[0071] [Fig.4] [Fig.4] is a view of a circuit schematically representing the principle of the calculation in memory of a MAC operation according to the invention;
[0072] [Fig.5] [Fig.5] recalls the equation of the scalar product of two vectors of 4 bits each;
[0073] [Fig.6] [Fig.6] shows the time evolution of the terms of the scalar product of the [Fig.5] ;
[0074] [Fig.7] [Fig.7] represents the curve of the temporal evolution of the voltage at level of an accumulation line during a 100 MAC calculation;
[0075] [Fig.8] [Fig.8] illustrates a first example of time mask chronograms;
[0076] [Fig.9] [Fig.9] represents a second example of mask timing diagrams temporal;
[0077] [Fig. 10] Figure 10 schematically illustrates an example of a first logical sub-block for the temporal masking of a 4-bit binary vector Xj;
[0078] [Fig. 11] Figure 11 schematically represents an example of a second logical sub-block for the temporal masking of a 4-bit binary vector Wjj;
[0079] [Fig. 12] Figure 12 is analogous to Figure 8 and shows the details of a part of the scalar product of said vectors Xj and Wjj from the time masks;
[0080] [Fig. 13] [Fig. 13] shows the timing diagram of the control signal resulting from a simulation of the scalar product of the vectors X and W given as an example in Figures 10 to 12 ;
[0081] [Fig.14] [Fig.14] schematically illustrates the architecture of a circuit according to the invention;
[0082] [Fig. 15] [Fig. 15] schematically represents a first embodiment of current sources which can be used in the circuit of [Fig. 14];
[0083] [Fig. 16] [Fig. 16] schematically illustrates a second embodiment of current sources that can be used in the circuit of [Fig. 14];
[0084] [Fig. 17] [Fig. 17] schematically represents an example of a switch that can be used in the circuit of [Fig. 14];
[0085] [Fig. 18] [Fig. 18] schematically illustrates an accumulation line that can be used in the circuit of [Fig. 14];
[0086] [Fig. 19] [Fig. 19] schematically represents an accumulation line and a secondary accumulation line which can be used in the circuit of [Fig. 14];
[0087] [Fig.20] [Fig.20] is similar to [Fig. 19] additionally representing switches connected to the accumulation lines;
[0088] [Fig.21] [Fig.21] schematically represents a first embodiment of an analog-digital converter which can be used in the circuit of [Fig. 14];
[0089] [Fig.22] [Fig.22] schematically illustrates a second embodiment of an analog-to-digital converter that can be used in the circuit of [Fig. 14];
[0090] [Fig.23] [Fig.23] schematically represents a third embodiment of an analog-to-digital converter that can be used in the circuit of [Fig.14];
[0091] [Fig.24] [Fig.24] schematically illustrates an example of architecture of a set of circuits according to the invention with the time weighting approach;
[0092] [Fig.25] [Fig.25] schematically represents an example of a circuit according to the invention operating according to the current weighting approach;
[0093] [Fig.26] [Fig.26] illustrates a first example of timing mask diagrams that can be used for the circuit of [Fig.25];
[0094] [Fig.27] [Fig.27] schematically represents a first example of architecture of a set of circuits according to the invention with the current weighting approach;
[0095] [Fig.28] [Fig.28] illustrates a second example of time mask timing diagrams that can be used for the circuit of [Fig.25]; and
[0096] [Fig.29] [Fig.29] schematically represents a second example of architecture of a set of circuits according to the invention with the current weighting approach. Detailed description
[0097] Figures 1 to 3 refer to examples of the prior art and have been described above.
[0098] [Fig.4] schematically illustrates a circuit illustrating the principle of the calculation in memory of a MAC operation according to the invention. This principle is based on the accumulation of charges on a capacitive line AL by means of current sources CS. The example illustrated relates to a ternary multiplication of a bit X by a bit of weight W taking into account the respective sign bits with which they are associated. Such a ternary multiplication is implemented in a logic block LG whose output is a control signal which, depending on the sign of the result of the ternary multiplication, activates a switch SW which allows a current source CS to either charge or discharge the capacitive line AL. The result is then digitized by an analog-to-digital converter ADC.
[0099] This principle of accumulation of charges on an accumulation line will be extrapolated to a multi-bit multiplication.
[0100] [Fig.5] recalls the equation of the scalar product of two vectors, a vector X and a weight vector W, of 4 bits each.
[0101] y .pj^”r22cX •)' with Hx ~ and nw = 'cs vectors having no sign bit in this example.
[0102] A control signal, representing this scalar product, will be used to control current sources.
[0103] For this purpose, the current sources will be activated for a reference unit time T weighted by the weights of the terms of the scalar product of [Fig.5].
[0104] [Fig.6] shows such a signal corresponding to the time evolution of the terms of the scalar product of [Fig.5].
[0105] Depending on the sign of the scalar product, the accumulation line will be charged or discharged using the current sources.
[0106] The method according to the invention exploits the time required to charge an accumulation line for a long period with a low current, thanks to the recent advanced CMOS technology. If we take a reference unit time T of 20 ns, the total time for a 5-bit MAC operation is 5 qs. For an accumulation line capable of receiving the result of 100 MACs, the throughput is 50 GOPS (Giga Operations Per Second). This throughput is high enough for many applications such as audio applications, in particular the identification of keywords or "Key Word Spotting" (KWS) in English where new data enters every 10 ms.
[0107] [Fig.7] represents the curve of the time evolution of the voltage at the level of an accumulation line during a calculation of 100 MAC. The line is initialized at 0.5 V (half the dynamic range of a circuit powered at 1 V) in order to be able to represent negative values.
[0108] To obtain the control signal representing the scalar product of the multi-bit vectors Xj and Wjj, a logic block LG receiving as input these vectors and associated time masks PX and PW is implemented.
[0109] [Fig.8] illustrates a first example of time mask timing diagrams where the generation of these signals begins with the lowest weighting.
[0110] A time mask PX, PW is generated for each bit of the elements of the vector X and the weight matrix W.
[0111] In the case of a 4-bit multiplication, 4 time masks are required. There are therefore 8 time masks in total: PX=[PX(0),..., PX(3)] associated with the bits of the elements of the vector X and PW=[PW(0),..., PW(3)] associated with the bits of the elements of the weight matrix W.
[0112] The time masks PW(d) are preferably generated so as to be activated successively for a duration Td substantially equal to 2^^^73 Y being the reference unit time corresponding to the activation duration of the time mask associated with the least significant bit Xj, the time masks PX(c) being generated so as to be activated successively for each T d for a duration substantially equal to 2c+rf* T; {c, d] e {0,..., 3}.
[0113] It is also possible for the generation of the PX and PW signals to start with the highest weighting, as illustrated in the example of [Fig.9], or in any order, as long as the PX and PW time masks are synchronized.
[0114] The logic block LG receiving as input the vectors X and W and the associated time masks PX and PW preferably comprises a first logic sub-block for the time masking of the bits X and a second logic sub-block for the time masking of the bits of weight W, these two sub-blocks being cascaded so that the outputs of the first sub-block are at the input of the second sub-block.
[0115] Figure 10 schematically illustrates an example of a first logical sub-block LGX for the temporal masking of a 4-bit binary vector Xj.
[0116] The first logic sub-block LGX comprises two stages of logic gates: a first stage 10 and a second stage 11.
[0117] In this example, the first stage 10 comprises 4 levels of logic gates, the levels being connected in cascade so that the signal or signals at the output of one level are the signal or signals at the input of the next level until they reach a single output signal 101. Each bit Xj[0],..., Xj[3] and the associated time mask PX(0),..., PX(3) are at the input of a first level logic gate, in this case a 2-input NAND gate.
[0118] The signal 101 at the output of the first stage therefore represents the following value:
[0119] The second stage 11 comprises in this example two levels of logic gates: a level of two-input NAND gates and a level of NOT gates. The signal 101 at the output of the first stage and each time mask PW(0),PW(3) are at the input of a logic gate of the NAND gate level.
[0120] The NOT gate level gives four output signals denoted XW(0),XW(3) which will be injected into the input of the second logic sub-block.
[0121] Figure 11 schematically represents an example of a second logical sub-block LGW for the temporal masking of the 4-bit weight vector Wjj.
[0122] The second sub-block LGW receives as input the output signals of the first sub-block XW(0),..., XW(3) and the weight bits Wj,i[0],..., Wjj[3] and comprises several levels of logic gates, in this case 4 levels, these levels being connected in cascade so that the output signal or signals of one level are the input signal or signals of the following level until they reach a single output signal corresponding to the control signal SC. Each first-level logic gate (a two-input NAND gate in this example) receives as input an output signal XW(0),..., XW(3) of the first sub-block LGX and the corresponding weight bit Wj.i[0],..., Wj,i[3]. The signal SC therefore represents the following value: SC = (eLp^^I^wcI) The second LGW sub-block also includes an XNOR gate receiving as input the sign bit Xj.wgn] of the vector Xj (X;[4]) and the sign bit of the vector Wjj (Wy([4])- The output of the XNOR gate gives a signal of polarity SP = Depending on its value (+ or -) representative of the sign of the scalar product of the vectors Xj and Wjj, SP activates one of the switches SW to charge or discharge the accumulation line AL.
[0123] Figure 12 reproduces the timing diagrams of Figure 8, showing the details of part of the scalar product of said vectors X / and Wjj from the associated time masks.
[0124] Figure 13 shows the timing diagram of the control signal SC resulting from a simulation of the scalar product of the vectors Xj and Wjj given as an example in Figures 10 to 12. The vectors X7 and Wjj are as follows: Xy=[H010] and W? / =
[10110] , the sign bit being the one on the far right.
[0125] The small peaks seen on the SC signal in [Fig. 13] simply indicate a transition to the same state, but this does not occur in reality.
[0126] [Fig. 14] schematically illustrates the architecture of a circuit 1 according to the invention.
[0127] This circuit includes a generator PG of time masks PX, PW, sub- LGX, LGW logic blocks, CS current sources, SW switches, AL accumulation lines and ADC analog-to-digital converters.
[0128] The generator PG is preferably located outside memory and generates time masks PX, PW for all the logical sub-blocks LGX associated with the vectors Xb ..., Xa.
[0129] The XW outputs of the LGX logic sub-blocks are distributed over several lines of a weight matrix W: Wi.i,..., Wm.i, Wi.2,..., Wm>2,(...), Wi.n,..., Wm>n, the indices in second position ranging from 1 to n corresponding to the columns of the matrix.
[0130] The values of the vectors X and the matrix W are preferably stored near the logic blocks in order to reduce the cost of accessing the data, by using Flip-Flop registers or any other memory capable of storing a value without consuming too much energy, such as an SRAM memory. Each bit is preferably stored individually to be used as an input to the logic block.
[0131] In this example, each column of the matrix W is associated with an accumulation line AL and an analog-to-digital converter ADC.
[0132] At each line of the weight matrix, the outputs XW are brought back to the input of the corresponding logic sub-block LGW, to obtain a control signal and a polarization signal (not shown in [Fig. 14]). As explained above, the polarization signal makes it possible to control a switch SW which will activate the current source CS which is connected to it to charge or discharge the accumulation line AL in accordance with the time evolution of the control signal.
[0133] A first embodiment of current sources CS that can be used in the circuit of [Fig. 14] is illustrated in [Fig. 15].
[0134] The current sources CS are created using current mirrors. A current mirror composed of PMOS transistors 14 charges the accumulation line AL and another mirror composed of NMOS transistors 15 discharges the accumulation line AL. [Fig. 15] shows an accumulation line AL where the outputs of the current mirrors are connected to the accumulation line AL via switches SW. The transistors used can be of standard type, but when the number of transistors in a current mirror increases, thick oxide transistors can be used to reduce gate leakage.
[0135] As the voltage of the accumulation line AL can vary, it is possible to increase the output impedance of the current mirrors by using a cascoded architecture as illustrated in [Fig. 16]. Such an architecture makes it possible to have a more stable current value over the entire dynamic range [Vdd - 0], Vdd being the supply voltage of the circuit.
[0136] Since a current source must be activated and deactivated at several voltage levels of the accumulation line, an electronic switch SW of the transmission gate type ("pass gate switch" in English) is preferably used.
[0137] [Fig. 17] schematically represents an example of such an electronic switch SW comprising two transistors, an NMOS 17 and a PMOS 16, whose drains are connected together and the sources too. The signal applied to the gate of one of these transistors is the binary complement of that applied to the gate of the other.
[0138] This switch allows to obtain good performances along the dynamics.
[0139] The accumulation line AL comprises a metal line. Its capacitance is that of the metal line and of all the parasitic capacitances 18 of the electronic switches SW which are connected to the accumulation line, as schematically illustrated in [Fig. 18].
[0140] The line value is easily estimated in simulations and the current source value can be set to match the capacitance of the accumulation line.
[0141] The capacity of the accumulation line can be increased by adding additional capacitors to the output of the switches.
[0142] In an embodiment illustrated in [Fig. 19], a secondary accumulation line DAL is added to perform current steering. A follower operational amplifier 19 may be connected between the two accumulation lines AL and DAL so that each remains at the same voltage level in order to counteract the charge sharing effect when the switches SW open and close. [Fig.20] is similar to [Fig. 19] additionally showing switches connected to the accumulation lines.
[0143] To reduce charge sharing, the secondary accumulation line DAL can be manufactured with a lower capacity than the accumulation line AL, so that the charges are shared proportionally to the ratio between the capacities of the two lines.
[0144] The analog-to-digital converter ADC is connected either directly to the accumulation line AL in the absence of a secondary accumulation line DAL, or to the secondary accumulation line DAL in particular by means of a follower operational amplifier 20 as illustrated in [Fig. 19].
[0145] [Fig.21] schematically represents a first embodiment of an analog-to-digital converter ADC which can be used in the circuit of [Fig.14],
[0146] In this example, the analog-to-digital converter ADC comprises two analog comparators 22 and 23, one comparing to +LSB and the other to -LSB the value of the charge of the accumulation line 21 to which the analog-digital converter ADC is connected. The analog-digital converter also comprises a counter 24 configured to increment or decrement depending on the result of the comparison of the analog comparators 22 and 23, and whose output 30 is that of the converter. The analog-digital converter ADC also comprises a charge injector block 25 configured to return the charge of the accumulation line 21 connected to the converter to its initial value in the case where the counter indicates that this line has been discharged, and a discharger block 26 configured to discharge the accumulation line 21 connected to the converter and return it to its initial value in the case where the counter 24 indicates that it has been charged.
[0147] Figures 22 and 23 schematically illustrate two variants of the converter of [Fig.21] in which the converter includes a feedback loop 32 for calibration purposes.
[0148] Indeed, a feedback loop 32 can be introduced either to modify the value of an LSB as a function of the output value of the counter 24 (case illustrated in [Fig.22]), or to modify the unit time T in the time mask generator PG (case illustrated in [Fig.23]). In the case of [Fig.23], the feedback loop 32 can be used to modify, instead of the unit time T, a unit current I of an elementary current source, as will be explained later in relation to FIGS. 25 to 27.
[0149] In the variants of figures 22 and 23, the analog-digital converter ADC comprises a look-up table LUT receiving as input the output 30 of the counter 24 and being configured to apply as output a substitution value corresponding to an activation function or a correction for calibration purposes. A feedback loop 32 applies the output of the LUT to the values of the reference inputs +LSB and -LSB of the analog comparators 22 and 23 and / or is used to adjust the unit time and / or the unit current. In the example shown in [Fig.22], feedback loop 32 applies the output of the LUT lookup table to the +LSB and -LSB reference inputs of comparators 22 and 23.
[0150] In the example illustrated in [Fig.23], said loop 32 applies the output of the look-up table LUT to the time mask generator PG to modify the unit time T or to the reference current source to modify the reference current I.
[0151] The correspondence table can thus make it possible to calibrate the circuit by correcting non-linearities or to obtain a predefined non-linear response which can correspond to an activation function in artificial intelligence.
[0152] [Fig.24] schematically illustrates the architecture of a set of circuits for the carrying out MAC multiplication and accumulation operations in memory, in particular for a neural network, configured to determine the result of the matrix product of a vector [XB ..., Xm] by a weight matrix ([Wi.i,..., Wm.i],..., [Wi>n, •••- Wm>n]), comprising a plurality of circuits according to the invention placed in parallel with each other and sharing a single PG time mask generator block, located outside memory, and operating for example at 50 MHz.
[0153] [Fig.25] schematically represents a circuit according to the invention operating according to the current weighting approach.
[0154] In relation to the previous figures, we have seen that it was possible to use a current source CS by weighting the charge / discharge time of the accumulation line AL. The voltage across the accumulation line AL for a multiplication and accumulation (MAC) operation is equal to:
[0155] C being the value of the capacitance of the accumulation line and I the intensity of the current of the current source CS.
[0156] It is possible to obtain the same result R of scalar product by weighting the current and reducing the charge / discharge time.
[0157] According to the current weighting approach, there are several current sources (k being the number of current sources), weighted by a power of two and controlled by control signals SC representative of the values of the bits associated with this weighting.
[0158] This advantageously increases the calculation speed by reducing the duration of the control signals.
[0159] In the example illustrated in [Fig.25], there are two control signals SC0,j,i and SCijj respectively activating the current sources of intensities I and 21*!, I being the elementary reference current associated with the least significant bits of index 0 to (k-1). In this example, k=2, since there are only two control signals.
[0160] The generation of PX temporal masks is the same as in the case of temporal weighting seen previously, but the generation of PW temporal masks differs.
[0161] Indeed, if we take the example of a 4-bit multiplication (without the sign bit), the time masks PX and PW are represented in [Fig.26].
[0162] The calculation takes half the time as if it were a time weighting, the signals of PW(0) and PW(2) are identical, as well as PW(1) and PW(3). Indeed, the same signals are used, but on different current sources.
[0163] In the example of the timing diagrams shown in [Fig.26], the 4 bits are distributed equally between the two sources, so that the control signals are such that : [01641 SC 0JJ = (E^PW(d) (c) *X14 SC w = (E^PW( d - 2 ) * ^PX( C) %F])'
[0165] [Fig.27] is analogous to [Fig.24] representing a set of circuits according to the invention, with the current weighting variant, in the case of a 4-bit multiplication with equitable distribution of the bits between the 2 current sources.
[0166] The reference current for charging is denoted Iref P and the reference current for discharging is denoted Iref N, with Iref_P= Iref_N- The transistors of the corresponding current sources are sized to provide such a current intensity. The W / L ratio is related to the transconductance gm which is defined as the ratio between the variation of the drain current and the variation of the gate-source voltage. Thus, for a given gate-source voltage, a higher W / L ratio results in a higher current, W being the width of the conductive channel and L its length. The transistors of the higher weighting current sources have a 4W / L ratio.
[0167] [Fig.28] illustrates a second example of time mask timing diagrams that can be used for the circuit of [Fig.25]. In this example, the 4 bits are distributed unequally between the 2 sources, so that the control signals are such that: 101681sc^=(Ow( d ) ) *xld) 101691 SC = pw(o) ^14)
[0170] [Fig.29] is analogous to [Fig.27] representing a set of circuits according to the invention, with the current weighting variant, in the case of a 4-bit multiplication with unequal distribution of the bits between the 2 current sources.
[0171] The transistors of the higher weighting current sources have a ratio of 8W / L. Indeed, since the second current source is only controlled by the bit Wy([3] and since PW(3) is equivalent to PW(0), the current carries the weighting 23. It is therefore necessary to modify the size of the transistor by the same factor.
[0172] Among the applications of the invention, we can cite the optimization of calculations in preprocessors, perceptrons and classifiers.
[0173] The invention is not limited to the exemplary embodiments described above. For example, the logic block may be implemented with different logic gates, typically NOR gates instead of NAND gates. The frequency of the time mask generator block may be modified to meet the needs of the target application.
Claims
Claims
1. Method of performing multiplication and accumulation operations (MAC) in memory, in particular for a neural network, in which we determine the result of the matrix product O of a vector X = [ X{ X7 ... Xm ] composed of m elements each represented on ('Y- 1') bits and a matrix W = W composed of mxn elements each represented on (nw - 1) bits: O=[XiXï W', IV ” mji = [O,O2... O,J each element of the matrix product O being the sum of products scalars Em with i G {1,...,n}, j G {1,...,m}, with the bit vector \ Jv ph „v|nx - 2]] and the vector of weight bits ^=[^[0], ....WyK-a]]. the multiplication and accumulation operations leading to the scalar product Rji by applying to a set of logic gates the values of the vectors Xj and Wjj and those of associated time masks, the time masks PX — [PX( 0 ), ..., PX(nx - 2)] being associated with the bits of the elements of the vector X / and the time masks PW = [PW( 0 ), .... PW(nw - 2)] being associated with the bits of the elements of the vector so as to generate at least one control signal SC
2. of at least one current source, the control signal SC activating said at least one current source so as to accumulate capacitive charges in an accumulation line, in particular in a manner proportional to the scalar product Rjj. Method according to the preceding claim, at least one of the two vectors is signed, X (y / gn] being the sign bit of the vector X; and / or being the sign bit of the vector of bits of weight 1V,„ the set of logic gates generating a signal of polarity SP = X Wjcharging or discharging the accumulation line (AL).
3. Method according to one of the two preceding claims, the time masks PX and PW being generated in the form of pulses.
4. Method according to any one of the preceding claims, the time masks PX and PW being generated by a time mask generator block located outside memory.
5. Method according to any one of the preceding claims, the time masks PW (d) being generated so as to be activated successively for a duration Td substantially equal to 2^^-^9^7^ being a unit time corresponding to the activation duration of the time mask associated with the least significant bit of the vector Xj, the time masks PX^c) being generated so as to be activated successively for each Td for a duration substantially equal to 2c+d* T.
6. Method according to the preceding claim, the control signal SC controlling a single current intensity value in absolute value, by charging or discharging the accumulation line, such that: sc = SC,J=(EXT pwW (iXopx < 7 *xld)'with ' .....
7. Method according to any one of claims 1 to 4, at least two current sources (CS) being present, each being controlled by a respective control signal SC, the time masks PW(d) being generated so as to activate successively for a duration Td substantially equal to Td = with d £ {0, ..., max(^ -1^)}, the (nw - 1) bits of PW being distributed in k intervals [ aj, ], k being the number of current sources or control signals, with 1 e {0,..., k - 1}, h; > at, e [0, (nw-2) ], b / being chosen so that the set of intervals [ , b / ] represents a partition of the interval [0 , nw - 2], the longest interval being [0 , max( - a{ ) ] , T being a unit time corresponding to the activation duration of the time mask associated with the least significant bit of the vector Xj, the time masks PX(c) being generated so as to activate successively for each Td for a duration substantially equal to 2t+f / * T.
8. Method according to the preceding claim, the control signals being such that: =(Et, PW ( J - «à *^14 (Et2™ ( 1 " {°' ■ ki).
9. Method according to any one of the preceding claims, the voltage Vp across the accumulation line (AL) for a multiplication and accumulation operation being equal to: v -1222. c C being the value of the capacitance of the accumulation line (AL) and I the intensity of the current of the current source (CS) of the least significant (ie for 6l=0).
10. Circuit (1) for performing multiplication and accumulation (MAC) operations in memory, in particular for a neural network, configured to determine the result of the matrix product 0 of a vector X = [Xj X? ... XH;] composed of m elements each represented on (nx~ 1) bits and a matrix W = Wij ■■■ Kd ■■■ composed of axn elements each represented on (nw - 1) bits: '^Ll ^1 O=[XfX2... X„] X : : = [O,O2...O„] W,».! each element of the matrix product 0 being the sum of scalar products 0, = with = 1 ' {1,..., n], je {1,..., with the bit vector Xj = [X} [ 0 ], ..., Xj [ nx - 2] ] and the vector of weight bits Wj4 = [Wj4 [ 0 ], ..., Wjj [ nw - 2 ] pc circuit comprising: - a logic block (LG) comprising logic gates to which the values of the vectors Xj and Wjj and those of associated time masks, the time masks PX = [PX( 0 ), ..., PX{nx - 2)] being associated with the bits of the elements of the vector Xj and the time masks PW — [PW ( 0 ), ..., PW(nw - 2)] being associated with the bits of the elements of the vector W jj, the logic block being configured to generate at output at least one control signal SC intended to activate at least one source. current source (CS); and - said at least one current source (CS) configured to be activated by the control signal SC so as to accumulate capacitive charges in an accumulation line (AL).
11. Circuit according to the preceding claim, at least one of the two vectors being signed, being the sign bit of the vector Aj and / or Wy|i7gz?] being the sign bit of the vector of weight bits Wÿ, said accumulation line (AL) being configured to be charged or discharged according to a polarity signal SP = Xjszgfî] X Wj ^ / g / z] this signal SP being in particular obtained at the output of an XOR or XNOR logic gate receiving as input the sign bit Xjszgn] and the sign bit ^y^zg / z]-
12. Circuit according to one of the two preceding claims, the logic block (LG) comprising NAND and / or NOR logic gates.
13. Circuit according to the preceding claim, the logic block (LG) comprising a first logic sub-block (LGX) for the temporal masking of the bits A; and a second logic sub-block (LGW) for the temporal masking of the bits of weight Wy, these two sub-blocks being cascaded so that the outputs of the first sub-block are at the input of the second sub-block.
14. Circuit according to the preceding claim, the first sub-block (LGX) receiving as input the time masks PX and PW as well as the bits Xj(): nx - 2], and comprising two stages of logic gates: - a first stage (10) comprising several levels of logic gates, the levels being connected in cascade so that the signal or signals at the output of one level are the signal or signals at the input of the following level until reaching a single output signal, first stage in which each bit X, [0: nx - 2] and the associated time mask PX are at the input of a first level logic gate, in particular a NAND or NOR gate; and - a second stage (11) comprising at least one level of logic gates, in which the signal at the output of the first stage and each time mask PW are at the input of a logic gate of said level, in particular a NAND or NOR gate.
15. Circuit according to the preceding claim, the second sub-block (LGW) receiving as input the output signals of the first sub-block (10) and the bits of weight: Hw - 2] ct comprising several levels (12) of logic gates, the levels being connected in cascade so that the signal or signals at the output of one level are the signal or signals at the input of the following level until reaching a single output signal corresponding to the control signal SC, each first level logic gate receiving as input an output signal from the first sub-block (LGX) and the corresponding bit of weight: flw- 2].
16. Circuit according to the preceding claim in its attachment to claim 11, comprising an XOR or XNOR gate receiving as input the sign bit of the bit vector Xj and the sign bit of the bit vector W ÿ.
17. Circuit according to any one of claims 10 to 16, said at least one current source (CS) being created by current mirrors using CMOS transistors (14, 15).
18. Circuit according to the preceding claim, the current mirrors comprising PMOS transistors (14) for charging the accumulation line (AL) and NMOS transistors (15) for discharging said line.
19. A circuit according to claim 17 or 18, the current mirrors having a cascoded architecture.
20. Circuit according to any one of claims 10 to 19, comprising at least one electronic switch (SW), being in particular a transmission gate, controlled by the control signal SC for the activation of said at least one current source (CS).
21. Circuit according to any one of claims 10 to 20, comprising at least one capacitor at the output of said at least one electronic switch (SW) to increase the capacity of the accumulation line (AL).
22. A circuit according to any one of claims 10 to 21, comprising a secondary accumulation line (DAL) connected to the accumulation line (AL) via a follower amplifier (19).
23. Circuit according to the preceding claim, comprising a follower amplifier (19) connected between the accumulation line (AL) and the secondary accumulation line (DAL).
24. A circuit according to claim 22 or 23, comprising an analog-to-digital converter (ADC) connected either directly to the line accumulation (AL) in the absence of a secondary accumulation line (DAL), or to the secondary accumulation line in particular by means of a follower amplifier (20).
25. Circuit according to the preceding claim, the analog-to-digital converter (ADC) comprising: - two analog comparators (22, 23), one comparing to +LSB and the other to -LSB the value of the charge of the accumulation line (21) to which the analog-to-digital converter is connected; - a counter (24) configured to increment or decrement depending on the result of the comparison of the comparators, and whose output (30) is that of the converter; - a charge injector block (25) configured to return the charge of the accumulation line (21) connected to the converter to its initial value in the case where the counter indicates that this line has been discharged; and - a discharger block (26) configured to discharge the accumulation line (21) connected to the converter and return it to its initial value in the case where the counter (24) indicates that it has been charged.
26. Circuit according to the preceding claim, the analog-digital converter (ADC) comprising a look-up table (LUT) receiving as input the output (30) of the counter (24) and being configured to apply as output a substitution value chosen during calibration and / or corresponding to an activation function.
27. Circuit according to the preceding claim, in the case where the values of the look-up table (LUT) are chosen during calibration, the output of the look-up table acts via a feedback loop (32) on the values of the reference inputs +LSB and -LSB of the analog comparators (22, 24) and / or on the unit time (T) and / or on the unit current (I).
28. Circuit according to any one of claims 10 to 27, comprising a time mask generator block (PG) located outside memory.
29. Circuit according to the preceding claim, the generator block (PG) being configured to generate the time masks so that the time masks PW(d) are activated successively for a duration Td substantially equal to and that the time masks PX(c) are activated successively for each Td for a duration
30.
31. substantially equal to 2c+d* T, T being a unit time corresponding to the activation duration of the time mask associated with the bit of the least significant vector. Circuit according to claim 28, comprising at least two current sources (CS), each being controlled by a respective control signal SC, the generator block (PG) being configured to generate the time masks so that the time masks PW(d) are activated successively for a duration Td substantially equal to *7^' aveC of •••' max(^-«z)}, the (nw- 1) bits of PW being distributed into k intervals, fy], k being the number of current sources or control signals, with 1 e {0,..., k- 1], b( > ah { bt} e [0, ( / îw - 2) ], ai and being chosen such that the set of intervals b[] represents a partition of the interval [(), nw - 2], the longest interval being [0 , mâx(bj - a;)] , T being a unit time corresponding to the activation duration of the time mask associated with the least significant bit of the vector Xj, and that the time masks PX(c) are activated successively for each Td for a duration substantially equal to 2c+d* T. Set of circuits for carrying out multiplication and accumulation operations (MAC) in memory, especially for a neural network, configured to determine the result of the matrix product of a vector X — [.Y, À^... Xm] by a weight matrix W= r^Lt comprising a plurality of circuits (1) according to one W , ■■■ W mi ; any of claims 10 to 30 placed in parallel with each other and sharing a single generator block (PG) of time masks located outside memory.