Binary neural network
The circuit optimizes binary neural network operations in near-memory architectures by using two-value arithmetic and masking, addressing performance and energy challenges, thereby improving efficiency.
Patent Information
- Application Number
- FR2024004299
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2025-10-31
AI Technical Summary
Existing binary neural networks implemented in near-memory architectures face challenges in terms of performance, energy consumption, and area efficiency due to the computational expense of operators like scalar products and gating mechanisms, particularly when using real number spaces and point-by-point multiplications.
A circuit design that includes memory elements, computing circuits, and logic gates to perform binary arithmetic operations efficiently, utilizing two-value arithmetic functions and masking operations to enhance performance and reduce energy consumption, while maintaining flexibility in arithmetic configurations across network layers.
The proposed circuit improves the performance and reduces energy consumption of binary neural networks by optimizing binary arithmetic operations and masking mechanisms, enhancing the efficiency of near-memory implementations.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Binary neural network technical field
[0001] The present description relates generally to circuits configured to execute binary neural networks, and in particular to the near-memory implementation of such networks. Previous technique
[0002] Operators, such as scalar products or Hadamard products, are generally involved in the operation of binary neural networks. These operators are, for example, used on each layer of the network, and the operator outputs are binarized or requantized.
[0003] When the neural network is implemented in hardware using a so-called "near-memory" architecture, the operators are placed directly at the memory output, for example, at the edge of a SRAM (Static Random Access Memory) memory tile. To save space, it is desirable to increase the activity rate of these operators during network execution. Similarly, the arithmetic used—that is, the definition of the useful number space, its relationships and properties, and the elementary mathematical operations that can be performed—is the same for all layers of the network.
[0004] Furthermore, some networks incorporate gating mechanisms to promote the emergence of contextualization of the processing carried out within the network (attention-like mechanism). These mechanisms are generally implemented via activation functions working from the real number space to the real number space, such as, for example, softmax and / or sigmoid functions, followed by a point-by-point multiplication stage, which is computationally expensive in terms of area.
[0005] There is a need to improve the architectures of near-memory binary neural networks, particularly in terms of performance, energy consumption and area. Summary of the invention
[0006] One embodiment provides a circuit comprising: - a first memory element configured to store a first piece of data; - a second memory element configured to store a first weight matrix in association with a first layer of an artificial binary neural network; - a computing circuit configured to: a) receive the first data and the k-th row, of the first weight matrix; b) receive a first control signal, indicating the nature of each of the first and second reading functions, from among at least the first and second reference functions, the first and second reference functions each being associated with a two-value arithmetic; c) generate a first vector by applying the first reading function to the k-th row of the first weight matrix and a second vector by applying the second reading function to the first data point; and d) generate the same component of a first output vector based on the first and second vectors.
[0007] According to one embodiment, the first reference function has values in {_ 1,1| and the second reference function has values in ]}.
[0008] According to one embodiment, the above circuit is configured to, following the generation of the same component of the first output vector, control the storage of the same component in the first memory element.
[0009] According to one embodiment, the above circuit further comprises: - a first logic gate configured to apply an AND operation between the same component of the first output vector, transmitted by the calculation circuit, and a component of a binary vector; - a second logic gate configured to apply an XNOR type operation between the same component of the first vector, transmitted by the calculation circuit, and the component of the binary vector; - a first multiplexer configured to select the output of the first logic gate, or the output of the second logic gate, based on a control signal, indicating the nature of a third read function among the first and second reference functions; - a second multiplexer configured to generate the same component of a second output vector by selecting the same component of the first vector, provided by the calculation circuit, or the output of the first multiplexer, based on a modulation signal.
[0010] According to one embodiment, the above circuit further includes a scheduler circuit configured to: - receive a component of a vector stored in the first memory element; and - based on the value of the component, command the reading, in the second memory element, of the Mth line of the first weight matrix.
[0011] According to one embodiment, the calculation circuit comprises: - a number N of multiplier circuits, each configured to receive the first and second read functions and each configured to generate a scalar of output based on a component of a vector and a component of a weight matrix; - an accumulator circuit configured to sum the output scalars provided by the plurality of multiplying circuits; and - a converter configured to convert the value generated by the accumulator circuit into a binary value.
[0012] According to one embodiment, the accumulator includes an adder tree, comprising a plurality of shift circuits and configured to generate a scalar, corresponding to the dot product between the first and second vectors, augmented by a power gain of 2.
[0013] According to one embodiment, the first data is of length being an integer, and is stored contiguously by vectors of length K being a divisor of the value N, and in which the ^-th row, of the first matrix is of length N and in which the calculation circuit comprises an L-NtK number of shift registers connected to the first memory element and is configured to, following the receipt of a sequence of vectors of the first data, convert said first data into a vector of size N, by concatenation of vectors of size K, in which the shift registers are, for example, further configured to perform the concatenation of the first data of size K with a sequence of NK bits each equal to 1, or to 0.
[0014] According to one embodiment, each of the multiplier circuits includes an NXOR and / or AND type logic gate configured to multiply a component of the first data and an element of the first weight matrix associated with the first layer.
[0015] According to one embodiment, the first memory element is further configured to store a masking vector, the scheduler circuit being configured to: - if the fc-th component of the masking vector is equal to 0, command the storage of the ^-th component of the first data as the ^-th component of the first output vector; and - if the fc-th component of the masking vector is equal to 0, command the execution of steps a) to c).
[0016] According to one embodiment, the first memory element is a memory configured for the implementation of masking operations.
[0017] According to one embodiment, the second memory element further stores a third weight matrix associated with a second layer of the neural network, and in which the computing circuit is further configured to: - select, following the receipt of a second control signal, from the fifth and sixth reading functions, among the first and second reference functions; - generate a fifth vector by applying the fifth read function to a row, or column, of the third weight matrix and a sixth vector by applying the sixth read function to an output vector from a previous layer, stored in the first memory element; - generate a third output value based on the fifth and sixth vectors.
[0018] According to one embodiment, the first reference function is defined by g -.u-^u |a second reference function is defined by h : u -* lu -1.
[0019] One embodiment provides a method comprising: - the provision of a first piece of data stored in a first memory element of a circuit and the ^-th line of a first weight matrix associated with a layer of a binary artificial neural network to a first calculation circuit of the circuit; - the provision of an indication, by means of a control signal, of the nature of a first and a second reading function, from among a first and a second reference function; - the generation, by the first calculation circuit, of a first vector by applying the first function to the k-th row of the weight matrix and of a second vector by applying the second function to the first data point, - the generation of a k-th component of a first output vector based on the first and second vectors, the first and second reference functions each associated with two-valued arithmetic.
[0020] According to one embodiment, the first computing circuit comprises a plurality of multiplier circuits and an accumulator configured to generate a scalar by performing a dot product between the first and second vectors, wherein the first computing circuit comprises, for example, a converter configured to convert the scalar into a binary value.
[0021] According to one embodiment, the above process further comprises: - reading, by a second calculation circuit, the value of a first component of the first data, the first output value, and a first masking value, stored in the first memory element; and - the command, based on the first masking value and by the second calculation circuit, to write a first masked value, in the first memory element, corresponding either to the first component of the first data, or to the k-th component of the first output vector.
[0022] According to one embodiment, the writing of the first masked value comprises: - the removal of the first component of the first data point; and - writing the first masked value to the address of the first component of the first data in the first memory element.
[0023] According to one embodiment, the above process further comprises, after writing the first masked value: - the generation of a k+l-th output component of the first output vector and its storage in the first memory element; - the reading, by the second calculation circuit, of the value of a k+l-th component of the first data, of the k+l-th component of the first output vector and of a second masking value, stored in the first memory element; and - the command, on the basis of the second masking value and by the second calculation circuit, of the writing of a second masked value, in the first memory element, corresponding either to the k+l-th component of the first data, or to the k+l-th component of the output vector. Brief description of the drawings
[0024] These features and advantages, as well as others, will be described in detail in the following description of particular embodiments, given by way of non-limiting example, in relation to the accompanying figures, among which:
[0025] [Fig.1] is a block diagram illustrating a rebinarized scalar product in a layer of a neural network;
[0026] [Fig.2] is a block diagram illustrating an example of a near-memory architecture of a fully connected network, according to an embodiment of the present description;
[0027] [Fig.3] is a diagram illustrating an example of implementation of a requantized dot product operator;
[0028] [Fig.4A] is a diagram illustrating an example of an adding tree, according to an embodiment of the present description;
[0029] [Fig.4B] is a diagram illustrating an example of a programmable bit-shifting circuit of the adder tree of [Fig.4A], according to an embodiment of the present description;
[0030] [Fig.5] is a block diagram illustrating a rebinarized scalar product in a layer of a neural network followed by a modulation operation;
[0031] [Fig.6] is a block diagram illustrating an example of a fully connected one-layer circuit integrating a modulation mechanism on an output vector, according to an embodiment of the present description;
[0032] [Fig.7A] schematically illustrates a sequence of operations enabling the realization of a binary masking mechanism;
[0033] Figure 7B schematically illustrates another example of a sequence of read-write operations for a memory circuit enabling a masking mechanism; and
[0034] [Fig.8] is a block diagram illustrating an example of a fully connected one-layer circuit integrating a modulation mechanism on a vector from an intermediate calculation and enabling a masking mechanism to be implemented by a conditional update of the output vector, according to an embodiment of the present description. Description of the implementation methods
[0035] The same elements have been designated by the same reference numerals in the different figures. In particular, the structural and / or functional elements common to the different embodiments may have the same reference numerals and may have identical structural, dimensional and material properties.
[0036] For the sake of clarity, only the steps and elements useful for understanding the described embodiments have been shown and are detailed. In particular, the operation and implementation of artificial neural networks, and especially binary neural networks, is known to those skilled in the art and is not described in detail.
[0037] Unless otherwise specified, when referring to two elements connected together, this means directly connected without intermediate elements other than conductors, and when referring to two elements connected (in English "coupled") together, this means that these two elements can be connected or linked through one or more other elements.
[0038] In the following description, when reference is made to absolute position qualifiers, such as the terms "front", "back", "top", "bottom", "left", "right", etc., or relative position qualifiers, such as the terms "above", "below", "superior", "inferior", etc., or to orientation qualifiers, such as the terms "horizontal", "vertical", etc., reference is made, unless otherwise specified, to the orientation of the figures.
[0039] Unless otherwise specified, the expressions "approximately", "roughly", and "in the order of" mean within 10%, preferably within 5%.
[0040] The [Fig. 1] is a block diagram illustrating a rebinarized scalar product in a layer of a neural network.
[0041] By way of example, a vector e — (e[1], e[2], , e[A / ]), of size N, where N is an integer greater than or equal to 1, is an input vector for a layer of the neural network. Each component e[n], n ∈ [1, , N], is a binary value, equal to 0 or 1. Depending on the type of binary arithmetic chosen, the binary values 0 and 1 respectively quantify either values equal to 0 and 1, or values equal to -1 and 1, or two other values for example equal to 0 and 2 or -2 and 2, etc.
[0042] By way of example, a layer operation allows the generation of a layer output vector y = ( J'[ 1], j[ 2], , jfA'] ), of size where N is a multiple of K. Each component & E- (f . K}^ of the layer output vector is then equal to the scalar product 100 between the input vector e and a row VFy[k] of a weight matrix VKj binarized, for example via an activation function 102.
[0043] By way of example, the dot product 100 is a component j[&] defined as the sum of the pointwise multiplication of the vector e with the vector WjW. In other words, j _ jy? [ ] [ n ] e [ nj, °ù Vf\ [ £ ] [ n ] is the coefficient on the k-th row, n-th column of the matrix Wy. By way of example, each weight of the weight matrix Wv is a binary value, the value of which 1 quantifies a value equal to 1, and the value 0 quantifies a value equal to 0 or, for example, -1, etc., depending on the arithmetic used.
[0044] The component y [A:] is then obtained, for example, by applying a binarization function b to the value y [k], so as to transform the value of j [k] into a binary value. In other words, y [k] — (J' [A]) • As an example, the function b is defined by:
[0045] . , , [ 1 sz x > 0 " l0«x < 0'
[0046] In another example, a bias value is added to the dot product 100, and in this case, the value supplied to the function b is equal to 3W [«]«[„] + Fwhere BiasW is a bias value for the k-th component. As an example, for all k E [1, , K], Biais[k] is a constant value, not depending on the value of the index k,
[0047] By way of example, the binarization function b is implemented in hardware by a comparator, or by an ADC-lb converter (from the English "Single bit Analog to Digital Converter").
[0048] Each component y[&] of the layer output vector is therefore a binary value, belonging to A. For example, depending on the arithmetic considered, the value 0 quantifies a value equal to 0, or a value equal, for example, to -1.
[0049] According to one embodiment, in order to take into account the arithmetic(s) considered, the dot product 100 is calculated from read functions. By way of example, the read functions are functions that take as input a binary value, equal to 0 or 1, and are configured to output a value belonging, by for example, to the set [q. 1} or to the set { _ 1, [}. In other words, the functions of
[0050] reading allows the transformation of the binary value encoding an input, weight, or output value into a value in another representation. As an example, the read functions are functions between a function s of sign, defined by g ■ u 2m -1, and a function * called of Heaviside type and defined by h : u u. Thus, the use of read functions allows us to have arithmetic in {- 1, 1} for the function s and / or in {0. ]} for the function. In the following description, the function X^O will represent the read function applied to the matrix J y of weight of the current layer for calculating the output y, and the quantity fiW'( ) represents a matrix, of the same size as the weight matrix Wy, and, for all ,K], f^ÇWy) [&] [m] = f'W\ [w] ) ' Similarly,' the function yW will represent the reading function applied to the input vector € of the current layer for calculating the output, and the quantity ^Çg) represents a vector, of the same size as the vector e, and, for all n G {1, , A7}, ( g ) [ « ] — ( ¢[^] ) • For each layer, and Therefore, for each layer output y, each of the functions is, in this example, either equal to the function £ or to the function h. Thus, the reading functions X^D and f allow the dot products on each layer to take place, by J p J y for example, in the sets {Ojfx {-l,lf' {-IJ^x {0,l}'V and / or (-1 fx {-1 1} V' These sets are given as examples; other arithmetics are of course conceivable, and in this case the read functions will need to be adapted by a person skilled in the art to take values in the desired sets. In the following description, for any vector u appearing in a layer operation, read functions / W and Xe' are read functions among the J uu functions £ and h used in the calculation of the vector u. The dot product between the two vectors yj) ( e} and yW ( ) is then denoted ( ey ( w).
[0051] According to one embodiment, considering that each of the components of the input vector as well as each of the coefficients of the weight matrix VFy are binary values, in {q, 1], the layer output vector y is such that y _ ' °ù f(W\ Wy) f^e>(e ) is a vector of size K, of which each component is equal to the dot product between the vector e and a line of the Wy weight matrix. The output vector is therefore also binary in
[0052] By way of example, the output vector J' of a layer for the current layer is an input vector for a subsequent layer. For example, for this subsequent layer, the arithmetic representing the input vector is different from that representing the input vector e for the current layer. Similarly, in one example, the weight matrix for this subsequent layer has a different representation than that of the matrix of the current layer. In another example, the representations of the vector ε, and / or the weight matrix for this subsequent layer, are the same as the representations of the vector e and / or the matrix Wy for the current layer. The reading functions used from one layer to another can therefore vary.
[0053]
[0054]
[0055] In one embodiment, one or more layers of the neural network are configured to apply an additional masking operation (also known as "gating" or "attention-like"). The masking operation, for example, hides certain components of the output vector from layer 3. A masked vector is then provided as input to the next layer of the network. Masking operations allow for the selection and / or modulation of outputs based on context or prior classification. According to one embodiment, an output vector 4, obtained by masking the output vector J, is defined by: 5 = 2'x ' A where the operator " " represents a Hadamard product, or product point by point, and where the vector x with values in is derived from the input vector e, and / or from internal and / or intermediate states calculated previously in the current layer, or in a previous layer. In particular, the operator " " represents a modulation operation. The vector ' is a binary vector, with values in , and is a masking vector. The vector Z corresponds to the vector which, when summed with the vector ', results in a vector composed only of 1s. Thus, the vector 4 includes some of its components from the vector x and others from its components from the output vector J7. The aforementioned masking operation is a "multiplexing" type masking operation. Other so-called masking operations can be performed by applying another function, for example by performing a modulation operation between two vectors with a vector exhibiting "on / off" properties, for example with values of "0". According to one embodiment, 4 is such that s = z • x + zy^ where, for example, W(e) being a weight matrix, x_(IV)(e))' being another weight matrix, and where y is the output vector described previously, or J = e. The functions and / ri are reading functions, respectively associated with the vectors z and x, and X for example, equal to the function and / or to the function A. Thus defined, the vectors z, x, and y are derived from projections of the vector e, or are directly equal to the vector e. The
[0056]
[0057]
[0058]
[0059] The masking thus performed allows us to adapt and take into account values obtained by projecting the past, the past being defined by the layer input vector. As an example, masking operations are generally performed in so-called recurrent neural networks (RNNs) where feedback is applied in the calculation of the output vector s. Indeed, when processing a data vector, at each index z of this vector, the equations involved in recurrent neural networks generally use versions prior to the vectors calculated at that time—that is, for example, vectors sm and htA that were calculated at time t-1. This type of operation allows, in particular, the memorization of past states. Recurrent structures are, for example, analogous to infinite impulse response (IIR) filtering. According to one embodiment, a vector dt in Jq ^j2^ is defined as the concatenation of an input vector and in {g of a layer of an RNN and a vector ht_} in {g, i.e., dt = concat ( et, ) and a vector dt in |g is defined as the concatenation of the input vector and in {g and an output vector in n* of a layer of the RNN, i.e., d't = concat(e(, stA). Several Variants of recurrent structures can be defined, from reading functions as described previously. As an example, the following table presents the different vectors involved in layer operations of a long short-term memory (LSTM) network. In particular, each vector (r, o, y, xs, h) has a binary value and is associated with functions reading that allows its description in another arithmetic. To do this, each The matrix in the following table has a binary value in 2W and is associated with reading functions allowing its description in another arithmetic. [Tables 1] zr=b( / H) (w / j (d,)) r <= b( / 7' (Wz) / / ' (d,)} y^b^cWyy^td,)} x^Hf^y,) s t = z t -x t + z t -s tA In this example, each reading function can be chosen from the functions g and 2i0 implementations are possible.
[0060] The following table presents the different vectors involved in layer operations of an LSTM type network.
[0061] [Tables2] z,=b(^ (wj (d,)) r<=b(^ (wj / ;* (d,)) ut = concat({lA, , 1};rt) y=b(fy(dtYfy(utY) ^b^W,)^)) st = zt-xt + zt-sf_} ht = b(fh(st) In this example, each reading function can be chosen from the functions £ and h, so 210 implementations are possible.
[0062] The following table presents the different vectors involved in layer operations of a Gated Recurrent Unit (GRU) network in a canonical variant.
[0063] [Tables3] z, = b(^ (WJ (d'j) (W^) Y ( d 'Y ut = concat ({1,1, , 1}; rt ) yt = h(fy(d't)-fy(ut)) x, = b(j^Wx) / ^yi)} st = zt-xt + zt-stA In this example, each reading function can be chosen from the functions £ and b, so 27 implementations are possible.
[0064] The following table presents the different vectors involved in layer operations of a minimal gated unit (MGU) network in a canonical variant.
[0065] [Tables4] (wa / ” n( = œnœ / ({l,l, , l};zj y=b(fym-fy(ut)) S^Zt-Xt + Zf-S^
[0066] In this example, each reading function can be chosen from the functions and A, so 25 implementations are possible.
[0067] Each of Tables 1 to 4 above presents examples of algorithms in which the vectors st and / or ht are obtained by calculating several dot products between intermediate vectors. The calculations of the vectors st and / or ht are, for example, performed in several steps. Initial steps include the generation, for example sequentially, of one or more intermediate vectors. In some examples, intermediate vectors, such as the vector , are obtained by modulation, denoted by the operator " " in Tables 2 to 4. In some examples, vectors such as the vector \ are obtained by masking operations in Tables 1 to 4.
[0068] Figure 2 is a block diagram illustrating an example of a near-memory architecture of a fully connected network, according to an embodiment of the present description. In particular, Figure 2 represents a circuit 200 comprising two memory elements 202 (DATA) and 204 (WEIGHTS).
[0069] By way of example, memory elements 202 and 204 are two separate memories of circuit 200. By way of example, memory element 202 is configured to store data of length, or in other words, to store a word of K bits. In particular, memory element 202 is configured to store data computed during layer operations, such as, for example, the input vector e, the vectors x, y, s, h, u, r, 0, and any other vector of length K, or a multiple of K, involved in a layer operation. The length of these vectors stored in memory element 202 is preferably a multiple of K, such that N = L * K, where L is a positive integer. Advantageously, a length K corresponding to the size of a smallest vector among those manipulated during the operations will be chosen, as will be explained in more detail below.
[0070] By way of example, during the execution of the neural network, the data stored in memory element 202 are dynamically deleted and written. For example, an input vector of layer e is deleted following the calculation of an output vector y of that same layer, and the vector y is stored in memory 202 as an input vector for the next layer. Similarly, each vector x, y, s, h, u, r, z, 0 calculated during an operation for a layer of the network is, for example, replaced by its new version for the next layer.
[0071] By way of example, modulation and / or masking operations are carried out live and dynamically, that is to say without prior storage of one and / or the other of the intermediate vectors involved in the operation.
[0072] By way of example, memory element 204 has a width of N. Memory element 204 is configured, for example, to store one or more weight matrices, such as the matrices Wy, Wz, Wr, Wo, Wx, etc., for each computation stage and for each layer of the neural network. The matrices stored in memory element 204 have rows with a length preferably equal to N, or 2N depending on the example, so that the entire row of the weight matrix stored in memory 204 can be obtained directly. In other cases, it will also be possible to use a matrix with a row size that is a submultiple of N, and in this case, as with the data, an accumulator register and sequential reading of the different parts of a row of the matrix will be necessary to reconstruct a word corresponding to all the bits of a matrix row.
[0073] The circuit 200 further includes a computing circuit 206 configured, for example, to generate the output vector components T of the layer, from the data contained in the memory 202 and from one or more matrices contained in the memory 204. The output vector of the layer is then a binary vector with value in {g .
[0074] By way of example, the calculation circuit 206 includes a circuit 208 (HADAMARD) configured to perform an operation such as a Hadamard product, from a vector included in memory 202 and a matrix, or a row or column of a matrix, included in memory 204.
[0075] According to one embodiment, the circuit is further configured to receive, for example via a 2-bit signal and during the execution of each layer, information indicating which read functions are to be applied to the data and weights before applying the dot product operation. For example, for the same layer, the intermediate layer operations, resulting in the output vector y, call upon different read functions, for example, the functions f'W' Ae) ri111 etc. The read functions used during the operations may differ from one another. For example, the functions Âe\ JO Àè) and Ae) are not all identical. J z J y
[0076] According to one embodiment, in the case where the integer N is strictly greater than the integer K and is an integer multiple of the value, the computing circuit 206 further comprises one or more shift registers 210 (L-BITSHIFT REG). In particular, considering the integer L, such that A = L x K, the computing circuit 206 comprises L shift registers, each shift register being configured to shift the received data by bits. The L shift registers allow for the rearrangement of non-contiguous data in memory 202 in order to provide the circuit 208 with data of dimension A appropriate to the operation performed with the weight matrix(ies). The L shift registers also allow for the manipulation of tensors in the case of two-dimensional (2D) convolutional layers.The register(s) 210 also allow for concatenations, for example, generating vectors e of length A resulting from the concatenation of several lines of memory 202, each line comprising the data from K channels of a data tensor at a spatial position, or pixel, of an image I of dimension H x VF x K, where H and VF represent the spatial dimensions respectively, vertical for the value H and horizontal for the value VF. Adapting memory 202 read schemes to convolutions of a dimension other than two is within the capabilities of a person skilled in the art.
[0077] The circuit 208 includes, for example, A number of XNOR gates, configured to perform the operation when the arithmetics of the weight vector and matrix are both in [-fj}. In other words, the A XNOR gates are configured to perform the operation when the indication received by the circuit 208 indicates that the read functions are both the function. The circuit 208 further includes, for example, one or more AND gates configured to hide inputs when one of the two arithmetics is in {0,1}. In other words, one or more AND gates allow, in combination with the N XNOR gates, the operations to be performed when at least one of the read functions is the Heaviside function h.
[0078] The calculation circuit 206 further includes an accumulator 212 (ACCUMULATOR), connected to the circuit 208 and configured to add the results of point-by-point multiplications performed by the circuit 208. According to one embodiment, the accumulator 212 is a signed adder tree.
[0079] The calculation circuit 206 further includes a converter 214 (1-bit converter) configured to binarize the value generated by the accumulator 212. In the following description, the term "binarization" refers to a quantization operation of a signal into two levels, represented by the values "0" and "1". For example, converter 214 is a comparator configured to receive digital values as input. In another example, converter 214 is an analog-to-digital converter (ADC) configured to receive analog values as input. For example, converter 214 is configured to apply the binarization function b to the value generated by the accumulator 212. Converter 214 is, for example, further configured to write the binarized value to memory 202.
[0080] According to one embodiment, the calculation circuit 206 is configured to perform operations, such as dot products, on binary values while respecting a configuration in a given arithmetic. The arithmetic is not fixed and is indicated to the circuit 206 during the execution of each layer operation. In particular, the arithmetic can vary between each layer operation. For example, a first dot product is then performed in the set { r * (g) and a second dot product is then performed in the set {g |g jp', etc. In a case of nominal use, each layer can use its own arithmetic and will be parameterized only by and J v J y
[0081] According to one embodiment, the circuit 200 comprises, or is connected to, a scheduler circuit 216 (SCHEDULER) configured to control the execution of layer operations such as, for example, those described in relation to any one of Tables 1 through 4. The scheduler circuit 216 is further configured to control memory access for reading and writing. In particular, the scheduler circuit 216 is configured to control the writing, for example sequentially, of each of the intermediate vectors, and of the x and / or T vectors, to memory 202. By way of example, the scheduler circuit 216 is a coprocessor in a system encompassing the circuit 200 and the scheduler circuit 216. The circuit The scheduler 216 also allows the sequencing of the different layer operations performed by the different elements of circuit 200.
[0082] In operation, the scheduler circuit 216 is configured, for example, to command the calculation circuit 206 to perform several calculation rounds and thus generate the vectors corresponding to the processing to be carried out, such as one or more of the processing operations shown in Tables 1 to 4 below. The vectors generated in each round are, for example, stored in memory 202, and at the end of the processing, the output vector -1' is, for example, available in this memory 202.
[0083] Figure 3 is a diagram illustrating an example of a circuit 300 implementing a requantized dot product operator. In particular, the circuit 300 illustrates an example of an embodiment of the circuit 208, the accumulator 212, and the converter 214 of Figure 2 in the form of a calculation performed on two channels. These two channels are separated, for example, so as to be able to perform calculations regardless of the arithmetic used, that is, regardless of the read functions AA and J. AA) used. For example, circuit 208 includes a number N of sub-circuits 208_l y to 208_N.
[0084] In the example illustrated in Figure 3, the circuit 300 is configured to generate a component J[k] of the output vector based on an input vector e stored in memory 202 and the k-th row of a weight matrix Wy, stored in memory 204, and the indication of two read functions AA and AM. In particular, the circuit 300 is configured to perform the operation J y J y
[0085] Each circuit 208_i, i ? ^y], is configured to receive the i-th component e[i] of the input vector e and the element. These values are, for example supplied to 302_i and 304_i multiplexers. As an example, the 302_i and 304_i multiplexers are configured to select one configuration, from configurations 00, 01, 10, and 11, based on the indication of the AA and A^XA read functions. J y J y As an example, configuration 00 represents the configuration in which j^A _ and j-AM _ configuration 01 represents the configuration in which j^A _ and / W _ „, configuration 10 represents the configuration in which AA _ and J yo J y & yW _ and configuration 11 represents the configuration in which j^A = g and yW_g. The values e[i] and JE?[k][i] are further provided to a 305_i XOR gate configured to perform the XOR Wy[£] [t] operation as well as an AND 306_i gate configured to perform the ¢[ / ] AND [ / ] operation.
[0086] By way of example, in configuration 00, multiplexer 302_i is configured to select a value equal to 1 and multiplexer 304_i is configured to select the output of AND gate 306_i. In configuration 01, multiplexer 302_i is, for example, configured to select the value Wv[A:][i] and multiplexer 304_i is, for example, configured to select the component e[i]. In configuration 10, multiplexer 302_i is, for example, configured to select the component e[i] and multiplexer 304_i is, for example, configured to select the element Wy[k][i]. In configuration 11, the 302_i multiplexer is, for example, configured to select a value equal to 1 and the 304_i multiplexer is, for example, configured to select the output of the 305_i XOR gate.
[0087] By way of example, the multiplexer 302_i is configured to transmit the selected value to one input of an AND gate 307_i and to one input of an AND gate 308_i. The multiplexer 304_i is, for example, configured to transmit the selected value to the other input of gate 308_i and to the input of an inverter 310_i. The inverter 310_i is then configured to invert the value, that is, to generate the value 0 when the supplied value is equal to 1 and vice versa, and to supply it to the other input of gate 307_i.
[0088] Gate 307_i is then configured to apply the AND operation to the values transmitted by multiplexer 302_i and inverter 310_i. Gate 308_i is then configured to apply the AND operation to the values transmitted by multiplexers 302_i and 304_i.
[0089] Each of the gates 307_i and 308_i of each of the subcircuits 208_l to 208_N is configured to provide the generated value to the accumulator 212. In particular, the gates 307_l to 307_N are configured to provide the generated values to an adder circuit 312 and the gates 308_l to 308_N are configured to provide the generated values to an adder circuit 314. The adders 312 and 314 are respectively configured to generate a value v2, vi, corresponding, for example, to the sum of the values provided by the gates 307_l to 307_N and 308_l to 308_N.
[0090] Adders 312 and 314 are then configured to provide the values vi and v2 to comparator 214. As an example, comparator 214 is configured to generate the binary value 1 if the value vi is greater than the value v2, and to generate the binary value 0 otherwise. The generated binary value corresponds to the
[0091] Figure 4A is a diagram illustrating an example of an adder tree 400, according to an embodiment of the present description. In particular, the adder tree dot product 400 is an example of an implementation of the adder circuits 312 and 314. According to one embodiment, the adder tree 400 comprises a log?(N) number of stages, each stage comprising one or more programmable bit shift circuits 402 and the same number of adders 404.
[0092] By way of example, the 400 adder tree is configured to receive an N-bit word (INPUT) and to generate an output value (OUTPUT) based on a vector c = (c[0], c[1], , c[log2(N) - 1]) of length log2(N). The word received by the 400 circuit corresponds to the N bits transmitted by the 208_1 to 208_N circuits. By way of example, in a first stage, the 400 adder tree includes A / 2 of 402-bit programmable shift circuits. By way of example, the 400 adder is configured to supply each odd-numbered bit, or even-numbered bit, to a 402 shift circuit. In other words, every other bit of the input word is supplied to a 402 shift circuit. In this first stage, each 402 shift circuit therefore receives one bit as input.
[0093] Fig. 4B is a diagram illustrating an example of a 402 programmable bit shift circuit, according to an embodiment of the present description.
[0094] Each 402 shift circuit is configured to receive a word comprising xp bits, for example, a number p of bits, where p is an integer greater than or equal to 1. The 402 circuit includes a 405 (0-PADDING) circuit configured to add one bit to the xp word. As an example, the 405 circuit is configured to generate a word of p+1 bits whose most significant bit is set to 0 and whose p least significant bits correspond to the xp word. As an example, the 405 circuit is configured to apply the so-called "0 Maximum Padding" method to the received word.
[0095] Each 402 shift circuit includes, for example, a 406 shift register configured to generate a word *p, corresponding to the multiplication by two of an unsigned integer represented in so-called "big-endian" binary by the word xp, shifting the input word xp by one bit and setting the least significant bit to 0. Each 402 shift circuit further includes, for example, a 408 multiplexer. The 408 multiplexer is, for example, configured to select a word from the word xp and the word based on a one-bit value c[z'] of the word c at index '. For example, if the bit c[z'] has a value of 0, the 408 multiplexer is configured to select the word xp and if the bit c [z] has a value of 1, the 408 multiplexer is configured to select the word Xp.
[0096] By way of example, the 402 shift circuits of the first stage of the 400 tree in Figure 4A are thus configured to generate 2-bit words. Each 402 shift circuit is further configured to provide the generated word to an adder 404. By way of example, each adder 404 of the first stage of the 400 tree is in Furthermore, it is configured to receive a bit from the input word that was not provided to a 402 shift circuit. For example, for each bit of position n in the input word provided to a 402 shift circuit, the bit of position n - 1 is provided to the 404 adder receiving the word generated by said shift circuit, starting from the bit of position n. Each 404 adder is configured to generate a word by adding the two received values. For example, the word generated by a 404 adder in the first stage of the 400 tree is of length 2. Indeed, the maximum attainable value on the first stage, in integer representation, is 2 + 1 = 3 and can be represented in binary by the word 11, which can therefore be encoded on 2 bits. From the next stage onward, each 404 adder is configured to generate a word containing one additional bit beyond the one provided by the 402 shift circuit.Indeed, on each of these stages, the maximum achievable value exceeds the dynamic range associated with the number of bits in the word provided by the 402 shift circuit.
[0097] As an example, following the processing of the input word by the first stage of the 400 tree, N / 2 2-bit words are provided to a second stage of the 400 tree.
[0098] The operation of the second stage of the 400 tree is, for example, identical to that of the first stage. Each 404 adder of the second stage is then configured to transmit the generated word to, depending on its position in the 400 tree, a 402 shift circuit of the second stage or a 404 adder of the second stage.
[0099] The last stage of the 400 tree, that is, the log-n stage, then comprises a single 402 shift circuit and a single 404 adder. This last 404 adder is configured to generate an output word (OUTPUT) from the 400 tree, comprising, for example, 21og (2V) bits. As with the preceding stages, the 404 adder of the last stage is configured to generate the output word by adding the word generated by the 402 shift circuit of the last stage with a value provided by one of the two 404 adders on the penultimate stage. In particular, the penultimate stage comprising two 404 adders, one is configured to provide the word it generates to the 402 shift circuit of the last stage, and the other is configured to provide the word it generates to the 404 adder of the last stage. The 402 shift circuit of the last stage is then configured to perform a selection based on the value of the bit c log^ ( N ) - IJ.
[0100] The [Fig.5] is a block diagram showing the rebinarized scalar product, described in relation to the [Fig.1], followed by a modulation operation.
[0101] According to one embodiment, the binarized output j[k] is provided to a modulation block 500 (MODULATION). By way of example, the modulation block 500 is configured to receive a vector, such as the intermediate data r. The modulation block 500 is then configured to perform a modulation operation between each output coefficient y [fc] and each coefficient r[k] of the vector r. As an example, block 500 generates the coefficient x [k], corresponding for example to y [k] when the value of r[Æ] is equal to 1, or corresponding to the value 0 when r[Æ] is equal to 0. The vector x then corresponds to the point-by-point product between the vectors y and r.
[0102] In the example described in relation to Table 1, block 500 is configured to generate each coefficient of the vector x by applying the modulation operation between / {y[k]) and f ( r [k ] . The generated vector x is then equal to / (y)-f (r ). XXX
[0103] As an example, each coefficient of the vector y is provided to block 500 directly and without being previously stored in memory.
[0104] An example of an implementation of block 500 is described in detail in relation to [Fig.6].
[0105] Figure 6 is a block diagram illustrating a circuit 600 configured to perform layer operations including, among other things, modulation operations, according to an embodiment of this description. By way of example, the circuit 600 illustrates embodiments of the circuit 200 and includes elements 202 to 214 described in relation to Figure 2. The circuit 600 further includes, or is connected to, the scheduler circuit 216 configured to control and drive the execution of layer operations in relation to, for example, one of the tables 1 to 4.
[0106] By way of example, the calculation circuit 206 of the circuit 600 is further configured to calculate the components of one or more corresponding "intermediate" vectors, for example, to the vectors y, z, r and o described in Table 1. By way of example, the circuit 208 of the circuit 600 is configured to perform a point-by-point product between the input vector d and a row of a weight matrix Wy. In particular, circuits 208 and 212 of circuit 600 are configured to perform the dot product between the input vector and a row of a weight matrix Wy based on the indication of read functions and / 3D among the functions and h. The calculation circuit 206 sequentially provides the components X&1 to a first input of each of the gates 602 and 604. As an example, circuit 206 is further configured to provide each component y[&] to a first input of a multiplexer 606.
[0107] Gates 602 and 604 are further configured to sequentially receive, at a second input, components of a binary vector r. Gate 602 is then, for example, configured to perform, upon each receipt of a component, an AND operation between the component [jffc] and the component [jffc]. Gate 604 is then, for example, configured to perform, upon each receipt of a component, an XNOR operation between the component [jffc] and the component [jffc]. The values generated by gates 602 and 604 are then transmitted to a multiplexer 608 configured to select the output of gate 602, or the output of gate 604, based on a read function f. The multiplexer 608 is then configured to transmit the selected value to a second input of the multiplexer 606. The multiplexer 606 is then configured to generate the output coordinate x[k] of the binary vector x by selecting the component or output of the multiplexer 608, based on a MODULATION ON signal. In particular, gates 602 and 604 and multiplexers 606 and 608 define an example implementation of the 500 modulation block.
[0108] In an example, when the layer operation performed incorporates a modulation corresponding to a masking, different from the "multiplexing" masking defined above, the function f chosen will be the Heaviside function. When the layer operation performed incorporates a modulation corresponding to a binary modulation, the function f chosen will be the function
[0109] Fig. 7A schematically illustrates a sequence of operations 700 enabling the realization of a binary masking mechanism possibly using a multiplexing circuit 704, but which can also be realized without a multiplexing circuit.
[0110] Fig. 7B schematically illustrates another example of a 702 sequence of operations, based on readings and writings of a memory circuit which also enables the implementation of a masking mechanism without additional hardware.
[0111] In particular, Figures 7A and 7B represent operations of a binary masking mechanism that can be integrated into the circuit 200 and / or 600 and in connection with the memory element 202 described in relation to Figures 2 and / or 6. Figures 7A and 7B illustrate sequences of operations 700 or 702, advantageously implemented by the scheduler circuit 216.
[0112] By way of example, the masking mechanism acts on two data vectors simultaneously, for example on the vectors x and T described in Table 1 and stored in memory 202. The masking mechanism is further controlled by a masking vector z, stored in memory 202.
[0113] By way of example, a vector s = z'x + z' as described in Table 1, corresponds to the masking of the vectors x and \ controlled by the vector z. As an example, the calculation of the vector is carried out either according to one of the sequence of operations 700 of [Fig.7A], or according to the sequence of operations 702 of [Fig.7B].
[0114] According to an embodiment illustrated in [Fig. 7A], the sequence of operations 700 orchestrated by the scheduler circuit 216 is as follows when a multiplexing circuit is used, located on the periphery of the memory element 202. The multiplexing circuit is, in practice, composed of K multiplexers, each with 2 inputs, driven by a selection bit. To begin, a read operation is performed on the memory element 202 of the x vector, then of the y vector. The read values are placed in registers, not shown, located at the input of the multiplexing circuit 704. A read of the z vector is then performed in memory element 202. The values of the z vector are used to control the multiplexing circuit 704. The data present at the output of the multiplexing circuit, following this selection based on the z vector, are then written to memory element 202 as forming the s vector.
[0115] According to another alternative embodiment, also illustrated in principle in Figure 7A, the sequence of operations 700 orchestrated by the general control device is performed without the use of a multiplexer. To do this, the vectors x, y, and z are read from memory. Then, depending on the composition of the computing resources of the general control device, the equivalent of the multiplexing function described above is implemented in hardware or software. It should be noted that it is potentially possible to reduce the number of reads from memory 202 by first reading the vector z, and then reading all or part of the vectors x and y. Thus, if, for example, all the bits of the vector z have the same value, only one complete read of the vector x or y can be performed, the other read being unnecessary.Finally, we do as before, writing the result to memory element 202 at the storage location of the vector s.
[0116] The operations described above are then repeated, for each value k between 1 and K. The scheduler circuit 216 is, for example, configured to command the storage of each component of the vector x in a buffer memory, and then, once the vector is fully generated, to command its writing into memory 202. In another example, the scheduler circuit 216 is configured to command the writing of each component of the vector, one by one and as soon as they are generated, into memory 202.
[0117] According to one embodiment, the circuit 704 is configured to write the vectors into memory 202. The circuit 704 includes, for example, K parallel multiplexers. By way of example, if the vector x is not reused in a subsequent calculation, the vector 4 is directly written into memory 202 to replace the vector x.
[0118] According to one embodiment, the circuit 702 includes the memory 202 configured so that signals from the data read (rdata) and write control (bwrite) execute a scheduling to perform the masking mechanism in a reduced number of clock strokes and using standard memory tile control signals.
[0119] In the example illustrated in [Fig. 7B], memory 202 is a memory configured for implementing masking operations such as, for example, described in more detail in US patent 6075721. The signals resulting from reading the data, from write control, and data to be written, then execute a two-step scheduling 708 and 710 successive.
[0120] The sequence of operations 708 is driven by the scheduler circuit 216. In this example, memory 202 stores, for instance, the vectors x and z, as well as the vector 5 corresponding, for example, to the previous turn. In order to generate the vectors for the current turn, memory 202 is configured to perform the masking S — J ■ Z + X • z. To do this, memory 202 is configured to read the vector x (rdata) and copy it as write data (wdata) by providing it on a so-called "data" input of memory 202 in place of the vectors from the previous turn.
[0121] The sequence of operations 710, following sequence 708, includes reading the vectors T etz (rdata). The vectors etz are, for example, temporarily stored in one or more buffer memories external to the matrix 202. The vector ' is provided on a so-called "mask" input of memory 202. Memory 202 then reads component by component the vectors y etz stored in the external buffer memories and rewrites the components s[k] into memory 202 such that s[&] = J'[fc] when — [ and s[k] = 0 when — Q- As an example, the component-by-component reading by memory 202 is performed in parallel. Thus, following operations 708 and 710, the vector stored in memory 202 corresponds to the mask S = y ■ z + X ■ z.
[0122] According to one embodiment, the calculation of the vector takes place following the generation of the vector y, or sequentially, following the generation of each component j[fc]. The writing of binary values, describing the components of the vectors stored in memory 202, is then carried out sequentially, and not in parallel.
[0123] Fig. 8 is a block diagram illustrating an example of a fully connected one-layer 800 circuit integrating a modulation mechanism on a vector from an intermediate calculation and enabling a masking mechanism to be achieved by a conditional update of the output vector, according to an embodiment of the present description.
[0124] According to one embodiment, the circuit 800 includes the scheduler circuit 216. By way of example, the scheduler circuit 216 is configured to control the calculation of a vector, resulting from the masking mechanism, such as S = Z - b^^b^X )yfx(r))+z-SA as an example, the vector * represents a layer operation illustrated in relation to Table 1. In the example shown in Figure 7B, the scheduler circuit 216 is configured to receive the vector z and, based on the values of each component z[&] of the vector z, to directly control the reading, or not, of the k-icmc row of a matrix of weight Wy, stored in memory 204. This allows us to execute only the calculations necessary to update the vector using the components of the vector x corresponding to the components of the vector z being zero, in their corresponding positions in the vector, that is to say, for each index k such that 4^ = 0, we have ) ) - AUd) ) ■In particular, as described in relation to Figure 6, the calculation of vector 5 is carried out in several turns of calculation by circuit 206. The scheduler circuit 216 is then configured to reduce the calculation latency and energy consumption of circuit 800 by avoiding the execution of unnecessary calculations. dot product between the vector m
[0125] In the example where the scheduler is configured to control the generation of the vectors described in relation to Table 1, memory 202 initially includes, for example, the vectors and. Memory 204 includes, for example, the weight matrices Wz, Wo and. As an example, the vector rt is entirely calculated by circuit 206. The scheduler circuit 216 is then configured to provide the indication of the read functions and to circuit 208 for the implementation of the and the lines r » To The vector rt fr is then, for example, stored in memory 202. The scheduler circuit 216 is further configured to control the calculation of a component of the vector zt, for example z[1] based on the read functions and As an example, the scheduler circuit 216 is configured to provide the indication of the read functions yW and to circuit 208 for the realization of the dot product between the vector (dt) and the line [1] ) • The component A [1] is then, by For example, stored in memory 202. In another example, the K components of the vector zt are calculated and the vector zt is stored in memory 202. The circuit Scheduler 216 is further configured to control the calculation, via circuit 206, and by providing the indication of the read functions and the component J [ 1 ]. The scheduler circuit 216 is then configured to activate the MODULATION ON signal in order to activate block 500. Block 500 is then configured to generate the component xr[ 1 ] by performing a modulation between f (yjl]) Ct / (rjl]). The component U is then not stored in memory 202 following its calculation by circuit 206. Memory 202 is then, for example, configured to perform masking, for example described in relation to Figure A or 7B. The The value of zf[1] is read and the value of component [1] is overwritten; in the case where Z / [1] = 0, P31 becomes component xt[1]. The value of component st[1] for the current vector is then equal to xt[1]. In the case where component zt[1] is equal to 1, component vJ[1] is not overwritten and becomes the value of component sz[1] for the current vector St.
[0126] The scheduler circuit 216 is further configured to control the calculation, by the calculation circuit 206, of the component oj 1] on the basis of the read functions and the scheduler circuit 216 is then configured to activate the JOJO signal MODULATION ON in order to activate the 500 block. The oj 1] component is, for example, directly supplied to the 500 modulation block and the latter is configured to generate the ht [ 1 ] component by performing a modulation between fh ( st [ 1 ] ) and fh( ot [ 1 ] ). The component is not then stored in memory 202 following its calculation by circuit 206. The vector ht is then updated in memory 202 following the storage of the component ht[ 1].
[0127] The circuit 800 then performs Kl other calculation turns, in order to calculate all the components of the vectors S( and ht.
[0128] Advantageously, the sequences of operations described above execute a scheduling that allows the masking mechanism to be carried out in a reduced number of clock strokes and using standard control signals for a memory tile by the scheduler circuit 216. In addition, intermediate data is used directly, for example for modulation operations, without being stored beforehand, which allows for a gain in space and time.
[0129] An advantage of the described embodiments is that they allow for near-memory layer operations, and in particular intermediate operations, by varying the arithmetics used, between each layer and between each operation, for the weight matrices and vectors manipulated.
[0130] An advantage of a near-memory implementation also lies in the fact that the canonical control signals of the memory plane can be used efficiently for all or part of the variants described above.
[0131] Various embodiments and variations have been described. A person skilled in the art will understand that certain features of these various embodiments and variations could be combined, and other variations will become apparent to a person skilled in the art.
[0132] Finally, the practical implementation of the embodiments and variants described is within the reach of a person skilled in the art, based on the functional indications given above.
Claims
Demands
1. Circuit comprising: - a first memory element (202) configured to store a first data point (e'x- s- 0 M r); - a second memory element (204) configured to store a first weight matrix (Wj?) in association with a first layer of an artificial binary neural network; - a computing circuit (206) configured to: a) receive the first data point and the ^-th row ( Wj, [ k ] ), of the first weight matrix; b) receive a first control signal, indicating the nature of each of the first and second read functions ( f^1 ), J y J y among at least the first and second reference functions ( g, h ), the first and second reference functions each being associated with a two-valued arithmetic;c) generate a first vector (Æ] ) ) by applying the first reading function ( f W) to the ^-th row of the first weight matrix and a second vector ( by applying the second reading function ( f W) to the first data ; and J yd) generate a ^-th component (jj#]) of a first output vector (^) on the basis of the first and second vectors.;
2. Circuit according to claim 1, wherein the first reference function takes values in {-1,1} and the second reference function takes values in [0,1}-
3. Circuit according to claim 1 or 2, configured to, following the generation of the ^-th component (jpt]) of the first output vector (^), control the storage of the k-icmc component in the first memory element (202).
4. Circuit according to any one of claims 1 to 3, further comprising: - a first logic gate (602) configured to apply an AND operation between the *-th component (jfjt]) of the first output vector (^), transmitted by the calculation circuit (206), and a component (y[jt]) of a binary vector (r); - a second logic gate (604) configured to apply an XNOR operation between the Ά-th component (j[#]) of the first vector, transmitted by the calculation circuit (206), and the component (r[jt]) of the binary vector (r); - a first multiplexer (608) configured to select the output of the first logic gate (602), or the output of the second logic gate (604), on the basis of a control signal, indicating the nature of a third read function among the first and second reference functions (g, &); - a second multiplexer (606) configured to generate a ^-th component (x[£J) of a second output vector (x) by selecting the *-th component (j[jt]) of the first vector, provided by the calculation circuit, or the output of the first multiplexer, on the basis of a modulation signal.
5. Circuit according to any one of claims 1 to 4, further comprising a scheduler circuit (216) configured to: - receive a component of a vector stored in the first memory element (202); and - based on the value of the component, command the reading, in the second memory element, of the ^-th line (Wj[k]) of the first weight matrix.
6. Circuit according to any one of claims 1 to 5, wherein the calculation circuit (206) comprises: - a number N of multiplier circuits (208_1, 208_N) each configured to receive the first and second read function and each configured to generate an output scalar on the basis of a component of a vector and a component of a weight matrix; - an accumulator circuit (212) configured to sum the output scalars supplied by the plurality of multiplier circuits; and - a converter (214) configured to convert the value generated by the accumulator circuit into a binary value.
7. Circuit according to claim 6, wherein the accumulator (212) comprises an adder tree (500), comprising a plurality of shift circuits (502) and configured to generate a scalar, corresponding to the dot product between the first and second vectors, augmented by a power gain of 2.
8. A circuit according to any one of claims 1 to 7, wherein the first data is of length N, N being an integer, and is stored contiguously by vectors of length K, K being a divisor of the value N, and wherein the ^-th row k ), of the first matrix is of length N and wherein the computing circuit (206) comprises an L = N / K number of shift registers (210) connected to the first memory element (202) and is configured to, upon receiving a sequence of vectors of the first data, convert said first data into a vector of size N, by concatenation of vectors of size K, wherein the shift registers (210) are, for example, further configured to perform the concatenation of the first data of size K with a sequence of NK bits each equal to 1, or to 0.
9. Circuit according to claim 8, wherein each of the multiplier circuits (208_l, 208_N) comprises an NXOR and / or AND logic gate configured to multiply a component of the first data and an element of the first weight matrix associated with the first layer.
10. Circuit according to claim 5, wherein the first memory element (202) is further configured to store a masking vector (z), the scheduler circuit (216) being configured to: - if the fc-th component of the masking vector is equal to 0, command the storage of the ^-th component of the first data as the ^-th component of the first output vector; and - if the same component of the masking vector is equal to 0, command the execution of steps a) to c).
11. Circuit according to claim 10, wherein the first memory element (202) is a memory configured for the implementation of masking operations.
12. Circuit according to any one of claims 1 to 11, wherein the second memory element (204) further stores a third weight matrix associated with a second layer of the neural network, and wherein the computing circuit (206) is further configured to: - select, following the reception of a second control signal, from the fifth and sixth read functions, among the first and a second reference functions; - generate a fifth vector by applying the fifth read function to a row, or column, of the third weight matrix and a sixth vector by applying the sixth read function to an output vector from a previous layer, stored in the first memory element (202); - generate a third output value based on the fifth and sixth vectors.
13. Circuit according to any one of claims 1 to 12, wherein the first reference function (^) is defined by # : w M and the second reference function (A) is defined by h : u 2w -1.
14. Method comprising: - the provision of a first data (e>x' £ o, K r) stored in a first memory element (202) of a circuit (200, 600, 700, 702) and of the ^-th line ( Wy [ k ] ) of a first weight matrix associated with a layer of a binary artificial neural network to a first computing circuit (206) of the circuit; - the provision of an indication, by means of a control signal, of the nature of a first and a second read function ( f^, among a first and a second reference function J y J y ($, A);- the generation, by the first calculation circuit, of a first vector ( [ & jp by applying the first function ( to the i-th row of the weight matrix and a second vector (p by applying the second function ( on the first data, - the generation of a k-th component (jf#]) of a first output vector (^) on the basis of the first and second vectors, the first and second reference functions each associated with two-valued arithmetic.;
15. A method according to claim 14, wherein the first computing circuit (206) comprises a plurality of multiplier circuits (208_l, 208_N) and an accumulator (212) configured to generate a scalar by performing a dot product between the first and second vectors, wherein the first computing circuit (206) comprises, for example, a converter (214) configured to convert the scalar into a binary value.
16. A method according to claim 14 or 15, further comprising: - the reading, by a second calculation circuit (704) of the value of a first component of the first data (x[à]), of the first output value (jfjfcj) and of a first masking value ($(&]), stored in the first memory element (202); and - the command, on the basis of the first masking value and by the second calculation circuit, of the writing of a first masked value (s[&]), in the first memory element, corresponding either to the first component of the first data, or to the k-th component of the first output vector.
17. A method according to claim 16, wherein the writing of the first masked value comprises: - the deletion of the first component of the first data; and - the writing of the first masked value to the address of the first component of the first data in the first memory element.
18. A method according to claim 16 or 17, further comprising, after writing the first masked value: - the generation of a k+l-th output component (j(jt + 1]) of the first output vector and its storage in the first memory element (202); - the reading, by the second calculation circuit, of the value of a k+l-th component of the first data (x[£ + 1]), of the k+l-th component (^(^ + 1]) of the first output vector and of a second masking value (^jt+ 1]), stored in the first memory element (202); and - the command, on the basis of the second masking value and by the second calculation circuit, of the writing of a second masked value 4. ] |), in the first memory element, corresponding either to the k+l-th component of the first data, or to the k+l-th component of the output vector.
Citation Information
Patent Citations
Random access memory having bit selectable mask for memory writes
US6075721A