Training of Artificial Neural Networks

By allocating weight storage bits in digital memory and analog multiplication-accumulation units, combined with the weight update of the digital processing unit, the calculation-intensive and time-intensive problems of the ANN training system are solved, and an efficient and accurate training method is achieved.

CN113826122BActive Publication Date: 2025-07-22INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080034604.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-05-16
Filing Date
2020-05-12
Publication Date
2025-07-22
Estimated Expiration
2040-05-12

AI Technical Summary

Technical Problem

Existing artificial neural network training systems are computationally intensive and time-intensive, requiring reduced complexity while maintaining training accuracy.

Method used

The least significant bits of each N-bit weight are stored in the digital memory, and the next n-bit parts are stored in the analog multiplication-accumulation unit, and the weight update calculation is performed in conjunction with the digital processing unit, and the memory and the multiplication-accumulation unit are periodically reprogrammed to store the update weight.

Benefits of technology

It reduces training complexity and power consumption, improves training accuracy, and realizes an efficient ANN training method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113826122B_ABST
    Figure CN113826122B_ABST
Patent Text Reader

Abstract

Methods and apparatus are provided for training an artificial neural network having a series of neuron layers with interposed synaptic layers, each synaptic layer having a corresponding set of N-bit fixed-point weights {w} for weighting signals propagating between its adjacent neuron layers via an iterative cycle of signal propagation and weight update computational operations. Such a method includes, for each synaptic layer, storing a plurality of p least significant bits of each N-bit weight w in a digital memory, and storing a next n-bit portion of each weight w in an analog multiply-accumulate unit including an array of digital memory elements. Each digital memory element includes n binary storage cells for storing respective bits of the n-bit portion of the weight, where n≥1 and (p + n + m) = N, where m≥0 corresponds to a defined number of most significant zero bits in the synaptic layer weights.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The present invention generally relates to the training of artificial neural networks.

[0002] Artificial Neural Networks (ANNs) have been developed to perform computational tasks in a manner inspired by the biological structure of the nervous system. These networks are based on the fundamental principles of biological systems, whereby neurons are interconnected via synapses that transmit weighted signals between neurons. ANNs are based on a logical structure comprising a series of neuron layers with intervening synaptic layers. The synaptic layers store weights for weighting the signals propagated between neurons in their adjacent neuron layers. A neuron n in a given layer i can be connected to one or more neurons n in the next layer j , and different weights w ij can be associated with each neuron-neuron connection n i -n j for weighting the signal sent from n i to n j . Each neuron generates an output signal based on its accumulated weighted input, whereby the weighted signals can propagate across successive layers of the network.

[0003] ANNs have been successfully applied to a variety of complex analysis tasks, such as speech and image recognition, e.g., the classification of handwritten digits based on the MNIST (Modified National Institute of Standards and Technology) dataset. ANNs undergo a training phase in which the set of weights for each synaptic layer is determined. During an iterative training process, the network is exposed to a training dataset (e.g., image data of handwritten digits), during which the weights are repeatedly updated as the network "learns" from the training data. Training involves an iterative loop of signal propagation and weight update computational operations, in which the network weights are gradually updated until a convergence condition is reached. The resulting trained network with weights defined via the training operations can then be applied to new (unseen) data to perform the inference task of the application under discussion.

[0004] Training of ANNs can be computationally and time intensive, with multiple neuron layers and millions of synaptic weights. Training methods using analog multiply-accumulate units based on memristive synaptic arrays have been proposed to alleviate these problems, where synaptic weights are stored in the analog conductance values of memristive devices such as PCM (Phase Change Memory) devices. These units employ crossbar arrays of memristive devices that are connected between row and column lines for applying signals to the devices, where each device implements a synaptic with a weight corresponding to the (variable) device conductance. The parallel computing capabilities of these multiply-accumulate arrays can be utilized to perform inexpensive vector-matrix calculations (such as generating the accumulated-weighted signals propagated across the synaptic layer as needed) in the analog domain with O(1) computational complexity. Such a training method of accumulating updates to synaptic weights during training in a high-precision digital accumulator is known in the art. An analog multiply-accumulate unit is also known in the art where 1-bit weights are stored digitally in binary SRAM (Static Random-Access Memory) cells for neural network inference calculations.

[0005] There is still a need for further neural network training systems to reduce complexity while maintaining training accuracy. Summary of the Invention

[0006] According to at least one embodiment of the present invention, there is provided a method for training an artificial neural network having a series of neuron layers with interposed synaptic layers, each synaptic layer having a corresponding set of N-bit fixed-point weights {w} for weighting signals propagated between its adjacent neuron layers via an iterative cycle of signal propagation and weight update computational operations. The method includes, for each synaptic layer, storing the plurality of p least significant bits of each N-bit weight w in digital memory and storing the next n-bit portion of each weight w in an analog multiply-accumulate unit including an array of digital memory elements. Each digital memory element includes n binary storage cells for storing respective bits of the n-bit portion of the weight, where n≥1 and (p + n + m) = N, where m≥0 corresponds to a defined number of most significant zero bits in the synaptic layer weights. The method further includes performing a signal propagation operation by providing the signal to be weighted by the synaptic layer to the multiply-accumulate unit to obtain an accumulated weighted signal depending on the n-bit portion of the stored weights, and performing a weight update computational operation in a digital processing unit (operably coupled to the digital memory and the multiply-accumulate unit) to calculate updated weights of the synaptic layer based on the signals propagated by the neuron layers. The method further includes periodically reprogramming the digital memory and the multiply-accumulate unit to store bits of the updated weights.

[0007] In the training method embodying the present invention, weights are defined in an N-bit fixed-point format with a desired precision of the training operation. For each N-bit weight w, at least p of the least significant bits of the weight are stored in a digital memory. The next n-bit portion (i.e., the n next-to-most significant bits) is stored digitally in binary storage cells of digital memory elements of an analog multiply-accumulate unit. This n-bit portion corresponds to a reduced-precision weight value of the weight w. During signal propagation operations, these reduced-precision weights are used to perform multiply-accumulate operations. In weight update operations, the N-bit weights of the updated synaptic layer are computed in a digital processing unit. Thus, weight update computations are performed with digital precision, and the digital memory and the multiply-accumulate unit are reprogrammed periodically to store the appropriate bits of the updated weights (i.e., the p least significant bits and the n-bit portion, respectively). By using the N-bit fixed-point weights stored in a combination of a digital memory and digital elements of a multiply-accumulate array, the method combines the accuracy advantages in weight update operations with fast, low-complexity vector matrix computations for signal propagation. Performing vector matrix operations with reduced-precision weights reduces the complexity, and thus the power and on-chip area of the multiply-accumulate unit. Accordingly, embodiments of the present invention provide a fast, efficient ANN training method based on a multiply-accumulate array.

[0008] For a synaptic layer, the parameter m can be defined as m = 0, regardless of the actual number of most significant zero bits in the weights of any given layer. This gives a simple implementation where (p + n) = N. In other embodiments of the present invention, the initial value of m for a synaptic layer can be defined according to the number of most significant zero bits in the weights {w} of the synaptic layer, and then, as the number of most significant zero bits in the weight set {w} changes, the value of m can be adjusted dynamically during training. In these embodiments of the present invention, at least p = (N - n - m) of the least significant bits of the weight w are stored in a digital memory, and as the value of m is adjusted during training, the n-bit portion stored in the multiply-accumulate unit is redefined and reprogrammed dynamically. This provides a more optimal definition of reduced-precision weights for various network layers, thereby improving training accuracy.

[0009] In some embodiments of the present invention, only p of the least significant bits of each N-bit weight are stored in a digital memory. The digital memory can be distributed in the multiply-accumulate unit such that each N-bit weight is stored in a unit cell that includes p bits of the digital memory (storing the p least significant bits of the weight) and a digital memory element (storing the n-bit portion of the weight). This provides an area-efficient implementation of a combined digital / analog storage unit based on unit cells with a small footprint.

[0010] In other embodiments of the present invention, all N bits of each N-bit weight can be stored in a digital storage unit providing digital memory. This provides an efficient operation for performing weight updates in digital memory, allowing infrequent updates of weights with reduced precision in the multiply-accumulate unit. For example, weights with reduced precision may only be updated after the network has processed many batches of training examples. To further improve the efficiency of the weight update operation, the n-bit portion of the updated weight can be copied from the digital memory to the multiply-accumulate unit only when a bit overflow occurs at the (N - p)-th bit during the weight update in the digital memory during training.

[0011] In embodiments of the present invention where the N-bit weights of all synaptic layers are stored in digital memory, during signal propagation through the network, the multiply-accumulate unit can be reused for weights with reduced precision in different layers. As successive sets of synaptic layers become active for signal propagation, the n-bit portions of the weights of those layers can be dynamically stored in an array of digital memory elements.

[0012] At least one other embodiment of the present invention provides an apparatus for implementing an artificial neural network in an iterative training loop of signal propagation and weight update computational operations. The apparatus includes a digital memory storing the p least significant bits of each N-bit weight w of each synaptic layer, and an analog multiply-accumulate unit for storing the next n-bit portion of each weight w of the synaptic layer. As described above, the multiply-accumulate unit includes an array of digital memory elements, each digital memory element including n binary storage units. The apparatus further includes a digital processing unit operatively coupled to the digital memory and the multiply-accumulate unit. During a signal propagation operation, the digital processing unit is adapted to provide a signal to be weighted for each synaptic layer to the multiply-accumulate unit to obtain an accumulated weighted signal depending on the n-bit portion of the stored weights. The digital processing unit is further adapted to perform a weight update computational operation to calculate updated weights for each synaptic layer based on signals propagated by the neuron layer, and to control periodic reprogramming of the digital memory and the multiply-accumulate unit to store the appropriate bits of the updated weights.

[0013] According to one aspect, there is provided a method for training an artificial neural network having a series of neuron layers with inserted synaptic layers, each synaptic layer having a corresponding set of N-bit fixed-point weights {w} for weighting signals propagated between its adjacent neuron layers via an iterative cycle of signal propagation and weight update calculation operations. The method includes, for each synaptic layer: storing multiple p least significant bits of each N-bit weight w in a digital memory; and storing the next n-bit portion of each weight w in an analog multiply-accumulate unit including an array of digital memory elements, each digital memory element including n binary storage units for storing respective bits of the n-bit portion of the weight, where n≥1 and (p + n + m) = N, where m≥0 corresponds to a defined number of most significant zero bits in the synaptic layer weights; performing the signal propagation operation by providing the signal to be weighted by the synaptic layer to the multiply-accumulate unit to obtain an accumulated weighted signal depending on the stored n-bit portion of the weight; and performing a weight update calculation operation in a digital processing unit (operably coupled to the digital memory and the multiply-accumulate unit) to calculate updated weights of the synaptic layer based on signals propagated by the neuron layers; and periodically reprogramming the digital memory and the multiply-accumulate unit to store bits of the updated weights.

[0014] According to another aspect, there is provided an apparatus for implementing an artificial neural network having a series of neuron layers with inserted synaptic layers, each synaptic layer having a corresponding set of N-bit fixed-point weights {w} for weighting signals propagated between its adjacent neuron layers via an iterative cycle of signal propagation and weight update calculation operations. The apparatus includes: a digital memory storing multiple p least significant bits of each N-bit weight w of each synaptic layer; an analog multiply-accumulate unit storing the next n-bit portion of each weight w of the synaptic layer, the multiply-accumulate unit including an array of digital memory elements, each digital memory element including n binary storage units for storing respective bits of the n-bit portion of the weight, where n≥1 and (p + n + m) = N, where m≥0 corresponds to a defined number of most significant zero bits in the synaptic layer weights; and a digital processing unit operably coupled to the digital memory and the multiply-accumulate unit, the digital processing unit being adapted to: in a signal propagation operation, obtain an accumulated weighted signal depending on the stored n-bit portion of the weight by providing the signal to be weighted by each synaptic layer to the multiply-accumulate unit; perform a weight update calculation operation to calculate updated weights of each synaptic layer based on signals propagated by the neuron layers; and control periodic reprogramming of the digital memory and the multiply-accumulate unit to store the bits of the updated weights. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Embodiments of the present invention will now be described in more detail by way of illustrative and non - limiting examples with reference to the accompanying drawings:

[0016] Figure 1 is a schematic diagram of an exemplary ANN;

[0017] Figure 2 is a schematic block diagram of a device for implementing an ANN in a training operation according to an embodiment of the present invention;

[0018] Figure 3 indicates the bit structure of the weights w of an ANN layer according to an embodiment of the present invention;

[0019] Figure 4 shows according to an embodiment of the present invention Figure 2 the structure of an array of digital memory elements in the multiply - accumulate unit of the device;

[0020] Figure 5 indicates according to an embodiment of the present invention by Figure 2 the steps of a training method performed by the device;

[0021] Figure 6 shows Figure 2 the structure of a memory device in an embodiment of the device;

[0022] Figure 7 shows according to an embodiment of the present invention Figure 6 a more detailed structure of an analog SRAM multiply - accumulate array in the device;

[0023] Figure 8 shows according to an embodiment of the present invention Figure 7 the structure of an SRAM unit cell in the array;

[0024] Figure 9 shows according to an embodiment of the present invention Figure 2 a memory device in another embodiment of the device;

[0025] Figure 10 shows according to an embodiment of the present invention Figure 9 a more detailed structure of a combined digital / analog SRAM cell in the device;

[0026] Figure 11 shows Figure 2 a memory device in yet another embodiment of the device; and

[0027] Figure 12 shows another embodiment of the analog SRAM multiply - accumulate array of the device. Detailed Description

[0028] Figure 1Shows the logical structure of an example of a fully connected ANN according to an embodiment of the present invention. ANN1 includes a series of neuron layers with inserted synaptic layers. In the simple example shown, the network has three neuron layers: the first layer N1 of input neurons that receive the network input signal; the last layer N3 of output neurons that provide the output signal of the network; and an intermediate ("hidden") layer N2 of neurons between the input layer and the output layer. The neurons in layer N1 are represented by n 1i (1 ≤ i ≤ l1), the neurons in layer N2 are represented by n 2j (1 ≤ j ≤ l2), and the neurons in layer N3 are represented by n 3k (1 ≤ k ≤ l3), where l x is the number of neurons in layer N x . As shown, all neurons in each layer are connected to all neurons in the next layer, thereby transmitting the neuron activation signal from one layer to the neurons in the next layer. Synaptic layers S1 and S2 are inserted together with the neuron layers and have respective weight sets {w ij} and {w jk} for weighting the signals propagated between their adjacent neuron layers. The weight w 1i defined for each connection between N1 neuron n 2j and N2 neuron n ij , whereby the signal propagated from n ij to n 1i is weighted according to the corresponding weight w 2j of this neuron pair. Thus, as shown, the weight set {w ij} of synaptic layer S1 can be represented by a matrix W with l2 rows and l1 columns having weights w ij . The signal propagated from N2 neuron n 2j to N3 neuron n 3k is similarly weighted by the corresponding weight w jk of synaptic layer S2, and the weight set {w jk} of synaptic layer S2 can be represented by a matrix with l3 rows and l2 columns having weights w jk .

[0029] The input layer neurons can simply transmit the input data signals they receive as the activation signal of layer N1. For the subsequent layers N2 and N3, each neuron n 2j , n 3k generates an activation signal depending on its cumulative input, i.e., the cumulative weighted activation signals from the neurons it is connected to in the previous layer. Each neuron applies a non-linear activation function f to the result A of this cumulative operation to generate its neuron activation signal for forward transmission. For example, the cumulative input A 2j of neuron n j is calculated by the dot product is given, where x 1i is the activation signal from neuron n 1i Thus, the cumulative input 2j to neuron n for vector A can be represented by the matrix-vector multiplication Wx of the weight matrix W ij and the activation signal 1i from neuron n for vector x. Then, each N2 neuron n 2j generates its activation signal x 2j as x 2j = f(A j ) to propagate to the N3 layer.

[0030] Although Figure 1 a simple example of a fully connected network is shown, generally neurons in any given layer can be connected to one or more neurons in the next layer, and the network can include one or more (usually up to 30 or more) consecutive layers of hidden neurons. A layer of neurons can include one or more bias neurons (not shown) that do not receive input signals but transmit bias signals to the next neuron layer. Other computations can also be associated with some ANN layers. In some ANNs, such as Convolutional Neural Networks (CNNs), a layer of neurons can include a three-dimensional volume of neurons with an associated three-dimensional weight array in the synaptic layer, although the signal propagation computations can still be expressed in terms of matrix-vector operations.

[0031] ANN training involves an iterative loop of signal propagation and weight update calculation operations in response to a collection of training examples provided as network inputs. For example, in the supervised learning of handwritten digits, training examples from the MNIST dataset (where the labels (here, the digit classes from 0 to 9) are known) are repeatedly input into the network. For each training example, the signal propagation operation includes a forward propagation operation and a backward propagation operation. In the forward propagation operation, signals propagate forward from the first neuron layer to the last neuron layer, and in the backward propagation operation, the error signal propagates backward through the network from the last neuron layer. In the forward propagation operation, the activation signal x is weighted layer by layer and propagated through the network as described above. For each neuron in the output layer, the output signal after forward propagation is compared with the expected output (based on the known label) of the current training example to obtain the error signal ε of that neuron. The error signal of the output layer neurons propagates backward through all layers of the network except the input layer. The error signal propagated backward between adjacent neuron layers is weighted by the appropriate weights of the inserted synaptic layers. Thus, backward propagation results in the calculation of the error signal for each neuron layer except the input layer. Then, the update of the weights of each synaptic layer is calculated based on the signals propagated through the neuron layers in the signal propagation operation. In general, the weight update can be calculated for some or all of the weights in a given iteration. For example, the update Δw ij of the weight w ij between neuron i in one layer and neuron j in the next layer can be calculated as:

[0032] Δw ij = ηx i ε j

[0033] where x i is the forward propagation activation signal from neuron i; ε j is the backward propagation error signal of neuron j; and η is a predefined learning parameter of the network. Thus, the training process gradually updates the network weights until a convergence condition is reached, and the resulting network with the trained weights can be applied to ANN inference operations.

[0034] Figure 2Shows an apparatus for implementing ANN 1 in a training operation according to a preferred embodiment. Apparatus 2 includes a memory device 3 and a digital processing unit 4, which is operably coupled to the memory device 3 herein via a system bus 5. The memory device 3 includes a digital memory (schematically indicated as 6) and an analog multiply-accumulate (MAC) unit 7. As further described below, the MAC unit 7 includes at least an array of digital memory elements based on binary storage units. A memory control device, indicated as memory controller 8, controls the operation of the digital memory 6 and the MAC unit 7. The digital processing unit 4 includes a central processing unit (CPU) 9 and a memory 10. The memory 10 stores one or more program modules 11, and the program modules 11 include program instructions executable by the CPU 9 to implement the functional steps of the operations described below.

[0035] The DPU 4 controls the operation of the apparatus 2 during the iterative training process. The DPU is adapted to generate activation signals and error signals propagated by the neuron layers during forward and backward propagation operations, and perform weight update calculations for the training operation. The set of weights {w} of each synaptic layer of the network is stored in the memory device 3. The weight w is defined in an N-bit fixed-point format, where N is selected according to the precision required for a particular training operation. In this embodiment of the present invention, N = 32 gives high-precision 32-bit fixed-point weights. However, in other embodiments of the present invention, N can be set differently, for example, N = 64.

[0036] In the operation of the apparatus 2, the N-bit weights w of the synaptic layers are stored in a combination of the digital memory 6 and the digital memory elements of the MAC unit 7. Specifically, referring to Figure 3 , for each synaptic layer, at least a plurality of p least significant bits (LSBs) of each weight w are stored in the digital memory 6. The next n-bit portion of each weight w (i.e., the (p + 1)-th bit to the (p + n)-th bit) is stored in the MAC unit 7 at least when required for signal propagation calculations. Specifically, each n-bit portion is stored in the array of digital memory elements in the MAC unit 7. Each of these digital memory elements includes (at least) n binary storage units for storing the respective bits of the n-bit portion of the weight. For different synaptic layers, the value of n can be different. However, generally, n≥1 and (p + n + m) = N, where m≥0 corresponds to a defined number of most significant zero bits in the synaptic layer weights. The value of m can thus vary between synaptic layers or can be defined as m = 0 for any given layer as described below, in which case (p + n) = N. Thus, it can be seen that the n-bit portion of each weight w defines a weight value with reduced precision of that weight, denoted herein as W.

[0037] According to one embodiment, Figure 4 The logical structure of a digital memory element array storing the reduced-precision weights W of the synaptic layer in the MAC unit 7 is shown. The array 15 can be conveniently implemented as a cross array of digital memory elements 16 (with associated analog circuits described below) connected between row lines and column lines as shown. This example shows storing Figure 1 the reduced-precision weights {W ij} of the synaptic layer S1 in the ANN ij in a cross array. Each element 16 in the array stores the indicated corresponding reduced-precision weight W ij for n bits. The elements 16 are arranged in logical rows and columns, and each device is connected between a specific row line r i and column line c j for applying signals to the device. The row lines and column lines are connected to the controller 8 of the memory device 3 via row and column digital-to-analog / analog-to-digital converters (not shown), which convert the array input / output signals between the digital domain and the analog domain.

[0038] In the signal propagation operation of the synaptic layer, the signal generated by the DPU 4 is provided to the memory device 3 via the bus 5, where the controller 8 provides the signal to the array 15 storing the reduced-precision weights W ij In the forward propagation operation, the controller 8 provides the activation signal x 1i to the row lines r i of the array 15. The output signal obtained on the column line c j corresponds to the accumulated weighted signal ∑ i W ij x 1i returned by the control 8 to the DPU 4. The backward propagation calculation of the synaptic layer can be similarly performed by applying the error signal ε j to the column lines of the array to obtain the accumulated weighted signal ∑ j (W ij ε j ) on the row lines. Thus, the array 15 implements the matrix-vector calculations required for signal propagation across the synaptic layer.

[0039] Although exemplary embodiments of the apparatus 2 are described, the DPU 4 may include one or more CPUs that may be implemented by one or more microprocessors. The memory 10 may include one or more data storage entities and may include a main memory, such as DRAM (Dynamic Random Access Memory) and / or other memories physically separated from the CPU 9, as well as caches and / or other memories local to the CPU 9. Generally, the DPU 4 may be implemented by one or more (general-purpose or special-purpose) computer / programmable data processing devices, and the functional steps of the processing operations performed by the DPU 4 may generally be implemented by hardware or software or a combination thereof. The controller 8 may also include one or more processors that can be configured by software instructions to control the memory device 2 to perform the functions described herein. In some embodiments of the present invention, the DPU 4 and / or the controller 8 may include electronic circuits, such as programmable logic circuits, Field-Programmable Gate Arrays (FPGAs), or Programmable Logic Arrays (PLAs), for executing program instructions to implement the described functions. In the case of describing embodiments of the present invention with reference to a flowchart, it will be understood that each block of the flowchart and / or combinations of blocks in the flowchart may be implemented by computer-executable program instructions. Program instructions / program modules may include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. The blocks or combinations of blocks in the flowchart may also be implemented by a dedicated hardware-based system that performs the specified functions or actions or executes a combination of dedicated hardware and computer instructions.

[0040] The system bus 5 may include one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, and a processor or local bus using any of various bus architectures. By way of example and not limitation, these architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MCA) buses, Enhanced ISA (EISA) buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.

[0041] The binary storage cells in the memory element 16 of the MAC unit may include SRAM cells, DRAM (Dynamic RAM) cells, MRAM (Magnetoresistive RAM) cells, floating gate cells, RRAM (Resistive RAM) cells, or more generally any binary cell for digitally storing the respective bits of weights with reduced precision. An exemplary implementation of an analog MAC array based on SRAM cells is described in detail below. Generally speaking, the MAC unit 7 may include one or more analog multiply-accumulate arrays, each of which may include one or more crossbar arrays of digital memory elements. At any time, the MAC unit 7 may store all or a subset of the weights W with reduced precision of one or more synaptic layers. In some embodiments of the present invention, all the weights W of each synaptic layer may be stored in the corresponding array of the MAC unit. In other embodiments, the MAC unit may store only the weights W of the set of synaptic layers (one or more) that are currently active in the signal propagation operation. However, for each synaptic layer S, the training method implemented by the device 2 involves Figure 5 the basic steps indicated in the flowchart of

[0042] As Figure 5 indicated in step 20 of Figure 4As explained, for forward propagation, the activation signal x is provided to the MAC array to obtain an accumulated weighted signal depending on the precision-reduced weights W. In a subsequent backpropagation operation, the error signal ε is provided to the array to obtain an accumulated weighted signal depending on the weights W. The signals generated in these multiply-accumulate operations are returned to the DPU 4. In step 23, the DPU 4 calculates the updated N-bit weights w of the synaptic layer. Here, based on the signals propagated by the neuron layer as described above, a weight update Δw is calculated for the corresponding weights w, and each weight is updated to w = w + Δw. In step 24, the DPU 4 determines whether a predetermined convergence condition of the training operation has been reached. (Convergence can be defined in various known ways, and the specific convergence condition is orthogonal to the operations described herein). If not reached ("N" at step 24), the operation proceeds to step 25, where the DPU 4 controls the reprogramming of the weights w in the memory device 4. As further explained below, in any given iteration, depending on the implementation, this step may involve reprogramming the bits of the weights stored in the digital memory 6 or stored in both the digital memory 6 and the MAC unit 7. However, both the digital memory 6 and the MAC unit 7 are periodically reprogrammed (at the same or different times) during training to store the appropriate bits of the updated weights w. Then, the operation returns to step 22 for the next training sample. This process iterates until convergence is detected ("Y" at step 24), whereupon the training operation terminates.

[0043] Using the above method, the weight update can be calculated with high precision (here 32-bit precision) in the DPU 4 to ensure the accuracy of ANN training. In addition, using the precision-reduced weights W stored digitally in the analog MAC unit, the multiply-accumulate calculations for signal propagation can be performed efficiently. Using the precision-reduced weights here reduces the complexity, power consumption, and on-chip area of the MAC unit. The value of n can vary between synaptic layers, providing the weights W of the required precision for each layer to optimize training. For example, on a layer-by-layer basis, the value of n can be set to 1 ≤ n ≤ 8. According to a preferred embodiment, the method embodying aspects of the present invention thus provides an efficient training of an artificial neural network.

[0044] Figure 6 is in the first embodiment Figure 2Schematic representation of the structure of the memory device 3. For this embodiment, for all synaptic layers, the parameter m is defined as m = 0, whereby (p + n) = N, where N = 32 in this example. In the memory device 30 of the present embodiment, the digital memory is provided by digital memory (here SRAM) cells 31, which store only the p = (32 - n) least significant bits (LSBs) of each 32-bit weight w of the synaptic layer. The reduced-precision weights W defined by the remaining highest n most significant bits (MSBs) of the respective weights w are stored in the digital memory elements 32 of the SRAM analog MAC array 33 of the MAC unit. A global memory controller 34 shared by the digital storage cells 31 and the MAC unit 7 acts on the storage cells and the programming of the weights for the input / output of the signals to / from the MAC array 33 for signal propagation. After each weight update calculation ( Figure 5 step 23), the controller 34 stores the updated 32-bit weight w + Δw by reprogramming the p LSBs of the weight in the digital SRAM 31 and the n-bit portion of the weight in the MAC array 33.

[0045] Figure 7 is a more detailed illustration of an embodiment of the analog MAC array 33. The array 33 includes rows and columns of SRAM unit cells 35. Each row provides a digital memory element 32 for storing the n-bit reduced-precision weight W. Each of the n bits is stored in a corresponding unit cell 35 of the element. Each unit cell 35 contains a binary SRAM cell of the digital memory element 32 and an analog circuit for implementing the analog MAC array. The structure of these unit cells (hereinafter referred to as "analog" SRAM cells) is as Figure 8 shown. Each unit cell 35 includes a binary SRAM cell 38, a capacitor 39, and switches 40, 41a, and 41b connected as shown. The size of the capacitor 39 in each unit cell depends on the power of 2 corresponding to the bit stored in the connected binary cell 38. Figure 7 The first column of the unit cells 35 in [reference] stores the LSBs of each n-bit weight. If the capacitors 39 in these cells have a capacitance C, then: the capacitors 39 in the second column of the unit cells have a capacitance of (2 1 × C); the capacitors in the third column have a capacitance of (2 2 × C); and so on, up to the nth column, where the capacitor 39 has a capacitance of (2 n-1 × C). The rows of the cells 35 are connected to a word line control circuit 42, and the columns of the cells 35 are connected to a bit line control circuit 43. The control circuit includes standard SRAM circuits, such as an input voltage generator, a line driver / decoder circuit, a sense amplifier, and an ADC / DAC circuit, for addressing and reprogramming the cells and inputting / outputting signals as needed.

[0046] In the multiply-accumulate operation in the array 32, the SRAM cells 38 of the elements 32 are connected to Figure 4 the appropriate row lines r of the array i . The input voltage generator applies different analog voltages to each row, where each voltage corresponds to the value of the input signal x for that row. By closing the switch 41a, all the capacitors in the analog SRAM cell 35 are charged to that value. Then, the input voltage is turned off and the switch 41a is opened, so that the SRAM cells 38 in the analog cell 35 then discharge their adjacent capacitors based on whether these cells store "0" or "1". Specifically, if the cell 38 stores "0", the switch 40 is closed to discharge the capacitor. If the cell 38 stores "1", the switch 40 remains open, as Figure 8 shown. This step effectively multiplies the SRAM cell values by the input voltage. Subsequently, the switch 41b in the SRAM cells connected to the same column line c j is closed to short-circuit all the capacitors in the same column, and an analog addition and averaging operation is performed by redistributing the charge on these capacitors. Different bit powers are accommodated by exponential sizing of the capacitors. Thus, the resulting output voltage on the capacitors in the column line corresponds to the result of the multiply-accumulate operation and is obtained via the ADC.

[0047] Figure 9 Another embodiment of the memory device is shown. In this memory device 45, digital memories are distributed in the MAC cell array. Each N-bit weight w of the synaptic layer is stored in the unit cell 46 of the combined digital / analog SRAM MAC array 47. Each unit cell 46 includes p=(32 - n) bits of digital SRAM, storing the p least significant bits of the weight, and n analog SRAM cells 35, which correspond to the rows of the MAC array 33 described above. The binary SRAM cells 38 of the n analog SRAM cells provide n-bit digital memory elements 32 that store the weight W with reduced n-bit precision. The memory controller 48 controls access to the analog SRAM cells 35 of the unit cell 46 for the multiply-accumulate operation as described above, and access to the digital SRAM of the unit cell 46. According to one embodiment of the present invention, the structure of the combined MAC array 47 is shown in more detail in Figure 10 . The combined digital / analog unit cell of this embodiment provides a small on-chip footprint for a high area efficiency implementation.

[0048] Figure 11 The structure of yet another embodiment of the memory device is shown. Compared with Figure 6The corresponding components are indicated by like reference numerals. The memory device 50 includes digital memory (here SRAM) cells 51 that store all bits of each 32-bit weight w of the synaptic layer. As described above, the reduced n-bit precision weights W are stored in the digital memory elements 32 of the SRAM analog MAC array 33. As previously described, a standard SRAM controller 52 controls the digital SRAM cells 51 and a MAC controller 53 controls the MAC array 33. In this embodiment, after each weight update calculation ( Figure 5 step 23), the SRAM controller 52 reprograms the 32-bit weight w stored in the digital cells 51 to the updated weight w+Δw. Thus, weight updates are accumulated in the digital SRAM 51. The SRAM controller 52 periodically (e.g., after weight update operations have been performed for a batch of training examples) copies the reduced n-bit precision weights W from the cells 51 to the MAC cells via the MAC controller 53. Thus, the n-bit portion of the updated 32-bit weights is copied to the digital memory elements 32 of the MAC array 33 that store the corresponding reduced precision weights. In the embodiments of the invention described herein, the memory controller 52 may be adapted to copy the n-bit portion of the updated weight w to the MAC cells only when a bit overflow occurs in the (N-p)th bit during the update of the weight during the batch weight update operation. This reduces the number of programming operations for updating the reduced precision weights and thus reduces the data transfer between the SRAM 51 and the MAC cells.

[0049] In Figure 6 and Figure 9 memory devices, the reduced precision weights W of each synaptic layer are stored in the corresponding arrays of the MAC cells. With Figure 11 the memory structure, as the signal propagation proceeds, a given MAC array can be reused for the weights W of different synaptic layers. Specifically, under the control of the DPU 4, as the signal propagation operations proceed and different layers become active, the SRAM controller 52 can dynamically store the n-bit portions of the weights w of consecutive active synaptic layer sets (one or more) in the MAC array. The MAC array can be used to perform propagation on the active layers for a batch of training examples and can then be reprogrammed with the reduced precision weights for subsequent active layer sets.

[0050] In Figure 11 a modification of the embodiment, if the weight matrix of a synaptic layer is too large for the MAC array to store all the weights W of that layer, then the signal propagation operation can be performed by storing blocks (valid submatrices) of the weights W consecutively in the MAC array, performing multiply-accumulate operations for each block, and then accumulating the resulting signals for all blocks in the DPU 4.

[0051] Figure 12Shows another embodiment of the analog MAC array in the memory device 2. Except that all the analog SRAM cells 56 of the array 55 include capacitors having the same capacitance C, the components of the array 55 generally correspond to Figure 7 's components. The array control circuit of this embodiment includes the digital shift-add circuit indicated by 57. This circuit performs a shift-add operation on the outputs on the column lines during the multiply-accumulate operation to accommodate different powers of 2 of the bits stored in different columns of the cells 56. After the digitization of the column line outputs, the circuit 57: shifts the digital output value of the nth column by (n - 1) bits; shifts the digital output value of the (n - 1)th column by (n - 2) bits; and so on. Then, the results from all n columns are added in the circuit 57 to obtain Figure 4 the result of the multiply-accumulate operation of the n-bit weights in the column of memory elements in the logical array configuration of Figure 10 . The MAC array 55 can also be integrated in a combined digital / analog array structure corresponding to the

[0052] According to the network, the weights in different synaptic layers can span different ranges, and it may not be optimal to use the same n bits of N-bit weights to represent weights W with reduced precision. This can be solved by defining the initial value of each synaptic layer parameter m according to the number of most significant zero bits in the weights of each synaptic layer (see Figure 3 ). Specifically, if all the weights w in the layer have M (M > 0) most significant zero bits, then m can be set to the initial value m = M. Then, (at least) a plurality of p LSBs stored in the digital memory are defined as p = (N - n - m). During training, under the control of the memory controller 8, the value of m is adjusted according to the change in the number of most significant zero bits in the weights of the synaptic layer. In response to adjusting the value of m, the n-bit portion of the weights of the synaptic layer is redefined according to the adjusted value of m. Thus, Figure 3 the n-bit portion in

[0053] effectively "slides" along the N-bit weight value because m changes with the number of zero MSBs in the weights {w}. Then, when it is necessary to store the redefined n-bit portion of the weights, the memory controller 8 reprograms the MAC cells. For example, the redefined n-bit portion can be copied from the N-bit weights stored in the digital SRAM of the memory device. -mScaling. When a bit overflow of the (N - m)-th bit is detected during the weight update of the N-bit weights in the digital memory, the memory controller 8 can reduce the value of m for the layer. The memory controller can periodically read the current n-bit weights stored for the layer and increase m when the MSB of all n-bit weights is zero. This scheme gives a more optimal definition of the weights for the multiply-accumulate operation and improves the training accuracy.

[0054] Of course, many changes and modifications can be made to the exemplary embodiments of the present invention described. For example, although the multiply-accumulate operation is performed in the MAC unit 7 for the forward propagation and backpropagation operations above, embodiments of the present invention can be envisioned where the MAC unit 7 is used only for one of the forward propagation and backpropagation. For example, the forward propagation can be performed using the MAC unit 7, and the backpropagation calculation is completed in the DPU 4.

[0055] The steps of the flowchart can be implemented in an order different from the order shown, and some steps can be performed in parallel where appropriate. Generally, in cases where features are described herein with reference to methods embodying aspects of the present invention, corresponding features can be provided in an apparatus embodying aspects of the present invention, and vice versa.

[0056] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to a person of ordinary skill in the art without departing from the scope and spirit of the embodiments described herein. The terms used herein are chosen to best explain the principles of the embodiments of the present invention, the practical application, or the technical improvement of the technology found in the market, or to enable other persons of ordinary skill in the art to understand the embodiments of the present invention disclosed herein.

[0057] The present invention can be a system, a computer-implemented method, and / or a computer program product. The computer program product can include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.

[0058] A computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer floppy disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device (such as a punched card or raised structures in a groove record instructions thereon), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium shall not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (such as an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0059] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or an external storage device via a network (such as the Internet, a local area network, a wide area network, and / or a wireless network). The network can include a copper transmission cable, an optical transmission fiber, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the corresponding computing / processing device.

[0060] The computer-readable program instructions for carrying out operations of the present invention may be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages (such as Smalltalk, C++) and conventional procedural programming languages (such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), may execute the computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit to perform aspects of the present invention.

[0061] Aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0062] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture including instructions implementing aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0063] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0064] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart and combinations of blocks in the block diagrams and / or flowchart may be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0065] The description of the various embodiments of the present invention has been presented for purposes of illustration, but is not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the embodiments described herein. The terms used herein were chosen to best explain the principles of the embodiments, the practical application, or technical improvement of technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for training an artificial neural network, the artificial neural network having a series of neuron layers with inserted synaptic layers, each synaptic layer having a corresponding set of N-bit fixed-point weights {w} for weighting signals propagating between its adjacent neuron layers via an iterative loop of signal propagation operations and weight update calculation operations, the method comprising, for each synaptic layer: Storing multiple p least significant bits of each N-bit weight w in a digital memory; Storing the next n-bit portion of each weight w in an analog multiply-accumulate unit including an array of digital memory elements, each digital memory element including n binary storage units for storing respective bits of the n-bit portion of the weight, where n≥1 and (p + n + m) = N, where m≥0 corresponds to a defined number of most significant zero bits in the synaptic layer weights; Performing the signal propagation operation by providing the signal to be weighted by the synaptic layer to the multiply-accumulate unit to obtain an accumulated weighted signal that depends only on the stored n-bit portion of the weight; Performing the weight update calculation operation in a digital processing unit operatively coupled to the digital memory and the multiply-accumulate unit to calculate updated N-bit weights of the synaptic layer based on signals propagated by the neuron layers; And Periodically reprogramming the digital memory and the multiply-accumulate unit to store the updated weights.

2. The method according to claim 1, wherein For the synaptic layer, m is defined as m = 0, whereby (p + n) = N.

3. The method according to claim 2, wherein, Only the p least significant bits of each N-bit weight are stored in the digital memory.

4. The method according to claim 3, wherein The reprogramming is performed after the weight update calculation operation by reprogramming the p least significant bits of the weights in the digital memory and the n-bit portion of the weights in the multiply-accumulate unit.

5. The method according to claim 4, wherein, The digital memory is provided in digital storage units, and wherein the reprogramming is performed by a memory controller shared by the digital storage units and the multiply-accumulate unit.

6. The method according to claim 4, wherein, The digital memory is distributed in the multiply-accumulate unit such that each N-bit weight is stored in a unit cell that includes p bits of the digital memory storing the p least significant bits of the weight and the digital memory element storing the n-bit portion of the weight.

7. The method according to claim 2, including storing all N bits of each N-bit weight in the digital storage units providing the digital memory.

8. The method according to claim 7, wherein The reprogramming is performed by: After the weight update calculation operation, reprogramming the N-bit weights in the digital storage units to the updated weights; And Periodically copying the n-bit portion of the updated weights in the digital storage units to the digital memory elements that store the n-bit portion of the weights in the multiply-accumulate unit.

9. The method according to claim 8, including copying the n-bit portion of the updated weights to the digital memory elements after a batch of weight update calculation operations.

10. The method according to claim 9, including copying the n-bit portion of the updated weight to the digital memory element only when a bit overflow occurs in the (N - p)-th bit during the update of the weight in the batch weight update calculation operation.

11. The method according to claim 7, further including: Storing the N-bit weights of all synaptic layers in the digital storage unit; And Dynamically storing the n-bit portions of the weights of a set of consecutive synaptic layers in the digital memory element array to perform the signal propagation operation.

12. The method according to claim 1, further including: Defining an initial value of m for the synaptic layer according to the number of most significant zero bits in the weight of the synaptic layer; For the synaptic layer, defining the plurality of p as p = (N - n - m); Adjusting the value of m during the training according to the change in the number of most significant zero bits in the weight of the synaptic layer; And In response to adjusting the value of m, redefining the n-bit portion of the weight of the synaptic layer according to the adjusted value of m, and reprogramming the digital memory element array to store the redefined n-bit portion of the weight.

13. The method according to claim 1, wherein, Each of the signal propagation operations includes a forward propagation operation in which a signal propagates through the network from a first neuron layer, and a backward propagation operation in which a signal propagates backward through the network from a last neuron layer. The method includes, for each synaptic layer, providing the signal weighted by the synaptic layer to the multiply-accumulate unit in the forward propagation operation and the backward propagation operation.

14. The method according to claim 1, including defining a corresponding value of n for each synaptic layer.

15. The method according to claim 1, wherein, For each synaptic layer, N = 32 and n ≤ 8.

16. An apparatus for implementing an artificial neural network, the artificial neural network having a series of neuron layers with inserted synaptic layers, each synaptic layer having a corresponding set of N-bit fixed-point weights {w} for weighting signals propagated between adjacent neuron layers in an iterative training loop of signal propagation operations and weight update calculation operations. The apparatus includes: A digital memory storing the p least significant bits of each N-bit weight w of each synaptic layer; An analog multiply-accumulate unit storing the next n-bit portion of each weight w of the synaptic layer, the multiply-accumulate unit including an array of digital memory elements, each digital memory element including n binary storage units for storing respective bits of the n-bit portion of the weight, where n ≥ 1 and (p + n + m) = N, where m ≥ 0 corresponds to a defined number of most significant zero bits in the weight of the synaptic layer; And A digital processing unit operatively coupled to the digital memory and the multiply-accumulate unit, the digital processing unit being adapted to: In the signal propagation operation, obtain an accumulated weighted signal that depends only on the n-bit portion of the stored weight by providing the signal to be weighted by each synaptic layer to the multiply-accumulate unit; Perform the weight update calculation operation to calculate the updated N-bit weights for each synaptic layer based on the signals propagated by the neuron layer; and Control the periodic reprogramming of the digital memory and the multiply-accumulate unit to store the updated weights.

17. The apparatus according to claim 16, wherein, For the synaptic layer, m is defined as m = 0, whereby (p + n) = N.

18. The device according to claim 17, wherein, Only the p least significant bits of each N-bit weight are stored in the digital memory.

19. The apparatus according to claim 18, comprising a digital storage unit providing the digital memory, and a memory controller shared by the digital storage unit and the multiply-accumulate unit for performing the reprogramming.

20. The apparatus according to claim 18, wherein, The digital memory is distributed in the multiply-accumulate unit such that each N-bit weight is stored in a unit cell, the unit cell comprising p bits of the digital memory storing the p least significant bits of the weight, and the digital memory element storing the n-bit portion of the weight.

21. The device according to claim 17, wherein, Store all N bits of each N-bit weight in the digital storage unit providing the digital memory.

22. The apparatus according to claim 21, wherein, The N-bit weights of all synaptic layers are stored in the digital storage unit, and wherein the apparatus is adapted to dynamically store the n-bit portions of the weights of a set of consecutive synaptic layers in the digital memory element array to perform the signal propagation operation.

23. The apparatus according to claim 16, wherein, The multiply-accumulate unit includes a corresponding array of the digital memory elements storing the n-bit portions of the weights of each synaptic layer.

24. The apparatus according to claim 16, wherein, For each synaptic layer, an initial value of m is defined according to the number of most significant zero bits in the weight of the synaptic layer, and for the synaptic layer, the plurality of p is defined as p = (N - n - m), wherein the apparatus is adapted to: Adjust the value of m for the synaptic layer during the training according to the change in the number of most significant zero bits in the weight of the synaptic layer; and In response to adjusting the value of m, redefine the n-bit portion of the weight of the synaptic layer according to the adjusted value of m, and reprogram the digital memory element array to store the redefined n-bit portion of the weight.

25. The apparatus according to claim 16, wherein, The binary storage unit includes SRAM cells.

26. A computer program product comprising program code which, when run on a computer, causes the method according to any one of claims 1 to 15 to be performed.

Citation Information

Patent Citations

  • System for implementing a sparse coding algorithm

    US20160358075A1

  • Compute in memory circuits with multi-VDD arrays and / or analog multipliers

    US20190042199A1