Dynamic configuration of a readout circuit for different operations in an analog resistor crossbar array

By dynamically configuring the readout circuit with different signal bounds for forward and backward pass operations, the system optimizes signal processing and quantization, resolving the vanishing gradient issue and enhancing the accuracy of neuromorphic computing systems.

JP7710521B2Active Publication Date: 2025-07-18INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023536806
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-12-15
Filing Date
2021-10-25
Publication Date
2025-07-18
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

Existing neuromorphic computing systems face challenges in efficiently performing forward and backward pass operations due to mismatched signal ranges, leading to issues like vanishing gradients and inadequate quantization of output signals, particularly in the backward pass operation.

Method used

The system dynamically configures the readout circuit to have different signal bounds for forward and backward pass operations by adjusting integration capacitors and ADC resolutions, ensuring optimal signal processing and quantization for each operation mode.

Benefits of technology

This approach enhances the detection and quantization of output signals during both forward and backward pass operations, addressing the vanishing gradient issue and improving the accuracy and efficiency of neuromorphic computing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710521000007
    Figure 0007710521000007
  • Figure 0007710521000008
    Figure 0007710521000008
  • Figure 0007710521000009
    Figure 0007710521000009
Patent Text Reader

Abstract

A device comprising an array of resistive processing unit (RPU) cells, a first plurality of control lines extending across the array of RPU cells in a first direction, and a second plurality of control lines extending across the array of RPU cells in a second direction. A peripheral circuit comprising a readout circuit is coupled to the first and second control lines. A control system generates control signals to control the peripheral circuit to perform a first operation on the array of RPU cells and to perform a second operation on the array of RPU cells. The control signals include a first configuration control signal for configuring the readout circuit to have a first hardware configuration when the first operation is performed on the array of RPU cells; and a second configuration control signal for configuring the readout circuit to have a second hardware configuration different from the first hardware configuration when the second operation is performed on the array of RPU cells.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to an analog resistive processing system for neuromorphic computing and to techniques for controlling peripheral circuits of the analog resistive processing system to perform various operations on an array of one of a plurality of resistive processing units (RPUs) of the analog resistive processing system.

Background Art

[0002] Information processing systems, such as neuromorphic computing systems and artificial neural network (ANN) systems, are being used in various applications, such as machine learning and inference processing for cognitive recognition and computing. Such systems are generally hardware-based systems that include a number of highly interconnected processing elements (referred to as "artificial neurons") that operate in parallel to perform various types of calculations. The artificial neurons (e.g., pre-synaptic neurons and post-synaptic neurons) are connected using artificial synapse devices that provide synaptic weights representing the connection strength between the plurality of artificial neurons. The synaptic weights can be implemented using an array of one of a plurality of RPU cells having tunable resistive memory devices, and the conductance states of the plurality of RPU cells are encoded or otherwise mapped to the synaptic weights.

Summary of the Invention

Means for Solving the Problems

[0003] Embodiments of the present disclosure include an analog resistance processing system for neuromorphic computing, and a technique for dynamically configuring the hardware configuration of a peripheral circuit of the analog resistance processing system, such as a readout circuit, when performing different operations on one array of a plurality of RPU cells of the analog resistance processing system.

[0004] An exemplary embodiment includes a device having one array of RPUs, a plurality of first control lines extending in a first direction across the one array of a plurality of RPU cells, and a plurality of second control lines extending in a second direction across the one array of a plurality of RPU cells. Each RPU cell is connected at an intersection of one of the plurality of first control lines and one of the plurality of second control lines, where the peripheral circuit includes a readout circuit. A control system is operably connected to the peripheral circuit. The control system controls the peripheral circuit to generate control signals for performing a first operation on the one array of a plurality of RPU cells and for performing a second operation on the one array of a plurality of RPU cells. The control signals include a first configuration control signal for configuring the readout circuit to have a first hardware configuration when the first operation is performed on the one array of a plurality of RPU cells; and a second configuration control signal for configuring the readout circuit to have a second hardware configuration different from the first hardware configuration when the second operation is performed on the one array of a plurality of RPU cells.

[0005] Other embodiments will be described in the following detailed description of the invention of the exemplary embodiments to be read in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6

Figure 7A

Figure 7B

Figure 8

Figure 9A

Figure 9B

[0007] Here, embodiments of the present invention will be described in further detail with respect to an analog resistive processing system for neuromorphic computing, and a technique for dynamically configuring a peripheral circuit of the analog resistive processing system, such as a readout circuit, to perform different operations on one array of a plurality of resistive processing unit (RPU) cells of the analog resistive processing system. For example, in some embodiments, as will be described in further detail below, the hardware configuration of the readout circuit of the analog resistive processing system is dynamically set to have different operation signal ranges for different operations (e.g., forward pass operation and backward pass operation of a neural network training process) performed on one array of a plurality of RPU cells of the analog resistive processing system.

[0008] It should be understood that the various features shown in the accompanying drawings are schematic diagrams not drawn to scale. Moreover, the same or similar reference numerals are used throughout the drawings to indicate the same or similar features, elements, or structures, and thus, a detailed description of the invention for the same or similar features, elements, or structures will not be repeated for each of the drawings. Further, the term "exemplary" as used herein means "serving as an example, instance, or illustration." Any embodiment or design described herein as "exemplary" should not be construed as preferred or advantageous over other embodiments or designs.

[0009] Exemplary embodiments of the present disclosure include in-memory computing systems or computing memory systems that utilize an array of a plurality of RPU cells for two purposes: storing data and processing the data to perform some computational tasks. In some embodiments, the plurality of RPU cells implement resistive memory devices, such as resistive random-access memory (ReRAM) devices, phase-change memory (PCM) devices, etc., which have a tunable conductance (G) with variable conductance states ranging from a minimum conductance (Gmin) to a maximum conductance (Gmax). As described above, neuromorphic computing systems and ANN systems are a type of in-memory computing system where artificial neurons are connected using artificial synapse devices to provide synaptic weights representing the strength of the connection between two artificial neurons. The synaptic weights can be implemented using a plurality of tunable resistive memory devices, where the variable conductance states represent the synaptic weights and are used to perform computations (e.g., vector-matrix multiplication). The conductance states of the analog resistive memory devices are encoded or otherwise mapped to the synaptic weights.

[0010] Various types of artificial neural networks, such as deep neural networks (DNNs) and convolutional neural networks (CNNs), implement neuromorphic computing architectures for machine learning applications, such as image recognition, object recognition, and speech recognition. In-memory computing associated with such neural networks includes, for example, training computations where synaptic weights of a plurality of resistive memory cells are optimized by processing a training data set, and forward inference computations where a trained neural network is used to process input data for purposes such as classifying the input data or predicting an event based on the input data.

[0011] Training of DNNs generally relies on the backpropagation algorithm, which includes the following three iterative cycles: a forward cycle, a backward cycle, and a weight update cycle, which are repeated many times until a convergence criterion is met. The forward cycle and the backward cycle mainly involve calculating vector-matrix multiplications in the forward and backward directions. This operation can be performed on one 2D array of a plurality of analog resistive memory cells. In the forward cycle, the stored conductance values of the resistive memory devices in the one 2D array form a matrix, and an input vector is transmitted as a voltage pulse through each input column of the 2D array. In the backward cycle, a voltage pulse is supplied as an input from the columns, and the vector-matrix product is calculated with the transpose of the matrix. The weight update includes calculating a vector-vector outer product composed of a multiplication operation and an incremental weight update to be locally performed in each resistive memory cell within the one 2D array.

[0012] A probabilistically trained DNN comprising an array of a plurality of RPU cells can have synaptic weights implemented using tunable resistive devices. In some embodiments, each RPU cell in computing system 100 comprises a resistive element having a conductance value representing a matrix element or weight of the RPU cell 110. In some embodiments, the plurality of resistive elements of the plurality of RPU cells 110 are resistive devices, such as resistive switching devices (interface or filamentary switching devices), ReRAM devices, memristor devices, PCM devices, etc., and other types of devices having a tunable conductance (or a tunable resistance level) that can be programmatically adjusted within a range of a plurality of different conductance levels for adjusting the weights of the RPU cell 110. In some embodiments, the variable conductance elements of the plurality of RPU cells 110 can be implemented using ferroelectric devices, such as ferroelectric field effect transistor devices. Moreover, in some embodiments, the plurality of RPU cells 110 can be implemented using an analog CMOS-based framework in which each RPU cell 110 comprises a capacitor and a read transistor. In this framework, the capacitor functions as a memory element of the RPU cell 110 and stores a weight value in the form of a capacitor voltage, where the capacitor voltage is applied to a gate terminal of the capacitor voltage to modulate the channel resistance of the read transistor based on the level of the capacitor voltage, where the channel resistance of the read transistor represents the conductance of the RPU cell 110 and correlates with the level of the read current generated based on the channel resistance.

[0013] To properly train a DNN and achieve high accuracy, a strict set of specifications for acceptable RPU device parameters where a given DNN algorithm can be tolerated without significant error penalties must be met by the operating characteristics of tunable resistive devices. These specifications include, for example, variations in the switching characteristics of resistive memory devices, such as the minimum incremental conductance change (±Δg min ) due to a single potential difference, symmetry in the up and down conductance changes, the tunable range of the conductance values, etc.

[0014] FIG. 1 schematically illustrates a computing system comprising an array of one of a plurality of resistive processing unit cells according to an exemplary embodiment of the present disclosure. More specifically, FIG. 1 schematically shows a computing system 100 (e.g., a neuromorphic computing system) comprising a two-dimensional (2D) crossbar array of a plurality of RPU cells 110 arranged in a plurality of rows R1, R2, R3, …, Rm and a plurality of columns C1, C2, C3, …, Cn. The plurality of RPU cells 110 in each row R1, R2, R3, …, Rm are commonly connected to respective row control lines RL1, RL2, RL3, …, RLm (collectively referred to as row control line RL). The plurality of RPU cells 110 in each row RL1, RL2, RL3, …, RLm are commonly connected to respective column control lines CL1, CL2, CL3, …, CLn (collectively referred to as column control line CL). Each RPU cell 110 is connected at (and between) each cross point (i.e., intersection) of one of the row control line and the column control line. In one exemplary embodiment, Computing system 100 comprises a 4,096×4,096 array of a plurality of RPU cells 110.

[0015] Computing system 100 further includes a peripheral circuit 120 connected to row control lines RL1, RL2, RL3, …, RLm, and a peripheral circuit 130 connected to column control lines CL1, CL2, CL3, …, CLn. Further, the peripheral circuit 120 is connected to a data input / output (I / O) interface block 125, and the peripheral circuit 130 is connected to a data I / O interface block 135. The computing system 100 further includes a control signal circuit 140 that includes various types of circuit blocks, such as power, clock, bias, and timing circuits, for providing power distribution and control signals and clocking signals for the operation of the peripheral circuits 120 and 130 of the computing system 100.

[0016] In a neuromorphic computing application, a plurality of RPU cells 110 include artificial synapses that provide weighted connections between pre-neurons and post-neurons. A plurality of pre-neurons and post-neurons are connected via a 2D crossbar array of the plurality of RPU cells 110, which naturally represents a fully connected neural network. In some embodiments, the computing system 100 is configured to perform DNN or CNN computations, where the conductance of each RPU cell 110 represents a matrix element or weight w ij which can be updated or accessed through the operation of the peripheral circuits 120 and 130 (where w ij(represents the weight value for the i-th row and j-th column in the one array of the plurality of RPU cells 110). As described above, DNN training generally relies on a backpropagation process that includes three iterative cycles, namely, a forward cycle, a backward cycle, and a weight update cycle. The computing system 100 can be configured to execute all three cycles of the backpropagation process in parallel, thus potentially providing a significant acceleration to DNN training using lower power and reduced computing resources. The computing system 100 can be configured to execute vector-matrix multiplication operations in the analog domain in parallel.

[0017] The row control line RL and the column control line CL are each shown as a single line in FIG. 1 for ease of illustration, but it should be understood that the control lines for each row and each column can include more than two control lines connected to the plurality of RPU cells 110 in their respective rows and columns, depending on the implementation and the particular architecture of the plurality of RPU cells 110. For example, in some embodiments, each row control line RL can comprise a complementary pair of word lines for a given RPU cell 110. Additionally, each column control line CL can comprise a plurality of control lines comprising, for example, one or more source lines (SL: source line) and one or more bit lines (BL: bit line).

[0018] Peripheral circuits 120 and 130 include various circuit blocks connected to respective rows and columns in a 2D array of a plurality of RPU cells 110, and the circuit blocks execute vector-matrix multiply functions, matrix-vector multiply functions, and outer product update operations to implement forward operations, backward operations, and weight update operations of a backpropagation process (for neural network training), and inference processing using a trained neural network. For example, in some embodiments, to support read operations / sensing operations of RPU cells (e.g., reading weight values of a given RPU cell 110), peripheral circuits 120 and 130 generate pulse-width modulation (PWM) read pulses in response to input vector values (read input values) received during a forward cycle / backward cycle and apply the pulse-width modulation (PWM) read pulses to a plurality of RPU cells 110, and include a pulse-width modulation (PWM) circuit and a read pulse drive circuit.

[0019] More specifically, in some embodiments, peripheral circuits 120 and 130 include digital-to-analog (D / A) conversion circuits, and the digital-to-analog (D / A) conversion circuits are configured to receive a digital input vector (to be applied to a row or column) and convert the digital input vector into an analog input vector value represented by an input voltage whose pulse width varies. In some embodiments, a time encoding scheme is used when the input vector is represented by a fixed amplitude Vin = 1V pulse having a tunable duration (e.g., the pulse duration is a multiple of 1 nanosecond (ns) and is proportional to the value of the input vector). The input voltage applied to a row (or column) generates an output vector value represented by an output current, where the weights of a plurality of RPU cells 110 are read by measuring the output current.

[0020] Peripheral circuits 120 and 130 further include a current integration circuit and an analog-to-digital (A / D) conversion circuit that integrate the read current (I READ ) output and accumulated from a plurality of connected RPU cells 110, and then convert the integrated current into a digital value (read output value) for subsequent calculations. In particular, the currents generated by the plurality of RPU cells 110 are combined on a column (or row), and this total current is integrated over a measurement time T meas by the current readout circuits of peripheral circuits 120 and 130. The current readout circuit includes an integrating ammeter and an analog-to-digital (A / D) converter. In some embodiments, each integrating ammeter includes an operational amplifier that integrates the current output from a given column (or row) on a capacitor (or the differential current from a pair of a plurality of RPU cells implementing positive and negative weights), and the analog-to-digital (A / D) converter converts the integrated current (e.g., an analog value) into a digital value.

[0021] Data I / O interfaces 125 and 135 are configured to interface with a digital processing core, where the digital processing core is configured to process inputs / outputs to the computing system 100 (neural core) and route data between different RPU arrays. Data I / O interfaces 125 and 135 are configured to (i) receive external control signals and data from the digital processing core and provide the received control signals and data to peripheral circuits 120 and 130, and (ii) receive digital readout output values from peripheral circuits 120 and 130 and transmit the digital readout output values to the digital processing core for processing. In some embodiments, the digital processing core implements a non-linear function circuit that calculates activation functions (e.g., sigmoid neuron function, softmax, etc.) and other arithmetic operations on data to be provided to the next or previous layer of a neural network.

[0022] As is known in the art, a fully connected plurality of DNNs comprises a stack of a plurality of fully connected layers through which a signal propagates from an input layer to an output layer through a series of linear and non-linear transformations. The entire DNN represents a single differentiable error function that maps input data to class scores at the output layer. Typically, a DNN is trained using a simple stochastic gradient descent (SGD) scheme in which the error gradient with respect to each parameter is calculated using the backpropagation algorithm. The backpropagation algorithm consists of three cycles, namely, a forward cycle, a backward cycle, and a weight update cycle, which are repeated many times until a convergence criterion is met. The forward cycle and the backward cycle mainly involve vector-matrix multiplication operations in the forward and backward directions using the 2D crossbar array of the RPU Se in row 110 of FIG. 1, including vector-matrix multiplication operations in the forward and backward directions.

[0023] In the computing system 100 of FIG. 1, the conductance value g ij is the weight value w in the 2D crossbar array of a plurality of RPU cells ijForm the matrix W. In the forward cycle (Figure 2A), the input vector (in the form of a voltage pulse) is transmitted through each of the input rows in the 2D crossbar array to perform a vector-matrix multiplication in a plurality of RPU cells 110. In the backward cycle (Figure 2B), the voltage pulses supplied from the columns are input into the plurality of RPU cells 110, and the vector-matrix product is calculated with the transpose of the values of the weight matrix W. In contrast to the forward and backward cycles, implementing weight updates on a 2D crossbar array of resistive devices requires calculating a vector-vector outer product that includes a multiplication operation and an incremental weight update that is locally executed at each cross-point RPU device in an array. Figures 2A, 2B, and 2C schematically show the forward path, backward path, and weight update operations, respectively, of the backpropagation algorithm that can be executed using the computing system 100 of Figure 1.

[0024] In the case of a single fully connected layer where N input neurons are connected to M output (or hidden) neurons, the forward path (Figure 2A) includes calculating the vector-matrix multiplication y = Wx, where the vector x of length N represents the activity of the input neurons, and the matrix W of size M×N stores the weight values between each pair of input and output neurons. The resulting vector y of length M is further processed by performing a non-linear activation on each of the elements and then passed to the next layer. Once the information reaches the last output layer, an error signal is calculated and backpropagated through the network. In the forward cycle, the conductance values stored in the crossbar array of the plurality of RPU cells 110 form a matrix, while the input vector is transmitted as a voltage pulse through each of the input rows R1, R2, R3, …, Rm.

[0025] The backward cycle (Figure 2B) on a single layer also involves the transpose of the weight matrix z = W TIncluding vector matrix multiplication at δ, where W represents the weight matrix, where the vector δ of length M represents the error calculated by the output neuron, and where the vector z of length N is processed using the derivative of the neuron's non-linearity and then passed to the previous layer. In the backward cycle, voltage pulses are supplied as input to a plurality of RPU cells 110 from columns CL1, CL2, CL3, …, CLn, and the vector matrix product is calculated for the transpose of the weight matrix W.

[0026] Finally, in the update cycle (Figure 2C), the weight matrix W is updated by performing the outer product of two vectors used in the forward cycle and the backward cycle. In particular, implementing weight updates locally and all in parallel on a 2D crossbar array of resistive devices, independent of the one array size, requires calculating a vector-vector outer product consisting of multiplication operations and incremental weight updates that should be performed locally at each cross point (RPU cell 110) in the computing system of Figure 1. As schematically shown in Figure 2C, the weight update process is calculated as follows: w ij ←w ij +ηx i ×δ j where w ij represents the weight value for the i-th row and j-th column (layer indices are omitted for clarity), where x i is the activity at the input neuron, δ j is the error value calculated by the output neuron, and where η represents the global learning rate.

[0027] In summary, all operations in the weight matrix W can be implemented using a 2D crossbar array of two-terminal RPU devices having M rows and N columns, where the conductance values stored in the crossbar array form the matrix W. In the forward cycle, the input vector x is transmitted as a voltage pulse to each of the rows, and the resulting vector y can be read as a current signal from the columns. Similarly, when a voltage pulse is supplied as an input from the columns in the backward cycle, the vector matrix product is calculated in the transpose of the weight matrix W T In the last, in the update cycle, voltage pulses representing the vectors x and δ are supplied simultaneously from the rows and columns. In the update cycle, each RPU cell 110 performs local multiplication and sum operations by processing the voltage pulses coming from the columns and rows, thus achieving an incremental weight update.

[0028] The x of the weight update cycle i vector and δ j To obtain the product of the vector and the δ vector, the probabilistic translator circuits of the peripheral circuits 120 and 130 are utilized to generate probabilistic bitstreams representing the input vectors x i and δ j . The probabilistic bitstreams for the vectors x i and δ j are supplied through the rows and columns of the 2D crossbar array of a plurality of RPU cells, where the conductance of a given RPU cell is the probabilistic pulse stream of x i input to the given RPU cell and δ jwill vary depending on its coincidence with the probabilistic pulse stream. The vector cross product operation for the weight update operation is implemented based on the known concept that the coincidence detection (using AND logic gate operation) of the probabilistic stream representing real numbers is equivalent to the multiplication operation. With the three operation modes described above, the plurality of RPU cells forming the neural network can be active in all three cycles, thus enabling a very efficient implementation of the backpropagation algorithm for calculating the updated weight values of the plurality of RPU cells during the learning process of the DNN.

[0029] During the forward pass operation and the backward pass operation in which the vector matrix multiplication is executed on the RPU array, the digital input vector (x or δ) is converted into an analog input vector that is transmitted as voltage pulses having a fixed amplitude and tunable durations across rows and columns. In some embodiments, the maximum pulse duration represents unity with respect to the integration time (T meas →1), where all pulse durations are scaled appropriately depending on the value of x i or δ j . This scheme functions optimally in the forward cycle where all x i in x are within a given range, for example, [-1, 1]. However, this scheme becomes problematic in the case of the backward cycle because there is no guarantee for the range of the error signal values in δ. For example, as the training process progresses and the classification error becomes increasingly smaller, all δ j in δ can become significantly smaller than unity (δ << 1). In this case, as a result of the low signal strength of the input error signal δ of the input error vector δ j , the signal strength of the output signal generated during the backward pass operation may be too small to fit within the full-scale operation range of the readout circuit.

[0030] In some embodiments, based on the maximum input error signal δ j of the input error vector δ, the input error signal δj is the input signal x of the input vector applied to the RPU array during the forward pass operation so as to fall, for example, between [-1, 1] i and an input error signal δ similar to the full input range of j A noise management system can be implemented within the RPU array so that δ is scaled. However, in some cases, the process of matching the input signal ranges for the forward pass operation and the backward pass operation may be insufficient to match the output signal ranges resulting from the forward pass operation and the backward pass operation, for example, by the manner in which a neural network (e.g., a DNN) is trained.

[0031] In particular, in the case of the forward pass operation, the distribution of Wx becomes large. Because the values of the input vector x and the weight values of the weight matrix W typically correlate (e.g., achieved through DNN learning intended to enhance such correlation), while W for the backward pass operation T in the case of δ, the error vector δ and the matrix W T are typically uncorrelated vectors and thus have results close to zero. As a result, the output signal strengths of the forward pass operation and the backward pass operation are significantly different. Moreover, another reason for weak output signals being generated during the backward pass operation is that many elements δ of the input error vector δ for many DNNs j are typically zero or very close to zero. In this example, even when the input range of the value δ of the input error vector δ matches the input range of the input signal x of the input vector x for the forward pass operation (e.g., [-1, 1]), scaling the value δ of the input error vector δ j may not significantly increase the signal strength of the output signal generated by the backward pass operation. i of the input vector x for the forward pass operation j

[0032] ​In at least the case of the reasons described above, when the read circuit (shared for the forward path operation and the backward path operation) is configured to have a fixed output signal bound b (e.g., an operation signal range) that is more optimal for the range of output signals generated during the forward path operation, the output signal generated as a result of the vector matrix multiplication operation executed during the backward path operation may be too small and, in some cases, may not be easily detectable or quantizable. For example, the fixed output signal bound b is the result of the current integration circuit of the shared read circuit having an integration capacitor of a fixed size, or the ADC circuit of the shared read circuit having a fixed ADC resolution, etc. In such a case, a relatively small (e.g., close to zero) analog output signal will be quantized to zero due to the finite ADC resolution.

[0033] This effect is particularly severe when the shared ADC circuit is configured to have a relatively low resolution (e.g., 3-bit, 4-bit, 5-bit resolution), where the ADC bin size (i.e., the least significant bit (LSB) voltage) is relatively large. The use of a shared readout circuit (for forward and backward path operations) with a fixed output bound b (e.g., an integration capacitor of the same size or the same ADC resolution or a combination thereof, etc.) may be sufficient to effectively read out and quantize the voltage signal generated during the forward path operation, but it is insufficient to effectively read out and quantize the voltage signal generated during the backward path operation (which represents the error signal propagated through the neural network layer during the backward path operation). Since the backward propagating error signals have a much lower signal strength than the forward propagating signals, the fixed signal output bound b of the shared readout circuit inadequately quantizes the backward propagating signals, thereby resulting in an effect referred to as the "vanishing gradients".

[0034] As is known in the art, an ADC circuit converts an analog signal (a continuous-time or continuous-amplitude analog signal) into a digital signal through a quantization process. The resolution of an ADC refers to the number of discrete values that the ADC can generate over an acceptable range of analog input values. For example, an ADC with an 8-bit resolution can convert an analog input signal into 256 different levels (2 8It can be encoded into one of [0, 255] (i.e., unsigned integers) or [-128, 127] (i.e., signed integers), which provides a 256:1 dynamic range, depending on the application. The resolution of the ADC is also indicated by the least significant bit (LSB) voltage (also referred to as "voltage resolution"). The LSB of the ADC represents the smallest detectable interval. For an 8-bit ADC, the LSB is approximately 1 / 256, i.e., 3.9x10 -3 is such. The LSB voltage (or voltage resolution) refers to the change in voltage required to guarantee a change in the output code level. The voltage resolution of the ADC is equal to the allowable range of the analog input voltage value (e.g., full-scale operating voltage range) divided by the number of discrete intervals.

[0035] For example, assuming that the full-scale operating voltage range of an ADC is from -10V to +10V (where V can be in microvolts (μV), millivolts (mV), etc.), the voltage resolution of an 8-bit ADC is 20 / 256, which is approximately 0.078V. If the analog voltage value for a given operation (e.g., the backward path) is within a much smaller voltage range, such as [-0.1V to +0.1V], compared to the ADC's full-scale operating voltage range [-10V to +10V], the effective voltage resolution for the smaller voltage range is significantly limited because the LSB is 0.078V, and thus any analog voltage signal with an absolute value less than 0.078V will be quantized as zero. In other words, in the smaller voltage range [-0.1V to +0.1V], an 8-bit ADC cannot ideally resolve voltage differences smaller than 0.078V. In this scenario, for an 8-bit ADC, it would be more desirable to increase the voltage resolution of the 8-bit ADC to more discrete levels (e.g., 256 levels) to more effectively quantize the voltage values within the smaller voltage range (-0.1V to +0.1V). In some embodiments, a lower bit resolution with a smaller full-scale range can be implemented so that it can be implemented to measure smaller voltages.

[0036] Exemplary embodiments of the present disclosure implement a bound management technique implemented in the analog domain by dynamically changing the configuration (e.g., hardware configuration) of a shared readout circuit to provide different signal bounds for different operation modes of an analog RPU crossbar array, such as a forward pass operation and a backward pass operation. Various techniques can be used to dynamically configure the shared readout circuit to have a first configuration during the forward pass operation and a second configuration during the backward pass operation in order to enhance the output signal for the backward pass operation when (δ << 1). In particular, in the forward pass operation, the shared readout circuit operates with a first output signal bound b1 (e.g., a first operating signal range) of the readout circuit, and in the backward pass operation, the shared readout circuit is configured to have a first configuration where it operates with a second output signal bound b2 (e.g., a second operating signal range), where b2 < b1.

[0037] FIG. 3A schematically illustrates a method for constructing a computing system comprising an array of one of a plurality of resistive processing unit cells to perform a forward pass operation in accordance with an exemplary embodiment of the present disclosure. In particular, FIG. 3A schematically shows a computing system 300 comprising a crossbar array of a plurality of RPU cells 305, where each RPU cell 310 in one array 305 comprises an analog non-volatile resistive element (represented as a variable resistor having a tunable conductance G) at the intersection of each row (R1, R2,..., Rm) and each column (C1, C2,..., Cn). As illustrated in FIG. 3A, the one array of the plurality of RPU cells 305 has a conductance value G of each of the plurality of RPU cells 310 ij (where i represents the row index and j represents the column index) encoded synaptic weight W ij is mapped to a matrix of conductance values G ijA matrix is provided. One array of a plurality of RPU cells 305 represents one array of artificial programmable synaptic elements that connect the nodes (artificial neurons) of the upstream layer (e.g., the input layer or the intermediate layer) of an artificial neural network to the nodes (artificial neurons) of the next downstream layer (intermediate layer or output layer) of the artificial neural network.

[0038] In the case of the forward pass operation, the multiplexer in the peripheral circuit of the computing system 300 is activated to selectively connect the row line drive circuit 320 to the row lines R1, R2,..., Rm. The row line drive circuit 320 is composed of a plurality of DAC circuit blocks 322-1, 322-2,..., 322-m (collectively referred to as DAC circuit blocks 322) connected to the respective row lines R1, R2,..., Rm. In addition, the multiplexer in the peripheral circuit of the computing system 300 is activated to selectively connect the readout circuit 330 to the column lines C1, C2,..., Cn. The readout circuit 330 includes a plurality of readout circuit blocks 330-1, 330-2,..., 330-n connected to the respective column lines C1, C2,..., Cn. The readout circuit blocks 330-1, 330-2,..., 330-m each include a respective current integration circuit 332-1, 332-2,..., 332-m and a respective ADC circuit 334-1, 334-2,..., 334-m.

[0039] The forward pass operation in a neural network is performed to calculate the neuron activation of the downstream layer based on the neuron activation of the upstream layer (e.g., the input layer or the hidden layer) and the synaptic weights that connect the neurons of the upstream layer to the neurons of the downstream layer (e.g., the hidden layer or the output layer). In FIG. 3A, the upstream neuron generates a digital signal x = [x1, x2,..., x m (referred to as the input vector), where the digital signals x1, x2,..., x mare input into respective DAC circuit blocks 322-1, 322-2, ..., 322-m, thereby generating an analog voltage V(t) (voltage as a function of time) signal on the row lines R1, R2, ..., Rm, which is proportional to the excitation of the upstream neuron, i.e., x = [x1, x2, ..., x m is proportional to.

[0040] In some embodiments, each of the DAC circuit blocks 322-1, 322-2, ..., 322-m includes a pulse width modulation circuit and a drive circuit configured to generate pulse width modulation (PWM) readout pulses V1, V2, ..., V m applied to respective row lines R1, R2, ..., Rm. More specifically, in some embodiments, the DAC circuit blocks 322-1, 322-2, ..., 322-m are configured to perform a digital-to-analog conversion process using a time encoding scheme represented by fixed amplitude pulses (e.g., V = 1V) having a tunable duration for the input vector, where the pulse duration is a multiple of a predefined time period (e.g., 1 nanosecond) and is proportional to the value of the input vector. For example, a given digital input value of 0.5 can be represented by a voltage pulse of 4 nanoseconds, and a digital input value of 1 can be represented by a voltage pulse of 80 nanoseconds (e.g., a digital input value of 1 can be encoded into an analog voltage pulse having a pulse duration equal to the integration time T meas ). As shown in FIG. 3A, the resulting analog input voltages V1, V2, ..., V m (e.g., readout pulses) are applied to the row lines R1, R2, ..., Rm.

[0041] During the forward pass operation, the analog input voltages V1, V2, ..., V m (e.g., readout pulses) are applied to the row lines R1, R2, ..., Rm, where each RPU cell 310 has a corresponding readout current I READ = V i × G ijis generated based on Ohm's law, where V i represents the analog input voltage applied to a given RPU cell 310 in a given row i, and G ij represents the conductance value of a given RPU cell 310 (in a given row i and a given column j). As shown in FIG. 3A, the read currents generated by a plurality of RPU cells 310 on each column j are summed (based on Kirchhoff's current law), and respective currents I1, I2,..., I n are generated at the outputs of respective columns C1, C2,..., Cn. In this manner, the resulting column currents I1, I2,..., I n represent the result of a vector matrix multiplication operation performed in the forward path operation, where, as shown in FIG. 3A, the input voltage vector [V1, V2,..., V m is multiplied by the conductance matrix G (of the conductance values G ij ) to generate the output current vector [I1, I2,..., I n . In particular, a given column current I j is calculated by the following formula. [Number] For example, the column current I1 of the first column C1 is determined as I1 = (V1G 11 + V2G 21 +,... + V m G m1 ).

[0042] The aggregate read currents I1, I2,..., I n obtained as a result at the outputs of respective columns C1, C2,..., Cn are input to respective read circuit blocks 330-1, 330-2,..., 330-n of the read circuit 330. The aggregate read currents I1, I2,..., I nare integrated by respective current integration circuits 332-1, 332-2, ..., 332-n to generate respective output voltages, which are then quantized by respective ADC circuits 334-1, 334-2, ..., 334-n to generate respective digital output signals y1, y2, ..., y n of the output vector y. The digital output signals y1, y2, ..., y n are processed and transmitted to the next downstream layer to continue the forward propagation. When data propagates forward through the neural network, vector matrix multiplications are performed, where hidden neurons / nodes receive inputs, perform non-linear transformations, and then send the results to the next weight matrix. This process continues until the data reaches the output layer with output neurons / nodes. The output neurons / nodes evaluate the classification error and generate a classification error signal δ, which is then backpropagated through the neural network using the backward pass operation. The error signal δ can be determined as the difference between the result of the forward inference classification (the estimated label) and the correct label at the output layer of the neural network.

[0043]

[0044] Figure 3B schematically shows a method for configuring a computing system comprising an array of one of a plurality of resistive processing unit cells to perform a backward pass operation, in accordance with an exemplary embodiment of the present disclosure. In particular, Figure 3B schematically shows the configuration of the computing system 300 when performing a backward pass operation, where the multiplexer in the peripheral circuitry of the computing system 300 is activated to selectively connect the column line driving circuit 340 to column lines C1, C2, ..., Cn and to connect the shared readout circuit 330 to row lines R1, R2, ..., Rm. The column line driving circuit 340 comprises a plurality of DAC circuit blocks 342-1, 342-2, ..., 342-n (collectively referred to as DAC circuit blocks 342) connected to respective column lines C1, C2, ..., Cn.In some embodiments, as described above and as illustrated in FIGS. 3A and 3B, the rows and columns do not share a DAC circuit such that each row and each column has a dedicated DAC circuit block. On the other hand, in some embodiments, the rows and columns share a readout circuit 330 (e.g., row R1 and column C1 share readout circuit block 330-1, and row R2 and column R2 share readout circuit block 330-2, etc.). In some embodiments, when the number of rows is the same as the number of columns (i.e., n = m), the number of readout circuit blocks of the shared readout circuit 330 is equal to n = m, while when the number of rows and columns is not the same (i.e., n ≠ m), the number of readout circuit blocks of the readout circuit 330 will be equal to the larger of n and m.

[0045] As shown in FIG. 3B, the backward pass operation performed on one array of the plurality of RPU cells 305 of the computing system 300 is performed in a manner similar to the forward pass operation (FIG. 3A), except that the computing system 300 receives a vector of error signals δ = [δ 1, δ 2, ..., δ n backpropagated from the downstream layer of the neural network. The digital error signals δ 1, δ2,..., δ n are input to respective DAC circuit blocks 342-1, 342-2,..., 342-n connected to respective columns C1, C2,..., Cn. The DAC circuit blocks 342-1, 342-2,..., 342-n generate analog voltages V1, V2,..., Vn using the same or similar time encoding techniques as described above, and each digital error signal δ 1, δ 2, ..., δ nGenerates a pulse - modulated voltage pulse (same amplitude but tunable pulse width) corresponding to the value of. As described in more detail below, in some embodiments, for the backward - pass operation, the analog voltage signals V1, V2, ..., Vn are the integration time T1 used for the forward - pass operation meas Greater integration time T2 meas Based on which it is generated.

[0046] During the backward - pass operation, the analog voltage signals V1, V2, ..., Vn (e.g., read - out pulses representing error signals) are applied to the column lines C1, C2, ..., Cn, where each RPU cell 310 generates a corresponding read - out current I READ =V j ×G ij (Based on Ohm's law), where V j Represents the analog input voltage applied to a given RPU cell 310 on a given column j, and G ij Represents the conductance value of a given RPU cell 310 (at a given row i and column j). As shown in Figure 3B, the read - out currents generated by the plurality of RPU cells 310 on each row i are summed (based on Kirchhoff's current law), and the row currents I1, I2, ..., I m Are generated at the outputs of each row R1, R2, ..., Rm respectively. In this manner, the row currents I1, I2, ..., I n Represent the result of the matrix - vector multiplication operation performed in the backward - pass operation, as shown in Figure 3B, where the conductance matrix G (having conductance values G ij ) is multiplied by the input voltage vector [V1, V2, ..., V n to generate the output current vector [I1, I2, ..., I m . In particular, a given row current I i Is calculated by the following formula.

Equation

[0047] The aggregated read currents I1, I2, …, I obtained as a result of the outputs of the respective rows R1, R2, …, Rm m are input to the respective read circuit blocks 330-1, 330-2, …, 330-m of the shared read circuit 330. The aggregated read currents I1, I2, …, I m are integrated by the respective current integration circuits 332-1, 332-2, …, 332-m to generate respective output voltages, whereby they are quantized by the respective ADC circuits 334-1, 334-2, …, 334-m, and the respective digital output signals z 1, z 2, ..., z m of the output vector z are generated. The digital output signals z 1, z 2, ..., z m are then processed and transmitted to the next upstream layer to continue the backward propagation operation. This process continues until an error signal reaches the input layer.

[0048] After the backward pass operation is completed in one array of the plurality of RPU cells 305 of the computing system 300, the forward propagated digital signals x1, x 2, …, x m and the backward propagated digital error signals δ , δ 2, …, δ nBased on this, a weight update process is performed to adjust the conductance values of a plurality of RPU cells 310 (which represent the conductance matrix G of one array of the plurality of RPU cells 305). Once the error signal value (i.e., the delta value) is integrated for a given neuron layer, that layer is ready for weight update. The update process executed on one array of the plurality of RPU cells 305 of the computing system 300 can be pipelined with the backward propagation of the error vector δ through additional upstream layers of the computing system 300. In some embodiments, backward propagation from the first hidden layer back to the input layer neurons is performed, but this is not required as the input neurons do not have upstream synapses, and thus the highest layer that uses the δ error value is the first hidden layer.

[0049] As further shown in FIGS. 3A and 3B, the mode control system 350 is configured to generate control signals to dynamically change the configuration (e.g., the hardware configuration) of the shared read circuit 330 to have different output signal bounds for different operations executed on one array of the plurality of RPU cells 305 of the computing system 300, such as forward pass and backward pass operations. For example, as shown in FIG. 3A, the mode control system 350 dynamically configures the shared read circuit 330 to generate one or more control signals to have a first configuration, CONFIG_1, optimized for the forward pass operation. In particular, in the forward pass operation, the shared read circuit 330 is configured to operate using a first output signal bound b1 (e.g., a first operating signal range), which is determined based on the expected range (e.g., signal strength) of the voltage and current signals received and generated by the RPU Cell 305 The one of array and Row line drive circuit 320 and Readout circuit 330 during the forward pass operation.

[0050] Moreover, as shown in FIG. 3B, the mode control system 350 dynamically configures the shared readout circuit 330 to generate one or more control signals to have a second configuration, i.e., CONFIG_2, optimized for the backward pass operation. In particular, in the backward pass operation, the shared readout circuit 330 is configured to operate with a second output signal bound b2 (e.g., a second operating signal range), which is determined based on the expected signal strength ranges of the voltage and current signals received and generated by the RPU Cell 305 The one of array and Column line drive circuit 340 and Readout circuit 330. As described above, in the backward pass operation, the error signal δ can have a low signal strength (e.g., δ << 1), where the output signal bound b1 (e.g., a first operating signal range) of the shared readout circuit 330 for the forward pass operation is insufficient to properly process and quantize the aggregated current signal output from one array of multiple RPU cells 310, as during the backward pass operation. As will be described in more detail below, during the backward pass operation, the RPU Cell 305 The one of To enhance the signal processing and quantization of the aggregated current signal output from multiple RPU cells 310 of the array, various techniques can be implemented to dynamically configure the peripheral circuit (e.g., the shared readout circuit 330) to have a second operating signal range.

[0051] The mode control system 350 shown in FIGS. 3A and 3B (as well as the exemplary mode control systems shown in FIGS. 6, 7A, 7B, 8, and 9A) generally includes, in some embodiments, various types of control circuits, processors, and associated functionality (implemented in software, firmware, hardware, or combinations thereof) for controlling the peripheral circuits of the neural core, collectively referred to as a control system (e.g., the peripheral circuits 120 and 130 of the computing system 100 (FIG. 1), the computing system 300'sRow line drive circuit 320, Readout circuit 330, and Column line drive circuit 340, etc.) are controlled, thereby controlling various operations (e.g., forward operations, backward operations, and update operations) performed by the computing system. For example, the mode control system 350 includes a control signal circuit 140 (FIG. 1), as well as control circuits and processors of one or more digital processing cores operably connected to the neural core. Here, the one or more digital processing cores provide data and control signals to instruct operations performed by the neural core.

[0052] FIGS. 3A and 3B schematically illustrate an exemplary method for generating aggregated columns and row currents during forward pass operations and backward pass operations. However, other techniques can be implemented to generate the aggregated columns and row currents using differential current techniques that enable "weight with sign". For example, FIGS. 4A and 4B schematically illustrate a method for configuring a computing system comprising an array of one of a plurality of resistive processing unit cells to perform a forward pass operation using a weighted value with sign according to an exemplary embodiment of the present disclosure. Additionally, FIGS. 5A and 5B schematically illustrate a method for configuring a computing system comprising an array of one of a plurality of resistive processing unit cells to perform a backward pass operation using a weighted value with sign according to an exemplary embodiment of the present disclosure.

[0053] More specifically, FIG. 4A schematically illustrates a method for generating an aggregated column current during a forward pass operation using a reference current (I REF ) generated by a reference current circuit 400 to enable "weight with sign". For ease of illustration, FIG. 4A shows only the first column C1 of the shared read circuit 330 and the associated read circuit block 330-1. FIG. 4A shows that the aggregated column current I COL1 input to the read circuit block 330-1 is such that I COL1 = I1 - I REFSchematically shows the differential readout scheme determined as. In this differential scheme, I COL1 The magnitude of represents the weight value, and the weight code will depend on whether I1 is greater than, equal to, or less than the reference current I REF . A positive sign (I COL1 > 0) will be obtained when I1 > IREF 。 A zero value (I COL1 = 0) will be obtained when I1 = I REF . A negative sign (I COL1 < 0) will be obtained when I1 < I REF . The reference current circuit 400 is generally illustrated in FIG. 4A, but the reference current circuit 400 can be implemented using known techniques. For example, in some embodiments, the reference current circuit 400 includes a fixed current source configured to generate a reference current I REF having a known fixed magnitude selected for a given application.

[0054] Next, FIG. 4B schematically shows a method for generating an aggregate column current I + during a forward pass operation using differential column currents I1 - and I1 + and I1 - from two separate RPU arrays 410-1 and 410-2 corresponding columns C1 COL1 , where the conductance is determined as (G + - G - ). More specifically, in the exemplary embodiment of FIG. 4B, each RPU cell 310 (artificial synapse) includes two units RPU cells 310-1 and 310-2 having respective conductance values G ij + and G ij - , where the conductance value of a given RPU cell is the difference between the respective conductance values, i.e., G ij = G ij + - Gij - is determined as such, where i and j are indices within the 2D array of synapses. Thus, negative and positive weights can be easily encoded using only positive conductance values.

[0055] In other words, since the conductance values of the RPU devices can only be positive, the difference scheme of FIG. 4B encodes the positive (w ij + ) weight values and the negative (w ij - ) weight values by implementing a pair of identical RPU device arrays. Here, the weight value (wij) is proportional to the difference between two corresponding conductance values stored within two corresponding devices (G ij + -G ij - ) at the same position in a pair of RPU arrays 410-1 and 410-2 (where the two RPU arrays 410-1 and 410-2 can be stacked on top of each other in the back-end-of-line metallization structure of the chip). In this example, a single RPU tile is considered a pair of RPU arrays with peripheral circuitry that supports parallel operation of one of the arrays in all three cycles.

[0056] As shown in FIG. 4B, positive voltage pulses (V1, V2,..., Vm) and corresponding negative voltage pulses (-V1, V2,..., Vm) are separately supplied to a plurality of RPU cells 310-1 and 310-2 in corresponding rows in the same RPU arrays 410-1 and 410-2 used to encode positive and negative weights. The aggregated column currents I1 + and C1 - output from the corresponding first columns C1 + and I1 - are combined to form a differential aggregate current I COL1is generated, and the differential aggregate current I COL1 is input to a read circuit block 330-1 connected to the corresponding first column C1 + and C1 - .

[0057] FIGS. 5A and 5B are similar to FIGS. 4A and 4B, but schematically illustrate an alternative embodiment that generates the output current of a row during a backward pass operation. More specifically, FIG. 5A schematically shows a method for generating an aggregate row current during a backward pass operation using a reference current (I REF ) generated by a reference current circuit 500 to enable "weight with sign". For ease of illustration, FIG. 5A shows only the first row R1 of the shared read circuit 330 and the associated read circuit block 330-1. FIG. 5A shows that the aggregate row current I ROW1 input to the read circuit block 330-1 is determined as I ROW1 = I1 - I REF in a differential read scheme. In this difference scheme, the magnitude of I ROW1 represents the weight value, and the weight sign will be determined by whether I1 is greater than, equal to, or less than the reference current I REF . A positive sign (I ROW1 > 0) will be obtained when I1 > IREF 。 A zero value (I ROW1 = 0) will be obtained when I1 = I REF . A negative sign (I ROW1 < 0) will be obtained when I1 < IREF 。 The reference current circuit 500 is generally illustrated in FIG. 5A, but the reference current circuit 500 can be implemented using known techniques. For example, in some embodiments, the reference current circuit 500 includes a fixed current source configured to generate a reference current I REF having a known fixed magnitude selected for a given application.

[0058] Next, FIG. 5B schematically shows a method for generating an aggregated row current I during a backward pass operation using differential row currents I1 + and I1 - from two separate RPU arrays 410-1 and 410-2 corresponding rows R1 + and I1 - . Also, in the exemplary embodiment of FIG. 5B, each RPU cell 310 (artificial synapse) comprises two units RPU cells 310-1 and 310-2 having respective conductance values G row1 where the conductance value of a given RPU cell 310 is the difference between the respective conductance values, i.e., G ij + and G ij - is determined as, where i and j are indices within the respective RPU arrays 410-1 and 410-2. In this way, negative and positive weights can be easily encoded using only positive conductance values. ij =G ij + -G ij -

[0059] As shown in FIG. 5B, positive voltage pulses (V1, V2,..., Vm) and corresponding negative voltage pulses (-V1, V2,..., Vm) are separately applied to a plurality of RPU cells 310-1 and 310-2 in corresponding columns in the same RPU arrays 410-1 and 410-2 used to encode positive and negative weights. The aggregated row currents I1 + and I1 - output from the corresponding first rows R1 + and I1 - in the respective RPU arrays 410-1 and 410-2 are combined to generate a differential aggregated current I COL1 , and the differential aggregated current I COL1 is input to a readout circuit block 330-1 connected to the corresponding first columns C1 + and C1 - .

[0060] ​Here, exemplary embodiments are related to FIGS. 6, 7A, 7B, 8, 9A, and 9B, which schematically show a bound management technique implemented in the analog domain to dynamically change the configuration (e.g., hardware configuration) of a shared readout circuit so as to provide different signal bounds for different operation modes of one analog RPU crossbar array, such as forward pass operation and backward pass operation. In particular, FIGS. 6, 7A, 7B, 8, 9A, and 9B schematically show various techniques for dynamically configuring a peripheral circuit (e.g., the shared readout circuit 330) to have a configuration (e.g., an asymmetric hardware configuration) specialized for different operation modes (e.g., forward pass operation and backward pass operation), thereby optimizing the signal processing and quantization of the integrated current signals output from the rows and columns of the analog RPU crossbar array for different operation modes. This is in contrast to conventional schemes where the peripheral circuits of the analog RPU crossbar array have a "symmetric" configuration for performing readout operations for backward and forward pass operations.

[0061] FIG. 6 schematically shows a system for dynamically configuring a readout circuit for different operations performed on one array of a plurality of resistive processing unit cells, according to an exemplary embodiment of the present disclosure. More specifically, FIG. 6 schematically shows a system 600 for dynamically configuring an integration capacitor of a current integration circuit to change the gain of the current integration circuit depending on the operation mode (e.g., forward pass operation or backward pass operation) of the analog RPU crossbar array. The system 600 includes a mode control system 610 and a readout circuit block 620. The mode control system 610 is configured to generate various control signals 612 that are applied to all readout circuit blocks of a shared readout circuit (e.g., the shared readout circuit 330, FIGS. 3A and 3B) utilized between forward pass operation and backward pass operation. In an exemplary embodiment, the control signal 612 is

Number

[0062] For ease of illustration, FIG. 6 schematically illustrates a given read circuit block 620 representing the i-th read circuit block of a plurality of read circuit blocks of the shared read circuit 330. In the exemplary embodiment of FIG. 6, each read circuit block of the shared read circuit 330 has the same circuit configuration as the read circuit block 620, and it is assumed that each read circuit block of the shared read circuit 330 receives a control signal 612 output from the mode control system 610 during forward path operations and backward path operations performed on the associated analog RPU crossbar array.

[0063] As schematically shown in FIG. 6, the read circuit block 620 includes a multiplexer circuit 630, a current integration circuit 640, and an ADC circuit 650. The current integration circuit 640 includes an operational amplifier 642, a first integration capacitor 644-1, a second integration capacitor 644-2, a first switch 646-1, and a second switch 646-2. The operational amplifier 642 includes a non-inverting input connected to a ground (GND) voltage, an inverting input (denoted as node N1) connected to the output of the multiplexer circuit 630, and an output (denoted as node N2) connected to the input of the ADC circuit 650. A first integration capacitor 644-1 and a first switch 646-1 are serially connected between node N1 and node N2, and a second integration capacitor 644-2 and a second switch 646-2 are serially connected between node N1 and node N2.

[0064] As further shown in FIG. 6, the multiplexer circuit 630 has a first input and a second input connected to corresponding row line ROW(i) and column line COL(i) of the RPU crossbar array. The multiplexer circuit 630 comprises a control input for receiving control signals F_mode and B_mode. The control signal F_mode includes a control signal output from the mode control system 610 when the RPU crossbar array is performing a forward pass operation, and the control signal B_mode includes a control signal output from the mode control system 610 when the RPU crossbar array is performing a backward pass operation. In some embodiments, the multiplexer circuit 630 is configured to (i) connect the row line ROW(i) to the input node N1 of the current integration circuit 640 in response to the assertion of the B_mode control signal, and (ii) connect the column line COL(i) to the input node N1 of the current integration circuit 640 in response to the assertion of the F_mode control signal.

[0065] In this configuration, the multiplexer circuit 630 is configured to control the sharing of the read circuit block 620 for forward pass operations and backward pass operations performed by the RPU crossbar array. For example, in the exemplary embodiments of FIGS. 3A and 3B, the multiplexer circuit 630 and the associated control signals F_mode and B_mode (implemented in each of the read circuit blocks of the shared read circuit 330) are configured to selectively connect the column lines C1, C2,..., Cn of the RPU Cell 305 The one of array to the shared read circuit 330 during forward pass operations (FIG. 3A), and the row lines R1, R2,..., Rm of the RPU Cell 305 The one of array to the shared read circuit 330 during backward pass operations (FIG. 3B).

[0066] The current integration circuit 640 has an integration period (T measExecute an integration function over a period, and convert the input current at the input node N1 of the current integration circuit 640 into an analog voltage V at the output node N2 of the current integration circuit 640. OUT After the end of the integration period, the ADC circuit 650 latches the output voltage V OUT and quantizes the output voltage V OUT to generate a digital signal corresponding to the analog output voltage V OUT The input current can be (i) an aggregated column current output from a column line COL(i) selectively connected (via the operation of the multiplexer circuit 630) to the input node N1 of the current integration circuit 640 during a forward path operation, or (ii) an aggregated row current output from a row line ROW(i) selectively connected (via the operation of the multiplexer circuit 630) to the input node N1 of the current integration circuit 640 during a backward path operation.

[0067] The current integration circuit 640 is configured as an operational transconductance amplifier (OTA) with selectable capacitive feedback provided by one of a first integration capacitor 644-1 and a second integration capacitor 644-2 to convert an input current (aggregated row current or aggregated column current) into an output voltage V on the output node N2 of the current integration circuit 640. In the exemplary configuration of FIG. 6, the first integration capacitor 644-1 includes a capacitance value C OUT 1, and the second integration capacitor 644-2 includes a capacitance value C INT 2, where C INT 1 is larger than C INT 1 is larger than C INT 2. The first integration capacitor 644-1 and the second integration capacitor 644-2 are selectively connected to the feedback path of the operational amplifier 642 for different operation modes (e.g., forward path operation and backward path operation), change the amount of the feedback capacitance, and thus change the gain of the current integration circuit 640 according to the operation mode of the current.

[0068]

Number

[0069]

Number

[0070] In the exemplary embodiment of FIG. 6, the current integration circuit 640 has a configurable gain that can be dynamically switched between a high-gain configuration and a low-gain configuration by selectively connecting one of the first current integration capacitor 644-1 and the second current integration capacitor 644-2 in the capacitive feedback path of the operational amplifier 642. For example, during forward path operation, the current integration circuit 640 can be configured to have a low-gain configuration to generate an output voltage V on the output node N2, and during backward path operation, the current integration circuit 640 can be configured to have a high-gain configuration to generate an output voltage V on the output node N2. OUT on the output node N2, and during backward path operation, the current integration circuit 640 can be configured to have a high-gain configuration to generate an output voltage V OUT on the output node N2.

[0071] Generally, the output voltage V generated by the current integration circuit 640 is determined by the following equation. OUT

Number

Number

[0072] In an exemplary embodiment of the system configuration shown in FIG. 6, the capacitance value C INT 1 of the first integrating capacitor 644-1 is greater than the capacitance value C INT 2 of the second integrating capacitor 644-2. In this case, during the forward path operation, the first integrating capacitor 644-1 is selectively connected to the feedback path between the input node N1 and the output node N2 of the operational amplifier 642 to configure the current integrating circuit 640 to have a low gain configuration. On the other hand, during the backward path operation, considering that the pulse width of the input voltage is expected to be of a significantly shorter duration than the pulse width of the input voltage for the forward path operation, the second integrating capacitor 644-2 is selectively connected to the feedback path between the input node N1 and the output node N2 of the operational amplifier 642 to configure the current integrating circuit 640 to have a high gain configuration, thereby generating an output voltage V OUT of sufficient magnitude for processing by the ADC circuit 650. The capacitance values C INT 1 and C INT 2 of the first integrating capacitor 644-1 and the second integrating capacitor 644-2 are selected to provide sufficient gain during the forward path operation and the backward path operation while preventing or otherwise minimizing the possibility of saturating the operational amplifier 642 during the current integration operation.

[0073] Figures 7A and 7B schematically show a system for dynamically configuring a readout circuit for different operations performed on one array of a plurality of resistive processing unit cells, in accordance with an exemplary embodiment of the present disclosure. More specifically, FIGS. 7A and 7B schematically show an alternative exemplary embodiment of a system for dynamically configuring a shared readout circuit to select two different ADC circuits having different resolutions depending on the operation mode (forward path or backward path) of the RPU crossbar array. Referring to FIG. 7A, system 700 includes a mode control system 710 and a readout circuit block 720. The mode control system 710 is configured to generate various control signals 712 that are applied to all readout circuit blocks of a shared readout circuit (e.g., shared readout circuit 330, FIGS. 3A and 3B) utilized during forward path and backward path operations. In an exemplary embodiment, the control signals 712 include control signals shown as F_mode, B_mode, ADC_H, and ADC_L.

[0074] For ease of illustration, FIG. 7A schematically illustrates a given readout circuit block 720 that represents the i-th readout circuit block of a plurality of readout circuit blocks of the shared readout circuit 330. In the exemplary embodiment of FIG. 7A, each readout circuit block of the shared readout circuit 330 has the same circuit configuration as readout circuit block 720, and it is assumed that each readout circuit block of the shared readout circuit 330 receives control signals 712 output from the mode control system 710 during forward path and backward path operations performed on the associated analog RPU crossbar array.

[0075] As schematically shown in FIG. 7A, readout circuit block 720 includes a multiplexer circuit 730, a current integration circuit 740, a selection circuit 760, a first ADC circuit 750-1, and a second ADC circuit 750-2. Current integration circuit 740 includes an operational amplifier 742 and an integration capacitor 744 connected between the input node N1 and the output node N2 of the operational amplifier. Multiplexer circuit 730 has a control input that receives control signals F_mode and B_mode. Multiplexer circuit 730 is configured to operate in the same manner as multiplexer circuit 630 of FIG. 6, and its details will not be repeated.

[0076] Current integration circuit 740 performs an integration function over an integration period (T MEAS ), and converts the input current at the input node N1 of current integration circuit 740 into an analog voltage V OUT at the output node N2 of current integration circuit 740. The input current can be (i) an aggregated column current output from column line COL(i) selectively connected to the input node N1 of current integration circuit 740 (via the operation of multiplexer circuit 730) during a forward path operation, or (ii) an aggregated row current output from row line ROW(i) selectively connected to the input node N1 of current integration circuit 740 (via the operation of multiplexer circuit 730) during a backward path operation. In the exemplary embodiment of FIG. 7A, current integration circuit 740 has a fixed gain based on the capacitance value C INT 1 of integration capacitor 744. In some embodiments, the capacitance value C INT 1 of integration capacitor 744 is selected for output signal bounding for a forward path operation, e.g., to provide sufficient gain to current integration circuit 740 to generate an output voltage V OUT on output node N2 during a forward path operation while preventing or otherwise minimizing the possibility of saturating operational amplifier 742 during the current integration operation.

[0077] As further shown in FIG. 7A, the selection circuit 760 includes (i) an input connected to the output node N2 of the current integration circuit 740, (ii) a control input that receives control signals ADC_H and ADC_L, and (iii) a first output and a second output respectively connected to the inputs of the first ADC circuit 750-1 and the second ADC circuit 750-2. In some embodiments, the first ADC circuit 750-1 includes a high-resolution ADC circuit that is utilized during a backward pass operation to digitize the output voltage V OUT generated at the output node N2 of the current integration circuit 740, and the second ADC circuit 750-2 includes a low-resolution ADC circuit that is utilized during a forward pass operation to digitize the output voltage V OUT generated at the output node N2 of the current integration circuit 740.

[0078] The control signal ADC_L includes a control signal output from the mode control system 710 when the RPU crossbar array is performing a forward pass operation, and the control signal ADC_H includes a control signal output from the mode control system 710 when the RPU crossbar array is performing a backward pass operation. In some embodiments, the selection circuit 760 is configured to (i) selectively connect the output node N2 of the current integration circuit 740 to the input of the first ADC circuit 750-1 in response to an assertion of the control signal ADC_H, and (ii) selectively connect the output node N2 of the current integration circuit 740 to the input of the second ADC circuit 750-2 in response to an assertion of the control signal ADC_L.

[0079] In some embodiments, the first ADC circuit 750-1 has a first resolution, and the second ADC circuit 750-2 has a second resolution, where the first resolution is higher than the second resolution. In some embodiments, the first ADC circuit 750-1 and the second ADC circuit 750-2 have the same bit resolution, where the first ADC circuit 750-1 and the second ADC circuit 750-2 each include an n-bit ADC resolution, for example, n = 3, 4, 5, 6, 7, 8, but the first ADC circuit 750-1 and the second ADC circuit 750-2 are configured to have different least significant bit (LSB) voltage resolutions for a given n-bit resolution.

[0080] For example, assume that the first ADC circuit 750-1 and the second ADC circuit 750-2 are 8-bit resolution ADCs, but the first ADC circuit 750-1 is configured to have a full-scale operating voltage range of, for example, [-0.1V to +0.1V], and the second ADC circuit 750-2 is configured to have a full-scale operating voltage range of, for example, [-10V to +10V] (where V can be in microvolts (μV), millivolts (mV), etc.). In this case, the voltage resolution (the first resolution) of the first ADC circuit 750-1 (8-bit ADC) is 0.2 / 256, which is approximately 0.000078V, while the voltage resolution (the second resolution) of the second ADC circuit 750-2 (8-bit ADC) is 20 / 256, which is approximately 0.078V. In this example, the first voltage resolution (0.000078V) of the first ADC circuit 750-1 provides a higher voltage resolution than the second voltage resolution (0.078V) of the second ADC circuit 750-2 (8-bit ADC). The enhanced voltage resolution of the first ADC circuit 750-1 is effective for digitizing the low-level output voltage V OUT generated at the output node N2 of the current integration circuit 740 during the backward path operation, and the voltage resolution of the second ADC circuit 750-2 is for the high-level output voltage V OUTwould be effective for digitization.

[0081] In other embodiments, the first ADC circuit 750-1 and the second ADC circuit 750-2 can be configured to have different gains, where the first gain of the first ADC circuit 750-1 is greater than the second gain of the second ADC circuit 750-2. For example, in some embodiments, the first ADC circuit 750-1 and the second ADC circuit 750-2 have the same bit resolution, the same voltage resolution, and the same operating voltage input range, which are designed for the expected voltage range of the output voltage V OUT generated by the forward path operation. However, the first ADC circuit 750-1 includes an analog front end, and the analog front end includes an amplifier or a level shift circuit or a combination thereof configured to provide an appropriate gain and level shift for the low-level output voltage V OUT generated during the backward path operation to match the operation input range of the ADC conversion circuit.

[0082] Next, FIG. 7B schematically shows a system 701 for dynamically configuring a read circuit shared to select between two different ADC circuits having different resolutions depending on the operation mode (forward path or backward path) of the RPU crossbar array, according to an exemplary embodiment of the present disclosure. As described above, the system 701 of FIG. 7B is an alternative embodiment of the system 700 of FIG. 7A. Referring to FIG. 7B, the system 701 includes a mode control system 711 and a read circuit block 721. The mode control system 711 is configured to generate a plurality of control signals 713 applied to all read circuit blocks of a shared read circuit (e.g., the shared read circuit 330, FIGS. 3A and 3B) utilized during forward path operations and backward path operations. In an exemplary embodiment, the control signals 713 include control signals shown as F_mode and B_mode.

[0083] For ease of illustration, FIG. 7B schematically illustrates a given read circuit block 721 representing the i-th read circuit block of a plurality of read circuit blocks of the shared read circuit 330. In the exemplary embodiment of FIG. 7B, each read circuit block of the shared read circuit 330 has the same circuit configuration as the read circuit block 721, and it is assumed that each read circuit block of the shared read circuit 330 receives a control signal 713 output from the mode control system 711 during the forward pass operation and the backward pass operation executed on the associated analog RPU crossbar array.

[0084] As schematically shown in FIG. 7B, the read circuit block 721 includes a selection circuit 731, a first current integration circuit 740-1, a second current integration circuit 740-2, a first ADC circuit 750-1, and a second ADC circuit 750-2. The first current integration circuit 740-1 has an input connected to the output of the selection circuit 731 and an output connected to the input of the first ADC circuit 750-1. The second current integration circuit 740-2 has an input connected to the output of the selection circuit 731 and an output connected to the input of the second ADC circuit 750-2. In some embodiments, the first current integration circuit 740-1 and the second current integration circuit 740-2 are similar to the current integration circuit 740 in FIG. 7A, and thus their details are not repeated. Further, in some embodiments, the first ADC circuit 750-1 and the second ADC circuit 750-2 implement the same circuit architecture and function as those described above in the exemplary embodiment of FIG. 7A, and thus their details are not repeated.

[0085] The selection circuit 731 is configured to operate in a manner similar to the multiplexer circuits 630 and 730 of FIGS. 6 and 7A. However, the selection circuit 731 has two outputs, and in some embodiments, the selection circuit 731 is configured to (i) selectively connect the row line ROW(i) to the input of the first current integration circuit 740-1 in response to the assertion of the B-mode control signal, and (ii) selectively connect the column line COL(i) to the input of the second current integration circuit 740-2 in response to the assertion of the F-mode control signal. In this configuration, the first current integration circuit 740-1 and the first ADC circuit 750-1 receive, integrate, and quantize the aggregated row current signal output from the row line ROW(i) during the backward path operation, and the second current integration circuit 740-2 and the second ADC circuit 750-2 receive, integrate, and quantize the aggregated column signal output from the column line COL(i) during the forward path operation.

[0086] In an alternative embodiment of FIG. 7B, the first current integration circuit 740-1 and the second current integration circuit 740-2 are configured to have different but fixed gains, and the first ADC circuit 750-1 and the second ADC circuit 750-2 are configured to have the same resolution. In this embodiment, the first current integration circuit 740-1 and the second current integration circuit 740-2 each have a single but different-sized integration capacitor in respective feedback paths of respective operational amplifiers such that the first current integration circuit 740-1 has a fixed first gain and the second current integration circuit 740-2 has a second fixed gain that is smaller than the first fixed gain. Using this configuration, the first current integration circuit 740-1 has a higher gain to generate an output voltage V OUT having levels in a voltage range corresponding to the operating input voltage range of the first ADC circuit 750-1 and thus will be designed with an output signal bound for the backward path operation.

[0087] FIG. 8 schematically shows a system for dynamically configuring a readout circuit for different operations executed on one array of a plurality of resistive processing unit cells, according to another exemplary embodiment of the present disclosure. More specifically, FIG. 8 schematically shows a system 800 for dynamically configuring an ADC circuit to have different resolutions or different gains depending on the operation mode (forward path or backward path) of an RPU crossbar array. The system 800 includes a mode control system 810 and a readout circuit block 820. The mode control system 810 is configured to generate various control signals 812 that are applied to all the readout circuit blocks of a shared readout circuit (e.g., the shared readout circuit 330, FIGS. 3A and 3B) utilized during forward path operations and backward path operations. In an exemplary embodiment, the control signals 812 include control signals shown as F_mode, B_mode, ADC_H, and ADC_L.

[0088] Also, for ease of illustration, FIG. 8 schematically shows a given readout circuit block 820 representing the i-th readout circuit block of a plurality of readout circuit blocks of the shared readout circuit 330. In the exemplary embodiment of FIG. 8, it is assumed that each readout circuit block of the shared readout circuit 330 has the same circuit configuration as the readout circuit block 820, and that each readout circuit block receives the control signals 812 output from the mode control system 810 when forward and backward path operations are executed on the associated analog RPU crossbar array.

[0089] As schematically shown in FIG. 8, the readout circuit block 820 includes a multiplexer circuit 830, a current integration circuit 840, and a configurable ADC circuit 850. The multiplexer circuit 830 and the current integration circuit 840 have the same configuration as the multiplexer circuit 630 in FIG. 6 and the current integration circuit 740 in FIG. 7A, and perform the same or similar functions, the details of which will not be repeated. As further shown in FIG. 8, the configurable ADC circuit 850 includes a control input that receives control signals ADC_H and ADC_L.

[0090] In some embodiments, the configurable ADC circuit 850 has a configurable LSB voltage resolution that is dynamically adjusted in response to the control signals ADC_H and ADC_L. For example, in response to the assertion of the control signal ADC_H during the backward path operation, the configurable ADC circuit 850 is dynamically configured to increase the voltage resolution to a level sufficient to effectively quantize the low-level output voltage V OUT generated on the output node N2 of the current integration circuit 840 during the backward path operation. Further, in response to the assertion of the control signal ADC_L during the forward path operation, the configurable ADC circuit 850 is dynamically configured to decrease the voltage resolution to a level sufficient to effectively quantize the high-level output voltage V OUT generated on the output node N2 of the current integration circuit 840 during the forward path operation. In this regard, the configurable ADC circuit 850 is dynamically configured to have a first resolution for the backward path operation and dynamically configured to have a second resolution for the forward path operation, where the first resolution is greater than the second resolution.

[0091] The configurable ADC circuit 850 having the ADC resolution can be implemented using appropriate time-based ADC conversion circuits and techniques. For example, the configurable ADC circuit 850 can be implemented using a single-slope or dual-slope integrating ADC architecture, where the ADC conversion is based on the output voltage V OUT , the reference voltage, or the output voltage V OUTand based on the integration of a reference voltage. In this exemplary embodiment, the LSB voltage resolution can be dynamically changed by changing one or more operating parameters of the integrating ADC, including but not limited to, for example, the level of the reference voltage, the integration time of the integrating ADC, etc.

[0092] In other embodiments, the configurable ADC circuit 850 includes a configurable gain that is dynamically adjusted in response to control signals ADC_H and ADC_L. For example, the configurable ADC circuit 850 may have a fixed (non-configurable) resolution designed for output signal bounding for forward path operation, yet still, to dynamically adjust the gain of a programmable in amplifier depending on the operating mode (e.g., forward path operation or backward path operation), Configurable a programmable gain amplifier and a level shift circuit can be implemented in the analog front end of the ADC circuit 850.

[0093] For example, in some embodiments, in response to the assertion of control signal ADC_H during backward path operation, the output voltage V OUT is Configurable amplified to a level that falls within the high operating input voltage range of the conversion circuit of the ADC circuit 850 to have a higher gain so that Configurable the front end analog circuit of the ADC circuit 850 is dynamically configured, whereby a low level output voltage V OUT is Configurable enabled to be more accurately quantized based on the fixed resolution of the ADC circuit 850. Further, in response to the assertion of control signal ADC_H during forward path operation, Configurable the front end analog circuit of the ADC circuit 850, the output voltage V OUT (which is latched into the ADC circuit 850 from the output node of the current integration circuit 840, Configurable is ConfigurableIt is dynamically configured to have a low gain (e.g., unity gain of 1) so as to be maintained at a level within the operating input voltage range of the conversion circuit of the ADC circuit 850.

[0094] Figures 9A and 9B schematically show a system for dynamically configuring a peripheral circuit comprising a readout circuit for different operations executed on one array of a plurality of resistive processing unit cells, according to another exemplary embodiment of the present disclosure. More specifically, Figures 9A and 9B schematically show a system for dynamically configuring the peripheral circuit of an analog RPU crossbar array to have different operation signal ranges for different operation modes of the RPU crossbar array, according to another exemplary embodiment of the present disclosure. Figures 9A and 9B schematically show a system 900 for dynamically configuring peripheral circuits (e.g., shared readout circuit 330 and column line drive circuit 340 (DAC circuit block 342)) to provide different integration times according to the operation mode (forward path or backward path) of the RPU crossbar array.

[0095] Referring to Figure 9A, system 900 includes a mode control system 910 and a readout circuit block 920. The mode control system 910 is configured to generate various control signals 912 that are applied to all readout circuit blocks of a shared readout circuit (e.g., shared readout circuit 330, Figures 3A and 3B) utilized during forward path operations and backward path operations. In an exemplary embodiment, the control signals 912 are F_mode, B_mode, T1 MEAS _mode, and T2 MEASIt includes a control signal indicated as a _ mode. Also, for ease of illustration, FIG. 9A schematically shows a given read circuit block 920 representing the i-th read circuit block of a plurality of read circuit blocks of the shared read circuit 330. In the exemplary embodiment of FIG. 9A, it is assumed that each read circuit block of the shared read circuit 330 has the same circuit configuration as the read circuit block 920, and that when forward path operations and backward path operations are performed on the associated analog RPU crossbar array, each read circuit block receives a control signal 912 output from the mode control system 910.

[0096] As schematically shown in FIG. 9A, the read circuit block 920 includes a multiplexer circuit 930, a current integration circuit 940, and an ADC circuit 950. In some embodiments, the multiplexer circuit 930 has the same configuration as the multiplexer circuit 630 of FIG. 6 and performs the same or similar functions, and thus its details will not be repeated. Additionally, in some embodiments, the ADC circuit 950 has the same configuration as the ADC circuit 650 of FIG. 6 and performs the same or similar functions, and thus its details will not be repeated. In this exemplary embodiment, the ADC circuit 950 is designed to have a fixed configuration with a resolution and operating signal range that effectively digitizes the expected high-level output voltage V OUT generated at the output node N2 of the current integration circuit 940 during forward path operations.

[0097] The current integration circuit 940 includes an operational amplifier 942, an integration capacitor 944 connected between the input node N1 and the output node N2 of the operational amplifier 942, and a control circuit 946 including an integration time counter and a reset control circuit. The current integration circuit 940 integrates the input current at the input node N1 of the current integration circuit 940 into an analog voltage V OUTwhere the current integration operation is performed over a first integration time T1 depending on the operation mode (e.g., forward pass operation or backward pass operation) of the RPU crossbar array. MEAS or the second integration time T2 MEAS The second integration time T2 is performed over a configurable integration period that can be dynamically adjusted to have a second integration time T3. MEAS is the first integration time T1 MEAS For example, in some embodiments, the first integration time T1 MEAS is 80 nanoseconds (ns), while the second integration time T2 MEAS is greater than 80ns (for example, T2 MEAS =2×T1 MEAS ).

[0098] In the exemplary embodiment of FIG. 9A, the current integrator circuit 940 includes an integrating capacitor 944 with a capacitance value C INT 1. In some embodiments, the capacitance value C of the integration capacitor 944 INT 1 provides an output signal bound for forward path operation, while preventing or otherwise minimizing the possibility of saturating opamp 942 during current integration operation, for example, during integration period T1. MEAS During the forward pass operation, the output voltage V OUT Current integrator circuit 940 is selected to provide sufficient gain to generate

[0099] Meanwhile, during backward path operation, current integrator circuit 940 provides a larger output voltage V while preventing or otherwise minimizing the possibility of saturating opamp 942 during the current integration operation for the backward path operation. OUT A second integration period T2 is then established on the output node N2 of the current integrator circuit 940. MEAS 9A, the integration time of the current integrator circuit 940 is dynamically configured to integrate the input current over a period of time. As shown diagrammatically in FIG. 9A, the integration time of the current integrator circuit 940 is controlled by an integration time control signal T1 input to a control circuit 946. MEAS _Mode and T2MEAS It is dynamically configured by control circuit 946 in response to the mode. For example, during forward path operation, mode control system 910 asserts control signal T1 MEAS _mode, which instructs control circuit 946 to configure current integration circuit 940 to perform a current integration operation over a first integration time T1 MEAS . On the other hand, during backward path operation, mode control system 910 asserts control signal T2 MEAS _mode, which instructs control circuit 946 to configure current integration circuit 940 to perform a current integration operation over a second integration time T2 MEAS .

[0100] Control circuit 946 is generally illustrated in FIG. 9A. However, to perform its functions, control circuit 946 may use various control circuit architectures and techniques, for example, controlling the integration period of current integration circuit 940, sending a control signal (e.g., ADC_EN) to latch the output voltage V OUT on output node N2 at the end of the current integration period to the ADC circuit 950, resetting the current integration circuit, for example, resetting the current integration circuit by setting the voltage on the output node to an initial voltage level (e.g., 0V) to initialize current integration circuit 940 in preparation for the next current integration process, etc. In this regard, control circuit 946 can be configured in various manners depending on the specific circuit framework of current integration circuit 940 and the hardware interface and configuration between current integration circuit 940 and ADC circuit 950.

[0101] In an exemplary embodiment, control circuit 946 includes an integration time counter that counts the number of clock pulses input to the integration time counter, where the integration time correlates to a specific count of received clock pulses as understood by those skilled in the art. In this regard, in some embodiments, the integration time of current integration circuit 940 is determined by integration time control signals T1 MEAS _mode and T2 MEASIn response to the _ mode, (i) the first integration time T1 MEAS corresponding to the first counting process, and (ii) the second integration time T2 MEAS is adjusted by configuring the integration time counter circuit of the control circuit 946 to execute either the second counting process corresponding to.

[0102] In addition to increasing the integration time (e.g., T2 MEAS ) of the current integration circuit 946 for the backward path operation, in some embodiments, the DAC circuit block of the line drive circuit (e.g., the DAC circuit block 342 of the column line drive circuit 340, FIG. 3B) is a digital error signal δ 1, δ 2, ..., δ n dynamically configured to increase the pulse duration of the pulse - modulated voltage pulses generated for by an amount proportional to the increase in the integration time from T1 MEAS to T2 MEAS . For example, FIG. 9B schematically shows a process for increasing the pulse duration of the analog voltages V1, V2,..., Vn generated by the DAC circuit block during the backward path operation and applied to each column C1, C2,..., Cn.

[0103] In the exemplary embodiment of FIG. 9B, it is assumed that the second integration time T2 MEAS for the backward path operation is twice the first integration time T1 MEAS for the forward path operation. In this regard, the analog voltages V1, V2,..., Vn generated for the backward path operation with the increased integration time T2 MEAS (which correspond to the values of the respective digital error signals δ 1, δ 2, ..., δ n ) have respective pulse widths W1_2, W2_2,..., Wn_2, which are the same integration time T1 MEASWhen used, it has a magnitude that is twice that of each of the pulse widths W1_1, W2_1, ..., Wn_1 that would be generated (in a conventional scheme) for the backward path operation. In this regard, as the pulse widths of the analog voltages V1, V2, ..., Vn for each of the digital error signals δ 1, δ 2, ..., δ n increase and the integration time T2 MEAS increases, an increase in the output voltage V OUT generated on the output node N2 of the current integration circuit of the shared read circuit for the backward path operation results. With this scheme, even though the current integration circuit of the shared read circuit has an integration capacitor of a fixed size that provides a fixed gain of the current integration circuit capacitor, the output voltage V OUT is increased and the reading of low-level digital error signals δ 1, δ 2, ..., δ n is enhanced, which is optimal for the forward path operation.

[0104] The descriptions of the various embodiments of the present disclosure are presented for purposes of illustration and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen in order to best explain the principles of the embodiments, the practical application, or a technical improvement found in the marketplace, or to enable those skilled in the art to understand the embodiments disclosed herein. One embodiment of the present invention is as follows. [Item 1] A device, the device comprising: One array of a plurality of resistive processing unit (RPU) cells; A plurality of first control lines extending in a first direction across the one array of a plurality of RPU cells, and a plurality of second control lines extending in a second direction across the one array of a plurality of RPU cells, wherein each RPU cell is connected at an intersection of one of the plurality of first control lines and one of the plurality of second control lines; Peripheral circuitry connected to the plurality of first control lines and to the plurality of second control lines, wherein the peripheral circuitry comprises a readout circuit; and A control system operably connected to the peripheral circuitry, wherein the control system controls the peripheral circuitry to generate control signals for performing a first operation on the one array of a plurality of RPU cells and for performing a second operation on the one array of a plurality of RPU cells, Comprising Wherein the control signals are A first configuration control signal for configuring the readout circuit to have a first hardware configuration when the first operation is performed on the one array of a plurality of RPU cells; and A second configuration control signal for configuring the readout circuit to have a second hardware configuration different from the first hardware configuration when the second operation is performed on the one array of a plurality of RPU cells Including Said device. [Item 2] The peripheral circuitry further comprises a control line voltage drive circuit and a multiplexer circuit; and The control signal further has a multiplexer control signal for selectively connecting the readout circuit to the plurality of first control lines and the control line voltage driving circuit to the plurality of second control lines to cause the multiplexer circuit to perform the first operation, and for selectively connecting the readout circuit to the plurality of second control lines and the control line voltage driving circuit to the plurality of first control lines to cause the multiplexer circuit to perform the second operation. The device according to claim 1. [Claim 3] The first operation includes a forward pass operation of a neural network training process performed on the one array of a plurality of RPU cells, where the one array of a plurality of RPU cells comprises one array of artificial synapse elements connecting an upstream layer of artificial neurons and a downstream layer of artificial neurons; The second operation includes a backward pass operation of the neural network training process performed on the one array of a plurality of RPU cells; In the case of the first operation, the control line voltage driving circuit applies a first voltage pulse to one or more of the plurality of second control lines, where the first voltage pulse represents a digital input signal received from the upstream layer of the neural network, and the readout circuit having the first hardware configuration receives a first current signal generated by the RPU cell and output to the plurality of first control lines in response to the first voltage pulse, and generates a first digital output signal representing the first current signal output from the plurality of first control lines; In the case of the second operation, the control line voltage driving circuit applies a second voltage pulse to one or more of the plurality of first control lines, where the second voltage pulse represents a digital error signal received from the downstream layer of the neural network, and the readout circuit having the second hardware configuration receives a second current signal generated by the RPU cell and output to the plurality of second control lines in response to the second voltage pulse, and generates a second digital output signal representing the second current signal output from the plurality of second control lines. The device according to claim 2. [Claim 4] The readout circuit comprises a plurality of readout circuit blocks, where each readout circuit block a current integration circuit having an input node and an output node; and An analog-to-digital conversion circuit having an input connected to the output node of the current integration circuit and comprising: wherein the current integration circuit comprises: an operational amplifier having an inverting input terminal connected to the input node and an output terminal connected to the output node; a first integration capacitor and a first switch connected in series in a first feedback path between the input node and the output node of the current integration circuit, wherein the first integration capacitor has a first capacitance; and a second integration capacitor and a second switch connected in series in a second feedback path between the input node and the output node of the current integration circuit, wherein the second integration capacitor has a second capacitance smaller than the first capacitance, and comprising: wherein the first configuration control signal activates the first switch and deactivates the second switch to configure the current integration circuit to have a first gain based on the first capacitance of the first integration capacitor connected in the first feedback path between the output node and the input node of the current integration circuit; and wherein the second configuration control signal activates the second switch and deactivates the first switch to configure the current integration circuit to have a second gain greater than the first gain based on the second capacitance of the second integration capacitor connected in the second feedback path between the output node and the input node of the current integration circuit, The device according to claim 1. [Claim 5] The readout circuit comprises a plurality of readout circuit blocks, each readout circuit block comprising: a current integration circuit having an input node and an output node; a selection circuit connected to the output node of the current integration circuit; a first analog-to-digital conversion circuit connected to a first output of the selection circuit, wherein the first analog-to-digital conversion circuit has a first least significant bit (LSB) voltage resolution; and a second analog-to-digital conversion circuit connected to a second output of the selection circuit, wherein the second analog-to-digital conversion circuit has a second LSB voltage resolution greater than the first LSB voltage resolution, and comprising: Here, the first configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the first analog-to-digital conversion circuit; and, Here, the second configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the second analog-to-digital conversion circuit, The device according to claim 1. [Claim 6] The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block A current integration circuit having an input node and an output node; A selection circuit connected to the output node of the current integration circuit; A first analog-to-digital conversion circuit connected to a first output of the selection circuit, where the first analog-to-digital conversion circuit has a first gain; A second analog-to-digital conversion circuit connected to a second output of the selection circuit, where the second analog-to-digital conversion circuit has a second gain greater than the first gain, Comprising, Here, the first configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the first analog-to-digital conversion circuit; and, Here, the second configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the second analog-to-digital conversion circuit, The device according to claim 1. [Claim 7] The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block A selection circuit having a first input connected to one of the plurality of first control lines and a second input connected to one of the plurality of second control lines; A first current integration circuit having an input node and an output node; A first analog-to-digital conversion circuit connected to the output node of the first current integration circuit, where the first analog-to-digital conversion circuit has a first least significant bit (LSB) voltage resolution; A second current integration circuit having an input node and an output node; and, A second analog-to-digital conversion circuit connected to the output node of the second current integration circuit, where the second analog-to-digital conversion circuit has a second LSB voltage resolution greater than the first LSB voltage resolution, Comprising, Here, the first configuration control signal causes the selection circuit to selectively connect the one first control line to the input node of the first current integration circuit; and, Here, the second configuration control signal causes the selection circuit to selectively connect the one second control line to the input node of the second current integration circuit. The device according to claim 1. [Claim 8] The first current integration circuit includes a first operational amplifier having an inverting input terminal connected to the input node of the first current integration circuit and an output terminal connected to the output node of the first current integration circuit, and a first integration capacitor connected in a feedback path between the input node and the output node, where the first integration capacitor has a first capacitance; and, The second current integration circuit includes a second operational amplifier having an inverting input terminal connected to the input node of the second current integration circuit and an output terminal connected to the output node of the second current integration circuit, and a second integration capacitor connected in a feedback path between the input node and the output node, where the second integration capacitor has a second capacitance substantially equal to the first capacitance. The device according to claim 1. [Claim 9] The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block A selection circuit having a first input connected to one of the plurality of first control lines and a second input connected to one of the plurality of second control lines; A first current integration circuit having an input node and an output node; A first analog-to-digital conversion circuit connected to the output node of the first current integration circuit, where the first analog-to-digital conversion circuit has a first gain; A second current integration circuit having an input node and an output node; and, A second analog-to-digital conversion circuit connected to the output node of the second current integration circuit, where the second analog-to-digital conversion circuit has a second gain greater than the first gain. is provided, Here, the first configuration control signal causes the selection circuit to selectively connect the one first control line to the input node of the first current integration circuit; and, Here, the second configuration control signal causes the selection circuit to connect the one second control line to the input node of the second current integration circuit. The device according to claim 1. [Claim 10] The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; and an analog-to-digital conversion circuit connected to the first output of the selection circuit, where the analog-to-digital conversion circuit has a configurable least significant bit (LSB) voltage resolution including a first LSB voltage resolution and a second LSB voltage resolution greater than the first LSB voltage resolution. and includes Here, the first configuration control signal configures the analog-to-digital conversion circuit to have the first LSB voltage resolution; and Here, the second configuration control signal configures the analog-to-digital conversion circuit to have the second LSB voltage resolution. The device according to claim 1. [Claim 11] The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; and an analog-to-digital conversion circuit connected to the first output of the selection circuit, where the analog-to-digital conversion circuit has a configurable gain including a first gain and a second gain greater than the first gain. and includes Here, the first configuration control signal configures the analog-to-digital conversion circuit to have the first gain; and Here, the second configuration control signal configures the analog-to-digital conversion circuit to have the second gain. The device according to claim 1. [Claim 12] The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; and an analog-to-digital conversion circuit connected to the first output of the selection circuit and includes Here, the current integration circuit has a configurable integration period including a first integration period and a second integration period longer than the first integration period; Here, the first configuration control signal configures the current integration circuit to perform a current integration process having the first integration period; and Here, the second configuration control signal configures the current integration circuit to perform a current integration process having the second integration period. The device according to claim 11. [Claim 13] The peripheral circuit further includes a control line voltage drive circuit, where the control line voltage drive circuit generates a voltage pulse having a first pulse duration proportional to the first integration period in response to the first configuration control signal and generates a voltage pulse having a second pulse duration proportional to the second integration period in response to the second configuration control signal, and the device according to claim 1 includes a configurable pulse width modulation circuit. [Claim 14] Performing a first operation on one array of a plurality of resistance processing unit (RPU) cells; Configuring a readout circuit to have a first hardware configuration for performing the first operation on the one array of the plurality of RPU cells; Performing a second operation on the one array of the plurality of RPU cells; and Configuring the readout circuit to have a second hardware configuration for performing the second operation on the one array of the plurality of RPU cells, where the second hardware configuration of the readout circuit is different from the first hardware configuration of the readout circuit. A method including the above. [Claim 15] The first operation includes a forward pass operation of a neural network training process performed on the one array of the plurality of RPU cells, where the one array of the plurality of RPU cells includes an array of artificial synapse elements connecting an upstream layer of artificial neurons and a downstream layer of artificial neurons; The second operation includes a backward pass operation of the neural network training process performed on the one array of the plurality of RPU cells. The method according to claim 14. [Claim 16] Configuring the readout circuit to have a first hardware configuration includes configuring the current integration circuit of the readout circuit to have a first integration capacitance; and Configuring the readout circuit to have a second hardware configuration includes configuring the current integration circuit of the readout circuit to have a second integration capacitance smaller than the first integration capacitance. The method according to claim 14. [Claim 17] Configuring the readout circuit to have a first hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a first least significant bit (LSB) voltage resolution; and, Configuring the readout circuit to have a second hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a second LSB voltage resolution greater than the first LSB voltage resolution, The method according to item 14. [Item 18] Configuring the readout circuit to have a first hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a first gain; and, Configuring the readout circuit to have a second hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a second gain greater than the first gain, The method according to item 14. [Item 19] A system comprising a neuromorphic computing system, wherein the neuromorphic computing system An array of one of a plurality of resistive processing unit (RPU) cells, where a plurality of first control lines extend in a first direction across the one array of a plurality of RPU cells, and a plurality of second control lines extend in a second direction across the one array of a plurality of RPU cells, where each RPU cell is connected at an intersection of one of the plurality of first control lines and one of the plurality of second control lines, and the one array of a plurality of RPU cells includes an array of one of artificial synapse elements that store synaptic weights representing connection strengths between an upstream layer of artificial neurons of the neuromorphic computing system and a downstream layer of artificial neurons of the neuromorphic computing system, where the synaptic weights are encoded by conductance values of resistive devices of a plurality of RPU cells; Peripheral circuits connected to the plurality of first control lines and the plurality of second control lines, where the peripheral circuits include a readout circuit; and, A control system operably connected to the peripheral circuit, wherein the control system controls the peripheral circuit to generate control signals for performing a first operation on the one array of a plurality of RPU cells and for performing a second operation on the one array of a plurality of RPU cells. Comprising wherein the control signals are a first configuration control signal for configuring the read circuit to have a first hardware configuration when the first operation is performed on the one array of a plurality of RPU cells, wherein the first operation includes a forward pass operation of a neural network training process; and a second configuration control signal for configuring the read circuit to have a second hardware configuration different from the first hardware configuration when the second operation is performed on the one array of a plurality of RPU cells, wherein the second operation includes a backward pass operation of the neural network training process. The system including. [Article 20] The peripheral circuit further includes a control line voltage drive circuit and a multiplexer circuit; The control signal further has a multiplexer control signal for selectively connecting the read circuit to the plurality of first control lines and the control line voltage drive circuit to the plurality of second control lines to cause the first operation to be performed on the multiplexer circuit, and for selectively connecting the read circuit to the plurality of second control lines and the control line voltage drive circuit to the plurality of first control lines to cause the second operation to be performed on the multiplexer circuit; In the case of the first operation, the control line voltage drive circuit applies a first voltage pulse to one or more of the plurality of second control lines, wherein the voltage pulse represents a digital input signal received from the upstream layer, and the read circuit having the first hardware configuration receives a first current signal generated by the RPU cell and output to the plurality of first control lines in response to the first voltage pulse, and generates a first digital output signal representing the first current signal output from the plurality of first control lines; In the case of the second operation, the control line voltage drive circuit applies a second voltage pulse to one or more of the plurality of first control lines, where the second voltage pulse represents a digital error signal received from the downstream layer, and the readout circuit having the second hardware configuration receives a second current signal generated by an RPU cell and output to the plurality of second control lines in response to the second voltage pulse, and generates a second digital output signal representing the second current signal output from the plurality of second control lines. The system according to claim 19.

Claims

1. A device, the device comprising: an array of one of a plurality of resistance processing units (RPUs); a plurality of first control lines extending in a first direction across the one array of the plurality of RPU cells, and a plurality of second control lines extending in a second direction across the one array of the plurality of RPU cells, wherein each RPU cell is connected at an intersection of one of the plurality of first control lines and one of the plurality of second control lines; peripheral circuitry connected to the plurality of first control lines and to the plurality of second control lines, wherein the peripheral circuitry comprises a readout circuit shared by the plurality of first control lines and the plurality of second control lines; and a control system operably connected to the peripheral circuitry, wherein the control system controls the peripheral circuitry to generate control signals for performing a first operation on the one array of the plurality of RPU cells and for performing a second operation on the one array of the plurality of RPU cells, wherein: the control signals comprise: a first configuration control signal configured to cause the readout circuit to have a first hardware configuration when the first operation is performed on the one array of the plurality of RPU cells, enabling the reading of the signals with a first output signal bound b1; and a second configuration control signal configured to cause the readout circuit to have a second hardware configuration different from the first hardware configuration when the second operation is performed on the one array of the plurality of RPU cells, enabling the reading of the signals with a second output signal bound b2 wherein b2 < b1, the device.

2. the peripheral circuitry further comprising a control line voltage drive circuit and a multiplexer circuit; and the control signals further having a multiplexer control signal for selectively connecting the readout circuit to the plurality of first control lines and the control line voltage drive circuit to the plurality of second control lines to cause the first operation to be performed by the multiplexer circuit, and for selectively connecting the readout circuit to the plurality of second control lines and the control line voltage drive circuit to the plurality of first control lines to cause the second operation to be performed by the multiplexer circuit, The device according to claim 1.

3. The first operation includes a forward pass operation of a neural network training process performed on the one array of a plurality of RPU cells, where the one array of the plurality of RPU cells comprises one array of artificial synapse elements connecting an upstream layer of artificial neurons and a downstream layer of artificial neurons; The second operation includes a backward pass operation of the neural network training process performed on the one array of the plurality of RPU cells; In the case of the first operation, the control line voltage driving circuit applies a first voltage pulse to one or more of the plurality of second control lines, where the first voltage pulse represents a digital input signal received from the upstream layer of the neural network, and the readout circuit having the first hardware configuration receives a first current signal generated by the RPU cell and output to the plurality of first control lines in response to the first voltage pulse, and generates a first digital output signal representing the first current signal output from the plurality of first control lines; In the case of the second operation, the control line voltage driving circuit applies a second voltage pulse to one or more of the plurality of first control lines, where the second voltage pulse represents a digital error signal received from the downstream layer of the neural network, and the readout circuit having the second hardware configuration receives a second current signal generated by the RPU cell and output to the plurality of second control lines in response to the second voltage pulse, and generates a second digital output signal representing the second current signal output from the plurality of second control lines. The device according to claim 2.

4. The readout circuit comprises a plurality of readout circuit blocks, where each readout circuit block comprises a current integration circuit having an input node and an output node; and, an analog-to-digital conversion circuit having an input connected to the output node of the current integration circuit and is provided with, wherein the current integration circuit an operational amplifier having an inverting input terminal connected to the input node and an output terminal connected to the output node; A first integrating capacitor and a first switch connected in series in a first feedback path between the input node and the output node of the current integrating circuit, where the first integrating capacitor has a first capacitance; and, A second integrating capacitor and a second switch connected in series in a second feedback path between the input node and the output node of the current integrating circuit, where the second integrating capacitor has a second capacitance smaller than the first capacitance, comprising: where the first configuration control signal activates the first switch and deactivates the second switch to configure the current integrating circuit to have a first gain based on the first capacitance of the first integrating capacitor connected in the first feedback path between the output node and the input node of the current integrating circuit; and, where the second configuration control signal activates the second switch and deactivates the first switch to configure the current integrating circuit to have a second gain greater than the first gain based on the second capacitance of the second integrating capacitor connected in the second feedback path between the output node and the input node of the current integrating circuit, The device according to claim 1.

5. The readout circuit comprises a plurality of readout circuit blocks, and each readout circuit block A current integrating circuit having an input node and an output node; A selection circuit connected to the output node of the current integrating circuit; A first analog-to-digital conversion circuit connected to a first output of the selection circuit, where the first analog-to-digital conversion circuit has a first least significant bit (LSB) voltage resolution; and, A second analog-to-digital conversion circuit connected to a second output of the selection circuit, where the second analog-to-digital conversion circuit has a second LSB voltage resolution greater than the first LSB voltage resolution, comprising: where the first configuration control signal causes the selection circuit to selectively connect the output node of the current integrating circuit to the first analog-to-digital conversion circuit; and, Here, the second configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the second analog-to-digital conversion circuit. The device according to claim 1. **Claim 6** The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; a selection circuit connected to the output node of the current integration circuit; a first analog-to-digital conversion circuit connected to a first output of the selection circuit, where the first analog-to-digital conversion circuit has a first gain; a second analog-to-digital conversion circuit connected to a second output of the selection circuit, where the second analog-to-digital conversion circuit has a second gain greater than the first gain. The readout circuit is provided with Here, the first configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the first analog-to-digital conversion circuit; and Here, the second configuration control signal causes the selection circuit to selectively connect the output node of the current integration circuit to the second analog-to-digital conversion circuit. The device according to claim 1. **Claim 7** The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a selection circuit having a first input connected to one of the plurality of first control lines and a second input connected to one of the plurality of second control lines; a first current integration circuit having an input node and an output node; a first analog-to-digital conversion circuit connected to the output node of the first current integration circuit, where the first analog-to-digital conversion circuit has a first least significant bit (LSB) voltage resolution; a second current integration circuit having an input node and an output node; and a second analog-to-digital conversion circuit connected to the output node of the second current integration circuit, where the second analog-to-digital conversion circuit has a second LSB voltage resolution greater than the first LSB voltage resolution. The readout circuit is provided with Here, the first configuration control signal causes the selection circuit to selectively connect the one first control line to the input node of the first current integration circuit; and Here, the second configuration control signal causes the selection circuit to selectively connect the one second control line to the input node of the second current integration circuit. The device according to claim 1. **Claim 8** The first current integration circuit includes a first operational amplifier having an inverting input terminal connected to the input node of the first current integration circuit and an output terminal connected to the output node of the first current integration circuit, and a first integration capacitor connected in a feedback path between the input node and the output node, where the first integration capacitor has a first capacitance; and The second current integration circuit includes a second operational amplifier having an inverting input terminal connected to the input node of the second current integration circuit and an output terminal connected to the output node of the second current integration circuit, and a second integration capacitor connected in a feedback path between the input node and the output node, where the second integration capacitor has a second capacitance substantially equal to the first capacitance. The device according to claim 7. **Claim 9** The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block A selection circuit having a first input connected to one of the plurality of first control lines and a second input connected to one of the plurality of second control lines; A first current integration circuit having an input node and an output node; A first analog-to-digital conversion circuit connected to the output node of the first current integration circuit, where the first analog-to-digital conversion circuit has a first gain; A second current integration circuit having an input node and an output node; and A second analog-to-digital conversion circuit connected to the output node of the second current integration circuit, where the second analog-to-digital conversion circuit has a second gain greater than the first gain. and Here, the first configuration control signal causes the selection circuit to selectively connect the one first control line to the input node of the first current integration circuit; and Here, the second configuration control signal causes the selection circuit to connect the one second control line to the input node of the second current integration circuit. The device according to claim 1. **Claim 10** The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; a selection circuit connected to the output node of the current integration circuit; and an analog-to-digital conversion circuit connected to a first output of the selection circuit, where the analog-to-digital conversion circuit has a configurable least significant bit (LSB) voltage resolution including a first LSB voltage resolution and a second LSB voltage resolution greater than the first LSB voltage resolution, and includes where the first configuration control signal configures the analog-to-digital conversion circuit to have the first LSB voltage resolution; and where the second configuration control signal configures the analog-to-digital conversion circuit to have the second LSB voltage resolution, The device according to claim 1.

11. The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; a selection circuit connected to the output node of the current integration circuit; and an analog-to-digital conversion circuit connected to a first output of the selection circuit, where the analog-to-digital conversion circuit has a configurable gain including a first gain and a second gain greater than the first gain, and includes where the first configuration control signal configures the analog-to-digital conversion circuit to have the first gain; and where the second configuration control signal configures the analog-to-digital conversion circuit to have the second gain, The device according to claim 1.

12. The readout circuit includes a plurality of readout circuit blocks, and each readout circuit block includes a current integration circuit having an input node and an output node; a selection circuit connected to the output node of the current integration circuit; and an analog-to-digital conversion circuit connected to a first output of the selection circuit and includes where the current integration circuit has a configurable integration period including a first integration period and a second integration period longer than the first integration period; where the first configuration control signal configures the current integration circuit to perform a current integration process having the first integration period; and Here, the second configuration control signal configures the current integration circuit to perform a current integration process having the second integration period. The device according to claim 11.

13. The peripheral circuit further includes a control line voltage drive circuit, where the control line voltage drive circuit generates a voltage pulse having a first pulse duration proportional to the first integration period in response to the first configuration control signal and generates a voltage pulse having a second pulse duration proportional to the second integration period in response to the second configuration control signal, and the device according to claim 12 includes a configurable pulse width modulation circuit.

14. Performing a first operation on one array of a plurality of resistance processing unit (RPU) cells, where the one array of the plurality of RPU cells is connected to row lines and column lines that share a readout circuit; Configuring a readout circuit to enable reading of the signal at a first output signal bound b1 by having a first hardware configuration for performing the first operation on the one array of the plurality of RPU cells; Performing a second operation on the one array of the plurality of RPU cells; and Configuring the readout circuit to enable reading of the signal at a second output signal bound b2 by having a second hardware configuration for performing the second operation on the one array of the plurality of RPU cells, where the second hardware configuration of the readout circuit is different from the first hardware configuration of the readout circuit and b2 < b1. A method including the above.

15. The first operation includes a forward pass operation of a neural network training process performed on the one array of the plurality of RPU cells, where the one array of the plurality of RPU cells includes one array of artificial synapse elements connecting an upstream layer of artificial neurons and a downstream layer of artificial neurons; The second operation includes a backward pass operation of the neural network training process performed on the one array of the plurality of RPU cells. The method according to claim 14.

16. Configuring the readout circuit to have a first hardware configuration includes configuring the current integration circuit of the readout circuit to have a first integration capacitance; and Configuring the readout circuit to have a second hardware configuration includes configuring the current integration circuit of the readout circuit to have a second integration capacitance smaller than the first integration capacitance. The method according to claim 14. **Claim 17** Configuring the readout circuit to have a first hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a first least significant bit (LSB) voltage resolution; and Configuring the readout circuit to have a second hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a second LSB voltage resolution greater than the first LSB voltage resolution. The method according to claim 14. **Claim 18** Configuring the readout circuit to have a first hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a first gain; and Configuring the readout circuit to have a second hardware configuration includes configuring the analog-to-digital conversion circuit of the readout circuit to have a second gain greater than the first gain. The method according to claim 14. **Claim 19** A system comprising a neuromorphic computing system, wherein the neuromorphic computing system has an array of one of a plurality of resistive processing unit (RPU) cells, where a plurality of first control lines extend in a first direction across the one array of a plurality of RPU cells, and a plurality of second control lines extend in a second direction across the one array of a plurality of RPU cells, where each RPU cell is connected at an intersection of one of the plurality of first control lines and one of the plurality of second control lines, and the one array of a plurality of RPU cells has an array of one of artificial synapse elements storing synaptic weights representing connection strengths between an upstream layer of artificial neurons of the neuromorphic computing system and a downstream layer of artificial neurons of the neuromorphic computing system, where the synaptic weights are encoded by conductance values of resistive devices of a plurality of RPU cells; A peripheral circuit connected to the plurality of first control lines and the plurality of second control lines, wherein the peripheral circuit includes a read circuit shared by the plurality of first control lines and the plurality of second control lines; and, A control system operably connected to the peripheral circuit, wherein the control system controls the peripheral circuit to generate control signals for performing a first operation on the one array of a plurality of RPU cells and performing a second operation on the one array of a plurality of RPU cells, Comprising, Wherein the control signal is, When the first operation is performed on the one array of a plurality of RPU cells, configuring the read circuit to have a first hardware configuration to enable reading of the signal with a first output signal bound b1, wherein the first operation includes a forward pass operation of a neural network training process; and, When the second operation is performed on the one array of a plurality of RPU cells, configuring the read circuit to have a second hardware configuration different from the first hardware configuration to enable reading of the signal with a second output signal bound b2, wherein the second operation includes a backward pass operation of the neural network training process, and b2 < b1, Including, the system.

20. The peripheral circuit further includes a control line voltage drive circuit and a multiplexer circuit; The control signal further has a multiplexer control signal for selectively connecting the read circuit to the plurality of first control lines and the control line voltage drive circuit to the plurality of second control lines to cause the multiplexer circuit to perform the first operation, and selectively connecting the read circuit to the plurality of second control lines and the control line voltage drive circuit to the plurality of first control lines to cause the multiplexer circuit to perform the second operation; In the case of the first operation, the control line voltage drive circuit applies a first voltage pulse to one or more of the plurality of second control lines, where the first voltage pulse represents the digital input signal received from the upstream layer, and the readout circuit having the first hardware configuration receives the first current signal generated by the RPU cell and output to the plurality of first control lines in response to the first voltage pulse, and generates a first digital output signal representing the first current signal output from the plurality of first control lines; In the case of the second operation, the control line voltage drive circuit applies a second voltage pulse to one or more of the plurality of first control lines, where the second voltage pulse represents the digital error signal received from the downstream layer, and the readout circuit having the second hardware configuration receives the second current signal generated by the RPU cell and output to the plurality of second control lines in response to the second voltage pulse, and generates a second digital output signal representing the second current signal output from the plurality of second control lines. The system according to claim 19.

Citation Information

Patent Citations

  • Memristor Neuromorphic Circuit and Method for Training a Memristor Neuromorphic Circuit

    JP2018521400A

  • Resistive processing unit architecture with separate weight update and inference circuitry

    WO2019202427A1