DNN training algorithm with dynamically computed zero references
By employing RPU crossbar arrays and digital memory with a chopper mechanism, the method addresses the inefficiencies of symmetric RPU devices in DNN training, improving computational efficiency and accuracy.
Patent Information
- Application Number
- JP2025520120
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-19
- Publication Date
- 2025-10-09
AI Technical Summary
Existing DNN training methods require symmetric RPU devices, which are prone to misestimation and noise, leading to inaccurate gradient updates and inefficient computational resources.
Utilize a combination of RPU crossbar arrays and digital memory to dynamically calculate and store reference values, allowing for asymmetric RPU devices by incorporating a chopper mechanism to negate biases and improve computational efficiency and accuracy.
Enhances the efficiency and accuracy of DNN training by reducing computing resources and noise, even with asymmetric RPU devices, through dynamic reference value calculation and on-the-fly updates.
Smart Images

Figure 2025533921000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to deep learning, and more particularly to systems and methods for training deep neural networks using hardware elements. [Background technology]
[0002] A deep neural network (DNN) can be embodied in an analog cross-point array of resistive devices, such as resistive processing units (RPUs). The RPU device generally includes a first terminal, a second terminal, and an active area. The conductance state of the active area identifies a weight value for the RPU, which can be updated / adjusted by applying a signal to the first terminal or the second terminal.
[0003] DNN-based models have been used for various cognitive-based tasks, such as object and speech recognition and natural language processing. DNN training is important for providing high levels of accuracy when performing such tasks. Training large DNNs is a computationally intensive task. Most common methods for DNN training, such as backpropagation and stochastic gradient descent (SGD), require the RPU to be "symmetric" in order to operate accurately. Typical systems assume that symmetric points are correctly estimated and initially stored in a reference device array. These symmetric points can be misestimated or even misdescribed, including noise. Summary of the Invention
[0004] According to an embodiment of the present disclosure, a computer-implemented method includes performing a gradient update of a stochastic gradient descent (SGD) algorithm for a deep neural network (DNN) using a first set of hidden weights stored in a first matrix including a resistor processing unit (RPU) crossbar array. A second matrix including a second set of hidden weights is stored in a digital medium. A third matrix including a set of reference values is calculated during a transfer cycle of the first set of weights from the first matrix to a second matrix, taking into account a sign change (chopper). The third matrix is stored in a digital medium. A third set of weights for the DNN is updated from the second matrix when a threshold value for the second set of weights is reached in a fourth matrix including the RPU crossbar array. This device has the technical effect of improving the efficiency and accuracy of system computations for data used in the RPU system.
[0005] In one embodiment, which may be combined with the preceding embodiment, the second set of weights takes into account a previous set of reference values from a previous iteration of the transfer cycle, which allows for more efficient computational performance.
[0006] In one embodiment that may be combined with the preceding embodiments, a fifth matrix stored on a digital medium is configured to calculate a next set of reference values from values read from the first matrix during a chopper cycle, and the fifth matrix is configured to partially update the third matrix after a chopper cycle is completed, thereby enabling more accurate data manipulation.
[0007] In one embodiment, which may be combined with the preceding embodiments, the calculation of the SGD includes a fifth matrix containing the set of previous reference values and storing the fifth matrix on a digital medium, which allows for more efficient computational performance.
[0008] In one embodiment, which may be combined with the preceding embodiment, the assignment of a set of reference values to a previous set of reference values in the digital medium occurs at the time of chopper switching, which allows for more accurate calculation functions.
[0009] In one embodiment, which may be combined with the preceding embodiment, the resetting of the set of reference values to zero occurs when the chopper switches, which allows for more efficient computational performance.
[0010] In one embodiment, which may be combined with the preceding embodiments, the device is configured to switch the sign of the chopper when the chopper switches, which allows for more accurate data manipulation.
[0011] In one embodiment, which may be combined with the preceding embodiment, the RPU crossbar array is not used to store the set of reference values, which allows for more efficient use of space within the IC array.
[0012] In one embodiment, which may be combined with the preceding embodiment, the previous set of reference values is set to the most recent read weight vector, which allows for more efficient use of space within the IC array.
[0013] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium tangibly embodied with computer-readable program code having computer-readable instructions for solving machine learning tasks, the instructions, when executed, causing a computing device to perform a method. The method includes performing a gradient update of a stochastic gradient descent (SGD) of a deep neural network (DNN) using a first set of weights stored in a first matrix including a resistive processing unit (RPU) crossbar array. A second matrix including a second set of weights is stored in a digital medium. A third matrix including a set of reference values is calculated for the SGD during a transfer cycle of the first set of weights from the first matrix to the second matrix, taking into account a chopper. The third matrix is stored in a digital medium. A third set of weights of the DNN is updated from the second matrix when a threshold value for the second set of weights is reached in a fourth matrix including the RPU crossbar array. This device has the technical effect of improving the efficiency and accuracy of system computations of data used in the RPU system.
[0014] According to an embodiment of the present disclosure, a device including a first matrix includes a resistive processing unit (RPU) crossbar array including a first set of weights configured for gradient updates of stochastic gradient descent (SGD) of a deep neural network (DNN). The device includes a second matrix including a second set of weights stored in a digital medium. The device further includes a third matrix including a set of reference values calculated for SGD, stored in the digital medium, the set of reference values being calculated during a transfer cycle of the first set of weights from the first matrix to the second matrix, taking into account a chopper. The device may also include a fourth matrix including an RPU crossbar array storing a third set of weights for the DNN that are updated from the second matrix when a threshold value for the second set of weights is reached. This device has the technical effect of improving the efficiency and accuracy of system calculations of data used in the RPU system.
[0015] In one embodiment, which may be combined with the preceding embodiment, the second set of weights takes into account a previous set of reference values from a previous iteration of the transfer cycle, which allows for more efficient computational performance.
[0016] In one embodiment, which may be combined with the preceding embodiments, the set of reference values takes into account the switching frequency, which allows for more accurate data manipulation.
[0017] In one embodiment, which may be combined with the preceding embodiments, a fifth matrix containing a set of previous reference values calculated for the SGD is stored in a digital medium, which allows for more efficient computational performance.
[0018] In one embodiment, which may be combined with the preceding embodiment, the device assigns a set of reference values to a previous set of reference values in the digital medium when the chopper switches, which allows for more efficient computational performance.
[0019] In one embodiment, which may be combined with the preceding embodiments, the device resets the set of reference values to zero when the chopper switches, which allows for more efficient computational performance.
[0020] In one embodiment, which may be combined with the preceding embodiments, the device switches the sign of the chopper when the chopper switches, which allows for more precise data manipulation.
[0021] In one embodiment, which may be combined with the preceding embodiment, the RPU crossbar array is not used to store the set of reference values, which allows for more efficient use of space within the IC array.
[0022] In one embodiment, which may be combined with the preceding embodiment, the previous set of reference values is set to the most recent read weight vector, which allows for more efficient use of space within the IC array.
[0023] The techniques described herein may be implemented in a number of ways. An example implementation is shown below with reference to the following figure. [Brief explanation of the drawings]
[0024] The drawings are of exemplary embodiments. They do not illustrate all embodiments. Other embodiments may be used in addition or instead. Details that may be obvious or unnecessary may be omitted to save space or for a more effective illustration. Some embodiments may be practiced using additional components or steps and / or without all components or steps illustrated. When the same numerals appear in different drawings, the numerals refer to the same or similar components or steps.
[0025] [Figure 1] FIG. 1 is a schematic diagram showing a DNN with a weight matrix W, an A matrix, and a hidden matrix H. [Figure 2] FIG. 1 illustrates a DNN embodied in an analog cross-point array of an RPU device, according to an embodiment. [Figure 3] 1 is a process flow illustrating an exemplary methodology for training a DNN, according to an embodiment. [Figure 4A] FIG. 1 illustrates an interconnected array with digital memories used to estimate reference values on the fly. [Figure 4B] FIG. 1 illustrates an interconnected array with digital memories used to estimate reference values on the fly. [Figure 5] FIG. 10 illustrates a forward cycle y=Wx performed according to an embodiment. [Figure 6] FIG. 10 illustrates a reverse cycle z=WTδ performed according to an embodiment. [Figure 7] FIG. 10 illustrates an embodiment in which array A is updated with x propagated in the forward cycle and δ propagated in the backward cycle. [Figure 8]FIG. 10 illustrates a forward cycle y′=Aei performed on a weight matrix, according to an embodiment. [Figure 9] FIG. 10 shows the hidden matrix H being updated with values calculated in the forward cycle of the A matrix. [Figure 10] FIG. 9 is a schematic diagram of a hidden matrix H902 being selectively applied to a weight matrix W1010, according to an embodiment. [Figure 11] FIG. 1 illustrates an example of a one-hot encoding vector, according to an embodiment. [Figure 12] FIG. 10 illustrates an example of a detailed algorithm, according to an embodiment. [Figure 13] FIG. 10 illustrates an example of detailed sub-algorithms, according to an embodiment. [Figure 14] FIG. 1 illustrates an exemplary apparatus that can be used to perform one or more of the present techniques, according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0026] In the following detailed description, by way of example, numerous specific details are set forth in order to provide a thorough understanding of the relevant teachings. However, it will be apparent that the present teachings may be practiced without such details. In other instances, well-known methods, procedures, components, and / or circuits have been described in relatively broad terms, without detail, in order to avoid unnecessarily obscuring aspects of the present teachings.
[0027] This specification provides methods for training DNNs using asymmetric RPU devices. The DNN is trained using two adjustable resistive device arrays and two or three digital memory arrays. These methods may include using an RPU crossbar array to represent the weights of the DNN. An additional crossbar array may be used for each weight to calculate gradient updates, without requiring a third adjustable RPU array used to store references. Furthermore, updates to both RPU arrays may be performed according to the algorithms described herein.
[0028] In typical systems, the symmetry point for each device can be incorrectly estimated. The symmetry point is the conductance at which the conductance change response to a single pulse update in the positive direction is the same as in the negative direction, on average. The symmetry point can be incorrectly written to the reference device with noise, resulting in an incorrect value being subtracted during gradient value readout. The update device can be variable, causing its symmetry point to be unstable and move over time. Furthermore, inputs are often too sparse or there are too many devices, resulting in the symmetry point being reached slowly and leaving a temporary offset. Furthermore, adding a dedicated reference device array is expensive per unit of integrated circuit chip area. Embodiments overcome these limitations by using digital memory to store the indices used dynamically on the fly to estimate the reference.
[0029] Thus, one or more of the methodologies discussed herein may eliminate the need for time-consuming data processing by a user, which may have the technical effect of reducing computing resources used by one or more devices in the system. Examples of such computing resources include, but are not limited to, processor cycles, network traffic, memory usage, storage capacity, and power consumption.
[0030] It should be understood that aspects of the teachings herein are beyond the capabilities of the human mind. It should also be understood that various embodiments of the subject disclosure described herein may include information that is impossible to manually obtain by an entity such as a human user. For example, the voltage inputs and conductance storage values discussed herein are impossible for a human user to perform.
[0031] Referring now to the drawings, FIG. 1 shows a weight matrix W 102, an A matrix 112, and a μ past 1 is a schematic diagram showing a DNN 100 having a matrix 113 and a hidden matrix H 114. The weight matrix W 102 is connected to the A matrix 112, μ pastIt is iteratively trained using matrix 113 and hidden matrix 114. As highlighted above, weight matrix W102 can be implemented in an analog cross-point array of the RPU. See, for example, the schematic diagram shown in FIG. 2.
[0032] As shown in FIG. 2, each parameter (weight W ij ) is a single RPU device on the hardware (RPU ij ), i.e., a physical crosspoint array 104 of the RPU device. The crosspoint array 104 includes a series of conductive row wires 106 and a series of conductive column wires 108 that are oriented orthogonal to and intersect the conductive row wires 106. The crosspoints between the row wires 106 and the column wires 108 are separated by RPUs 110, forming the crosspoint array 104 of the RPU device. Each RPU 110 may include a first terminal, a second terminal, and an active area. The conduction state of the active area identifies a weight value for the RPU 110, which can be updated / adjusted by applying a signal to the first or second terminal. Furthermore, a three-terminal (or more) device can effectively function as a two-terminal resistive memory device by controlling the additional terminal.
[0033] Each RPU 110 (RPUij) is uniquely identified based on its location (i.e., its ith row and jth column) within the crosspoint array 104. For example, working from top to bottom and left to right through the crosspoint array 104, the RPU at the intersection of the first row wire 106 and the first column wire 108 is designated RPU11, the RPU at the intersection of the first row wire 106 and the second column wire 108 is designated RPU12, and so on. Furthermore, the same rule is followed when mapping the parameters of the weight matrix 102 to the RPUs of the crosspoint array 104. For example, a weight w i1 is the RPU of the crosspoint array 104 i1 and the weights w i2is mapped to RPUi2 in the crosspoint array 104, and so on.
[0034] The RPUs 110 of the cross-point array 104 actually function as weighted connections between neurons in the DNN. The conduction state (e.g., resistance) of the RPUs 110 can be changed by controlling the voltages applied between the row wires 106 and column wires 108, respectively. Data is stored by changing the conduction state of the RPUs. The conduction state of the RPUs 110 is read by applying a voltage and measuring the current passing through the target RPU 110. All operations, including weights, are performed by the RPUs 110 in fully parallel fashion.
[0035] In machine learning and cognitive science, DNN-based models are a family of statistical learning models inspired by biological neural networks in animals, specifically the brain. These models can be used to estimate or approximate systems and cognitive functions that depend on many generally unknown inputs and connection weights. DNNs are often embodied as so-called "neuromorphic" systems of interconnected processor elements that function as simulated "neurons" that exchange "messages" between each other in the form of electronic signals. The connections in a DNN that transmit electronic messages between simulated neurons are provided with numerical weights corresponding to the strength or weakness of a given connection. These numerical weights can be adjusted and tuned based on experience, allowing the DNN to adapt to and learn from inputs. For example, a DNN for handwriting recognition is defined by a set of input neurons that can be activated by pixels in an input image. The activation of these input neurons is weighted and transformed by a function determined by the network designer before being passed on to other downstream neurons. This process is repeated until an output neuron is activated, which determines which character has been read.
[0036] The DNN 100 shown in FIG. 1 calculates weight values W via the A matrix 112, as described in more detail below. ijand then summing the resulting output from the A matrix 112 into the hidden matrix 114 to update the elements of the hidden matrix 114 (i.e., H ij ) reaches a threshold. However, before and after the weight values are updated in the A matrix 112, the chopper 116 multiplies the input and output signals by a chopper value. The chopper value at a given time is equal to either a positive value (+1) or a negative value (-1). The chopper 116 randomly or periodically flips between chopper values so that updates are applied to the A matrix 114 with opposite signs during part of the training period. This sign flip by the chopper 116 means that any "bias" contributed to the weight values by the A matrix 112 has one sign (i.e., positive or negative) during some periods of the training time and the other sign (i.e., negative or positive) during other periods of the training time. The chopping period or switching probability may also be assigned by the user. Bias can be inherent in any analog system, including non-ideal RPUs that may be used in the DNN 100.
[0037] For training purposes, such an ideal device would fully implement the DNN training process of backpropagation and stochastic gradient descent (SGD). Backpropagation is a training process that runs in three cycles: a forward cycle, a backward cycle, and a weight update cycle, and is repeated multiple times until a convergence criterion is met. Stochastic gradient descent (SGD) uses backpropagation to iterate over each parameter (weight W ij ) error gradient.
[0038] To perform backpropagation, DNN-based models consist of multiple processing layers that learn representations of data at multiple levels of abstraction. For a single processing layer with N input neurons connected to M output neurons, the forward cycle involves computing vector-matrix multiplication (y = Wx), where the length-N vector x represents the activations of the input neurons and the size-MxN matrix W stores the weight values between each pair of input and output neurons. The resulting length-M vector y is further processed by performing nonlinear activation on each of the resistive memory elements and then passed to the next layer.
[0039] Once the information reaches the final output layer, the backward cycle involves computing the error signal and backpropagating it through the DNN. The backward cycle for a single layer involves vector-matrix multiplication (z=W T δ), where the vector δ of length M represents the error calculated by the output neuron, and the vector z of length N is further processed using the derivative of the neuron's nonlinearity and then passed to the previous layer.
[0040] Finally, in the weight update cycle, the weight matrix W is updated by performing the cross product of the two vectors used in the forward and backward cycles. The cross product of these two vectors is W ← W + η(δx T ), where η is the global learning rate.
[0041] All operations performed on the weight matrix W during this backpropagation process may be performed using the crosspoint array 104 of the RPU 110, which has a corresponding number of M rows and N columns, with the conductance values stored in the crosspoint array 104 forming the matrix W. In the forward cycle, an input vector x is transmitted as a voltage pulse through each of the column wires 108, and the resulting vector y is read as a current output from the row wires 106. Similarly, when a voltage pulse is provided as an input from the row wires 106 to the reverse cycle, the weight matrix W TFinally, in the update cycle, voltage pulses representing vectors x and δ are simultaneously fed from the column wires 108 and row wires 106. In this configuration, each RPU 110 performs local multiplication and addition operations by processing the voltage pulses coming from the corresponding column wires 108 and row wires 106, thus achieving incremental weight updates.
[0042] A symmetric RPU can perfectly perform backpropagation and SGD. That is, in such an ideal RPU, w ij ←w ij +ηΔw ij and here w ij is the weight value of the ith row and jth column of the cross point array 104.
[0043] FIG. 3 illustrates an exemplary method 300 for training a DNN, according to an embodiment. During training, weight updates are first accumulated in an A matrix, which is a hardware component made up of rows and columns of an RPU with symmetric behavior around the zero point. Weight updates from the A matrix are then selectively moved to a weight matrix W, which is also a hardware component made up of rows and columns of an RPU. The training process involves selecting parameters (weights W ) that maximize the accuracy of the DNN. ij ) is determined iteratively. The matrix W is initialized to randomly distributed values using common methods applied in DNN training. The hidden matrix H, stored in digital form, is initialized to zero.
[0044] During training, weight updates are performed on the A matrix. The information processed by the A matrix is then accumulated in a hidden matrix H (another matrix that effectively performs a low-pass filter). Values of the hidden matrix H that reach an update threshold are applied to the weight matrix W. The update threshold effectively minimizes noise generated within the A matrix hardware. However, for elements of the A matrix initialized with a bias, the update threshold is reached early because each iteration from the element performs a consistent update (either positive or negative) based on the bias, rather than based on the weight updates associated with training the DNN. The chopper value negates the bias by inverting its sign for a certain period, during which the bias is summed with the opposite sign to the hidden matrix H. Specifically, during some periods, a positive bias is added to the weight values and summed to the hidden matrix H, while during other periods, a negative bias is added to the weight values and summed to the hidden matrix H. The random flipping of the chopper value means that periods of positive bias tend to be evenly spaced with periods of negative bias. Thus, hardware bias and noise associated with a non-ideal RPU are tolerated (or absorbed into the H matrix), thus resulting in lower test error compared to standard SGD methods, hidden matrix H only, or other training methods using asymmetric devices, even for small numbers of states.
[0045] The method 300 initializes the A matrix, the digitally computed values μ, the hidden matrix H (also stored in a digital buffer), and the weight matrix W at block 302. Initializing the A matrix includes, for example, setting all values to zero. The array A can be implemented as a single interconnected array.
[0046] 4A-4B illustrate an interconnected array with digital memory used to estimate a reference value on the fly. As shown, μ represents the recent past of the gradient update matrix A. In some embodiments, the recent past μ is used in a difference calculation in digital storage or memory, resulting in the value ω used to update H. This creates a floating-point representation of the reference value, where the reference value changes over time according to method 300. This dynamic update and on-the-fly calculation of the reference value helps eliminate the bias of previous systems that used a hardware reference RPU matrix for the reference value.
[0047] Initializing the hidden matrix H involves zeroing out the current values stored in the matrix or allocating digital storage capacity on the connected computing device. Initializing the weight matrix W involves loading random values into the weight matrix W so that the training process of the weight matrix W can begin. ω is assigned based on reading from the A matrix for each column or row, where ω is the processed digital conversion value after using the ADC.
[0048] The digital H is a hidden matrix used to filter the gradient values calculated in A. ω is the readout of the analog A matrix, which can be read from each column or row by inputting a unit vector containing voltage (e.g., [1 0 0 0]). The weight in the current unit for that column is obtained and converted back to digital using an ADC. λ is the scaling factor, i.e., the learning rate. S is used to change the chopper value that switches between negative and positive.
[0049] When a threshold is reached in H, a pulse is then sent to the weight matrix W. In other words, the gradient is placed on the A crossbar RPU. Once placed, the gradient contains a lot of noise. The gradient is read again, a chopper is applied, a reference value is subtracted to remove any bias, and then added to the filter matrix to remove the noise. The gradient is then integrated over time, and the weights are updated when the gradient reaches a threshold. Thus, the weights W remain largely unchanged without any bias being applied. This significantly improves the noise performance and accuracy of prior art RPU algorithms.
[0050] Method 300 includes determining activation values by performing a forward cycle using a weight matrix W (block 304). Figure 5 illustrates the forward cycle performed, according to an embodiment. The forward cycle includes computing vector matrix multiplication (y = Wx), where activation values embodied as input vectors x represent the activity of input neurons, and weight matrix W stores the weight values between each pair of input and output neurons. Figure 5 illustrates that the vector matrix multiplication operation of the forward cycle is performed in the cross-point array 502 of the RPU device, and the conductance values stored in the cross-point array 502 form a matrix.
[0051] An input vector x is transmitted as a voltage pulse through each of the conductive column wires 512, and the resulting output vector y is read as a current output from the conductive row wires 510 of the crosspoint array 502. An analog-to-digital converter (ADC) 513 is used to convert the analog output vector 516 from the crosspoint array 502 to a digital signal.
[0052] The method 300 also includes determining an error value by performing a backward cycle on the weight matrix W (block 306). Figure 6 illustrates a backward cycle performed, according to an embodiment. In general, the backward cycle involves calculating an error value δ and backpropagating the error value δ to the weight matrix W via a vector-matrix multiplication on the transpose of the weight matrix W (i.e., z=W T δ, W T denotes the transpose of the matrix W), where the vector δ represents the error calculated by the output neuron, and the vector z is further processed using the derivative of the neuron's nonlinearity and then passed to the previous layer.
[0053] Figure 6 shows the reverse cycle vector-matrix multiplication operation being performed in crosspoint array 502. Error value δ is transmitted as a voltage pulse over each of the conductive row wires 510, and the resulting output vector z is read as a current output from conductive column wires 512 of crosspoint array 502. As voltage pulses are provided as inputs to the reverse cycle from row wires 510, a vector-matrix product is calculated on the transpose of weight matrix W. As also shown in Figure 6, ADC 513 is used to convert the (analog) output vector 518 from crosspoint array 502 to a digital signal.
[0054] The method 300 also includes applying a chopper value to the activation or error values (block 308). The chopper value may be applied by a chopper (e.g., chopper 116 of FIG. 1) included in each row wire and each column wire in the A matrix 502. In particular embodiments, the cross point array 502 may have choppers only on the column wires 506 or only on the row wires 504. After the chopper value is applied to the activation or error values, the method 300 also includes updating the A matrix using the activation values, the error values (input vectors x and δ), and the chopper value (block 310).
[0055] 7 illustrates an embodiment in which array A 502 is updated with x propagated in the forward cycle and δ propagated in the reverse cycle. Each row and each column has a chopper value 550 applied to its respective wire. The sign of chopper value 550 is represented by "+" for a positive chopper value (i.e., no change in the activation or error value) or "X" for a negative chopper value (i.e., a change in the sign of the activation or error value). In cross-point array 502, updates are performed by transmitting voltage pulses representing vector x (from the forward cycle) and vector δ (from the reverse cycle) simultaneously fed from conductive column wires 506 and conductive row wires 504, respectively. In this configuration, each RPU in cross-point array 502 performs local multiplication and addition operations by processing the voltage pulses coming from the corresponding conductive column wires 506 and conductive row wires 504, thus achieving incremental weight updates. The forward cycle (block 304), the backward cycle (block 306), and the update of the A matrix with input vectors from the forward and backward cycles (block 310) may be repeated multiple times to refine the updated values of the A matrix.
[0056] The method 300 calculates an input vector e i (i.e., y'=Ae i ) and the chopper value to read the i-th column of A by performing a forward cycle on the A matrix (block 312). At each time step, a new input vector e i is used, where the sub-index i indicates its time index. As will be explained in more detail below, according to an exemplary embodiment, the input vector e iis a one-hot encoding vector. For example, as known in the art, a one-hot encoding vector is a group of bits that has only combinations with a single high bit (1) and all other bits being low bits (0). Using a simple, non-limiting example for illustrative purposes, assuming a matrix of size 4x4, the one-hot encoding vector would be one of the vectors [1 0 0 0], [0 1 0 0], [0 0 1 0], and [0 0 0 1]. A new one-hot encoding vector is used at each time step, with the sub-index i indicating its time index. However, it is noted here that other methods for selecting the input vector ei are also contemplated herein.
[0057] FIG. 8 is a diagram illustrating reading the i-th column of A by performing a forward cycle y′=Aei on the A matrix using chopper values, where e i is the i-th unit vector. Alternatively, Y'=A T e i (Transpose read) can also be performed. i is transmitted as a voltage pulse through each of the conductive column wires 506, and the resulting output vector y' is read as a current output from the conductive row wires 504 of the crosspoint array 502. Each column wire 506 and row wire 504 is read with the same chopper value (i.e., positive or negative) as when the A matrix was updated. For example, in FIGS. 7 and 8, the first column wire 506 i1 has a positive chopper value (+), and in FIGS. 7 and 8, the second column wire 506 i2 has a negative chopper value (X), and in FIGS. 7 and 8, the first row wire 504 1i has a negative chopper value (X). A voltage pulse is provided from the column wire 506 as an input to this forward cycle, and then the vector-matrix product is calculated. The method 300 includes updating the hidden matrix H (block 314).
[0058] 9 shows a hidden matrix H 902 that is updated with values calculated in the forward cycle of the A matrix 904. The hidden matrix H 902 is a digital matrix, rather than a physical device like the A matrix and the weight matrix W, and each RPU (i.e., ij For each RPU in ij ) is stored. When the forward cycle is performed, the output vector y'e i T is generated, also referred to as ω. This output vector is used to calculate other digital matrices, as described in more detail below, and also to update the hidden matrix H. Thus, each time the output vector is read, the hidden matrix H 902 is modified. For an RPU with low noise levels, the H value 906 increases consistently. For a constant gradient and input, the increase can be positive or negative, depending on the value of the output vector ω. If the output vector ω contains significant noise, its value may be positive in one iteration and negative in another. This combination of positive and negative output vector ω values means that the H value 906 increases more slowly and consistently.
[0059] The values of the hidden matrix may be updated on the fly using digital storage to store and update the values of μ as follows:
[0060] The digital computation for each transfer cycle involves the following operations (for each element i of one read vector k): 1.μ ik ←(1-γ)μ ik +γω i 2.h ik ←h ik +s k λ′(ω i -μ ik past ),
[0061] where ω is the read weight vector, h ik is the digital buffer value, s kis the current chopper code, λ is the learning rate, and μ is a floating-point reference that changes over time and at different iterations. K is the number of n s It may wrap around and increase with each update.
[0062] Each time a vector k is read, after digital calculation, a buffer (with a threshold) can be written to the weight matrix W. γ is a user-defined parameter that can be positive or zero and is typically set to 2 / p, where p is the switching frequency assuming periodic switching.
[0063] As the H value 906 increases, the method 300 includes tracking whether the H value 906 has increased by more than a threshold value (block 316). ij If the H value 906 of the input vector e is not greater than the threshold (block 316, "No"), the method 300 repeats from performing the forward cycle (block 304) to updating the hidden matrix H (block 314) and possibly inverting the chopper value (blocks 320-322). If the H value 906 is greater than the threshold (block 316, "Yes"), the method 300 repeats from performing the forward cycle (block 304) to updating the hidden matrix H (block 314) and possibly inverting the chopper value (blocks 320-322). i to the weight matrix W, but only for the particular RPU (block 318). As mentioned above, the increase in H value 906 can be positive or negative, so the threshold value can also be positive or negative. Figure 10 is a schematic diagram of a hidden matrix H 902 being selectively applied to a weight matrix W 1010, according to an embodiment.
[0064] FIG. 10 shows a first H value 1012 and a second H value 1014 being transmitted to the weight matrix W 1010 beyond the threshold. The first H value 1012 reaches the positive threshold, and therefore a positive value of "1" is placed in the row of the input vector 1016. The second H value 1014 reaches the negative threshold, and therefore a negative value of "-1" is placed in the row of the input vector 1016. The remaining rows of the input vector 1016 are filled with zeros because their values (i.e., H values 906) have not increased significantly above the threshold. The threshold can be much larger than the value added to the hidden matrix H. For example, the threshold can be 10 or 100 times the expected strength of the value updated each cycle. Because the on-the-fly calculated reference value does not add bias to the H matrix, it is typically not necessary to make the threshold excessively large. A higher threshold reduces the frequency of updates performed on the weight matrix W. However, the filtering function performed by the H matrix reduces the error in the neural network's objective function. These updates can only be generated after processing a large number of data examples, thus increasing the confidence level of the updates. This technique allows neural networks to be trained on noisy RPU devices with only a limited number of states, even when the symmetry point shifts or becomes unstable. After the H value is applied to the weight matrix W, the H value 906 is reset to zero, and the iterations of method 300 continue.
[0065] Method 300 also includes inverting the sign of the chopper value at an inversion rate (block 320). In particular embodiments, the chopper value is inverted only after the chopper product is added to the hidden matrix H. That is, the chopper value is used twice: once when the activation and error values are written to the A matrix; and once when the forward cycle is read from the A matrix. The chopper value should not be inverted before the H matrix is updated. The inversion rate may be defined as a user preference so that after each chopper product is added to the hidden matrix H, the chopper has a percentage chance to invert the chopper value. For example, the user preference may be 50 percent so that half the time, the chopper value has a chance to change sign (i.e., from positive to negative or from negative to positive) after the chopper product is calculated. In other embodiments, for example, the chopper may be inverted every third or fourth time throughout the cycle.
[0066] If it is determined that the chopper is reversing (YES, block 320), the digital buffer value is further updated for on-the-fly reference estimation. For example, the following updates may occur:
[0067] μ ik past ←μ ik , μ past is updated with the current μ value of the i-th row and j-th column.
[0068] μ ik ←If 0, μ ik The value of is reset.
[0069] s k In the case of ←-s, the chopper value is inverted.
[0070] μ at reset past ik The final ω i Setting it to a value (e.g., γ=1 and omitting the moving average) can reduce memory space usage by a factor of two.
[0071] The chopper value is inverted and μ pastAfter is updated, method 300 continues by determining whether training is complete. If training is not complete, e.g., if certain convergence criteria are not met (block 324, "No"), method 300 begins again by performing a forward cycle y = Wx. For example, by way of example only, training may be considered complete when no further improvement is observed in the error signal. If training is complete (block 324, "Yes"), method 300 ends.
[0072] As highlighted above, according to an exemplary embodiment, the input vector e i is a one-hot encoding vector, which is a group of bits that have only combinations with a single high (1) bit and all other bits are low bits (0). See, for example, Figure 11. Assuming a matrix of size 4x4 as shown in Figure 11, the one-hot encoding vector will be one of the vectors [1 0 0 0], [0 1 0 0], [0 0 1 0], and [0 0 0 1]. At each time step, a new one-hot encoding vector is used, indicated by the sub-index i at that time index.
[0073] Figure 12 illustrates an example of a detailed algorithm according to an embodiment of the present disclosure. Figure 13 illustrates an example of a detailed sub-algorithm according to an embodiment of the present disclosure.
[0074] The present invention may be a system, a method, and / or a computer program product, which may include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to perform aspects of the present disclosure.
[0075] A computer-readable storage medium may be a tangible device that can hold and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer diskettes, hard disks, Random Access Memory (RAM), Read-Only Memory (ROM), Erasable Programmable Read-Only Memory (EPROM or flash memory), Static Random Access Memory (SRAM), portable Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disk (DVD), memory sticks, floppy disks, punch cards, or mechanically encoded devices such as ridge structures in grooves in which instructions are recorded, and any suitable combination of the foregoing. Computer-readable storage medium, as used herein, should not be construed as a transitory signal per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over electrical wires.
[0076] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to each computing / processing device or to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium in the respective computing / processing device for storage.
[0077] The computer-readable program instructions for carrying out the operations of the present disclosure may be source or object code written in any combination of one or more programming languages, including assembler instructions, Instruction-Set-Architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk®, C++, and the like, and conventional procedural programming languages such as the “C” programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, as a standalone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuitry to perform aspects of the present disclosure.
[0078] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0079] These computer-readable program instructions may be provided to a specially configured computer processor or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the computer processor or other programmable data processing apparatus, produce means for performing the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams. These computer-readable program instructions, which can instruct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, may also be stored on a computer-readable storage medium, such that the computer-readable storage medium having the instructions stored thereon comprises a product containing instructions for performing aspects of the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0080] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device and cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to generate a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowcharts and / or block diagrams.
[0081] 14, there is shown a block diagram of an apparatus 1400 for implementing one or more of the methodologies presented herein. By way of example only, the apparatus 1400 may be configured to control input voltage pulses applied to the array and / or process output signals from the array.
[0082] Apparatus 1400 includes a computer system 1410 and removable media 1450. Computer system 1410 includes a processor device 1420, a network interface 1425, memory 1430, a media interface 1435, and an optional display 1440. Network interface 1425 allows computer system 1410 to connect to a network, while media interface 1435 allows computer system 1410 to interact with media such as a hard drive or removable media 1450.
[0083] The processor device 1420 may be configured to implement the methods, steps, and functions disclosed herein. The memory 1430 may be distributed or local, and the processor device 1420 may be distributed or unitary. The memory 1430 may be implemented as electrical, magnetic, or optical memory, or any combination thereof, or other type of storage device. Furthermore, the term "memory" should be interpreted broadly enough to encompass any information that can be read from or written to an address within an addressable space accessed by the processor device 1420. Under this definition, because the processor device 1420 can obtain information from a network, information on a network accessible via the network interface 1425 is still within the memory 1430. Note that each distributed processor comprising the processor device 1420 generally includes its own addressable memory space. Note also that part or all of the computer system 1410 may be incorporated into an application-specific or general-purpose integrated circuit. The optional display 1440 is any type of display suitable for interaction with a human user of the apparatus 1400. Typically, the display 1440 is a computer monitor or other similar display. conclusion
[0084] The description of various embodiments of the present teachings is presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been selected to best explain the principles of the embodiments, practical applications, or technical improvements to technology found in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0085] While the above describes what is considered to be the best mode and / or alternative examples, it will be understood that various modifications may be made herein, that the subject matter disclosed herein may be embodied in various forms and examples, and that the teachings may be applied to many applications, only some of which are described herein. It is intended that the following claims claim all such applications, modifications, and variations as fall within the true scope of the present teachings.
[0086] The components, steps, features, objects, benefits, and advantages discussed herein are merely exemplary. None of them, nor the discussion related thereto, are intended to limit the scope of protection. While various advantages have been discussed herein, it will be understood that not all embodiments necessarily include all advantages. Unless otherwise specified, all measurements, values, ratings, positions, dimensions, sizes, and other specifications set forth in this specification, including the following claims, are approximate and not exact. They are intended to have a reasonable range consistent with the functions to which they relate and those customary in the technical field to which they pertain.
[0087] Many other embodiments are also contemplated, including embodiments having fewer, additional, and / or different components, steps, features, objects, benefits, and advantages, including embodiments in which the components and / or steps are arranged and / or ordered differently.
[0088] Aspects of the present disclosure are described herein with reference to call flow diagrams and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each step in the flowchart diagrams and / or block diagrams, and combinations of blocks in the call flow diagrams and / or block diagrams, can be implemented by computer-readable program instructions.
[0089] These computer-readable program instructions may be provided to a computer, a processor of a special purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed via the computer's processor or other programmable data processing apparatus, create means for performing the functions / acts specified in one or more blocks of the call flow processes and / or block diagrams. These computer-readable program instructions may also be stored in a computer-readable storage medium that can instruct a computer, programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having the instructions stored thereon comprises a product containing instructions that implement aspects of the functions / acts specified in one or more blocks of the call flow diagrams and / or block diagrams.
[0090] The computer-readable program instructions may also be loaded into a computer, other programmable data processing apparatus, or other device and cause the computer, other programmable apparatus, or other device to perform a series of operational steps to generate a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the call flow process and / or block diagram.
[0091] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a call flow process or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions shown in the blocks may occur out of the order shown in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functionality involved. It should also be noted that each block of the block diagrams and / or call flow diagrams, and combinations of blocks in the block diagrams and / or call flow diagrams, can be implemented by a special-purpose hardware-based system that performs the specified functions or operations or executes a combination of special-purpose hardware and computer instructions.
[0092] While the above is described in conjunction with exemplary embodiments, it is understood that the term "example" is intended as an example only, not as best or optimal. Except as immediately stated, nothing described or illustrated is intended to, or should be construed to, convey to the public any elements, steps, features, objects, benefits, advantages, or equivalents, whether claimed or not.
[0093] It will be understood that the terms and expressions used herein have the ordinary meanings ascribed to such terms and expressions with respect to the corresponding respective fields of inquiry and study, unless a specific meaning is otherwise stated herein. Relationship terms such as first and second may be used solely to distinguish one entity or action from another, without necessarily requiring or implying any actual relationship or order between such entities or actions. The terms "comprises," "comprising," or any other variation thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus comprising a list of elements may include not only those elements, but also other elements not expressly listed or other elements inherent to such process, method, article, or apparatus. The use of "a" or "an" preceding an element does not, without further constraints, exclude the presence of additional identical elements in a process, method, article, or apparatus comprising the element.
[0094] The Abstract of the Disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Moreover, in the foregoing Detailed Description, it may be seen that various features are grouped together in various embodiments for the purpose of streamlining the disclosure. This method of disclosure is not to be interpreted as reflecting an intention that the claimed embodiments have more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive subject matter lies in less than all features of a single disclosed embodiment. Accordingly, the following claims are hereby incorporated into the Detailed Description, with each claim standing on its own as separately claimed subject matter.
Claims
1. a first matrix having a resistive processing unit (RPU) crossbar array including a first set of hidden weights configured for gradient updates of a stochastic gradient descent (SGD) of a deep neural network (DNN); a second matrix having a second set of hidden weights of the DNN stored in a digital medium; a third matrix having a set of reference values stored in the digital medium, wherein the set of reference values is calculated during a transfer cycle of the first set of weights from the first matrix to the second matrix, taking into account a code change (chopper); and a fourth matrix having an RPU crossbar array that stores a third set of weights for the DNN that is updated from the second matrix when a threshold value for the second set of weights is reached; 1. A device comprising:
2. a fifth matrix stored on the digital medium and configured to calculate a next set of reference values from values read from the first matrix during a chopper cycle, the fifth matrix configured to partially update the third matrix after the chopper cycle is completed. The device of claim 1 further comprising:
3. The device of claim 1 , wherein the second set of weights takes into account a previous set of reference values from a previous iteration of the transfer cycle.
4. a fifth matrix used to calculate a next set of reference values to be used in a next chopper cycle based on readings from the first matrix stored in the digital medium; The device of claim 1 further comprising:
5. The device of claim 4 , wherein the device is configured to assign the set of reference values to the previous set of reference values in the digital medium upon chopper switching.
6. The device of claim 5 , wherein the device is configured to reset the set of reference values to zero when the chopper switches.
7. The device of claim 6 , wherein the device is configured to switch the sign of the chopper at a switching time of the chopper.
8. The device of claim 1 , wherein an RPU crossbar array is not configured to store the set of reference values.
9. The device of claim 1 , wherein the device is configured to copy a set of previous reference values into a recent read weight vector.
10. performing a stochastic gradient descent (SGD) gradient update of the deep neural network (DNN) using a first set of hidden weights stored in a first matrix comprising a resistive processing unit (RPU) crossbar array; storing, in a digital medium, a second matrix including a second set of hidden weights for the DNN; calculating a third matrix containing a set of reference values during a transfer cycle of the first set of hidden weights from the first matrix to the second matrix, taking into account a sign change (chopper); storing the third matrix on the digital medium; and updating a third set of weights of the DNN from the second matrix when a threshold value for the second set of weights is reached in a fourth matrix comprising an RPU crossbar array; 1. A computer-implemented method comprising:
11. calculating a next set of reference values from values read from said first matrix during a chopper cycle; and storing in the digital medium a next set of reference values in a fifth matrix, wherein the fifth matrix is configured to partially update the third matrix after the chopper cycle is completed. The method of claim 10 further comprising:
12. The method of claim 10 , wherein the second set of weights takes into account a previous set of reference values from a previous iteration of the transfer cycle.
13. calculating a fifth matrix containing a set of previous reference values for the SGD; and storing said fifth matrix on said digital medium. The method of claim 10 further comprising:
14. assigning said set of reference values to said previous set of reference values in said digital medium upon switching of said chopper; The method of claim 13 further comprising:
15. resetting the set of reference values to zero when the chopper switches The method of claim 14 further comprising:
16. switching the sign of the chopper at the time of the switching The method of claim 15 further comprising:
17. The method of claim 11 , wherein an RPU crossbar array is not configured to store the set of reference values.
18. copying the set of previous reference values into the most recent read weight vector; The method of claim 11 further comprising:
19. A non-transitory computer-readable storage medium tangibly embodied with computer-readable program code having computer-readable instructions for solving a machine learning task, the instructions, when executed, causing a computing device to: performing a stochastic gradient descent (SGD) gradient update of the deep neural network (DNN) using a first set of hidden weights stored in a first matrix comprising a resistive processing unit (RPU) crossbar array; storing, in a digital medium, a second matrix including a second set of hidden weights; calculating a third matrix containing a set of reference values during a transfer cycle of the first set of weights from the first matrix to the second matrix, taking into account a sign change (chopper); storing the third matrix on the digital medium; and updating a third set of weights of the DNN from the second matrix when a threshold value for the second set of weights is reached in a fourth matrix comprising an RPU crossbar array; A non-transitory computer-readable storage medium for performing a method comprising:
20. 20. The non-transitory computer-readable storage medium of claim 19, wherein the second set of weights takes into account a previous set of reference values from a previous iteration of the transfer cycle.