Zero-reference DNN training algorithm with dynamic calculation

By using two weight matrices and choppers to calculate the reference value in the RPU cross-switch array, the problem of degradation in the calculation efficiency and accuracy when training deep neural networks in the prior art is solved, and more efficient and accurate calculations are achieved.

CN120019387APending Publication Date: 2025-05-16INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074057.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-20
Filing Date
2023-10-19
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Prior art When training deep neural networks, it is easy to cause computational efficiency and accuracy to be reduced due to incorrect estimation of symmetric points and write noise, especially when using resistive processing unit (RPU) systems.

Method used

By using two weight matrices in the RPU cross-switch array, one stored in the digital medium and the other stored in the RPU array, the reference value is calculated in the weight transfer cycle using choppers, improving computational efficiency and accuracy.

Benefits of technology

It realizes improving the efficiency and accuracy of data calculation in the RPU system, reducing the impact of noise, and improving the overall performance of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019387A_ABST
    Figure CN120019387A_ABST
Patent Text Reader

Abstract

A computer-implemented method includes performing a gradient update of stochastic gradient descent (SGD) of a deep neural network (DNN) using a first set of hidden weights stored in a first matrix including a resistive processing unit (RPU) crossbar array. A second matrix including a second set of hidden weights is stored in the digital medium. A third matrix including a set of reference values is calculated during a transfer cycle of the first set of weights from the first matrix to the second matrix in consideration of a symbol change (chopper). The third matrix is stored in a digital medium. When a threshold is reached for the second set of weights, in a fourth matrix including the RPU crossbar array, a third set of weights of the DNN is updated from the second matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to deep learning, and more particularly, to systems and methods for training deep neural networks using hardware elements. Background Art

[0002] A deep neural network (DNN) can be embodied in an analog cross-point array of a resistive device such as a resistive processing unit (RPU). An RPU device typically includes a first terminal, a second terminal, and an active region. The conductance state of the active region identifies a weight value of the RPU, which can be updated / adjusted by applying a signal to the first / second terminal.

[0003] DNN-based models have been used for a variety of different cognitive-based tasks, such as object and speech recognition and natural language processing. DNN training is remarkable in providing a high level of accuracy when performing such tasks. Training large DNNs is a computationally intensive task. The most popular methods for DNN training, such as backpropagation and stochastic gradient descent (SGD), involve the RPU being "symmetric" to work accurately. Typical systems assume that symmetric points are correctly estimated and initially stored to a reference device array. Symmetric points may be estimated incorrectly, and may also be written incorrectly, including by noise. Summary of the invention

[0004] According to an embodiment of the present disclosure, a computer-implemented method includes: performing gradient updates of stochastic gradient descent (SGD) of a deep neural network (DNN) using a first set of hidden weights stored in a first matrix including a resistor processing unit (RPU) crossbar switch array. A second matrix including a second set of hidden weights is stored in a digital medium. Considering a sign change (chopper), a third matrix including a set of reference values ​​is calculated during a transfer cycle of the first set of weights from the first matrix to the second matrix. The third matrix is ​​stored in a digital medium. When a threshold is reached for the second set of weights, a third set of weights of the DNN is updated from the second matrix in a fourth matrix including the RPU crossbar switch array. The device has a technical effect of improving the efficiency and accuracy of system calculations of data used in the RPU system.

[0005] In one embodiment, which may be combined with the previous embodiment, the second set of weights takes into account a set of previous reference values ​​from a previous iteration of the transfer loop. This allows for more efficient computational power.

[0006] In one embodiment that may be combined with the preceding embodiment, a fifth matrix stored in the digital medium is configured to calculate the next set of reference values ​​from the values ​​read from the first matrix during the chopper cycle. The fifth matrix is ​​configured to partially update the third matrix after the chopper cycle is completed. This allows for greater accuracy in data manipulation.

[0007] In one embodiment which may be combined with the preceding embodiment, the calculation of the SGD includes a fifth matrix containing a set of previous reference values, and storing the fifth matrix in a digital medium. This allows for more efficient computing power.

[0008] In one embodiment which may be combined with the preceding embodiment, the assignment of the set of reference values ​​to the set of previous reference values ​​in the digital medium occurs at a chopper switching time. This allows for a more accurate calculation capability.

[0009] In one embodiment, which may be combined with the previous embodiments, resetting the set of reference values ​​to zero occurs at the chopper switching time. This allows for more efficient computing power.

[0010] In one embodiment which may be combined with the before-previous embodiment, the device is configured to switch the sign of the chopper at the chopper switching time. This enables a higher accuracy of the data manipulation.

[0011] In one embodiment, which may be combined with the previous embodiment, the RPU crossbar array is not used to store the set of reference values. This enables more efficient use of space in the IC array.

[0012] In one embodiment, which may be combined with the previous embodiment, a set of previous reference values ​​is set as the most recently read weight vector. This enables more efficient use of space in the IC array.

[0013] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium tangibly embodies a computer-readable program code, the computer-readable program code having computer-readable instructions for solving a machine learning task, the instructions causing a computer device to perform a method when executed. The method includes performing a gradient update of a stochastic gradient descent (SGD) of a deep neural network (DNN) using a first set of weights stored in a first matrix including a resistor processing unit (RPU) crossbar array. A second matrix including a second set of weights is stored in a digital medium. Considering a chopper, a third matrix including a set of reference values ​​is calculated for SGD during a transfer cycle of the first set of weights from the first matrix to the second matrix. The third matrix is ​​stored in a digital medium. When a threshold is reached for the second set of weights, a third set of weights of the DNN is updated from the second matrix in a fourth matrix including the RPU crossbar array. The device has a technical effect of improving the efficiency and accuracy of system calculations of data used in an RPU system.

[0014] According to an embodiment of the present disclosure, a device including a first matrix includes a resistor processing unit (RPU) crossbar switch array having a first set of weights configured for gradient updates of stochastic gradient descent (SGD) of a deep neural network (DNN). The device includes a second matrix, which includes a second set of weights stored in a digital medium. In addition, the device includes a third matrix, which includes a set of reference values ​​calculated for SGD stored in a digital medium, wherein the set of reference values ​​is calculated during a transfer cycle of the first set of weights from the first matrix to the second matrix taking into account a chopper. The device may also include a fourth matrix, which includes an RPU crossbar switch array, which stores a third set of weights of the DNN, and when a threshold is reached for the second set of weights, the third set of weights is updated from the second matrix. The device has a technical effect of improving the efficiency and accuracy of system calculations of data used in the RPU system.

[0015] In one embodiment, which may be combined with the previous embodiment, the second set of weights takes into account a set of previous reference values ​​from a previous iteration of the transfer loop. This allows for more efficient computational power.

[0016] In one embodiment, which may be combined with the previous embodiments, the set of reference values ​​takes into account the switching frequency. This allows for a higher accuracy in data manipulation.

[0017] In one embodiment which may be combined with the previous embodiment, a fifth matrix comprising a set of previous reference values ​​calculated for SGD is stored in a digital medium. This allows for more efficient computing power.

[0018] In one embodiment which may be combined with the before-addressed embodiment, the device assigns the set of reference values ​​to the set of previous reference values ​​in the digital medium at the chopper switching time. This allows for more efficient computing power.

[0019] In one embodiment which may be combined with the before-previous embodiments, the device resets the set of reference values ​​to zero at the chopper switching time. This allows for more efficient computing power.

[0020] In one embodiment which may be combined with the previous embodiments, the device switches the sign of the chopper at the chopper switching time. This allows for a higher accuracy in data manipulation.

[0021] In one embodiment, which may be combined with the previous embodiment, the RPU crossbar array is not used to store the set of reference values. This enables more efficient use of space in the IC array.

[0022] In one embodiment, which may be combined with the previous embodiment, a set of previous reference values ​​is set as the most recently read weight vector. This enables more efficient use of space in the IC array.

[0023] The technology described herein can be implemented in a variety of ways. Example embodiments are provided below with reference to the following drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings are illustrative embodiments. They do not show all embodiments. Other embodiments may be used in addition or alternatively. Details that may be obvious or unnecessary may be omitted to save space or for more effective description. Some embodiments may be practiced with additional components or steps and / or without all components or steps shown. When the same number appears in different drawings, it refers to the same or similar components or steps.

[0025] Figure 1 is a schematic diagram showing a DNN having a weight matrix W, an A matrix and a latent matrix H;

[0026] Figure 2 is a diagram illustrating a DNN embodied in an analog cross-point array of an RPU device according to an embodiment;

[0027] Figure 3 is a process flow illustrating an example method for training a DNN according to an embodiment;

[0028] FIG. 4A to FIG. 4B is a diagram showing an interconnected array with digital memories for estimating reference values ​​on the fly;

[0029] Figure 5 is a diagram illustrating a forward cycle y=Wx performed according to an embodiment;

[0030] Figure 6 is a diagram showing the reverse cycle z=W performed according to an embodiment T δ graph;

[0031] Figure 7 is a diagram showing updating of an array A using x propagated in a forward loop and δ propagated in a reverse loop according to an embodiment;

[0032] Figure 8 is a diagram showing the forward loop y'=Ae performed on the weight matrix according to an embodiment i 's picture;

[0033] Fig. 9 is a diagram showing the updating of the hidden matrix H using the values ​​calculated in the forward loop of the A matrix;

[0034] Fig.10 is a schematic diagram of the hidden matrix H 902 being selectively applied back to the weight matrix W 1010 according to an embodiment;

[0035] Fig.11is a diagram illustrating example one-hot encoded vectors according to an embodiment;

[0036] Fig.12 is a diagram showing an example detailed algorithm according to an embodiment;

[0037] Fig.13 is a diagram showing an exemplary detailed sub-algorithm according to an embodiment;

[0038] Fig.14 is a diagram illustrating example devices that may be employed in performing one or more of the present techniques according to an embodiment. DETAILED DESCRIPTION

[0039] Overview

[0040] In the following detailed description, many specific details are set forth by way of example in order to provide a thorough understanding of the relevant teachings. However, it should be apparent that the present teachings can be practiced without these details. In other cases, well-known methods, processes, components and / or circuits have been described at a relatively high level without detailed description to avoid unnecessarily obscuring aspects of the present teachings.

[0041] A DNN training technique with an asymmetric RPU device is provided herein. The DNN is trained by using two tunable resistor device arrays and two or three digital memory arrays. The method may include using an RPU crossbar array to represent the weights of the DNN. An additional crossbar array for each weight may be used to calculate gradient updates without requiring a third tunable RPU array for storing references. In addition, updates to the two RPU arrays may occur according to the algorithm described herein.

[0042] In a typical system, the symmetry point of each device may be estimated incorrectly. The symmetry point is the conductance where the conductance change response to a single pulse update in the positive direction is, on average, the same as in the negative direction. The symmetry point may be incorrectly written to a noisy reference device, causing an erroneous value to be subtracted during gradient value readout. The update device may be variable, causing its symmetry point to be unstable and move over time. Additionally, often the inputs are too sparse or the number of devices is too large, causing the symmetry point to be reached only slowly and the transient offset to remain. Additionally, adding an array of dedicated reference devices is expensive in integrated circuit chip area. An embodiment overcomes these limitations by using a digital memory for storing metrics that are used dynamically on the fly to estimate the reference.

[0043] Therefore, one or more methods discussed herein can eliminate the need for time-consuming data processing by the user. This can have the technical effect of reducing the computing resources used by one or more devices within the system. Examples of such computing resources include, but are not limited to, processor cycles, network traffic, memory usage, storage space, and power consumption.

[0044] It should be understood that various aspects of the teachings herein are beyond the capabilities of the human mind. It should also be understood that various embodiments of the subject disclosure described herein may include information that is not possible to be manually obtained by an entity such as a human user. For example, the voltage input and conductance storage values ​​discussed herein are not possible for a human user to perform.

[0045] Now turning to the attached figure, Figure 1 1 is a schematic diagram showing a DNN 100 having a weight matrix W 102, an A matrix 112, and μ past matrix 113 and hidden matrix H 114. Using A matrix 112, μ past The matrix 113 and the hidden matrix 114 iteratively train the weight matrix W 102, as Figure 1 As indicated by the direction of the arrows shown in . As emphasized above, the weight matrix W 102 can be embodied in an analog cross-point array of the RPU. See, e.g. Figure 2 The schematic diagram shown in .

[0046] like Figure 2 As shown, each parameter (weight w ij ) is mapped to a single RPU device in hardware (RPU ij ), that is, a physical cross-point array 104 of the RPU device. The cross-point array 104 includes a series of conductive row lines 106 and a series of conductive column lines 108, which are oriented to be orthogonal to the conductive row lines 106 and intersect with the conductive row lines 106. The intersections between the row lines 106 and the column lines 108 are separated by RPUs 110, thereby forming the cross-point array 104 of the RPU device. Each RPU 110 may include a first terminal, a second terminal, and an active area. The conduction state of the active area identifies the weight value of the RPU 110, which can be updated / adjusted by applying a signal to the first terminal / second terminal. In addition, by controlling additional terminals, a three-terminal (or even more terminal) device can be effectively used as a two-terminal resistive memory device.

[0047] Each RPU 110 (RPU ij) is uniquely identified based on its location (i.e., row i and column j) in the cross-point array 104. For example, working from the top to the bottom and from the left to the right of the cross-point array 104, the RPU at the intersection of the first row line 106 and the first column line 108 is designated as RPU 11 , the RPU at the intersection of the first row line 106 and the second column line 108 is designated as RPU 12 , and so on. In addition, the mapping of the parameters of the weight matrix 102 to the RPUs of the cross-point array 104 follows the same convention. For example, the weight wi1 of the weight matrix 102 is mapped to the RPU of the cross-point array 104 i1 , the weight w of the weight matrix 102 i2 RPU mapped to the cross point array 104 i2 ,etc.

[0048] The RPUs 110 of the cross point array 104 actually serve as weighted connections between neurons in the DNN. The conduction state (e.g., resistance) of the RPU 110 can be changed by controlling the voltage applied between the row lines 106 and the column lines 108, respectively. Data is stored by changing the conduction state of the RPU. The conduction state of the RPU 110 is read by applying a voltage and measuring the current through the target RPU 110. All operations involving weights are performed by the RPU 110 completely in parallel.

[0049] In machine learning and cognitive science, DNN-based models are a family of statistical learning models inspired by biological neural networks in animals (especially the brain). These models can be used to estimate or approximate systems and cognitive functions that depend on many inputs and weights of connections that are usually unknown. DNNs are typically embodied as so-called "neuromorphic" systems of interconnected processor elements that act as simulated "neurons" that exchange "messages" between each other in the form of electronic signals. The connections in DNNs that carry electronic messages between simulated neurons are provided with digital weights that correspond to the strength of a given connection. These digital weights can be adjusted and tuned based on experience so that the DNN adapts to the input and is able to learn. For example, a DNN for handwriting recognition is defined by a set of input neurons that can be activated by pixels of an input image. After being weighted and transformed by a function determined by the network designer, the activation of these input neurons is then passed to other downstream neurons. The process is repeated until the output neuron is activated. The activated output neuron determines which character is read.

[0050] By updating the weight value W via the A matrix 112 ij And then the resulting outputs from the A matrix 112 are summed into the hidden matrix 114 until the element of the hidden matrix 114 (i.e., H ij) reaches the threshold to train Figure 1 DNN 100 is shown, as explained in detail below. However, before and after the weight values ​​are updated in A matrix 112, a chopper 116 multiplies the input and output signals by a chopper value. The chopper value at a given time is equal to positive 1 (+1) or negative 1 (-1). Chopper 116 flips between chopper values ​​randomly or periodically so that for a portion of the training cycle, updates are applied to A matrix 114 with opposite signs. This sign flipping of chopper 116 means that any "bias" that A matrix 112 contributes to the weight values ​​has one sign (i.e., positive or negative) during some cycles of training time and another sign (i.e., negative or positive) during other cycles of training time. Chopping cycles or switching probabilities can also be assigned by the user. Bias can be inherent in any analog system, including non-ideal RPUs that can be used in DNN 100.

[0051] For training purposes, such an ideal device perfectly implements the DNN training process of backpropagation and stochastic gradient descent (SGD). Backpropagation is a training process performed in three cycles: forward cycle, backward cycle, and weight update cycle, which is repeated many times until the convergence criteria are met. Stochastic gradient descent (SGD) uses backpropagation to calculate each parameter (weight w ij ) is the error gradient.

[0052] To perform backpropagation, the DNN-based model includes multiple processing layers that learn data representations with multiple levels of abstraction. For a single processing layer where N input neurons are connected to M output neurons, the forward loop involves computing a vector-matrix multiplication (y=Wx), where the vector x of length N represents the activity of the input neurons and the matrix W of size M×N stores the weight values ​​between each pair of input and output neurons. The resulting vector y of length M is further processed by performing a nonlinear activation on each of the resistive memory elements and then passed to the next layer.

[0053] Once the information reaches the final output layer, the backward loop involves computing the error signal and propagating it back through the DNN. The backward loop on a single layer also involves a vector-matrix multiplication on the transpose (interchanging each row and corresponding column) of the weight matrix (z=W T δ), where the vector δ of length M represents the error computed by the output neuron and the vector z of length N is further processed using the derivatives of the neuron’s nonlinearity before being passed down to the previous layers.

[0054] Finally, in the weight update loop, the weight matrix W is updated by performing the outer product of the two vectors used in the forward and backward loops. This outer product of two vectors is usually expressed as W←W+η(δx T ), where η is the global learning rate.

[0055] All operations performed on the weight matrix W during this back-propagation process can be implemented using the cross-point array 104 of the RPU 110 having a corresponding number of M rows and N columns, where the conductance values ​​stored in the cross-point array 104 form the matrix W. In the forward loop, the input vector x is sent as a voltage pulse through each column line 108, and the resulting vector y is read as a current output from the row line 106. Similarly, when a voltage pulse is provided from the row line 106 as an input to the backward loop, then the conductance values ​​stored in the weight matrix W are stored in the cross-point array 104. T The vector-matrix product is calculated on the transpose of . Finally, in the update cycle, voltage pulses representing vectors x and δ are provided simultaneously from the column line 108 and the row line 106. In this configuration, each RPU 110 performs local multiplication and summation operations by processing voltage pulses from the corresponding column line 108 and row line 106, thereby achieving incremental weight updates.

[0056] Symmetric RPUs can perfectly implement backpropagation and SGD. That is, for such an ideal RPU w ij ←w ij +ηΔw ij , where wij is the weight value of the i-th row and j-th column of the cross-point array 104.

[0057] Figure 3 is a diagram illustrating an example method 300 for training a DNN according to an embodiment. During training, weight updates are first accumulated on the A matrix. The A matrix is ​​a hardware component consisting of rows and columns of RPUs with symmetric behavior around zero. Then, weight updates from the A matrix are selectively moved to the weight matrix W. The weight matrix W is also a hardware component consisting of rows and columns of RPUs. The training process iteratively determines a set of parameters (weights wij) that maximize the accuracy of the DNN. The matrix W is initialized to randomly distributed values ​​using common practices applied to DNN training. The digitally stored hidden matrix H is initialized to zero.

[0058] During training, weight updates are performed on the A matrix. The information processed by the matrix is ​​then accumulated in the hidden matrix H (a separate matrix that effectively performs low-pass filtering). The values ​​of the hidden matrix H that reach the update threshold are then applied to the weight matrix W. The update threshold effectively minimizes the noise generated within the hardware of the A matrix. However, for elements of the matrix initialized with a bias, the update threshold will be reached prematurely because each iteration from the element carries a consistent update (positive or negative) based on the bias rather than on the weight updates associated with training the DNN. The chopper value counteracts the bias by flipping the sign of the bias during certain time periods during which the bias is summed to the hidden matrix H with the opposite sign. Specifically, certain time periods will sum the weight value plus the positive bias to the hidden matrix H, while other time periods will sum the weight value plus the negative bias to the hidden matrix H. The random flipping of the chopper value means that time periods with positive bias tend to be flush with time periods with negative bias. Therefore, hardware bias and noise associated with non-ideal RPUs are tolerated (or absorbed by the H matrix) and thus give less test error compared to standard SGD techniques, a separate hidden matrix H, or other training techniques using asymmetric devices, even with a smaller number of states.

[0059] The method 300 initializes the A matrix, the digital calculated values ​​μ, the hidden matrix H (also stored in the digital buffer), and the weight matrix W in block 302. Initializing the A matrix includes, for example, setting all values ​​to zero. The array A may be embodied in an interconnected array.

[0060] FIG. 4A to FIG. 4B 300 is a diagram showing an interconnected array with digital storage for estimating reference values ​​on the fly. As shown, μ represents the most recent past of the gradient update matrix A. In some embodiments, the most recent past μ can be used for difference calculations in a digital storage device or memory to produce a value ω for updating H. This creates a floating point representation of the reference value. In this case, the reference value changes over time according to method 300. This dynamic updating and on-the-fly calculation of the reference value helps eliminate bias in previous systems that use a hardware reference RPU matrix for the reference value.

[0061] Initialization of the latent matrix H includes zeroing the current values ​​stored in the matrix or allocating digital storage space on the connected computing device. Initialization of the weight matrix W includes loading the weight matrix W with random values ​​so that the training process of the weight matrix W can begin. ω is assigned based on reading from the A matrix of each column or row, where ω is the digital converted value processed after using the ADC.

[0062] The number H is the hidden matrix used to filter the calculated gradient values ​​onto A. ω is the readout of the analog A matrix, which can be read out by placing a unit vector (e.g. [1 0 0 0]) with voltages at each column or row. The weight of that column in current units is retrieved, which is turned back into a number by using the ADC. λ is the scaling factor or learning rate. S is the varying chopper value used to switch between negative and positive.

[0063] Once a threshold is met on H, a pulse is sent to the weight matrix W. In other words, the gradient is placed on the crossbar RPU. As placed, the gradient includes a lot of noise. The gradient is read again, a chopper is applied and a reference value is subtracted to remove any bias, then added to the filter matrix, filtering out the noise. The gradient is then integrated over time, and once the gradient reaches a threshold, the weights are updated. Therefore, the weights W are only rarely modified, without any bias applied. This greatly improves the noise properties and accuracy of the prior art RPU algorithm.

[0064] The method 300 includes determining activation values ​​by performing a forward loop using the weight matrix W (block 304 ). Figure 5 is a diagram showing a forward loop being executed according to an embodiment. The forward loop involves computing a vector-matrix multiplication (y=Wx), where the activation values ​​embodied as input vector x represent the activity of the input neurons, and the weight matrix W stores the weight values ​​between each pair of input and output neurons. Figure 5 The vector-matrix multiplication operation of the forward loop is shown to be implemented in the cross-point array 502 of the RPU device, where the conductance values ​​stored in the cross-point array 502 form a matrix.

[0065] Input vector x is transmitted as a voltage pulse through each of the conductive column lines 512, and the resulting output vector y is read as a current output from the conductive row lines 510 of the cross point array 502. Analog to digital converter (ADC) 513 is used to convert the analog output vector 516 from the cross point array 502 into a digital signal.

[0066] The method 300 also includes determining an error value by performing a back loop on the weight matrix W (block 306 ). Figure 6 is a diagram showing a backward loop being performed according to an embodiment. In general, the backward loop involves calculating an error value δ and backpropagating the error value δ through the weight matrix W (i.e., z=W) via a vector-matrix multiplication on the transpose of the weight matrix W. T δ, where W T indicates the transpose of the matrix W), where the vector δ represents the error computed by the output neuron, and the vector z is further processed using the derivatives of the neuron’s nonlinearity before being passed down to the previous layers.

[0067] Figure 6The vector-matrix multiplication operation of the backward loop is illustrated in the cross-point array 502. The error value δ is transmitted as a voltage pulse through each of the conductive row lines 510, and the resulting output vector z is read as a current output from the conductive column lines 512 of the cross-point array 502. When a voltage pulse is provided from the row line 510 as an input to the backward loop, the vector-matrix product is calculated on the transpose of the weight matrix W. Figure 6 As shown in , ADC 513 is used to convert the (analog) output vector 518 from the cross-point array 502 into a digital signal.

[0068] The method 300 also includes applying a chopper value to the activation value or error value (block 308). The chopper value may be generated by a chopper (e.g., from Figure 1 The chopper 116 is applied to the A matrix 502, which is included for each row line and each column line in the A matrix 502. In some embodiments, the cross-point array 502 may have choppers only on the column lines 506 or only on the row lines 504. After applying the chopper values ​​to the activation values ​​or the error values, the method 300 also includes updating the A matrix with the activation values, the error values ​​(the input vectors x and δ), and the chopper values ​​(block 310).

[0069] Figure 7 is a diagram showing updating of array A 502 using x propagating in a forward loop and δ propagating in a reverse loop according to an embodiment. Each row and each column has a chopper value 550 applied to the corresponding line. The sign of the chopper value 550 is represented as "+" for a positive chopper value (i.e., the activation value or error value has not changed), or as "X" for a negative chopper value (i.e., the sign of the activation value or error value has changed). Updates are implemented in the cross-point array 502 by transmitting voltage pulses, which represent vector x (from the forward loop) and vector δ (from the reverse loop) provided simultaneously from the conductive column line 506 and the conductive row line 504, respectively. In this configuration, each RPU in the cross-point array 502 performs local multiplication and summation operations by processing voltage pulses from the corresponding conductive column line 506 and conductive row line 504, thereby implementing incremental weight updates. The forward loop (block 304), the backward loop (block 306), and updating the A matrix with the input vectors from the forward and backward loops (block 310) may be repeated multiple times to improve the updated value of the A matrix.

[0070] The method 300 also includes using the input vector e i (ie, y'=Aei) and the chopper value perform a forward loop over the A matrix to read the i-th column of A (block 312). At each time step, a new input vector e i , and the sub-index i represents the time index. As will be described in detail below, according to an example embodiment, the input vector e iis a one hot encoded vector. For example, as is known in the art, a one hot encoded vector is a set of bits that has only those combinations containing a single high (1) bit and all other bits being low (0). To use a simple non-limiting example for illustrative purposes, assuming a matrix of size 4×4, the one hot encoded vector will be one of the following vectors: [1 0 0 0], [0 1 0 0], [0 0 1 0], and [0 0 0 1]. At each time step, a new one hot encoded vector is used, and the sub-index i represents that time index. However, it is worth noting that this article also contemplates the use of a method for selecting the input vector e i other methods.

[0071] Figure 8 is a diagram showing that according to an embodiment, by performing a forward loop on the A matrix with chopper values ​​y' = Ae i To read the graph of the i-th column of A, where e i is the i-th unit vector. Alternatively, we can perform Y'=A T e i (Transposed read). Input vector e i The resulting output vector y' is sent as a voltage pulse through each conductive column line 506 and read as a current output from the conductive row lines 504 of the cross-point array 502. Each column line 506 and row line 504 is read with the same chopper value (i.e., positive or negative) that updates the A matrix. For example, the first column line 506 i1 exist Figure 7 and Figure 8 The second column has a positive chopper value (+), and the wiring 506 i2 exist Figure 7 and Figure 8 has a negative chopper value (X), and the first row wiring 504 1i exist Figure 7 and Figure 8 , with a negative chopper value (X). When a voltage pulse is provided from column line 506 as input to this forward cycle, a vector-matrix product is calculated. Method 300 includes updating the implicit matrix H (block 314).

[0072] Fig. 9 904. The hidden matrix H 902 is a digital matrix rather than a physical device like the A matrix and the weight matrix W, which stores the values ​​of each RPU in the A matrix (i.e., the value at the A position). ij The H value 906 (i.e., H ij). When the forward loop is executed, an output vector y'eiT, or ω, is produced. This output vector is used to calculate other digital matrices as detailed below, and is also used to update the hidden matrix H. Therefore, each time the output vector is read, the hidden matrix H902 changes. For those RPUs with low noise levels, the H value 906 will grow consistently. For constant gradients and inputs, the growth in value can be in the positive or negative direction, depending on the value of the output vector ω. If the output vector ω includes significant noise, its value may be positive for one iteration and negative for another iteration. This combination of positive output vector ω values ​​and negative output vector ω values ​​means that the H value 906 will grow more slowly and less consistently.

[0073] The hidden matrix values ​​can be updated on the fly using a digital storage device that stores and updates the value of μ as follows:

[0074] For each digital calculation of the transfer cycle (for each element i of a readout vector k), perform:

[0075] 1. μ ik ← (1 - γ)μ ik + γω i

[0076] 2. h ik ← h ik + s k λ′(ω i - μ ik past ),

[0077] where ω is the readout weight vector, h ik is the digital buffer value, s k is the current chopper symbol, λ is the learning rate, and μ is a floating point reference that changes over time and in various iterations. K can be varied with each n s The update increases due to the wraparound on M.

[0078] Each time a vector k is read, after digital computation, a buffer (with thresholds) can be written into the weight matrix W. γ is a user-defined parameter and is either positive or zero, and is typically set to 2 / p, where p is the switching frequency, assuming regular switching.

[0079] As the H value 906 increases, the method 300 includes tracking whether the H value 906 has increased by more than a threshold value (block 316). If the H value 906 at a particular location (i.e., H ij) is not greater than the threshold (block 316 “No”), the method 300 repeats from executing the forward loop (block 304) to updating the hidden matrix H (block 314) and potentially flipping the chopper values ​​(blocks 320-322). If the H value 906 is greater than the threshold (block 316 “Yes”), the method 300 continues to convert the input vector e i is sent to the weight matrix W, but only for that particular RPU (block 318). As described above, the increase in the H value 906 can be in a positive or negative direction, so the threshold is also a positive or negative value. Fig.10 is a schematic diagram of the hidden matrix H 902 being selectively applied back to the weight matrix W 1010 according to an embodiment.

[0080] Fig.10 A first H value 1012 and a second H value 1014 are shown that have reached above the threshold and are sent to the weight matrix W 1010. The first H value 1012 reaches the positive threshold and therefore carries a positive one for its row in the input vector 1016: "1". The second H value 1014 reaches the negative threshold and therefore carries a negative one for its row in the input vector 1016: "-1". The remaining rows in the input vector 1016 carry zeros because those values ​​(i.e., H values ​​906) have not yet become greater than the threshold. The threshold can be much larger than the value added to the hidden matrix H. For example, the threshold can be ten or a hundred times the expected strength of the update value for each cycle. Because no bias is added to the H matrix due to the reference values ​​calculated on the fly, the threshold does not usually need to be too large. A higher threshold reduces the frequency of updates performed on the weight matrix W. However, the filtering function performed by the H matrix reduces the error of the objective function of the neural network. These updates can only be generated after processing many data examples, so the confidence level in the updates is also increased. This technique enables training of neural networks with noisy RPU devices that have only a limited number of states even with shifted or unstable symmetry points.After applying the H values ​​to the weight matrix W, the H values ​​906 are reset to zero and iterations of the method 300 continue.

[0081] The method 300 also includes flipping the sign of the chopper value at a flip percentage (block 320). In some embodiments, the chopper value is flipped only after the chopper product is added to the hidden matrix H. That is, the chopper value is used twice: once when the activation value and the error value are written to the A matrix; and once when the forward cycle is read from the A matrix. The chopper value should not be flipped before updating the H matrix. The flip percentage can be defined as a user preference so that after each chopper product is added to the hidden matrix H, the chopper has a percentage chance of flipping the chopper value. For example, the user preference can be fifty percent, so that half of the time, after calculating the chopper product, the chopper value has a chance to change sign (i.e., positive to negative or negative to positive). In other embodiments, the chopper can be flipped once every three or four times throughout the cycle, for example.

[0082] When the chopper is determined to be flipped (yes, block 320), then the digital buffer value is further updated for the instantaneous reference estimate. For example, the following updates may occur:

[0083] μ ik past ← μ ik, , μ past is updated with the current μ value in the i-th row and k-th column.

[0084] μ ik ←0, reset μ ik The value of .

[0085] s k ← -s, the chopper value is flipped.

[0086] When μ past ik During reset it is set to the last ω i When γ=1 is used (e.g., omitting the running mean with γ=1), the memory space usage can be reduced by a factor of 2.

[0087] After flipping the chopper value and updating μ past Thereafter, the method 300 continues by determining whether the training is complete. If the training is not complete, e.g., a certain convergence criterion is not met (block 324 "No"), the method 300 begins to repeat again by performing a forward loop y = Wx. For example, by way of example only, the training may be considered complete when no improvement in the error signal is seen. When the training is complete (block 324 "Yes"), the method 300 ends.

[0088] As emphasized above, according to an example embodiment, the input vector e i is a one-hot encoded vector, which is a set of bits that has only those combinations with a single high (1) bit and all other bits low (0). See e.g. Fig.11 .like Fig.11 As shown, given a matrix of size 4×4, the one-hot encoded vector will be one of the following vectors: [1 0 0 0], [0 1 0 0], [0 0 1 0], and [00 0 1]. At each time step, a new one-hot encoded vector is used, represented by the sub-index i at that time index.

[0089] Fig.12 is a diagram showing an example detailed algorithm according to an embodiment of the present disclosure. Fig.13 is a diagram illustrating an exemplary detailed sub-algorithm according to an embodiment of the present disclosure.

[0090] The present invention may be a system, method and / or computer program product. The computer program product may include (one or more) computer-readable storage media having computer-readable program instructions thereon for causing a processor to perform various aspects of the present disclosure.

[0091] A computer-readable storage medium may be a tangible device that can retain and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device (such as a punched card or a raised structure in a groove on which instructions are recorded), and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0092] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network). The network may include copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium in the corresponding computing / processing device.

[0093] The computer-readable program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" programming language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, electronic circuits including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions to personalize the electronic circuits by utilizing the state information of the computer-readable program instructions so as to perform various aspects of the present disclosure.

[0094] Various aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each block of the flowchart illustration and / or block diagram and the combination of blocks in the flowchart illustration and / or block diagram can be implemented by computer-readable program instructions.

[0095] These computer-readable program instructions may be provided to a processor of a specially configured computer or other programmable data processing device to produce a machine, such that instructions executed via the processor of the computer or other programmable data processing device create components for implementing the functions / actions specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may instruct a computer, a programmable data processing device, and / or other device to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0096] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operational steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0097] Now turn to Fig.14 , a block diagram of an apparatus 1400 for implementing one or more methods presented herein is shown. By way of example only, the apparatus 1400 may be configured to control input voltage pulses applied to an array and / or process output signals from the array.

[0098] The apparatus 1400 includes a computer system 1410 and removable media 1450. The computer system 1410 includes a processor device 1420, a network interface 1425, a memory 1430, a media interface 1435, and an optional display 1440. The network interface 1425 allows the computer system 1410 to connect to a network, while the media interface 1435 allows the computer system 1410 to interact with a medium such as a hard drive or removable media 1450.

[0099] Processor device 1420 can be configured to implement the methods, steps and functions disclosed herein. Memory 1430 can be distributed or local, and processor device 1420 can be distributed or single. Memory 1430 can be implemented as electrical, magnetic or optical memory, or any combination of these or other types of storage devices. In addition, the term "memory" should be interpreted broadly enough to cover any information that can be read or written from an address in an addressable space accessed by processor device 1420. With this definition, information on a network accessible by network interface 1425 is still in memory 1430 because processor device 1420 can retrieve information from the network. It should be noted that each distributed processor constituting processor device 1420 typically contains its own addressable memory space. It should also be noted that some or all of computer system 1410 may be incorporated into a dedicated or general purpose integrated circuit. Optional display 1440 is any type of display suitable for interacting with a human user of device 1400. Typically, display 1440 is a computer monitor or other similar display.

[0100] in conclusion

[0101] The description of various embodiments of the present teachings has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to technologies found in the marketplace, or to enable those of ordinary skill in the art to understand the embodiments disclosed herein.

[0102] Although the foregoing has described what is considered to be the best state and / or other examples, it should be understood that various modifications may be made therein, and the subject matter disclosed herein may be implemented in various forms and examples, and the teachings may be applied to many applications, only some of which are described herein. The appended claims are intended to claim any and all applications, modifications, and variations that fall within the true scope of the present teachings.

[0103] The components, steps, features, objects, benefits and advantages discussed herein are merely illustrative. None of them and the discussion related to them is intended to limit the scope of protection. Although various advantages have been discussed herein, it should be understood that not all embodiments necessarily include all advantages. Unless otherwise stated, all measurements, values, ratings, positions, sizes, dimensions and other specifications set forth in this specification (including in the appended claims) are approximate and not precise. They are intended to have a reasonable range consistent with the functions to which they are related and the practices in the field to which they belong.

[0104] Many other embodiments are also contemplated. These include embodiments with fewer, additional and / or different components, steps, features, objects, benefits and advantages. These also include embodiments in which components and / or steps are arranged and / or ordered differently.

[0105] Various aspects of the present disclosure are described herein with reference to call flow diagrams and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each step of the flowchart diagrams and / or block diagrams and combinations of blocks in the call flow diagram diagrams and / or block diagrams can be implemented by computer-readable program instructions.

[0106] These computer-readable program instructions may be provided to a processor of a computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create components for implementing the functions / actions specified in one or more blocks of the call flow and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium, which may instruct a computer, a programmable data processing apparatus, and / or other device to function in a particular manner, such that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the call flow and / or block diagram.

[0107] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the call flow process and / or block diagram.

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate possible implementations of the system, method, and computer program product according to various embodiments of the present disclosure, The architecture, functions, and operations. In this regard, each box in the call flow or block diagram may represent a module, segment, or portion of an instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions marked in the box may not occur in the order marked in the figure. For example, two boxes shown in succession may actually be executed substantially simultaneously, or the boxes may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each box in the block diagram and / or call flow diagram and the combination of boxes in the block diagram and / or call flow diagram can be implemented by a dedicated hardware-based system that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.

[0109] Although the foregoing has been described in conjunction with example embodiments, it should be understood that the term "example" is meant only as an example, not the best or optimal. Except as stated immediately above, nothing stated or shown is intended or should be construed as causing any component, step, feature, object, benefit, advantage, or equivalent to be dedicated to the public, whether or not recited in the claims.

[0110] It should be understood that the terms and expressions used herein have the ordinary meaning consistent with these terms and expressions with respect to the corresponding fields of investigation and research corresponding thereto, unless a specific meaning is otherwise set forth herein. Relational terms such as first and second etc. may be used only to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between these entities or actions. The terms "comprises", "comprising" or any other variation thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a list of elements includes not only those elements, but may also include other elements that are not explicitly listed or inherent to such process, method, article or device. In the absence of further constraints, an element beginning with "one" or "an" does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0111] An abstract of the present disclosure is provided to allow the reader to quickly determine the nature of the present technical disclosure. It should be understood that the abstract will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing specific embodiments, it can be seen that various features are grouped together in various embodiments for the purpose of simplifying the present disclosure. The method of the present disclosure should not be interpreted as reflecting the intention that the claimed embodiments have more features than the features explicitly recited in each claim. On the contrary, as reflected in the following claims, the subject matter of the invention lies in less than all the features of a single disclosed embodiment. Therefore, the attached claims are hereby incorporated into the specific embodiments, with each claim independently serving as a separately claimed subject matter.

Claims

1. A device comprising: a first matrix including a resistor processing unit (RPU) crossbar array having a first set of hidden weights configured for gradient updates of stochastic gradient descent (SGD) of a deep neural network (DNN); a second matrix comprising a second set of hidden weights of the DNN stored in the digital medium; a third matrix comprising a set of reference values ​​stored in the digital medium, wherein the set of reference values ​​is calculated during a transfer cycle of the first set of weights from the first matrix to the second matrix taking into account a sign change (chopper); as well as A fourth matrix includes an RPU crossbar switch array, wherein the RPU crossbar switch array stores a third set of weights of the DNN, and when a threshold is reached for the second set of weights, the third set of weights is updated from the second matrix.

2. The device according to claim 1, further comprising: A fifth matrix is ​​stored in the digital medium, the fifth matrix being configured to calculate a next set of reference values ​​based on values ​​read from the first matrix during a chopper cycle, and the fifth matrix being configured to partially update the third matrix after the chopper cycle is completed.

3. The device according to claim 1, wherein: The second set of weights takes into account a set of previous reference values ​​from a previous iteration of the transfer loop.

4. The apparatus according to claim 1, further comprising: A fifth matrix for calculating a next set of reference values ​​to be used in a next chopper cycle based on readings from the first matrix stored in the digital medium.

5. The device according to claim 4, wherein: The device is configured to assign the set of reference values ​​to the set of previous reference values ​​in the digital medium at a chopper switching time.

6. The device according to claim 5, wherein: The device is configured to set a reference value to zero at the chopper switching time.

7. The device according to claim 6, wherein: The device is configured to switch the sign of the chopper at the chopper switching time.

8. The device according to claim 1, wherein: No RPU crossbar array is configured to store the set of reference values.

9. The device according to claim 1, wherein: The device is configured to copy a set of previous reference values ​​to a most recently readout weight vector.

10. A computer-implemented method comprising: performing a gradient update of a stochastic gradient descent (SGD) of a deep neural network (DNN) using a first set of hidden weights stored in a first matrix comprising a resistive processing unit (RPU) crossbar array; storing in a digital medium a second matrix comprising a second set of hidden weights for the DNN; During the transfer cycle of the first set of hidden weights from the first matrix to the second matrix, a third matrix including a set of reference values ​​is calculated taking into account a sign change (chopper); storing the third matrix in the digital medium; as well as When a threshold is reached for the second set of weights, a third set of weights of the DNN is updated from the second matrix in a fourth matrix comprising an RPU crossbar array.

11. The method according to claim 10, further comprising: During a chopper cycle, calculating a next set of reference values ​​based on the values ​​read from said first matrix; as well as A next set of reference values ​​is stored in a fifth matrix in the digital medium, wherein the fifth matrix is ​​configured to partially update the third matrix after the chopper cycle is completed.

12. The method according to claim 10, wherein: The second set of weights takes into account a set of previous reference values ​​from a previous iteration of the transfer loop.

13. The method according to claim 10, further comprising: a fifth matrix comprising a set of prior reference values ​​for said SGD calculation; as well as The fifth matrix is ​​stored in the digital medium.

14. The method according to claim 13, further comprising: The set of reference values ​​is assigned to the set of previous reference values ​​in the digital medium at a switching time of the chopper.

15. The method according to claim 14, further comprising: The set of reference values ​​is reset to zero at the chopper switching time.

16. The method according to claim 15, further comprising: The sign of the chopper is switched at the switching time.

17. The method according to claim 11, wherein: No RPU crossbar array is configured to store the set of reference values.

18. The method according to claim 11, further comprising: Copies a set of previous reference values ​​to the most recently read weight vector.

19. A non-transitory computer-readable storage medium tangibly embodying a computer-readable program code having computer-readable instructions for solving a machine learning task, the instructions when executed causing a computer device to perform a method comprising: performing a gradient update of a stochastic gradient descent (SGD) of a deep neural network (DNN) using a first set of hidden weights stored in a first matrix comprising a resistive processing unit (RPU) crossbar array; storing in a digital medium a second matrix including a second set of hidden weights; calculating a third matrix comprising a set of reference values ​​taking into account sign changes (chopper) during a transfer cycle of the first set of weights from the first matrix to the second matrix; storing the third matrix in the digital medium; as well as When a threshold is reached for the second set of weights, a third set of weights of the DNN is updated from the second matrix in a fourth matrix comprising an RPU crossbar array.

20. The non-transitory computer-readable storage medium of claim 19, wherein: The second set of weights takes into account a set of previous reference values ​​from a previous iteration of the transfer loop.