Weight Iteration on the RPU Crossbar Array
By transforming the weight matrix from a rectangular to a nearly square configuration in an RPU array, the method enhances the signal strength of backward pass signals, addressing noise and accuracy issues in ANN training, and improving overall neural network performance.
Patent Information
- Application Number
- JP2023522777
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-02
- Filing Date
- 2021-10-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-10-28
AI Technical Summary
Existing artificial neural network (ANN) training methods using analog resistive processing unit (RPU) arrays face challenges due to analog noise and limited weight range caused by bounded analog-to-digital converter (ADC) resolutions, leading to weak backward pass signals and reduced training accuracy.
The method involves storing weight values in an array of RPU devices as resistance values, defining a weight matrix with an output dimension smaller than the input dimension, and transforming it from a rectangular to a nearly square configuration by repeating or concatenating weight elements, thereby enhancing the signal strength of backward pass signals.
This approach significantly increases the signal strength of backward pass signals, reduces the requirements for high-resolution ADCs, improves power efficiency, and enhances neural network training performance by averaging noise and increasing accuracy.
Smart Images

Figure 0007692477000028 
Figure 0007692477000029 
Figure 0007692477000030
Abstract
Description
Technical Field
[0001] The present invention generally relates to an artificial neural network (ANN) having an analog cross-point array of resistive processing unit (RPU) devices, and more particularly to enhancing signal strength by weight iteration on an RPU cross-point array.
Background Art
[0002] Machine learning is used to broadly represent the main function of an electronic system that learns from data. In machine learning and cognitive science, an ANN is a family of statistical learning models inspired by biological neural circuits and especially the brain. An ANN depends on many inputs and can generally be used to estimate or approximate unknown systems and functions. An ANN is often embodied as a so-called "neuromorphic" system of interconnected processor elements that function as simulated "neurons" and exchange messages with each other in the form of electrical signals. Similar to the so-called "plasticity" of synaptic neurotransmitter connections that carry messages between biological neurons, the connections in an ANN that carry electrical messages between simulated neurons are provided with numerical weights corresponding to the strength or weakness of a given connection. The weights can be adjusted and tuned based on experience, making the ANN adaptable and capable of learning with respect to inputs.
Summary of the Invention
[0003] According to an embodiment, a method for artificial neural network (ANN) training is provided. The method includes storing weight values in an array of resistive processing unit (RPU) devices, where the array of RPU devices represents a weight matrix W of an m-by-n ANN as weight values of the weight matrix W are stored in the array as resistance values of the RPU devices; defining the weight matrix W to have an output dimension smaller than an input dimension such that the weight matrix W has a rectangular configuration; during a forward cycle pass, copying an input of a repeated weight element, summing calculation results output from the repeated weight element that produces one output per row, and updating each of the repeated weight elements according to a backpropagated error, or alternatively, during an update pass, updating only one of the repeated weight elements by setting all forward values except from one to zero, and converting the weight matrix W from a rectangular configuration to a nearly square configuration by repeating or concatenating the rectangular configuration of the weight matrix W to enhance a signal strength of a backward path signal.
[0004] A non-transitory computer-readable storage medium including a computer-readable program for training an artificial neural network (ANN) is presented. When the computer-readable program is executed on a computer, it stores weight values in an array of resistive processing unit (RPU) devices, where the array of RPU devices stores the weight values of the weight matrix W of an m-by-n ANN as the resistance values of the RPU devices, thereby representing the weight matrix W; defines the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration; during the forward cycle path, copies the input of the repeated weight elements, sums the calculated results output from the repeated weight elements that produce one output per row, and updates each of the repeated weight elements according to the backpropagated error, or alternatively, during the update path, updates only one of the repeated weight elements by setting all forward values except from one to zero, and transforms the weight matrix W from a rectangular configuration to an approximately square configuration by repeating or concatenating the rectangular configuration of the weight matrix W to enhance the signal strength of the backward path signal, and causes the computer to perform the steps.
[0005] A system for training an artificial neural network (ANN) is presented. The system includes an array of resistive processing unit (RPU) devices for storing weight values, where the array of RPU devices stores the weight values of the weight matrix W of an m-by-n ANN as the resistance values of the RPU devices, thereby representing the weight matrix W, and a processor for controlling the voltage between the RPU devices in the array, where the processor defines the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration, and transforms the weight matrix W from a rectangular configuration to an approximately square configuration by repeating or concatenating the rectangular configuration of the weight matrix W to enhance the signal strength of the backward path signal.
[0006] Note that the exemplary embodiments are described with reference to various subjects. In particular, some embodiments are described with reference to method-type claims, while on the other hand, other embodiments have been described with reference to apparatus-type claims. Nevertheless, those skilled in the art will presume from the above and following descriptions that any combination of features belonging to one type of subject, as well as any combination between features regarding different subjects, especially between the features of method-type claims and the features of apparatus-type claims, is considered to be described in this document, unless otherwise notified.
[0007] These and other features and advantages will become apparent from the following detailed description of its exemplary embodiments, which is to be read in conjunction with the accompanying drawings.
[0008] The present invention provides details in the following description of preferred embodiments with reference to the following figures.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
[0010] Throughout the drawings, the same or similar reference numerals represent the same or similar elements.
[0011] Exemplary embodiments according to the present invention are provided for enhancing the signal strength of backward pass signals by weight replication on a resistive processing unit (RPU) crossbar array. In particular, the weight matrix is modified by repeating or replicating columns or rows or both of the RPU crossbar array. This is achieved by repeating some of the weight elements on the resistive device array and accumulating the results of the repeated weight elements (i.e., rows or columns or both) in the digital periphery to enhance the output signal strength during the backward pass cycle, which increases accuracy and reduces the requirements for a high-resolution analog-to-digital converter (ADC). As a result, using a lower-precision ADC helps improve the power efficiency of the analog hardware chip, which provides better noise and border management and improves neural network training performance.
[0012] A crossbar array, also known as a cross-point array or cross-wire array, is a high-density, low-cost circuit architecture used to form various electronic circuits and devices, including ANN architectures, neuromorphic microchips, and ultra-high density non-volatile memories. A basic crossbar array configuration includes a set of conductive row wires and a set of conductive column wires formed to intersect the set of conductive row wires. The intersections between the two sets of wires are separated by so-called cross-point devices, which can be formed from thin film materials.
[0013] Cross-point devices effectively function as the weighted connections of an ANN between neurons. To emulate highly energy-efficient synaptic plasticity, nanoscale two-terminal devices, such as memristors with "ideal" conductive state switching characteristics, are often used as cross-point devices. The conductive state (e.g., resistance) of an ideal memristor material can be changed by controlling the voltage applied between individual ones of the row and column wires. Digital data can be stored by changing the conductive state of the memristor material at the intersection to achieve a high-conductive state or a low-conductive state. The memristor material can further be programmed to maintain two or more distinct conductive states by selectively setting the conductive state of the material. The conductive state of the memristor material can be read by applying a voltage between the materials and measuring the current passing through the target cross-point device.
[0014] However, ANN training by analog resistive crossbar arrays such as analog RPU arrays can be difficult due to analog noise. Furthermore, the training process is limited by the bounded ranges of the analog-to-digital converters (ADCs) and digital-to-analog converters (DACs) used for the RPU array. The ADCs and DACs are used to convert digital inputs to the RPU to analog signals and the outputs from the RPU to digital signals, respectively. Analog noise can be reduced by a noise management approach that involves increasing the output signal for the rectangular weight matrix.
[0015] Exemplary embodiments of the present invention disclose a method and system for advantageously managing analog noise by repeatedly encoding the weights of a deep neural network (DNN) network on one physical analog crossbar array and increasing the output signal of the backward pass signal. Exemplary embodiments of the present invention further disclose a method and system for using a rectangular weight matrix n times to make the rectangular weight matrix more square and thus better suited for a physical square crossbar array. Exemplary embodiments of the present invention repeatedly encode the weights of a DNN (e.g., convolutional neural network (CNN)) network layer on a single analog crossbar array and then sum (average) or spread (copy) the outputs of the repeated rows or columns or both in the digital periphery to maintain the correct DNN / CNN network architecture. Exemplary embodiments of the present invention repeatedly encode the weights of a DNN / CNN network layer on a single analog crossbar array and then maintain the correct DNN / CNN network architecture and sum (average) or spread (copy) the outputs of the repeated rows or columns or both in the digital periphery to update just one (e.g., randomly selected) of the repeated weights or all of the repeated weights simultaneously or in parallel.
[0016] Accordingly, the exemplary method and system provide improved neural network training performance. In particular, the weak output signal of the backward pass signal is significantly increased, which reduces the requirements for an accurate ADC. Moreover, the noise of the analog elements is averaged by repetition, resulting in improved training accuracy.
[0017] The present invention is described with respect to a given exemplary architecture, but it should be understood that other architectures, structures, substrate materials, as well as process features and steps / blocks can be varied within the scope of the present invention. Note that not all specific features can be shown in all figures for clarity. This is not intended to be construed as a limitation on the scope of any particular embodiment, illustration, or claim.
[0018] Various exemplary embodiments of the present invention are described below. For clarity, not all features of an actual implementation are described herein. In the development of any such actual implementation, it will be appreciated that numerous implementation-specific decisions must be made to achieve the developer's specific goals, such as compliance with system-related and business-related constraints that will vary from one implementation to another. Moreover, such development efforts can be complex and time-consuming, but will nevertheless be routine work for those skilled in the art having the benefit of the present invention.
[0019] FIG. 1 is a diagram showing an artificial neural network (ANN) embodied in an analog cross-point array of a resistive processing unit (RPU) device according to an embodiment of the present invention.
[0020] As shown in FIG. 1, each parameter (weight w ij ) of the algorithmic (abstract) weight matrix 10 is a single RPU device (RPU) on hardware ij) That is, it is mapped to the physical cross-point array 12 of the RPU device. The cross-point array 12 has a series of conductive row wires 14 and a series of conductive column wires 16 that are oriented perpendicular to and intersect the conductive row wires 14. The intersections of the conductive row wires 14 and the column wires 16 are separated by RPU devices 18, forming the cross-point array 12 of the RPU device. Each RPU device 18 can include a first terminal, a second terminal, and an active region. The conductive state of the active region identifies the weight value of the RPU device 18, and the weight value can be updated / regulated by applying a signal to the first / second terminals. Further, a three-terminal (or in some cases, more terminals) device can effectively function as a two-terminal resistive memory device by controlling the additional terminals.
[0021] The m-by-n matrix W is typically mapped to an RPU array having M columns and N rows. Thus, the integration in the figure occurs along the columns of the RPU array, while the summation occurs mathematically along the rows of W. Therefore, the mathematical mapping of W to the RPU array is actually transposed. As a result, for ease of display, the mathematical rows of the matrix W are shown as the columns of the RPU array. For example, moving from the top to the bottom and from left to right of the cross-point array 12, the RPU device 18 at the intersection of the first conductive row wire 14 and the first conductive column wire 16 is the RPU 11 as represented, and the RPU device 18 at the intersection of the first conductive row wire 14 and the second conductive column wire 16 is the RPU 12 as represented, and so on. Usually, although the convention is to swap the columns and rows of the RPU array for display, the mapping of the weight parameters in the weight matrix 10 to the RPU devices 18 in the cross-point array 12 follows a similar convention. For example, the weight w i1 of the weight matrix 10 is mapped to the RPU 1i of the cross-point array 12, and the weight w i2 of the weight matrix 10 is mapped to the RPU 2iis mapped to, etc.
[0022] The RPU device 18 of the cross-point array 12 functions as a weighted connection between neurons in the ANN. The resistance of the RPU device 18 can be changed by controlling the voltage applied between the individual conductive row wires 14 and the conductive column wires 16. Changing the resistance is, for example, related to how data is stored in the RPU device 18 based on a high-resistance state or a low-resistance state. The resistance state of the RPU device 18 is read by applying a voltage and measuring the current passing through the target RPU device 18. All operations with weights are performed completely in parallel by the RPU device 18.
[0023] In machine learning and cognitive science, ANN-based models are a family of statistical learning models inspired by the biological neural circuits of animals and especially the brain. These models can be used to estimate or approximate systems and cognitive functions that depend on many inputs and generally unknown connection weights. The ANN is often embodied as a so-called "neuro-morphological" system of interconnected processor elements that function as simulated "neurons" that exchange "messages" with each other in the form of electronic signals (Figure 8). The connections in the ANN that carry the electronic messages between the simulated neurons are provided with numerical weights corresponding to the strength or weakness of a given connection (Figure 8). These numerical weights can be adjusted and tuned based on experience, making the ANN adaptable and capable of learning with respect to the input. For example, an ANN for handwritten recognition is defined by a set of input neurons that can be activated by the pixels of the input image. After being weighted and transformed by a function determined by the network designer, the activation of these input neurons is then transmitted to other downstream neurons. This process is repeated until the output neuron is activated. The activated output neuron determines which character has been read.
[0024] ANN can be trained in an incremental or stochastic gradient descent (SGD) process, in which the error gradient for each parameter (weight w ij ) is calculated using backpropagation. Backpropagation is carried out in three cycles: a forward cycle, a backward cycle, and a weight update cycle, and these cycles are repeated multiple times until the convergence criterion is met. A DNN-based model includes multiple processing layers that learn the representation of data at multiple levels of abstraction. For a single processing layer where N input neurons are connected to M output neurons, the forward cycle involves calculating a vector-matrix multiplication (y = Wx), where the vector x of length N represents the activities of the input neurons, and the matrix W of size M×N stores the weight values between each pair of input and output neurons. The resulting vector y of length M is further processed by performing a non-linear activation for each of the resistive memory elements and then passed to the next layer.
[0025] When the information reaches the final output layer, the backward pass cycle involves calculating the error signal and backpropagating the error signal through the ANN. The backward pass cycle for a single layer involves a vector-matrix multiplication with respect to the transpose (swapping each row with the corresponding column) of the weight matrix (z = W T δ), where the vector δ of length M represents the error calculated by the output neurons, and the vector z of length N is further processed using the derivative of the neuron non-linearity and then passed to the previous layer.
[0026] Finally, in the weight update cycle, the weight matrix W is updated by performing the outer product of two vectors used in the forward and backward pass cycles. This outer product of the two vectors is often expressed as W←W + η(δx T ), where η is the global learning rate.
[0027] During this backpropagation process, all the operations performed on the weight matrix W can be performed in the cross-point array 12 of the RPU device 18 having the corresponding number of m rows and n columns, where the conductance values stored in the cross-point array 12 form the matrix W. In the forward cycle, the input vector x is transmitted as a voltage pulse through each of the conductive column wires 16, and the resulting vector y is read as a current output from the conductive row wires 14. Similarly, when a voltage pulse is supplied from the conductive row wires 14 as an input to the backward-pass cycle, a vector-matrix product is calculated for the transpose of the weight matrix W T . Finally, in the update cycle, voltage pulses representing the vectors x and δ are supplied simultaneously from the conductive column wires 16 and the conductive row wires 14. Thus, each RPU device 18 performs local multiplication and summation operations by processing the voltage pulses coming from the corresponding conductive column wires 16 and conductive row wires 14, and thus achieves an incremental weight update.
[0028] The resistance values of the RPU devices are limited to a bounded range having limits and a limited and finite state resolution that limit the weight range that can be used for ANN training. Further, the operations performed on the RPU array are inherently analog and thus are prone to various noise sources. When the input values to the RPU array are small (as for the backward pass), the output signal y may be buried in noise and thus produce incorrect results. In the training phase, ANN training may involve an SGD process with backpropagation.
[0029] CNN training is carried out using batches. Therefore, a batch of input data that will be used for training is selected. Using the input map and the convolutional kernel, an output map is generated. The generation of the output map is usually called the "forward pass". Further, the method includes using the output map to determine how close or far the predicted character recognition is to the CNN. The degree of error for each of the matrices including the CNN is determined using, for example, gradient descent. The determination of the relative error is called the "backward pass". The method further includes modifying or updating the matrices in order to adjust the error. Adjusting the convolutional kernel based on the output error information and using this to determine the modification to each neural network matrix is called the "update pass".
[0030] FIG. 2 shows an exemplary analog vector matrix multiplication on an RPU array according to an embodiment of the present invention.
[0031] The analog vector matrix multiplication 100 is accompanied by a set of digital input values (δ) 110, where each of the digital input values (δ) 110 is represented by a respective analog signal pulse width 120. The analog signal pulse width 120 is provided as an input to the array, and the generated current signal is connected to the inverting input of the operational amplifier 131 and to a capacitor (C int ) 132 with the operational amplifier 131 to the input of the operational amplifier (op-amp) integrated circuit 130 having the operational amplifier 131. The non-inverting input of the operational amplifier 131 is connected to ground. The output of the operational amplifier 131 is also connected to the input of an analog-to-digital converter (ADC) 140. The ADC 140 outputs a signal y l representing the (digitized) result of the analog vector matrix multiplication 100 on the RPU array.
[0032] During the full integration time, analog noise is integrated in the operational amplifier 131. When the input value (δ) 110 becomes very small (e.g., as for the backward path), the output signal is buried by the noise integrated during the cycle (SNR~0), producing an incorrect result.
[0033] The actual pulse duration is much shorter than the full integration time, but the ADC 140 waits for the entire cycle to evaluate the analog output from the operational amplifier 131. The analog noise is desirably reduced. FIGS. 3-6 present methods and systems for managing noise by providing symmetric signal strengths between the forward path signal and the backward path signal.
[0034] FIG. 3 is a diagram showing an exemplary rectangular RPU array in the forward path in which a rectangular sub-region of the RPU array is used, according to an embodiment of the present invention.
[0035] Each circle 202 represents a separate digital input x to the RPU hardware system 200. For example, in the forward cycle path, the digital input x (i.e., 202) is provided to the m rows of the matrix W. The digital input 202, when received by the RPU array 200, is represented as the digital RPU input x' (i.e., 204). The digital RPU input 204 is fed into the noise / boundary management unit or component 210. The vector-matrix multiplication performed on the RPU array 225 is essentially analog and thus prone to various noise sources. Therefore, the noise / boundary management unit or component 210 performs a noise reduction operation. The digital-to-analog converter (labeled "DA converter 212") provides the digital RPU input x' (i.e., 204) as an input to the RPU array 225, such as the analog pulse width 215. The RPU array 225 includes a first region 230 and a second region 235. The first region 230 is the rectangular region being used, while the second region 235 is the unused region. The term "being used" means that the RPU is loaded with the conductance corresponding to its weight. The (analog) output 240 from the RPU array 230 is converted by an analog-to-digital converter (labeled "AD converter 250") into a vector of the digital RPU output y' (i.e., 260). The digital RPU output 260 is fed into another noise / boundary management unit or component 270. Further, the result of the vector-matrix multiplication is an analog voltage, and thus the result is bounded by the signal limits imposed by the circuit. Therefore, the noise / boundary management unit or component 270 performs a noise reduction operation to ensure that the result at the output of the RPU array 230 is always within an acceptable voltage amplitude range.
[0036] As a result, the output capacitor of the resistive crossbar device is of finite size (resulting in a finite output boundary b), and an analog output signal close to zero is set to zero due to the finite ADC resolution. Thus, when the analog output of the RPU array 230 is very small, the digital output may all be zero. This effect is not welcome when the ADC resolution is small (e.g., when the output boundary b remains unchanged and the ADC bin size increases). This effect is particularly unwelcome when the weighted matrix encoded on the RPU array 230 is not square (as shown in FIG. 3). Then, on average, the forward and backward directions have very different average signal strengths. For example, in a 10-class classification network, the last fully connected layer is typically on the order of 1000×10 in size. Thus, on average, the signal in the backward direction is at least [Number] times less. When the backward signal is very small (e.g., smaller than the minimum ADC bin size), the error is set to zero and learning fails. For a symmetric RPU (e.g., the same hardware specs for the forward and backward directions such as ADC resolution and output boundary), this effect is not desirable. FIG. 4 shows a solution for mitigating such problems.
[0037] FIG. 4 is a diagram showing an exemplary square RPU array in the forward path in which the entire square RPU array according to an embodiment of the present invention is used.
[0038] Elements similar to those in FIG. 3 are not described for clarity. The RPU hardware system 200’ differs from the RPU hardware system 200 (FIG. 3) in that the weight matrix W(225) is replicated k times so that additional weight elements are used to convert or modify the rectangular configuration of the weight matrix W(225) to a substantially or approximately or nearly square configuration. The RPU array (225) is noted to always be physically square. Nevertheless, when the weight matrix is rectangular and assuming the number of columns of the weight matrix matches the input dimension of the RPU array, only a rectangular sub-region of the physically square RPU array is used. As described above, the term “used” means that the RPU is loaded with the conductance corresponding to its weight. Thus, in FIG. 3, the second region 235 (including some rows or columns or both of the RPU) is not used. Nevertheless, in FIG. 4, the weight matrix is sized up by adding more rows or columns or both of the RPU 405 to convert or modify the rectangular configuration of the weight matrix W(225) to a substantially or nearly square configuration. The replicated or repeated rows or columns or both 405 make the rectangular weight matrix 225 more square. The replicated or repeated rows or columns or both are denoted as 280 after the noise / boundary management unit 270, summed (282), and thus produce one output per row. All or just one of the repeated weights 280 is updated. When just one is updated, the repeated weight 290 can be randomly selected or sequentially selected. Thus, in the second embodiment, only all subsets of the repeated weights are updated simultaneously or in parallel.
[0039] The system knows how many rows / columns should be added to achieve an approximately or substantially square configuration by calculating the number of iterations of all weight matrices. By calculating the number of iterations of all weight matrices, the system can simply divide the number of input dimensions N of matrix W (number of columns) by the output dimension m (number of rows) of matrix W and obtain the maximum integer, e.g., r = floor(N / M). This is the number of how many times all rows / columns are repeated. For example, when M = 250 and N = 512, r = 2, and the resulting weight matrix has size N = 512 and M = 250*r = 500. This method only makes the weight matrix approximately square.
[0040] Therefore, to solve the problem described above with respect to FIG. 3, rows / columns 405 are added, providing symmetry to the signal strength for both the forward pass and the backward pass. It is noted that the backward pass signal strength does not have to exactly match the forward pass signal strength. Still, assuming the input dimension is the same as the RPU array size, the backward pass signal strength is maximized. This is because since the physical RPU layout is substantially or approximately square, the system should prevent adding more rows than necessary.
[0041] In particular, assuming a square-sized RPU crossbar array, the exemplary embodiment enables the replication of rows or columns or both to make the rectangular weight matrix W more square. In other words, physically, the RPU array is always square. Still, as detailed in FIG. 3, only the rectangular sub-region of the RPU array is actually used. Therefore, the exemplary embodiment enables the use of more available crosspoints, and thus, by adding repeated columns together, output / input processing is added digitally.
[0042] For example, it is assumed that W has size m×n. In many DNN networks, m << n, e.g., the output dimension is much smaller than the input dimension.
[0043] The exemplary embodiments create or construct a larger matrix of size km×n
Number
Number
[0044] In the forward pass,
Number
Number
[0045] In the backward pass, it is described in detail below with reference to FIGS. 5 and 6.
[0046]
Number
Number
[0047] Accordingly, the corresponding new delta input is copied from the original delta.
[0048] Typically, the number of replications k is chosen such that mk≒n (while not exceeding the physical size limits of the crossbar array). Additionally, k is b / w maxis guaranteed to be smaller, where b is the output boundary (about 12) and w max is the maximum weight (about 0.6). Thus, k is maximized to achieve mk ≒ n, but not greater than b / w max (about 20).
[0049] In the update path, changes (except for possible learning rate adaptation) are also, alternatively, a random fraction of the error [Number] and is not set to zero either. [Number] Figure 5 is a diagram showing an exemplary rectangular RPU array in the backward path in which a rectangular sub-region of the RPU array according to an embodiment of the present invention is used.
[0050] Figure 5 is a diagram showing an exemplary rectangular RPU array in the backward path in which a rectangular sub-region of the RPU array according to an embodiment of the present invention is used.
[0051] In the RPU hardware system 500, the vector-matrix multiplication performed on the RPU array 525 is essentially analog and thus is prone to various noise sources. Therefore, the noise / boundary management unit 510 performs noise reduction operations. The digital-to-analog converter (labeled "DA converter 512") provides a digital RPU input x' (i.e., 560, FIG. 6) as an input to the RPU array 525, such as an analog pulse width 515. The RPU array 525 includes a first region 530 and a second region 535. The first region 530 is the rectangular region being used, while the second region 535 is the unused region. The term "being used" means that the RPU is loaded with a conductance corresponding to its weight. The (analog) output 240 from the RPU array 530 is converted by an analog-to-digital converter (labeled "AD converter 250") into a vector of digital RPU outputs y' (i.e., 260). The digital RPU output 260 is fed into another noise / boundary management unit 270. Further, the result of the vector-matrix multiplication is an analog voltage, and thus the result is bounded by the signal limits imposed by the circuit. Therefore, the noise / boundary management unit 270 performs noise reduction operations to ensure that the result at the output of the RPU array 530 is within an acceptable voltage amplitude range. The digital output 272 is output from the noise / boundary management unit 270.
[0052] As described above, the output capacitor of the resistive crossbar element has a finite size (resulting in a finite output boundary b), and an analog output signal close to zero is set to zero due to the finite ADC resolution. Therefore, when the analog output of the RPU array 530 is very small, the digital outputs may all become zero. This effect is not welcome when the ADC resolution is low (e.g., the ADC bin size increases when the output boundary b remains unchanged). This effect is particularly unwelcome when the weighted matrix encoded on the RPU array 530 is not square (as shown in FIG. 5). Therefore, on average, the forward and backward directions have very different average signal strengths. When the backward signal is very small (e.g., smaller than the minimum ADC bin size), the error is set to zero and learning fails. For a symmetric RPU (e.g., the same hardware specifications for the forward and backward directions such as ADC resolution and output boundary), this effect is undesirable. FIG. 6 shows a solution to mitigate such problems.
[0053] FIG. 6 shows an exemplary rectangular RPU array in the backward path according to an embodiment of the present invention, where the entire square RPU array is used and the output dimension m is half of the input dimension such that each column is repeated exactly once.
[0054] In the backward cycle path, the digital input x (i.e., 550) is provided to the n columns of the matrix W (525). The RPU hardware system 500’ differs from the RPU hardware system 500 (FIG. 5) in that the weight matrix W (525) is replicated k times so that additional weight elements are used to convert or modify the rectangular configuration of the weight matrix W (525) into a substantially or nearly square configuration. It is noted that the weight matrix W (525) is always physically square. Nevertheless, only a rectangular sub-region of the physically square RPU array is being used. As described above, the term “used” means that the RPU is loaded with the conductance corresponding to its weight. Thus, in FIG. 5, a second region 535 (including some rows or columns or both of the RPU) is not being used. Nevertheless, in FIG. 6, the weight matrix is sized up by adding more rows or columns or both of the RPU 605 to convert or modify the rectangular configuration of the weight matrix W (525) into a substantially or nearly square configuration. The replicated or repeated rows or columns or both 605 make the rectangular weight matrix 525 more square. The replicated or repeated rows or columns or both are denoted as 560 in front of the noise / boundary management unit 510.
[0055] FIG. 7 is a block / flow diagram of an exemplary equation 700 used in the forward path and the backward path according to an embodiment of the present invention.
[0056] In conclusion, the exemplary embodiments disclose a method and system for replicating / repeating / concatenating a rectangular weight matrix of a DNN on a physical analog crossbar and copying the input to the repeated weight elements during the forward pass. The exemplary embodiments further disclose a method and system for replicating / repeating / concatenating a rectangular weight matrix of a DNN on a physical analog crossbar and averaging or summing the output calculation results from the repeated weight elements to produce one output per row for the original weight matrix. The exemplary embodiments further disclose a method and system for replicating / repeating / concatenating a rectangular weight matrix of a DNN on a physical analog crossbar and updating each repeated weight according to the backpropagation error. The exemplary embodiments further disclose a method or system for replicating / repeating / concatenating a rectangular weight matrix of a DNN on a physical analog crossbar and updating only one of the repeated weight elements by setting all of the backward deltas (or forward values), excluding zero from one, during the update pass.
[0057] FIG. 8 depicts a block diagram of the components of a system 900 that includes a computing device 905. It should be understood that FIG. 8 provides only an illustration of one implementation and does not imply any limitation as to the environment in which different embodiments may be implemented. Many modifications to the depicted environment may be made.
[0058] Computing device 905 includes a communication fabric 902 that provides communication between computer processor 904, memory 906, persistent storage 908, communication unit 910, and input / output (I / O) interface 912. The communication fabric 902 can be implemented in any architecture designed to transfer data, control information, or both, between processors (such as microprocessors, communication and network processors), system memory, peripheral devices, and any other hardware components within the system. For example, the communication fabric 902 can be implemented with one or more buses.
[0059] Memory 906, cache memory 916, and persistent storage 908 are computer-readable storage media. In this embodiment, memory 906 includes random access memory (RAM) 914. In general, memory 906 can include any suitable volatile or non-volatile computer-readable storage media.
[0060] In some embodiments of the present invention, a deep learning program 925 is included and is operated by a neuromorphic chip 922 as a component of a computing device 905. In other embodiments, the deep learning program 925 is stored in a persistent storage 908 for execution by a neuromorphic chip 922 together with one or more of the respective computer processors 904 via one or more memories of a memory 906. In this embodiment, the persistent storage 908 includes a magnetic hard disk drive. As an alternative to, or in addition to, the magnetic hard disk drive, the persistent storage 908 can include a solid state hard drive, a semiconductor storage device, a read only memory (ROM), an erasable programmable read only memory (EPROM), a flash memory, or any other computer readable storage medium capable of storing program instructions or digital information.
[0061] The medium used by the persistent storage 908 can further be removable. For example, a removable hard drive can be used for the persistent storage 908. Other examples include optical and magnetic disks, thumb drives, and smart cards that are inserted into a drive for transfer to another computer readable storage medium that is also part of the persistent storage 908.
[0062] In some embodiments of the present invention, the neuromorphic chip 922 is included in the computing device 905 and is connected to a communication fabric 902.
[0063] In these examples, the communication unit 910 provides communication with other data processing systems or devices that include resources of the distributed data processing environment. In these examples, the communication unit 910 includes one or more network interface cards. The communication unit 910 can provide communication using either or both physical and wireless communication links. The deep learning program 925 can be downloaded to the persistent storage 908 through the communication unit 910.
[0064] The I / O interface 912 enables the input and output of data with other devices that can be connected to the computing system 900. For example, the I / O interface 912 can provide a connection to an external device 918, such as a keyboard, keypad, touch screen, or some other suitable input device, or a combination thereof. The external device 918 can further include portable computer-readable storage media, such as, for example, a thumb drive, portable optical or magnetic disk, and memory card.
[0065] The display 920 provides a mechanism for displaying data to a user and can be, for example, a computer monitor.
[0066] FIG. 9 is an exemplary block / flow diagram of a method for enhancing signal strength by weight iteration on an RPU crossbar array, according to an embodiment of the present invention.
[0067] In block 1010, replicate or iterate or concatenate the rectangular weight matrix of the DNN on the physical analog crossbar.
[0068] In block 1020, copy the input to the repeated weight elements during the forward pass.
[0069] In block 1030, average or sum the output calculation results from the repeated weight elements that produce one output per row for the original weight matrix.
[0070] In block 1040, update each repeated weight according to the backpropagation error, or alternatively, update only one of the repeated weight elements by setting all of the backward deltas (or forward values), excluding zero from one, during the update pass.
[0071] The terms "data", "content", "information", and similar terms as used herein are used interchangeably to refer to data that can be captured, transmitted, received, displayed, or stored, or a combination thereof, according to various example embodiments. Thus, the use of any such terms should not be understood to limit the spirit and scope of the present disclosure. Further, when a computing device is described herein as receiving data from another computing device, the data can be received directly from the other computing device or indirectly via one or more intermediate computing devices such as, for example, one or more servers, repeaters, routers, network access points, base stations, or the like, or a combination thereof.
[0072] To interact with a user, embodiments of the subject matter described in this specification may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, for the user to provide input to the computer. Similarly, other types of devices may be used to interact with the user, and for example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user may be in any form including acoustic, speech, or tactile input.
[0073] The present invention can be a system, a method, or a computer program product, or a combination thereof. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions for causing a processor to execute aspects of the present invention.
[0074] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction-executing device. A computer-readable storage medium can be, for example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing, but is not limited thereto. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium as used herein should not be construed as being a signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire, which are essentially transient signals.
[0075] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, or a wireless network, or a combination thereof. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. Each computing / processing device's network adapter card or network interface receives the computer-readable program instructions from the network and transfers the computer-readable program instructions for storage on a computer-readable storage medium of each computing / processing device.
[0076] Computer-readable program instructions for carrying out the operations of the present invention can be source code or object code written in any combination of one or more programming languages, including assembly instructions, instruction set architecture (ISA) instructions, machine language instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk(R), C++, or the like, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, or partly on the user's computer as a stand-alone software package, or partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, electronic circuit components including programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to customize the electronic circuit components to implement aspects of the present invention.
[0077] Aspects of the invention will be described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0078] These computer-readable program instructions may be provided to at least one processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks or modules of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, a programmable data processing apparatus, or other devices to function in a particular manner, such that the medium storing the instructions contains an article of manufacture including instructions for implementing the function / act aspects specified in one or more blocks or modules of the flowchart and / or block diagram.
[0079] The computer-readable program instructions may also be loaded onto a computer, other programmable apparatus, or other devices to cause a series of operational blocks / steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-executed process, such that the instructions which execute on the computer, other programmable apparatus, or other devices implement the functions / acts specified in one or more blocks or modules of the flowchart and / or block diagram.
[0080] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, segment, or portion of one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks shown in succession can in fact be executed substantially simultaneously, or the blocks can sometimes be executed in the reverse order depending on the functionality involved. It is further noted that each block of the block diagrams or flowcharts, or combinations of blocks in the block diagrams or flowcharts, or both, can be implemented by a dedicated hardware-based system that performs the specified function or acts, or by a combination of dedicated hardware and computer instructions.
[0081] References in this specification to "one embodiment" or "an embodiment" and other variations thereof mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment of the present principle. Thus, the appearances of the phrases "in one embodiment" or "in an embodiment" and other variations thereof in various places throughout this specification are not necessarily all referring to the same embodiment.
[0082] It should be recognized that any use of the following: " / ", "and / or", and "at least one of", for example, in the cases of "A / B", "A and / or B", and "at least one of A and B", is intended to encompass the selection of only the first listed option (A), or only the second listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to encompass the selection of only the first listed option (A), or only the second listed option (B), or only the third listed option (C), or the selection of only the first and second listed options (A and B), or the selection of only the first and third listed options (A and C), or the selection of only the second and third listed options (B and C), or the selection of all three options (A, B, and C). This can be extended, as will be readily apparent to those skilled in the art, to the same number of listed items.
[0083] Preferred embodiments of a system and method for enhancing signal strength by weight iteration on an RPU cross-point array have been described (which is intended to be illustrative and not limiting), but it is pointed out that modifications and changes can be made by those skilled in the art from the perspective of the above teachings. Therefore, it should be understood that changes may be made in the specific embodiments described within the scope of the present invention as outlined in the appended claims. Thus, while the aspects of the present invention, as particularly required by patent law, have been described in detail, what is claimed and desired to be protected by the patent is shown in the appended claims.
Claims
1. A method for training an artificial neural network (ANN) to be executed by a computer, comprising: storing weight values in an array of resistive processing unit (RPU) devices, wherein the array of RPU devices stores the weight values of a weight matrix W of the ANN having m rows and n columns as the resistance values of the RPU devices in the array, thereby representing the weight matrix W; defining the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration; copying the input of repeated weight elements during the forward cycle path; summing the output calculation results from the repeated weight elements that produce one output per row; and updating each of the repeated weight elements according to the backpropagated error transforming the weight matrix W from the rectangular configuration to a more square configuration by repeating or concatenating the rectangular configuration of the weight matrix W to enhance the signal strength of the backward path signal; and A method comprising:
2. The method according to claim 1, wherein the weight elements are the m rows and n columns of the weight matrix W.
3. The method according to claim 1 or claim 2, wherein the repeated weight elements added to the weight matrix W provide symmetric signal strength between the forward path signal and the backward path signal.
4. The method according to any one of claims 1 to 3, wherein in the forward cycle path, a digital RPU input is fed into a first noise / boundary measurement component and a digital-to-analog converter (DAC) before being received by the array of RPU devices.
5. The method according to claim 4, wherein in the forward cycle path, the output calculation results from the repeated weight elements are summed after being processed by an analog-to-digital converter (ADC).
6. In the forward cycle path, 【Number 1】 and 【Number 2】 wherein [Number 3] is the direct output signal vector from the modified RPU array / weight matrix, 【Number 4】 is the modified weight matrix stored in the RPU array, x is the input signal vector, yi is the i-th element of the original output vector, and k is the number of replications. The method according to claim 5.
7. In a backward cycle path, 【Number 5】 wherein, 【Number 6】 wherein, j = 0, ... , k - 1, 【Number 7】 is the transposed and modified weight matrix stored in the RPU array, 【Number 8】 is the modified error signal vector that is the input to the modified RPU array / weight matrix during the backward path, 【Number 9】 is the (i + jk)-th element of the modified error signal, the method according to any one of claims 1 to 6. **Claim 8** A non-transitory computer-readable storage medium including a computer-readable program for training an artificial neural network (ANN), wherein when the computer-readable program is executed by a computer, storing weight values in an array of resistive processing unit (RPU) devices, wherein the array of RPU devices stores the weight values of the weight matrix W of m rows and n columns of the ANN as the resistance values of the RPU devices in the array, thereby representing the weight matrix W, the storing; defining the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration; copying the input of repeated weight elements during the forward cycle path; summing the calculation results output from the repeated weight elements that produce one output per row, and updating each of the repeated weight elements according to the backpropagated error to enhance the signal strength of the backward path signal, converting the weight matrix W from the rectangular configuration to a more square configuration by repeating or concatenating the rectangular configuration of the weight matrix W; causing the computer to perform the steps of. A non-transitory computer-readable storage medium. **Claim 9** The non-transitory computer-readable storage medium according to claim 8, wherein the weight elements are the m rows and n columns of the weight matrix W. **Claim 10** The non-transitory computer-readable storage medium according to any one of claims 8 to 9, wherein the repeated weight elements added to the weight matrix W provide symmetric signal strength between the forward path signal and the backward path signal. **Claim 11** In the forward cycle path, before the digital RPU input is received by the array of the RPU devices, it is sent to a first noise / boundary measurement component and a digital-to-analog converter (DAC), the non-transitory computer-readable storage medium according to any one of claims 8 to 10.
12. In the forward cycle path, after the output calculation result from the repeated weight elements is processed by an analog-to-digital converter (ADC), it is summed, the non-transitory computer-readable storage medium according to claim 11.
13. In the forward cycle path, 【Number 10】 and 【Number 11】 where 【Number 12】 is a direct output signal vector from the modified RPU array / weight matrix, 【Number 13】 is the modified weight matrix stored in the RPU array, x is the input signal vector, yi is the i-th element of the original output vector, and k is the number of replications, the non-transitory computer-readable storage medium according to claim 12.
14. In the backward cycle path, 【Number 14】 where 【Number 15】 where j = 0,..., k - 1, 【Number 16】 is the transposed modified weight matrix stored in the RPU array, 【Number 17】 is the modified error signal vector which is the input to the modified RPU array / weight matrix during the backward path, 【Number 18】 is the (i + jk)-th element of the modified error signal, the non-transitory computer-readable storage medium according to claim 8.
15. A system for artificial neural network (ANN) training, comprising an array of resistive processing unit (RPU) devices for storing weight values, wherein the weight values of the weight matrix W of the m-by-n ANN are stored in the array as the resistance values of the RPU devices, representing the weight matrix W, the array of the RPU devices, and a processor for controlling the voltage between the RPU devices in the array, defining the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration, and converting the weight matrix W from the rectangular configuration to a more square configuration by repeating or concatenating the rectangular configuration of the weight matrix W to enhance the signal strength of the backward path signal performed by the processor A system comprising
16. wherein the signal strength is during the forward cycle path, copying the input of the repeated weight element, summing the calculated results output from the repeated weight element that produces one output per row, and updating each of the repeated weight elements according to the backpropagated error, or alternatively, updating only one of the repeated weight elements by setting all forward values except from one to zero during the update path The system according to claim 15, which is enhanced by
17. The system according to claim 16, wherein the repeated weight element added to the weight matrix W provides symmetric signal strength between the forward path signal and the backward path signal.
18. In the forward cycle path, the digital RPU input is fed to a first noise / boundary measurement component and a digital-to-analog converter (DAC) before being received by the array of RPU devices. The system according to claim 17.
19. In the forward cycle path, the calculated results output from the repeated weight element are summed after being processed by an analog-to-digital converter (ADC). The system according to claim 18.
20. The system according to claim 16, wherein the one updated repeated weight element is randomly selected.
21. A method for training an artificial neural network (ANN) executed by a computer, comprising storing weight values in an array of resistive processing unit (RPU) devices, wherein the array of RPU devices stores the weight values of the weight matrix W of the m-by-n ANN as the resistance values of the RPU devices, thereby representing the weight matrix W, said storing; defining the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration; during the forward cycle path, copying the input of the repeated weight element, summing the calculated results output from the repeated weight element that produces one output per row, and Updating only one of the repeated weight elements by setting all forward values except from one to zero during the update pass Converting the weight matrix W from a rectangular configuration to a more square configuration by repeating or concatenating the rectangular configuration of the weight matrix W to enhance the signal strength of the backward pass signal A method comprising the above. **Claim 22** The method according to claim 21, wherein the repeated weight element added to the weight matrix W provides symmetric signal strength between the forward pass signal and the backward pass signal. **Claim 23** The method according to claim 21 or claim 22, wherein in the forward cycle path, a digital RPU input is fed into a first noise / boundary measurement component and a digital-to-analog converter (DAC) before being received by an array of the RPU devices. **Claim 24** The method according to claim 23, wherein in the forward cycle path, the output calculation results from the repeated weight elements are totaled after being processed by an analog-to-digital converter (ADC). **Claim 25** A non-transitory computer-readable storage medium including a computer-readable program for training an artificial neural network (ANN), wherein when the computer-readable program is executed by a computer, Storing weight values in an array of resistive processing unit (RPU) devices, wherein the array of RPU devices stores the weight values of the weight matrix W of the ANN with m rows and n columns as the resistance values of the RPU devices in the array, thereby representing the weight matrix W, the storing; Defining the weight matrix W to have an output dimension smaller than the input dimension such that the weight matrix W has a rectangular configuration; During the forward cycle path, copying the input of the repeated weight elements; Summing the output calculation results from the repeated weight elements that produce one output per row, and Updating only one of the repeated weight elements by setting all forward values except from one to zero during the update pass By repeating or concatenating the rectangular configuration of the weight matrix W to enhance the signal strength of the backward path signal, converting the weight matrix W from the rectangular configuration to a more square configuration A non-transitory computer-readable storage medium that causes the computer to perform the step of doing so.
Citation Information
Patent Citations
Noise and bound management for RPU array
US20180293208A1
Training of artificial neural networks
WO2019082077A1