Enabling hierarchical data loading within the Resistor Processing Unit (RPU) array to reduce communication costs
By using hierarchical loading technology of baseline and differential data in RPU arrays, the problem of high communication cost in large-scale neural network training is solved, and efficient data transmission and training process is realized, which significantly reduces communication overhead.
Patent Information
- Application Number
- JP2023550331
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-16
- Filing Date
- 2022-02-17
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2042-02-17
AI Technical Summary
When training large-scale parallel neural networks, the communication cost of input data becomes a bottleneck, especially under high-speed RPU operation, the prior art is difficult to effectively reduce communication overhead.
Using a technology to implement layered data loading in an RPU array, two random pulse generators (baseline and differential) are used to process the input layer data. The baseline data is not frequently transmitted, while the differential data is transmitted with fewer bits, and combined with the control circuit to generate pulse training signals, reducing communication needs.
The communication cost of input data is significantly reduced, reaching a 75% reduction, and no additional computing resources are required, achieving efficient neural network training.
Smart Images

Figure 0007725173000002 
Figure 0007725173000003 
Figure 0007725173000004
Abstract
Description
[Technical Field]
[0001] The present invention relates to electrical, electronic and computer technology, and more particularly to electronic circuits suitable for implementing neural networks and the like. [Background technology]
[0002] Neural networks are becoming increasingly popular for a variety of applications. They are used to perform machine learning. Computers learn to perform certain tasks by analyzing training examples, which are typically pre-labeled by human experts. Neural networks comprise thousands or even millions of simple processing nodes that are densely interconnected. Training neural networks and inference using trained neural networks are computationally expensive. In fact, the datasets required to train massively parallel neural networks require much larger input data, exceeding the order of terabytes (TB) in size.
[0003] To address the computational challenges associated with neural networks, hardware-based techniques have been proposed. For example, resistive processing unit (RPU) devices have the potential to speed up neural network training by orders of magnitude while using much less power. However, even at high-speed RPU operation, the cost of communicating input data over wireless or off-chip interfaces remains a significant overhead burden in many applications, such as training deep neural networks. Summary of the Invention [Means for solving the problem]
[0004] The principles of the present invention provide techniques for enabling hierarchical data loading in a resistive processing unit (RPU) array to reduce communication costs. In one aspect, an exemplary electronic circuit includes a plurality of word lines; a plurality of bit lines intersecting the plurality of word lines at a plurality of lattice points; a plurality of resistive processing units disposed at the plurality of lattice points; a plurality of baseline stochastic pulse input units connected to the plurality of word lines; a plurality of differential stochastic pulse input units connected to the plurality of word lines; and a plurality of bit line stochastic pulse input units connected to the plurality of bit lines. Also included is a control circuit connected to the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units, the control circuit configured to cause each of the plurality of baseline stochastic pulse input units to generate a baseline pulse train using base input data, to cause each of the plurality of differential stochastic pulse input units to generate a differential pulse train using differential input data defining a difference from the base input data, and to cause each of the plurality of bit line stochastic pulse input units to generate a bit line pulse train using bit line input data.
[0005] In another aspect, a hardware description language (HDL) design construct is encoded on a machine-readable data storage medium and comprises elements that, when processed in a computer-aided design system, generate a machine-executable representation of a device, the HDL design construct comprising an electronic circuit as just described.
[0006] In yet another aspect, an exemplary method includes providing the electronic circuitry just described, wherein the control circuitry causes each of the plurality of baseline stochastic pulse input units to generate a baseline pulse train using base input data, causes each of the plurality of differential stochastic pulse input units to generate a differential pulse train using differential input data defining a difference from the base input data, and causes each of the plurality of bit line stochastic pulse input units to generate a bit line pulse train using bit line input data.
[0007] As used herein, "facilitating" an action includes performing the action, facilitating the action, assisting in performing the action, or causing the action to be performed. Thus, by way of example and not limitation, instructions executing on one processor may facilitate an action performed by instructions executing on a remote processor by sending appropriate data or commands to cause or assist the action. For the avoidance of doubt, when an actor facilitates an action in a way other than performing the action, the action is performed by some entity or combination of entities.
[0008] One or more embodiments of the present invention or elements thereof can be implemented in hardware, e.g., digital circuitry. This digital circuitry can then be used in a computer to train / execute machine learning software in a computationally efficient manner. The machine learning software can be implemented in the form of a computer program product comprising a computer-readable storage medium having computer-usable program code for performing the indicated method steps. The software can then be executed on a system (or device) comprising a memory and at least one processor coupled to the memory and operative to perform exemplary machine learning training and inference, where the processor can be configured as described herein.
[0009] The techniques of the present invention can provide substantial beneficial technical effects. For example, one or more embodiments provide: Dramatically reduce the communication cost of input data for training large-scale parallel neural networks; Unlike other compression techniques to reduce the amount of input data, no explicit decompression step is required; and / or Since consecutive frames have high similarity, the amount of input data is greatly reduced.
[0010] These and other features and advantages of the present invention will become apparent from the following detailed description of illustrative embodiments, which is to be read in connection with the accompanying drawings. [Brief explanation of the drawings]
[0011] [Figure 1] FIG. 1 illustrates a prior art technique for applying a probabilistic update rule to an RPU-based array. [Figure 2] FIG. 2 illustrates a technique for applying probabilistic update rules to an RPU-based array in accordance with one aspect of the present invention. [Figure 3] FIG. 3 shows two exemplary cycles of updating for the array of FIG. 2 in accordance with one aspect of the present invention. [Figure 4] FIG. 4 shows an example of baseline data and difference data according to one aspect of the present invention. [Figure 5] FIG. 5 illustrates exemplary data of sparsity versus threshold in accordance with one aspect of the present invention. [Figure 6] FIG. 6 illustrates an exemplary file (data) size versus threshold value in accordance with one aspect of the present invention. [Figure 7] FIG. 7 illustrates the array of FIG. 2 during an exemplary inference process in accordance with one aspect of the present invention. [Figure 8]FIG. 8 illustrates a computer system (also representative of a general-purpose computer capable of implementing a design process such as that shown in FIG. 9) using a coprocessor in accordance with an aspect of the present invention, suitable for accelerating the implementation of neural networks. [Figure 9] FIG. 9 is a flow diagram of a design process used in semiconductor design, manufacturing and / or testing. DETAILED DESCRIPTION OF THE INVENTION
[0012] As noted above, even at high-speed RPU operations, input data communication costs remain a major overhead burden. One or more embodiments advantageously enable similar accelerations on the order of 10,000 times compared to conventional digital accelerator hardware. One or more embodiments take advantage of the fact that many data sets of interest exist that exhibit high data-to-data similarities. Examples include, but are not limited to, frames in molecular dynamics (MD) simulation data and pattern recognition in video frames with continuously moving objects. One or more embodiments advantageously split the input data (during preprocessing) into baseline data and difference data to eliminate similarities and reduce input data size. While splitting input data into baseline data and difference data is known per se, one or more embodiments provide a further improvement by efficiently implementing it in the RPU context without the need to explicitly recover the original data.
[0013] However, instead of using one stochastic pulse generator, one or more embodiments use two stochastic pulse generators (baseline and differential) per input layer row. In one or more embodiments, baseline data is transmitted occasionally, while differential data is transmitted using a reduced number of bits. This advantageously allows for significant reductions in communication costs (up to 75% reduction compared to RPUs that do not apply this technique). Because the baseline data and differential data are typically not calculated in the digital domain, no additional computational hardware / cost is required.
[0014] Moreover, in this regard, Figure 1 illustrates a prior art technique for applying a stochastic update rule to an RPU-based array. During an update cycle, all RPUs in the array can be updated in parallel, regardless of array size. A stochastic bit stream is used to encode the numerical values. By superimposing one or more signals, the weights are updated in an incremental manner. Referring to Equation 101, the change in the weights is, on average, proportional to:
[0015]
number
[0016] (W ij is the weight value of the ith row and jth column, x i is the activity in the input neuron, and δ jis the error calculated by the output neuron. The update time is proportional to BL (the length of the probabilistic bit stream at the output of the probabilistic translator 199). Switching of RPU devices 103-1,1, 103-1,2, 103-2,1, and 103-2,2 in the matrix occurs only when a positive pulse coincides with a negative pulse. Thus, arrow 105 represents the first update of device 103-1,1; arrow 107 represents the first update of device 103-2,1; arrow 109 represents the first update of device 103-2,2; arrow 111 represents the second update of device 103-1,1; arrow 113 represents the second update of device 103-2,1; and arrow 115 represents the first update of device 103-1,2. The voltage V x1 and V x2 The probability that V will toggle high is 0.5 and 0.6, respectively. x2 For a total update period of 10, V x1 Similarly, the voltage V δ1 and V δ2 The probabilities that are low are 0.3 and 0.4, as labeled in decimal form on the right side of FIG. 1, and in fractional form as shown to the right of arrows 105-115.
[0017] 2 and 3 show an exemplary embodiment of the present invention in which dual pulse generators are used for each row in the first neural network layer. "Diff" data is received every cycle, while baseline data only needs to be received every nth cycle. As in FIG. 1, the weights are updated in an incremental manner by superimposing one or more signals. Referring to Equation 201, the weight change here is, on average, (x i _ ベース +x i _ 差分 )xδ j is proportional to (W ij is the weight value of the ith row and jth column, x i_ベース is the base part of the activity in the input neuron, xi _ 差分 is the differential part of the activity in the input neuron, and δ j is the error computed by the output neuron).
[0018] Note the word lines (WL) 701 (only two shown to avoid clutter) and bit lines (BL) 703 (only two shown to avoid clutter). In Figure 2, each bit line has a single stochastic translator 299, while each word line 701 has two stochastic translators: a "baseline" (B) stochastic translator 297 and a "differential" (D) stochastic translator 295.
[0019] The update time is proportional to BL (the length of the probabilistic bitstream at the output of the probabilistic translator 295). Referring again to FIG. 3, two consecutive update cycles are illustrated. Thus, in FIG. 3, the total time of BL is 20 (2×10). FIG. 3 also shows an example where the difference (diff) value changes over the cycle, and therefore the number of pulses also changes. Switching of RPU devices 203-1,1, 203-1,2, 203-2,1, and 203-2,2 in the matrix occurs only when the positive and negative pulses coincide. Thus, arrow 205 represents a first update of device 203-1,1 with differential data during a first cycle; arrow 207 represents a first update of device 203-2,1 with differential data during the first cycle; arrow 209 represents a first update of device 203-2,2 with differential data during the first cycle; arrow 211 represents a second update of device 203-1,1 with differential data during the first cycle; arrow 213 represents a second update of device 203-2,1 with differential data during the first cycle; and arrow 215 represents a first update of device 203-1,2 with differential data during the first cycle. As described elsewhere, the baseline data typically does not change from cycle to cycle, but the differential data does.
[0020] In FIG. 2 , boxes 265 having the numbers (0.45, 0.05, 0.55, 0.05, 0.3, 0.4) represent temporary storage devices, e.g., registers, where baseline number 0.45 and diff number 0.05 are stored. Optionally, the baseline number registers may have a relatively high bit precision (e.g., 16 bits), and the diff number registers may have a relatively small precision (e.g., 4 bits), since diff numbers are typically very small compared to baseline numbers. In one or more embodiments, rather than providing any special enable signals, these number registers will be updated periodically. In the example of FIG. 2 , the diff numbers will be updated every cycle, and the baseline numbers will be updated infrequently; the updates are managed by controller 279. Note the “Baseline Update” signal connected to boxes 265 having numbers 0.45 and 0.55, and the “Diff Update” signal connected to the other box 265 having number 0.05. Stochastic translators 295 and 297 convert the input data of register 265 into a stochastic stream of highs and lows. Stochastic pulse generators 705 and 707 then drive WLs 701 based on the output (high / low stream) of the stochastic translators; i.e., they no longer use the input data directly. Thus, the stochastic pulse generators do not use the baseline / differential input data themselves, but rather the output of the stochastic translators. In other words, as shown, data A and data B come into registers, the stochastic translator uses the values to create a stream of highs and lows, and then the stream is sent to the stochastic pulse generator, which drives multiple word lines (WLs).
[0021] Furthermore, in FIG. 3, arrow 205' represents the first update of device 203-1,1 with differential data during the second cycle; arrow 207' represents the first update of device 203-2,1 with differential data during the second cycle; arrow 209' represents the first update of device 203-2,2 with differential data during the second cycle; arrow 211' represents the second update of device 203-1,1 with differential data during the second cycle; arrow 213' represents the second update of device 203-2,1 with differential data during the second cycle; and arrow 215' represents the first update of device 203-1,2 with differential data during the second cycle.
[0022] Voltage V x1 and voltage V x2 is high during the first cycle is 0.5 (0.45 + 0.05) and 0.6 (0.55 + 0.05), and δ1 and voltage V δ2 The probability that V is low during the first cycle is 0.3 and 0.4. This is shown in decimal form in Figure 2 and in fractional form in Figure 3. x1 and voltage V x2 is high during the second cycle is 0.4 (0.45-0.05), 0.5 (0.55-0.05), and the voltage V δ1 and voltage V δ2 The probabilities that σ is low are 0.3 and 0.4 during the second cycle (see also the decimal notation in FIG. 3).
[0023] Also shown in FIG. 2 are a conventional voltage supply 278, control circuitry 279, on-chip memory 269 (discussed further below) connected to register 265, and external memory 267 (discussed further below). Some or all of the components other than external memory 267 may be implemented on integrated circuit chip 263 (the voltage supply and / or control circuitry may be off-chip, if desired). Control circuitry 279 performs the functions defined herein; given the teachings and explanations of the functions herein, known control circuitry techniques, e.g., multi-cycle or pipelined, hardwired or microprogrammed, using any appropriate technology family (e.g., 7 nm CMOS, 5 nm CMOS, etc.), may be used. For example, the identified functions may be instantiated in logic circuitry, as described below with respect to FIG. 9.
[0024] FIG. 4 shows an example of baseline data and differential data. Original data is shown at 401. The original data is divided into baseline data 403 and differential data 405. When the differential data is added to the baseline data, the original data is reconstructed at 407. Assuming that copying the baseline data is negligible, the total overhead for copying the original data (96 bytes) for an exemplary embodiment of the present invention would be 16 bytes ((16 / 96)*100=17%) by ignoring the one-time baseline data overhead. Moreover, at this point, the amount of training data in applications is typically very large; therefore, data typically needs to be moved from an external mass storage device, such as a solid-state drive (SSD), to on-chip memory (random access memory (RAM)) via an off-chip interface or wireless communication. This communication cost is very expensive; therefore, one or more embodiments seek to minimize the communication cost. SSD and RAM are non-limiting examples of such communication. Appropriately, in one or more embodiments, the baseline data is copied infrequently, and therefore the communication costs for the baseline data are negligible.
[0025] In one or more embodiments, the 8 bytes are assumed to represent a floating-point number. The baseline data can be determined, for example, using heuristic or statistical methods (e.g., mean, median). The baseline data should typically be some representative data over a long period of time. For example, if the total data over three cycles is 0.6 → 0.7 → 0.5, the baseline data = 0.6, and the difference data is 0 → 0.1 → -0.1. Here, 0.6 is the average value over the three values 0.6, 0.7, and 0.5. The "mean and median" are the most widely used indicators for finding a representative value among many values. The baseline data 403 is 32 bytes (4 * 8), while the difference data 405 is 16 bytes (2 * 8). The 32 bytes of baseline data can be amortized if n (the number of difference frames before the baseline changes, i.e., the baseline data only needs to be received every nth cycle) is sufficiently large. Therefore, when n is sufficiently large, it can be assumed that the overhead of the baseline data is negligible (in this figure, a value of n=3, i.e., 3 frames of differential data, is shown, but this is for convenience of explanation; in a practical case, this number should typically be much larger, for example, more than 24). If the overhead of the baseline data is 0, the data can be transmitted in 16 bytes instead of 96 bytes (leading to a value of 17%). Note that the value of 96 bytes is determined by the following formula: original data has three 2x2 arrays; 2*2=4, 4*3*8 bytes / floating-point number=96 bytes. Furthermore, differential data with sparse differential matrices, where only two elements in the differential matrices have non-zero values, is 8 bytes times per floating-point number, i.e., 16 bytes.
[0026] Figure 5 shows sparsity versus threshold, and Figure 6 shows file (data) size versus threshold. The input data is a molecular dynamics distance matrix, which represents the distance between proteins (183x183*8 bytes = 262KB per frame). The threshold (tolerable error rate) indicates how much error is allowed to be drawn (in the illustrated example, Angstroms). Sparsity is the number of zeros divided by the number of cells in the matrix. A larger threshold increases sparsity, while a higher sparsity reduces data (file) size. Therefore, the higher the sparsity, the smaller the file size and, therefore, the less data to load. In such applications, consecutive image frames have high similarity, leading to very small difference data with very high sparsity. Therefore, very low data communication can be achieved. Using a high threshold within the range of tolerable error can further reduce communication.
[0027] It will be appreciated that training large neural networks therefore requires large streams of input data, resulting in frequent raw data communications exceeding TB in size, e.g., video frames of moving objects and molecular dynamics. It is known to separate input data into baseline data and differential data to eliminate similarities, with the baseline data being transmitted infrequently and the smaller volume differential data being transmitted frequently. While this achieves significant reductions in communication costs, it requires an explicit recovery process by combining the baseline data and the differential data. Advantageously, one or more embodiments do not require or use such an explicit decompression process.
[0028] Additionally, one or more embodiments provide an efficient training hardware system that uses an array of resistive processor units (RPUs) for streamed input images with high frame-to-frame similarity. By having two separate pulse generators per word line for baseline data and differential data, baseline data and differential data can be applied to the array simultaneously during training. As described above, unlike conventional digital implementations, an architecture according to one or more embodiments does not require explicit recovery of the original data.
[0029] In one or more embodiments, the weights are stored in the array to be trained. As described elsewhere herein, stochastic translators 295 and 297 convert the input data in register 265 into a stochastic stream of highs and lows. Stochastic pulse generators 705 and 707 then simply drive WL 701 based on the output (high / low stream) of the stochastic translator; i.e., they do not use the input data directly any more. Thus, these pulse generators do not use the base / differential input data themselves, but rather the output of the stochastic translator. In other words, as shown, data A and B arrive at a register, and then the stochastic translator uses the values to generate a stream of highs and lows, which are then sent to a pulse generator, which drives multiple word lines (WL). Differential data is received every cycle, while baseline data is received every nth cycle. The learning rate can be controlled by modulating the pulse width (BL) of the update enable signal.
[0030] Accordingly, one or more embodiments provide a method and circuit for efficient hardware-based training of neural networks using a resistive processor unit (RPU) for streamed input image or similar data with high similarity between multiple frames.
[0031] While the exemplary use case is described in the context of the update process during training, it will be understood that an important computational kernel for forward and inference is also matrix multiplication. For example, it may be necessary to compute X*W between a (voltage) vector X and a (matrix, e.g., weight) W. In a non-limiting use case, during inference, a stream of video frames may be received to find the location of a target. In such a case, to save on communication costs, a similar technique may be used to convert a vector X into "X_W". ベース +X_ 差分 " can be used by breaking it down into "
[0032] An exemplary inference process is shown in Figure 7. During inference, the stored values of the RPU bit cells 203 are not changed. ベース +X_ 差分 is input to the word lines 701 from the voltage vector peripheral circuit 796, and the current at the bottom of the bit lines 703 is measured using an integrator (op-amp 753 and capacitor 755) and an analog-to-digital converter (ADC) 751.
[0033] Those skilled in the art will be familiar with conventional training and inference of RPU arrays, for example, from Gokmen T. and Vlasov Y., Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices: Design Considerations, Front.Neurosci. 10:333, doi:10.3389 / fnins.2016.00333, 21 July 2016 and Gokmen T., Onen M. and Haensch W., Training Deep Convolutional Neural Networks with Resistive Cross-Point Devices, Front.Neurosci. 11:538. doi:10.3389 / fnins.2017.00538, 10 October 2017.
[0034] In light of the foregoing discussion, it will be appreciated that, generally speaking, an exemplary electronic circuit according to one aspect of the present invention comprises a plurality of word lines 701; a plurality of bit lines 703 intersecting the plurality of word lines at a plurality of grid points; and a plurality of resistance processing units 203-1,1, ..., 203-2,2 disposed at the plurality of grid points. Also included are a plurality of baseline stochastic pulse input units 265 (i.e., their plurality of registers connected to a plurality of translators 297) 297 and 705 connected to the plurality of word lines; a plurality of differential stochastic pulse input units 265 (i.e., their plurality of registers connected to a plurality of translators 295) 295 and 707 connected to the plurality of word lines; and a plurality of bit line stochastic pulse input units 265 (i.e., their plurality of registers connected to a plurality of translators 299) 299 and 298 connected to the plurality of bit lines. The system further includes a control circuit 279 connected to the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units, the control circuit 279 being configured to cause each of the plurality of baseline stochastic pulse input units to generate one baseline pulse train using base input data, to cause each of the plurality of differential stochastic pulse input units to generate one differential pulse train using differential input data defining a difference from the base input data, and to cause each of the plurality of bit line stochastic pulse input units to generate one bit line pulse train using bit line input data.
[0035] In one or more embodiments, the control circuit controls the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units to store neural network weights in the plurality of resistive processing units. In one or more embodiments, whatever the application, the weights are simply numerical values obtained from an iterative training machine learning process to make an accurate final decision, such as whether an image is a human or a cat, or whether a cell is a cancer cell or a normal cell. Once the weights are obtained, they can be used for the application (i.e., during inference) by multiplying the numerical values with input image values during matrix multiplication or convolution. Thus, inference can be performed, for example, to recognize an image. Appropriate action, such as controlling an autonomous vehicle, a robotic surgical device, etc., can be taken based on the recognized image.
[0036] In one or more embodiments, the plurality of baseline stochastic pulse input units each include a baseline register (i.e., a plurality of registers connected to a plurality of translators 297) configured to store a corresponding portion of the baseline input data, a baseline stochastic translator 297 connected to the baseline register, and a baseline pulse generator 705 connected to the baseline stochastic translator and a corresponding one of the plurality of word lines, and the plurality of differential stochastic pulse input units each include a differential register (i.e., a plurality of registers connected to a plurality of translators 297) configured to store a corresponding portion of the baseline input data. and each of the plurality of bit line stochastic pulse input units comprises a bit line register (i.e., a plurality of registers connected to a plurality of translators 299) configured to store a corresponding portion of the bit line input data, and a bit line pulse generator 298 connected to the bit line stochastic translator and a corresponding one of the plurality of bit lines 703.
[0037] In one or more embodiments, the baseline stochastic translator 297 is configured to convert baseline data in the plurality of baseline registers into a plurality of baseline output stochastic streams of highs and lows, and the baseline pulse generator 705 is configured to drive the plurality of word lines based on the baseline outputs; the plurality of differential stochastic translators 295 are configured to convert differential data in the plurality of differential registers into a plurality of differential output stochastic streams of highs and lows, and the plurality of differential pulse generators 707 are configured to drive the plurality of word lines based on the differential outputs; and the plurality of bit line stochastic translators 299 are configured to convert bit line data in the plurality of bit line registers into a plurality of bit line output stochastic streams of highs and lows, and the plurality of bit line pulse generators 298 are configured to drive the plurality of bit lines based on the bit line outputs.
[0038] In one or more embodiments, the electronic circuitry is implemented as an integrated circuit chip 263 and further includes on-chip memory 269 (e.g., random-access memory (RAM), such as static RAM (SRAM)) connected to the registers 265 and providing an interface (I / F) to off-chip storage / external memory 267 (in a non-limiting example, a solid-state drive (SSD)).
[0039] In one or more embodiments, the differential data is received every cycle; while the baseline data is received every nth cycle. In one or more embodiments, whenever new input data is received from an external device (e.g., SSD), such data is typically stored in an on-chip temporary storage device 269, such as RAM (in a non-limiting example, SRAM). However, during (update / training) calculations, the data is moved closer to the RPU core, such as into a register. The values are stored in the register, and the base register is not updated frequently.
[0040] In update equation 201, the bit line data is "delta (δ)." The result of the delta comes from a previous backpropagation process and is stored in on-chip storage device 269 and moved to register 265 near the RPU core during the update process. However, in one or more embodiments, this delta is not updated frequently. The delta is reused for many new inputs through so-called "batch-based" processing.
[0041] As discussed elsewhere, the learning rate can be controlled by modulating the pulse width (BL) of the update enable signal. Thus, in one or more embodiments, the control circuitry is configured to control the update time by controlling the length of the probabilistic bitstream (see BL for the first and second cycles in FIG. 3) at the output of the plurality of differential probabilistic translators.
[0042] 7, the electronic circuit further includes a voltage vector peripheral circuit 796 connected to the word lines and a plurality of integrators (e.g., 751, 753, and 755) connected to the bit lines. The control circuit controls the voltage vector peripheral circuit and the integrators to perform inference on the resistance processing units with the neural network weights stored in the resistance processing units. Suitable integration techniques are known, for example, from the Gokmen et al. paper mentioned above.
[0043] In one or more embodiments, the control circuit controls the voltage vector peripheral circuitry to apply a voltage vector to the plurality of word lines as baseline data plus differential data, i.e., X_ ベース +X_ 差分 , and enter it as
[0044] Given the teachings herein, one skilled in the art can implement the circuits herein using known integrated circuit fabrication techniques. The weights are updated in an incremental manner by superimposing one or more signals, taking into account the various states of the cells and how they are programmed. In one or more embodiments, two independent random number generators for columns and rows are sufficient.
[0045] In another aspect, an exemplary method for training a computer-implemented neural network includes providing the electronic circuit described above (alternatively, instead of an explicit providing step, the circuit is the workpiece on which the method operates), the method including causing (e.g., using control circuitry) each of the plurality of baseline stochastic pulse input units to generate a baseline pulse train using base input data, each of the plurality of differential stochastic pulse input units to generate a differential pulse train using differential input data defining a difference from the base input data, and each of the plurality of bit line stochastic pulse input units to generate a bit line pulse train using bit line input data.
[0046] In one or more embodiments, the method further includes controlling (e.g., using the control circuit) the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units to store a plurality of neural network weights within the plurality of resistance processing units.
[0047] In one or more embodiments, in the providing (or alternatively, in a workpiece on which the method operates), the plurality of baseline stochastic pulse input units each include a baseline register configured to store a corresponding portion of the base input data, a baseline stochastic translator connected to the baseline register, and a baseline pulse generator connected to the baseline stochastic translator and a corresponding one of the plurality of word lines; the plurality of differential stochastic pulse input units each include a differential register configured to store a corresponding portion of the differential input data, a differential stochastic translator connected to the differential register, and a differential pulse generator connected to the differential stochastic translator and a corresponding one of the plurality of word lines; and the plurality of bit line stochastic pulse input units each include a bit line register configured to store a corresponding portion of the bit line input data, a bit line stochastic translator connected to the bit line register, and a bit line pulse generator connected to the bit line stochastic translator and a corresponding one of the plurality of bit lines. The method further includes the plurality of baseline stochastic translators converting baseline data in the plurality of baseline registers into a plurality of baseline output stochastic streams of highs and lows; the baseline pulse generator driving the plurality of word lines based on the baseline outputs; the plurality of differential stochastic translators converting differential data in the plurality of differential registers into a plurality of differential output stochastic streams of highs and lows; the plurality of differential pulse generators driving the word lines based on the differential outputs; the plurality of bit line stochastic translators converting bit line data in the plurality of bit line registers into a plurality of bit line output stochastic streams of highs and lows; and the plurality of bit line pulse generators driving the plurality of bit lines based on the bit line outputs.
[0048] One or more embodiments further include controlling the update time by controlling (e.g., using the control circuitry) the length of the probabilistic bitstream at the output of the plurality of differential probabilistic translators.
[0049] 7, in one or more embodiments, in the providing (or alternatively, in the workpiece on which the method operates), the electronic circuitry further includes a voltage vector peripheral circuit connected to the plurality of word lines and a plurality of integrators connected to the plurality of bit lines. A further step further includes controlling (e.g., with the control circuitry) the voltage vector peripheral circuitry and the plurality of integrators to perform inference in the plurality of resistance processing units with the plurality of neural network weights stored in the plurality of resistance processing units.
[0050] Still referring to FIG. 7, one or more embodiments further include controlling (e.g., using the control circuit) the voltage vector peripheral circuitry to input a voltage vector to the plurality of word lines as baseline data plus differential data.
[0051] Referring to Figure 8, some aspects of the present invention can be implemented as a hardware coprocessor 999 that uses specialized hardware techniques to accelerate training (and optionally inference) for neural networks and the like. Figure 8 illustrates a computer system 12 that includes such a hardware coprocessor. The computer system 12 includes, for example, one or more conventional processors or processing units 16, a system memory 28, and a bus 18 that connects various system components, including the system memory 28 and one or more hardware coprocessors 999, to the processor 16. The elements 999 and 16 can be connected to the bus using, for example, a suitable bus interface unit.
[0052] Bus 18 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0053] Computer system / server 12 typically includes a variety of computer-readable media, which may be any available media that can be accessed by computer system / server 12 and includes both volatile and nonvolatile media, removable and non-removable media.
[0054] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 or cache memory 32, or a combination thereof. The computer system / server 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium (not shown, typically referred to as a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk, such as a CD-ROM, DVD-ROM, or other optical medium, may be provided. In such cases, each may be connected to the bus 18 by one or more data medium interfaces. As further shown and described below, the system memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to execute software-implemented portions of, for example, a neural network or a digital filter.
[0055] A program / utility 40 having a set (at least one) of program modules 42 may be stored in system memory 28, by way of example and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include implementation of a network environment. The program modules 42 generally perform software-implemented functions or methodologies and combinations thereof.
[0056] The computer system / server 12 may also communicate with one or more external devices 14, such as a keyboard, pointing device, display 24, one or more devices that allow a user to interact with the computer system / server 12, or any device (e.g., a network card, modem, etc.) that allows the computer system / server 12 to communicate with one or more other computer devices, or a combination thereof. Such communication may occur via an input / output (I / O) interface 22. The computer system / server 12 may also communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof, via a network adapter 20. As shown, the network adapter 20 communicates with other components of the computer system / server 12 via a bus 18. While not shown, it should be understood that other hardware or software components, or combinations thereof, may be used with the computer system / server 12. Examples include, but are not limited to, microcode, device drivers, redundant processing units, and external disk drive arrays, RAID systems, tape drives, and data archival storage systems.
[0057] Still referring to FIG. 8 , attention is drawn to the processor 16, memory 28, and input / output interface 22 between the display 24 and one or more external devices 14, e.g., a keyboard, pointing device, etc. As used herein, the term “processor” is intended to encompass any processing device, such as a central processing unit (CPU) or other form of processing circuitry (e.g., 999), or the like, or a combination thereof. Additionally, the term “processor” may refer to one or more individual processors. The term “memory” is intended to encompass memory associated with a processor, i.e., a CPU, such as random access memory (RAM) 30, read only memory (ROM), fixed memory devices (e.g., hard drive 34), removable memory devices (e.g., diskettes), and flash memory. Additionally, as used herein, the phrase "input / output interface" is intended to contemplate, for example, an interface to one or more mechanisms for inputting data into a processing device (e.g., a mouse) and one or more mechanisms for providing results associated with the processing device (e.g., a printer). The processor 16, the coprocessor 999, the memory 28, and the input / output interface 22 may be interconnected, for example, via a bus 18, as part of the data processing unit 12. Suitable interconnections, for example, via the bus 18, may also be provided to a network interface 20, for example, a network card, which may be provided for interfacing with a computer network, and a media interface, for example, a diskette or CD-ROM drive, which may be provided for interfacing with suitable media.
[0058] Thus, computer software containing instructions or code to perform desired tasks may be stored in one or more associated memory devices (e.g., ROM, fixed or removable memory) and loaded in part or in whole (e.g., into RAM) and executed by the CPU when ready for use. Such software may include, but is not limited to, firmware, resident software, microcode, etc.
[0059] A data processing system suitable for storing and / or executing program code includes at least one processor 16 coupled directly or indirectly to memory elements 28 through a system bus 18. The memory elements may include local memory used during the actual implementation of the program code, bulk storage, and cache memory 32 that provides temporary storage of at least some of the program code to reduce the number of times the code must be retrieved from bulk storage during implementation.
[0060] Input / output devices or I / O devices (including but not limited to keyboards, displays, pointing devices, etc.) can be connected to the system either directly or through intervening I / O controllers.
[0061] Network adapters 20 may also be connected to the data processing system to enable the system to connect to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modems, and Ethernet cards are just a few of the currently available types of network adapters.
[0062] As used herein, including the claims, a "server" encompasses a physical data processing system (e.g., system 12 shown in FIG. 8) that executes a server program. It will be understood that such a physical server may or may not include a display and keyboard. Moreover, FIG. 8 also represents a conventional general-purpose computer (e.g., a general-purpose computer without a coprocessor 999) that can be used, for example, to implement aspects of the design process described below.
[0063] Exemplary Design Processes Used in Semiconductor Design, Manufacturing, or Test, or a Combination thereof
[0064] One or more embodiments of hardware according to aspects of the present invention can be implemented using techniques for semiconductor integrated circuit design simulation, testing, layout, or fabrication, or a combination thereof. In this regard, FIG. 9 illustrates a block diagram of an exemplary design flow 700 used, for example, in semiconductor IC logic design, simulation, testing, layout, and fabrication. Design flow 700 encompasses a process, machine, or mechanism, or combination thereof, for processing a design structure or device to generate a logically or other functionally equivalent representation of the design structure or device, or combination thereof, such as those described herein. The design structure processed, generated, or processed and generated by design flow 700 may be encoded on a machine-readable storage medium to include data or instructions, or a combination thereof, that, when executed on a data processing system or otherwise processed, generates a logical, structural, mechanical, or other functionally equivalent representation of a hardware component, circuit, device, or system. Machine encompasses, but is not limited to, any machine used in an IC design process, for example, designing, fabricating, or simulating a circuit, component, device, or system. For example, a machine may include a lithography apparatus, a machine or apparatus for generating masks, or a combination thereof (e.g., e-beam writers), a computer or apparatus for simulating a design structure, any apparatus used in a manufacturing or testing process, or any machine for programming a functionally equivalent representation of a design structure into any medium (e.g., a machine for programming a programmable gate array).
[0065] The design flow 700 may differ depending on the type of representation being designed. For example, a design flow 700 for building an application specific IC (ASIC) may differ from a design flow 700 for designing standard components, or for converting the design into a programmable array, such as an Altera 登録商標Inc. or Xilinx 登録商標 The design flow 700 may differ for instantiation into a programmable gate array (PGA) or field programmable gate array (FPGA), offered by Inc.
[0066] FIG. 9 illustrates a plurality of such design structures, preferably comprising input design structure 720, processed by design process 710. Design structure 720 may be a logic simulation design structure generated and processed by design process 710 to generate a logically equivalent functional representation of a hardware device. Design structure 720 may also, or alternatively, include data or program instructions, or a combination thereof, that, when processed by design process 710, generates a functional representation of the physical structure of a hardware device. Whether representing functional or structural design features, or a combination thereof, design structure 720 may be generated using electronic computer-aided design (ECAD), such as electronic computer-aided design implemented by a core developer / designer. When encoded, such as on a gate array or storage medium, design structure 720 may be accessed and processed by one or more hardware or software modules, or a combination thereof, in design process 710 to simulate or otherwise functionally represent an electronic component, circuit, electronic or logic module, apparatus, device, or system. Thus, design structure 720 may include files or other data structures comprising human- or machine-readable source code, or a combination thereof, compiled structures, and computer-executable code structures, that, when processed by a design or simulation data processing system, functionally simulate or otherwise represent a circuit or other level of hardware logic design. Such data structures may include hardware-description language (HDL) design entities or other data structures compliant with or compatible with low-level HDL design languages, e.g., Verilog and VHDL, or high-level design languages, e.g., C or C++, or a combination thereof.
[0067] Design process 710 preferably uses and incorporates hardware or software modules, or a combination thereof, for synthesizing, translating, or otherwise manipulating design / simulation functional equivalents of components, circuits, devices, or logic structures to generate netlist 780, which may comprise design structures, such as design structure 720. Netlist 780 may include, for example, compiled or otherwise manipulated data structures representing lists describing connections to other elements and circuits in an integrated circuit design, such as wires, discrete components, logic gates, control circuits, I / O devices, models, etc. Netlist 780 may be synthesized using an iterative process in which netlist 780 is resynthesized one or more times depending on the device's design specifications and parameters. As with the other design structure types described herein, netlist 780 may be recorded on a machine-readable data storage medium or programmed into a programmable gate array, which may be a non-volatile storage medium, such as a magnetic or optical disk drive, a programmable gate array, compact flash, or other flash memory. Additionally or alternatively, the medium may be system or cache memory, buffer space, or other suitable memory.
[0068] The design process 710 may include hardware and software modules for processing various input data structure types, such as those described above, including netlist 780. Such data structure types may reside, for example, in library elements 730 and may include a set of commonly used elements, circuits, and devices for a given manufacturing technology (e.g., different technology nodes, such as 32 nm, 45 nm, 90 nm, etc.), including models, layouts, and symbolic representations of the commonly used elements, circuits, and devices. The data structure types may further include design specifications 740, characterization data 750, verification data 760, design rules 770, and test data files 785, which may include input test patterns, output test results, and other test information. The design process 710 may also include standard mechanical design processes, such as stress analysis; thermal analysis; mechanical event simulation; and process simulation for operations, such as casting, molding, and die pressing. Those skilled in the art of mechanical design will appreciate the range of mechanical design tools and applications that may be used in design process 710 without departing from the scope of the present invention. Design process 710 may also include modules for performing standard circuit design processes, such as timing analysis, verification, design rule checking, place and route operations, etc.
[0069] Design process 710 uses and incorporates logical and physical design tools, such as HDL compilers and simulation model building tools, to process design structure 720 along with some or all of the illustrated supporting data structures, along with any additional mechanical design or data (if applicable), to generate second design structure 790. Design structure 790 resides on a storage medium or programmable gate array in a data format used for the exchange of mechanical device and structure data (e.g., information stored in IGES, DXF, Parasolid XT, JT, DRG, or any other suitable format for storing or rendering such mechanical design structures). Like design structure 720, design structure 790 preferably resides on a data storage medium and includes one or more files, data structures, or other computer-encoded data or instructions that, when processed by an ECAD system, generate a logical or other functionally equivalent form of one or more IC designs, such as those disclosed herein. In one embodiment, design structure 790 may include a compiled, executable HDL simulation model that functionally simulates a device disclosed herein.
[0070] Design structure 790 may also use data formats used for the exchange of integrated circuit layout data or symbolic data formats, or combinations thereof (e.g., information stored in GDSII (GDS2), GL1, OASIS, map files, or any other suitable format for storing such design data structures). Design structure 790 may include information such as symbol data, map files, test data files, design content files, manufacturing data, layout parameters, wires, metal levels, vias, shapes, data for routing through a manufacturing line, and any other data needed by a manufacturer or other designer / developer to manufacture the devices or structures described herein. Design structure 790 then proceeds to stage 795, where, for example, design structure 790 may proceed to tape-out, be released to manufacturing, be released to a mask house, be sent to another design house, or be sent back to the customer.
[0071] The description of various embodiments of the present invention has been presented for illustrative purposes and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used in this specification have been selected to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the market, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. An electronic circuit comprising: a plurality of word lines; a plurality of bit lines intersecting the plurality of word lines at a plurality of grid points; a plurality of resistance processing units arranged at the plurality of grid points; a plurality of baseline stochastic pulse input units connected to the plurality of word lines; a plurality of differential stochastic pulse input units connected to the plurality of word lines; a plurality of bit line stochastic pulse input units connected to the plurality of bit lines; and a control circuit connected to the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units, the control circuit configured to cause each of the plurality of baseline stochastic pulse input units to generate a baseline pulse train using base input data, cause each of the plurality of differential stochastic pulse input units to generate a differential pulse train using differential input data defining a difference from the base input data, and cause each of the plurality of bit line stochastic pulse input units to generate a bit line pulse train using bit line input data. The electronic circuit comprises:
2. 2. The electronic circuit of claim 1, wherein the control circuit controls the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units to store a plurality of neural network weights in the plurality of resistive processing units.
3. each of the plurality of baseline stochastic pulse input units comprising a baseline register configured to store a corresponding portion of the base input data, a baseline stochastic translator coupled to the baseline register, and a baseline pulse generator coupled to the baseline stochastic translator and to a corresponding one of the plurality of word lines; each of the plurality of differential stochastic pulse input units comprising a differential register configured to store a corresponding portion of the differential input data, a differential stochastic translator coupled to the differential register, and a differential pulse generator coupled to the differential stochastic translator and the corresponding one of the plurality of word lines; and each of the plurality of bit line stochastic pulse input units includes a bit line register configured to store a corresponding portion of the bit line input data; a bit line stochastic translator connected to the bit line register; and a bit line pulse generator connected to the bit line stochastic translator and to a corresponding one of the plurality of bit lines.
3. The electronic circuit of claim 2.
4. the plurality of baseline stochastic translators configured to convert baseline data in the plurality of baseline registers into a plurality of baseline output stochastic streams of highs and lows, and the baseline pulse generator configured to drive the plurality of word lines based on the baseline outputs; the plurality of differential stochastic translators configured to convert differential data in the plurality of differential registers into a plurality of differential output stochastic streams of highs and lows, and the plurality of differential pulse generators configured to drive the plurality of word lines based on the differential outputs; and the plurality of bit line stochastic translators configured to convert bit line data in the plurality of bit line registers into a plurality of bit line output stochastic streams of highs and lows, and the plurality of bit line pulse generators configured to drive the plurality of bit lines based on the bit line outputs.
4. The electronic circuit of claim 3.
5. 5. The electronic circuit of claim 4, wherein the electronic circuit is implemented as an integrated circuit chip and further comprises an on-chip random access memory connected to the register and having an interface to off-chip storage.
6. 6. The electronic circuit of claim 5, wherein the control circuitry is configured to control update times by controlling lengths of probabilistic bitstreams at the outputs of the plurality of differential probabilistic translators.
7. 3. The electronic circuit of claim 2, further comprising: a voltage vector peripheral circuit connected to the plurality of word lines; and a plurality of integrators connected to the plurality of bit lines, wherein the control circuit controls the voltage vector peripheral circuit and the plurality of integrators to perform inference in the plurality of resistance processing units with the plurality of neural network weights stored in the plurality of resistance processing units.
8. 8. The electronic circuit of claim 7, wherein the control circuit controls the voltage vector peripheral circuit to input voltage vectors to the plurality of word lines as baseline data plus differential data.
9. 1. A method for training a computer-implemented neural network, comprising: Providing an electronic circuit, said electronic circuit comprising: a plurality of word lines; a plurality of bit lines intersecting the plurality of word lines at a plurality of grid points; a plurality of resistance processing units arranged at the plurality of grid points; a plurality of baseline stochastic pulse input units connected to the plurality of word lines; a plurality of differential stochastic pulse input units connected to the plurality of word lines; a plurality of bit line stochastic pulse input units connected to the plurality of bit lines; and a control circuit connected to the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units; It is equipped with wherein the control circuit causes each of the plurality of baseline stochastic pulse input units to generate one baseline pulse train using base input data, causes each of the plurality of differential stochastic pulse input units to generate one differential pulse train using differential input data defining a difference from the base input data, and causes each of the plurality of bit line stochastic pulse input units to generate one bit line pulse train using bit line input data. The method.
10. 10. The method of claim 9, further comprising using the control circuitry to control the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units to store a plurality of neural network weights in the plurality of resistive processing units.
11. In the preparing, each of the plurality of baseline stochastic pulse input units comprises a baseline register configured to store a corresponding portion of the base input data, a baseline stochastic translator coupled to the baseline register, and a baseline pulse generator coupled to the baseline stochastic translator and to a corresponding one of the plurality of bit lines; each of the plurality of differential stochastic pulse input units comprising a differential register configured to store a corresponding portion of the differential input data, a differential stochastic translator coupled to the differential register, and a differential pulse generator coupled to the differential stochastic translator and to the corresponding one of the plurality of word lines; and each of the plurality of bit line stochastic pulse input units comprises: a bit line register configured to store a corresponding portion of the bit line input data; a bit line stochastic translator connected to the bit line register; and a bit line pulse generator connected to the bit line stochastic translator and to a corresponding one of the plurality of bit lines; Furthermore, the plurality of baseline stochastic translators converting baseline data in the plurality of baseline registers into a plurality of high and low baseline output stochastic streams; the baseline pulse generator driving the plurality of word lines based on the baseline output; said plurality of differential stochastic translators converting the differential data in said plurality of differential registers into a plurality of differential output stochastic streams of high and low; the plurality of differential pulse generators driving the plurality of word lines based on the differential outputs; the plurality of bit line stochastic translators converting bit line data in the plurality of bit line registers into a plurality of bit line output stochastic streams of highs and lows; and the plurality of bit line pulse generators drive the plurality of bit lines based on the bit line outputs; Further comprising: The method of claim 10.
12. 12. The method of claim 11, further comprising using the control circuitry to control update times by controlling lengths of probabilistic bitstreams at the outputs of the plurality of differential probabilistic translators.
13. 11. The method of claim 10, wherein in the providing, the electronic circuitry further comprises a voltage vector peripheral circuit connected to the plurality of word lines and a plurality of integrators connected to the plurality of bit lines, and further comprising controlling the voltage vector peripheral circuitry and the plurality of integrators by the control circuitry to perform inference in the plurality of resistance processing units with the plurality of neural network weights stored in the plurality of resistance processing units.
14. 14. The method of claim 13, further comprising using the control circuitry to control the voltage vector peripheral circuitry to input voltage vectors to the plurality of word lines as baseline data plus differential data.
15. 1. A method for simulating an electronic circuit comprising: providing a hardware description language (HDL) code encoded on a machine-readable data storage medium, the HDL code comprising: a plurality of word lines; a plurality of bit lines intersecting the plurality of word lines at a plurality of grid points; a plurality of resistance processing units arranged at the plurality of grid points; a plurality of baseline stochastic pulse input units connected to the plurality of word lines; a plurality of differential stochastic pulse input units connected to the plurality of word lines; a plurality of bit line stochastic pulse input units connected to the plurality of bit lines; and a control circuit connected to the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units, the control circuit configured to cause each of the plurality of baseline stochastic pulse input units to generate a baseline pulse train using base input data, cause each of the plurality of differential stochastic pulse input units to generate a differential pulse train using differential input data defining a difference from the base input data, and cause each of the plurality of bit line stochastic pulse input units to generate a bit line pulse train using bit line input data. The HDL code comprises:
16. each of the plurality of baseline stochastic pulse input units comprising a baseline register configured to store a corresponding portion of the base input data, a baseline stochastic translator coupled to the baseline register, and a baseline pulse generator coupled to the baseline stochastic translator and to a corresponding one of the plurality of word lines; each of the plurality of differential stochastic pulse input units comprising a differential register configured to store a corresponding portion of the differential input data, a differential stochastic translator coupled to the differential register, and a differential pulse generator coupled to the differential stochastic translator and the corresponding one of the plurality of word lines; and each of the plurality of bit line stochastic pulse input units comprises: a bit line register configured to store a corresponding portion of the bit line input data; a bit line stochastic translator connected to the bit line register; and a bit line pulse generator connected to the bit line stochastic translator and to a corresponding one of the plurality of bit lines.
16. The HDL code of claim 15.
17. the plurality of baseline stochastic translators configured to convert baseline data in the plurality of baseline registers into a plurality of baseline output stochastic streams of highs and lows, and the baseline pulse generator configured to drive the plurality of word lines based on the baseline outputs; the plurality of differential stochastic translators configured to convert differential data in the plurality of differential registers into a plurality of differential output stochastic streams of highs and lows, and the plurality of differential pulse generators configured to drive the plurality of word lines based on the differential outputs; and the plurality of bit line stochastic translators configured to convert bit line data in the plurality of bit line registers into a plurality of bit line output stochastic streams of highs and lows, and the plurality of bit line pulse generators configured to drive the plurality of bit lines based on the bit line outputs.
17. The HDL code of claim 16.
18. 20. The HDL code of claim 17, further comprising an on-chip random access memory coupled to the register and having an interface to off-chip storage.
19. 20. The HDL code of claim 18, wherein the control circuitry is configured to control update times by controlling lengths of probabilistic bitstreams at outputs of the plurality of differential probabilistic translators.
20. The method of claim 20, wherein the control circuit controls the plurality of baseline stochastic pulse input units, the plurality of differential stochastic pulse input units, and the plurality of bit line stochastic pulse input units to store a plurality of neural network weights in the plurality of resistive processing units; 17. The HDL code of claim 16, further comprising: a voltage vector periphery circuit connected to the plurality of word lines; and a plurality of integrators connected to the plurality of bit lines, wherein the control circuitry controls the voltage vector periphery circuitry and the plurality of integrators to perform inference on the plurality of resistor processing units with the plurality of neural network weights stored in the plurality of resistor processing units.
Citation Information
Patent Citations
Systems and methods for efficient generation of stochastic spike patterns in core-based neuromorphic systems
JP2017126332A
Differential Coding in Neural Networks
JP2017516192A
Resistive Processing Unit
JP2019502970A
Resistive processing unit array, method for forming a resistive processing unit array and method for hysteretic operation - Patents.com
JP2020514886A
Convolutional neural networks using resistive processing unit array
US20180075338A1