Analog error correction circuit, analog error detection and correction method, and non-transitory, computer-readable medium for analog error detection and correction in analog in-memory crossbars
The circuit implementation with three crossbar array sections addresses the inaccuracies in crossbar arrays by enabling single-cycle error detection and correction, enhancing the efficiency and accuracy of analog computing systems.
Patent Information
- Application Number
- DE102022109137
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-01-20
- Filing Date
- 2022-04-13
- Publication Date
- 2025-07-24
- Estimated Expiration
- 2042-04-13
AI Technical Summary
Crossbar arrays used in analog computing suffer from inaccuracies due to programming errors and noise, which can lead to significant computational errors that are not effectively detected and corrected in existing systems, limiting their adoption in applications requiring high accuracy.
A circuit implementation using three crossbar array sections - one for the target calculation matrix, one for the encoder matrix, and one for the decoder matrix - enables single-cycle analog error detection and correction, allowing for efficient detection and correction of errors in vector matrix multiplication without the need for a separate look-up table, and provides dynamic control over error detection to save energy and reduce latency.
The solution enhances the efficiency of analog error detection and correction processes by enabling single-cycle error detection, reducing chip area and power consumption, and improving the accuracy of computations in crossbar-based matrix multiplication accelerators.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
background
[0001] Vector matrix calculations are performed in many applications such as data compression, neural networks, encryption, etc. Hardware techniques for optimizing vector matrix calculations include application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), and, more recently, an analog dot product calculator based on a crossbar array. US 2021 / 0 209 456 A1 describes embodiments of an algorithm and apparatus for depositing an amount of charge on a floating gate of a non-volatile memory cell within a vector matrix multiplication array in an artificial neural network. It is an object of the invention to propose an analog error correction circuit, an analog error detection and correction method, and a non-transitory, computer-readable medium for analog error detection and correction in analog in-memory crossbars.This object is achieved by an analog error correction circuit according to the invention according to claim 1, an error detection and correction method according to the invention according to claim 9 and a non-transitory, computer-readable medium according to the invention according to claim 16. Brief description of the drawings
[0002] The present disclosure will be described in detail in accordance with one or more different aspects with reference to the following figures. The figures are for illustrative purposes only and represent merely typical or exemplary aspects of the disclosed technology. Fig. Figure 1 shows a circuit implementation for analog error detection and correction according to aspects of the disclosed technology. Fig. 2 schematically shows the acquisition of a measured syndrome vector as the output of a decoding matrix, wherein the values of the measured syndrome vector indicate whether a calculation error is present in a matrix multiplication output, according to aspects of the disclosed technology. Fig. 3 schematically shows an inverse matrix multiplication operation between a thresholded syndrome vector and a decoding matrix to obtain an inverse matrix multiplication output vector indicating a location of a computation error in a matrix multiplication output, according to aspects of the disclosed technology. Fig. 4 is a flowchart of an illustrative method for programming a crossbar-based analog error detection and correction circuit and using the programmed circuit to detect and correct computation errors in a matrix multiplication output in accordance with aspects of the disclosed technology. Fig. 5 is an example of a computer component that may be used to implement various features of the disclosed technology.
[0003] The figures are not exhaustive and do not limit the present disclosure to the precise form disclosed. Detailed description
[0004] Analog computing devices based on crossbar arrays have proven efficient for a number of applications. A crossbar array, as used herein, refers to an array comprising a set of row lines and a set of column lines that intersect the row lines to form an intersection at each intersection point, with a corresponding memory device connected to each intersection and programmed with an array value. The memory devices connected to the junctions can be memristors or devices using other non-volatile resistive memory technologies, such as flash memory, magnetic random access memory (MRAM) with spin-transfer torque (STT) or magnetic tunnel junctions (MTJ), phase-change memory, or similar.An input value along each row of the crossbar array is weighted by the matrix values in each column and accumulated as the output from each column to form a dot product. An analog memory crossbar array can therefore also be referred to as a dot product engine (DPE).
[0005] Although crossbar arrays are efficient for analog computations, inaccuracies in programming the memory devices within the crossbar array and noise when reading the output can affect the accuracy of the computations. Furthermore, inaccuracies can also occur at the junctions within the crossbar array if they become short-circuited, unprogrammable, or stuck in an open state. While memristor or other non-volatile resistive analog memory crossbar-based matrix multiplication accelerators are promising, their tendency to generate unexpected computational errors can limit their ability to replace traditional digital systems. While various analog computing applications tolerate small errors (e.g., neural networks), significant outliers can be problematic if not detected and corrected.
[0006] A technique for detecting / correcting errors in analog computations was described in commonly filed U.S. Application No. 16 / 429,983. This technique provides an error correction code (ECC) for analog computations that enables the detection and correction of computation errors when using DPEs for vector-matrix multiplication. Aspects of the technology disclosed herein relate to a circuit implementation for the analog ECC described in the above-referenced commonly filed application.According to aspects of the disclosed technology, a first crossbar array section is programmed with target calculation matrix values, a second crossbar array section is programmed with encoder matrix values for correcting calculation errors in a matrix multiplication result output of the first crossbar array section, and a third crossbar array section is programmed with decoder matrix values for detecting the calculation errors.
[0007] In some aspects, the circuit disclosed herein connects the first crossbar array portion encoding the target computation matrix and the second crossbar array portion encoding the encoder matrix to the third crossbar array portion encoding the decoder matrix, thus enabling single-cycle analog error detection. The first, second, and third crossbar array portions may be part of one or more crossbar arrays. For example, a first crossbar array of memristors may include the first and second crossbar array portions, and a second crossbar array connected to the first crossbar array may include the third crossbar array portion. Alternatively, the circuit may include a single crossbar array comprising the first, second, and third crossbar array portions.
[0008] Aspects of the disclosed technology also relate to an analog error correction circuit implementation that enables a reverse DPE operation, e.g., an inverse matrix multiplication involving the third crossbar array portion that encodes the decoder matrix. The reverse DPE operation enables the location of a computational error in a matrix multiplication output to be determined. More specifically, the third crossbar array portion can receive as input a matrix multiplication output from the first and second crossbar array portions and itself output a syndrome vector indicating whether computational errors are present in the matrix multiplication output.
[0009] As described in more detail later in this disclosure, in some aspects, each column in the decoder matrix may contain two non-zero values, such that a calculation error may be detected when the syndrome vector contains two out-of-threshold values, where an out-of-threshold value is a value above a first threshold or below a second threshold (the second threshold may be a negative of the first threshold). Further, in some aspects, the measured syndrome vector may be thresholded (i.e., converted to binary values) and fed back to the third crossbar array portion to perform the inverse matrix multiplication with the decoder matrix encoded in the third crossbar array portion. In some aspects, the position of a particular expected value in the resulting inverse matrix multiplication output vector may indicate the position of the calculation error.
[0010] The circuit implementation disclosed herein provides a technical solution to the technical problem of implementing the aforementioned analog ECC. This technical solution offers several technical improvements / advantages for analog error detection technology. These include the ability to perform analog error detection in a single cycle, which improves the efficiency of the analog error detection / correction process and is enabled by the circuit connection of the first and second crossbar array portions, which encode the encoding matrix and the calculation matrix, respectively, with the third crossbar array portion, which encodes the decoding matrix.Furthermore, the encoder and decoder matrices can be selectively disabled, reused, and / or rerouted between crossbar arrays to dynamically control the use of error detection and save power and / or reduce latency overhead when error detection is not desired or when error detection is only required for a portion of the data in a computation stack. Furthermore, an analog error correction implementation according to example aspects of the disclosed technology, incorporating a reverse DPE operation, offers the technical advantage of eliminating the need for a separate lookup table for error location calculation, resulting in further savings in chip area and power.
[0011] Fig. 1 shows an example implementation of an analog error detection circuit 100 according to aspects of the disclosed technology. The circuit 100 includes various crossbar array portions that are part of one or more analog crossbar arrays. The circuit 100 includes a first crossbar array portion including a circuit 102 programmed with analog values representing values of an actual target calculation matrix A' for performing a vector-matrix multiplication. The circuit 100 further includes a second crossbar array portion including a circuit 104 programmed with the values of an encoder matrix A" used to correct errors in a first output of the first crossbar array portion, the first output being a multiplication result of an input vector u with the calculation matrix A'.
[0012] More specifically, circuit 100 includes a crossbar array of ℓ row lines and n column lines that intersect to form crosspoints, with a memory device (e.g., a memristor) connected to each crosspoint and encoding a matrix value. Thus, the crossbar array includes ℓ × n memory locations. The ℓ rows of the crossbar array can receive an input signal as a vector of length ℓ. The n columns can then output an output signal as a vector of length n that is the dot product of the input signal and the matrix values ℓ × n memory locations encoded in the memory locations.
[0013] In some aspects, an ℓ × n matrix A can be programmed into the crossbar array containing the ℓ rows and n columns. The matrix A can have the following structure: A = (A'|A''), where A' is the target computation matrix (a ℓ × k matrix containing a first k columns of A) and A'' is the encoder matrix (a ℓ × m matrix containing the remaining m = n - k columns of the matrix A). More specifically, a set of k columns of the n columns (where k < n) of the crossbar array (i.e., the k columns that form the circuit 102) can be programmed with the values of the target computation matrix A'. Further, a second set of m (e.g., n - k) columns (i.e., the circuitry 104) can be programmed with continuous analog values corresponding to the values of the encoder matrix A''.Each row of the encoder matrix A'' can be determined from a corresponding row of the computation matrix A' so that computation errors above a threshold error value can be detected and corrected in an output vector y which is a result of matrix multiplication of the input vector u with the matrix A.
[0014] As in Fig. 1, each memory device connected to a row line and a column line at a respective crossover point within circuit 102 may be a memristor 106 or other suitable non-volatile resistive memory device. Similarly, each memory device connected to a row line and a column line at a respective crossover point within circuit 104 may be a memristor 108 or other suitable non-volatile resistive memory device. Thus, in example aspects, the conductance values of memristors 106 may be tuned to represent the values of the computation matrix A', and the conductance values of memristors 108 may be tuned to represent the values of the encoder matrix A''.
[0015] In operation, a set of digital-to-analog converters (DACs) (not shown) may be provided to convert an input signal representing an input vector of digital values into a set of corresponding voltages, which are then applied to the set of row lines shared between the first crossbar array portion (i.e., circuit 102) and the second crossbar array portion (i.e., circuit 104). More specifically, each element u i of an input vector u is fed into a corresponding DAC to generate a corresponding voltage level proportional to u iA first output signal from the first crossbar array portion may represent a first output vector c' = uA', which is a result of the desired matrix multiplication calculation between the input vector u and the calculation matrix A'. More specifically, each element of the first output vector c' may be the dot product of the input vector u with a respective one of the k columns of the circuit 102 encoding the calculation matrix A'.
[0016] Similarly, a second output signal from the second crossbar array portion may represent a second output vector c'' = uA'', which is the result of a matrix multiplication operation between the input vector u and the encoder matrix A''. More specifically, each element of the second output vector c'' may be the dot product of the input vector u with a respective one of the n-k columns of the circuit 104 that encodes the encoder matrix A''. The second output vector c'' may include redundancy symbols that allow correction of calculation errors in c'. The first and second output signals may be determined (i.e., the first and second output vectors may be calculated) by reading the currents at analog current measuring devices 110, which may be grounded column conductors such as transimpedance amplifiers.
[0017] In examples, the values of the encoder matrix A'' depend on the values of the computation matrix A', but not on the values of the input vector u. In further examples, the size of the encoding matrix A'' with n - k columns may depend on the alphabet size q (i.e., the number of levels or information bits programmed into each memristor cell 106 in the first crossbar array part that encodes the computation matrix A'); the desired number of errors τ that the encoder matrix A'' can correct; and a desired error correction capability (i.e., a threshold error value Δ to delimit a detectable / correctable error).
[0018] In general, the actually measured output vector y from the first and second crossbar array parts (i.e., circuits 102 and circuits 104) can be represented as follows: y = c + ε + e, where c represents an ideal calculation result vector of the matrix multiplication of the input vector u with the matrix values of both the calculation matrix A' and the encoder matrix A'' (e.g., c can be the concatenation of the first output vector c' and second output vector c''); ε is a tolerable analog inaccuracy; and e is the undesired error. More precisely: -δ < ε < δ, and thus can be a limited tolerable inaccuracy, while e represents a detectable and correctable error when e > Δ or e < -Δ and Δ > δ. In particular, e can be a vector whose non-zero entries are misses. In some aspects, a deviating error may be an error that is greater than δ or less than -δ. Such an error (ieAn error greater than δ or less than -δ) may be detected but not corrected if it does not exceed the threshold error value (i.e., is within [-Δ, Δ]). An aberrant error that exceeds the threshold (i.e., an error greater than Δ or less than -Δ), on the other hand, may be an intolerable error that must be corrected. In some aspects, both δ and Δ are preset thresholds, with the ratio between them being a tunable parameter. In some aspects, the ratio of δ to Δ (or vice versa) may be adjusted to reduce the difference between δ and Δ. In one example, the ratio may be 1, so that if an error is detectable, it is also correctable. However, this may increase the number of redundancy columns in the encoder matrix A''.
[0019] Various types of analog computation errors can occur, including transient errors (i.e., errors that will not recur predictably), permanent errors, and intrinsic errors. Examples of transient errors include, but are not limited to, read glitches from peripheral circuitry, memristor conductance fluctuations, and the like. While transient errors can be detected and corrected using circuit 100 implementing analog ECC, in some scenarios, as an alternative to correcting the error, the computation may simply be rerun because the error is unlikely to recur, being by definition a transient error. Permanent errors include, but are not limited to, conductance drift, open / shorted connections and / or wires, and the like. In examples, permanent errors are at least detected and, in some cases, can be corrected.In other scenarios, the DPE block experiencing the permanent error may instead be reprogrammed or replaced entirely. Intrinsic errors include, but are not limited to, inaccurate programming, line resistance, device IU nonlinearity, and the like. In examples, intrinsic errors may be treated as part of the "normal" operation of a crossbar-based analog matrix multiplication accelerator and ignored. In other examples, circuit 100 and the analog ECC it implements may be employed, where possible, to correct the errors to increase computational accuracy and / or intentionally tune the randomness of the result when intrinsic errors are detected.
[0020] As in Fig. 1, circuit 100 also includes a third crossbar array portion that encodes the values of a decoder matrix. More specifically, circuit 114 of the third crossbar array portion is programmed with the values of a decoder matrix / parity check matrix. Similar to the first and second crossbar array portions, the third crossbar array portion (i.e., circuit 114) includes intersecting row and column lines, with a corresponding memory device (e.g., a memristor 118) connected to each intersection between a corresponding row and column line. The first and second crossbar array portions (i.e., circuits 102 and 104) may be connected to the third crossbar array portion (i.e., circuit 114) to enable single-cycle analog error detection.That is, the current outputs of the first and second crossbar array portions (which include an output vector c comprising c' (the desired calculation result) and c'' (the redundancy symbols for correcting an error in the desired calculation result)) may be fed directly to the third crossbar array portion (i.e., circuit 114) without any processing, such as analog-to-digital conversion. In some aspects, the column lines of circuit 114 may be the same as the column lines of circuit 102 and circuit 104. Thus, circuits 102, 104, and 114—representing the first, second, and third crossbar array portions, respectively—may form a single crossbar array.
[0021] In examples, the values of the decoder / parity check matrix may conform to various rules imposed by the analog ECC depending on the number of errors to be detected and / or corrected. For example, to detect a single errant error, given positive integers r and n such that r ≤ n, then the decoder / parity check matrix H is an r×n matrix over {0,1} that satisfies the following properties: 1) each column in H is a unit vector, meaning each column contains exactly one 1, and 2) the number of 1s in each row is either the nearest integer below n / r or the nearest integer above n / r.On the other hand, to detect and correct a single error, for positive integers r and n such that n ≤ r(r - 1), then the decoder / parity check matrix H is an r × n matrix over {-1,0,1} that satisfies the following properties: 1) all columns of H are distinct, 2) each column in H contains exactly two non-zero entries, the first of which is a 1, and 3) the number of non-zero entries in each row of H is the nearest integer less than 2n / r or the nearest integer greater than 2n / r.
[0022] In example aspects, any value in the decoder / parity check matrix H that satisfies the properties described above for correcting a single, aberrant error can be encoded by the conductances of an adjacent pair of memristors 118 in the circuit 114 (i.e., adjacent memristors 118 in the same column). More specifically, a 1 value in the decoder matrix can be mapped to an adjacent pair of memristors 118 having the matched conductances (low resistance state (LRS), high resistance state (HRS)), while a -1 value in the decoder matrix can be mapped to (HRS, LRS), and a 0 value can be mapped to (HRS, LRS).
[0023] If the calculation error in the result of the matrix multiplication 112 (which is the result of a matrix multiplication of the input vector by the u with the calculation matrix A') is within a preset tolerance value (e.g., the threshold error value Δ), all outputs of the circuit 114 programmed with the decoder matrix values would be below a threshold determined by comparators between adjacent outputs. In examples, this is a result of the properties followed by the decoder / parity check matrix. In particular, as in Fig. 1, transimpedance amplifiers 120 may be provided that virtually ground the adjacent row wires in the decoder matrix circuit 114 and convert the currents into voltages. In particular, the voltage output of each amplifier 120 may be - I in * R ref where I represents the input current to amplifier 120 and R refis a reference voltage. The comparator 122 can then compare these converted voltages to determine whether a calculation error is within the error threshold Δ.
[0024] More specifically, the threshold error value Δ in comparator 122 can be implemented as a tunable parameter or as an input terminal (not shown), depending on the design. Although for the sake of clarity in Fig. 1, it should be understood that corresponding amplifiers 120 and a corresponding comparator 122 may be provided for each pair of adjacent row outputs of the decoder matrix circuit 114. That is, a different set of amplifiers 120 may be provided to convert the current outputs of rows 3 and 4 of the decoder matrix circuit 114 into corresponding voltages, which may then be input to another comparator 122 to determine whether a calculation error present in the matrix multiplication result 112 is within the error threshold Δ. The same may apply to rows 5 and 6 of the decoder matrix circuit 114, rows 7 and 8 of the circuit 114, and so on.
[0025] If a calculation error is within the error threshold Δ, the comparator 122 may output a low logic signal, which may then be inverted to a high logic signal and provided as an input to an AND gate 124. If the calculation error is outside the threshold error value (i.e., either greater than Δ or less than -Δ), the comparator 122 may output a high logic signal, which may be inverted to a low logic signal and provided to the AND gate 124. In some aspects, the result of the matrix multiplication 112 may be validated if each comparator 122 determines that the difference between the respective adjacent outputs of the decoder matrix circuit 114 that it is configured to compare is less than the threshold error value Δ. In this case, each comparator 122 outputs a logic high signal to the AND gate 124, and the result of the matrix multiplication 112 is validated.On the other hand, if one or more comparators 122 output a logic "low" signal indicating a calculation error that is outside the threshold (either greater than Δ or less than -Δ), then these comparators 122 would output a logic "low" signal to the AND gate 124, and the AND gate 124 would in turn output a logic "low" signal, resulting in the result of the matrix multiplication 112 not being validated.
[0026] As mentioned above, various types of analog computation errors may occur. In the case of a transient error, the error may be detected and corrected by the circuit implementation 100 disclosed herein and the analog ECC implemented by it, or alternatively, the computation error may be corrected by simply re-reading the computation results after the error has been detected based on the output of the decoder matrix circuit 114. On the other hand, if a computation error is detected during a first read operation and then during a second read operation (or a third read operation, where n ≥ 2), a device controller (which may be a particular implementation of a Fig. 5) may mark the calculation error as permanent, and the crossbar array may be replaced or reprogrammed. In some cases, the calculation error may undergo a certain number of corrections and be classified as permanent only if it undergoes a certain number of corrections again.
[0027] If the goal is to detect permanent errors (i.e., errors that occur every time after the first occurrence of an error), then in some examples, to reduce power consumption, the decoder matrix circuit 114 may be enabled only for a last vector input in a stack of inputs. This may be achieved by selectively applying a decoder signal 116 to enable the decoder matrix circuit 114. Furthermore, in some examples, if the computation matrix A' is duplicated into multiple copies (i.e., encoded into multiple crossbar array parts)—such as in the case of convolution kernels in convolutional neural networks (CNNs) for pixel-level parallelism—the decoder matrix circuit 114 may be divided among the different copies and enabled for one copy at a time.
[0028] Fig. 2 schematically illustrates the acquisition of a measured syndrome vector 212 as output from the decoder matrix A'', wherein values of the measured syndrome vector 212 indicate whether a calculation error is present in a matrix multiplication output 202, according to aspects of the disclosed technology. In example aspects, the matrix multiplication output 202 may include the Fig. 1, the currents read by the analog current measuring devices 110 may be the matrix multiplication operation between an input vector u and the matrix A, which comprises the calculation matrix A' and the encoder matrix A''. Fig. 2 shows a decoding matrix representation 204 containing a collection of pixels, where each pixel represents a conductance difference between two adjacent memristors 118 of the circuit 114 encoding the decoding matrix. In the decoding matrix representation 204, a pixel 206 may represent a conductance difference G HRS - G LRSwhich corresponds to the (HRS, LRS) representation of a -1 value in the decoder matrix; the pixel 208 can represent a conductance difference G LRS - G HRS which corresponds to the (LRS, HRS) representation of a 1-value in the decoder matrix; and the pixel 210 can represent a conductance difference G HRS - G HRS which corresponds to the (HRS, HRS) representation of a 0 value in the decoder matrix. Since the values of the decoder matrix are programmed into circuit 114 in the analog domain, it should be understood that the representation of the decoder matrix 204 may contain a range of conductance differences between adjacent memristors corresponding to a real number range between -1 and 1.
[0029] In examples of the disclosed technology where no calculation error lies outside the threshold error value Δ, all values of the output of the decoder matrix circuit—the syndrome vector 212—are below a threshold value. On the other hand, if a DPE output of the matrix multiplication output 202 is impaired (i.e., has a calculation error that exceeds the threshold error value Δ), the corresponding values of the decoder matrix are added to the syndrome vector 212. In example aspects, due to the predefined properties of the decoder / parity check matrix, each column of the decoder matrix is unique and contains exactly two non-zero entries. Thus, in the scenario where a single calculation error lies outside the threshold, there are exactly two values above the threshold in the syndrome vector 212. The absolute value of the calculation error that adds up to the syndrome vector 212 can be given by (G LRS - G HRS ) * Verr , where V err is the error to be corrected. The location of the detected calculation error can then be determined by an inverse matrix multiplication between a thresholded version of the syndrome vector 212 and the decoder matrix.
[0030] Fig. 3 schematically illustrates an inverse matrix multiplication between a thresholded syndrome vector 302 and the decoding matrix to obtain an inverse matrix multiplication output vector 306 indicating the location of a computation error in a matrix multiplication output (e.g., matrix multiplication output 202), according to aspects of the disclosed technology. The measured syndrome vector 212 may be thresholded to obtain the threshold syndrome vector 302 by converting the values of the syndrome vector 212 to binary values. In particular, any value in the syndrome vector 212 that is above a first threshold may be converted to a 1, and any value in the syndrome vector 212 that is below a second threshold may be converted to a -1. In the case of a single error correction, there are two such values. In some aspects, the first and second thresholds may be additively inverse.Any value in the syndrome vector 212 that lies between the first and second thresholds (which would be a zero value in the case of single error correction) becomes a binary 0.
[0031] In some aspects, the threshold syndrome vector 302 may be fed back to the third crossbar array portion (i.e., the decoder matrix circuit 114) to perform an inverse matrix multiplication operation. In particular, the values of the threshold syndrome vector 302 may be mapped to corresponding voltage pairs, and these voltage pairs may be applied to the row lines of the decoder matrix circuit 114 to generate an output signal representing the inverse matrix multiplication output vector 306. More specifically, a binary 1 in the threshold syndrome vector 302 (corresponding to a value in the syndrome vector 212 that is above the first threshold) may be mapped to (V r , -V r), while a binary -1 in the thresholded syndrome vector 302 (corresponding to a value in the syndrome vector 212 that is below the second threshold) is mapped to (-V r , V r ).
[0032] In example aspects, due to the properties of the decoder / parity check matrix when performing the inverse matrix multiplication, there is only one output (corresponding to a specific value of the inverse matrix multiplication output vector 306) whose current is equal to 2V r - (G LRS - G HRS ) and is output as the only matching element multiplication. In particular, (V r , -V r ) multiplied by (G LRS - G HRS ) and (-Vr, Vr) multiplied by (G LRS - G HRS) are the only results that yield a 1, and there is only one column in the decoder matrix with two matches that produce this result based on the properties of the decoder / parity check matrix.
[0033] As already mentioned, there is only one output for the multiplication of the inverse matrix, which generates a current equal to 2V r - (G LRS - G HRS ). Therefore, the position of a particular value in the output vector 306 of the inverse matrix multiplication, which for 2V r - (G LRS - G HRS) corresponds to the location of the calculation error. In particular, the position of this particular value in the output vector 306 corresponds to a particular column of the decoder matrix and thus to a particular column of the decoder matrix circuit 114. Due to the connection between the first and second crossbar array parts (i.e., circuits 102 and 104) and the third crossbar array part (i.e., circuit 114), the particular column of the decoder matrix circuit 114 corresponds to a column in the circuit 102 that contains the calculation error (e.g., it is the same column as this).
[0034] In example aspects, the measured syndrome vector 212 may be used to determine the actual value of the error to be corrected. In particular, in example aspects, the value of the error to be corrected may be determined by the mean absolute value of the two off-threshold values in the measured syndrome vector 212, which is approximately (G LRS- G HRS ) * V err . In further examples, the same decoder matrix circuit 114 may be used to perform error correction.
[0035] Fig. 4 is a flowchart of an illustrative method 400 for programming a crossbar-based analog error detection and correction circuit and using the programmed circuit to detect and correct computation errors in a matrix multiplication output, according to aspects of the disclosed technology. In some aspects, the method 400 may be performed in response to one or more processing units (e.g., Fig. 5, processor(s) 504, a controller, etc.) that execute machine / computer executable instructions stored in a storage device such as main memory 506, read-only memory (ROM) 512, and / or memory 514 ( Fig. 5). In some aspects, the method 400 may be implemented at least in part by hard-wired logic in a crossbar-based matrix multiplication hardware accelerator 508 ( Fig. 5). One or more operations of method 400 may be described below as being performed by a crossbar-based matrix multiplication hardware accelerator (e.g., hardware accelerator 508) that, for example, implements the Fig. 1 comprises circuit 100.
[0036] In block 402 of method 400, a first crossbar array portion may be programmed with analog values corresponding to the values of a target calculation matrix A' for a desired matrix multiplication with an input vector u. For example, each memristor 106 of circuit 102 ( Fig. 1) be programmed with a tuned conductance representing a corresponding value in the calculation matrix A'.
[0037] In block 404 of the method 400, a second crossbar array part can be programmed with analog values, the values of an encoder matrix A'' which is a redundancy symbol for performing error correction of a calculation result of a matrix multiplication of an input vector u with the calculation matrix A'. For example, each memristor 108 of the circuit 104 ( Fig. 1) be programmed with a tuned conductance value corresponding to a corresponding value in the encoder matrix A''.
[0038] In block 406 of method 400, a third crossbar array portion may be programmed with analog values representative of values in a decoder / parity check matrix H. As previously mentioned, the decoder / parity check matrix H may maintain various properties, which may include, without limitation, in the case of single error correction, the uniqueness of each column of the matrix H. Each memristor 118 of the circuit 114 ( Fig. 1) may be programmed with a tuned conductance value that is representative of a corresponding value in the decoder / parity check matrix H. In some embodiments, the first, second, and third crossbar array portions may be part of the same crossbar array or multiple crossbar arrays that are interconnected.
[0039] In block 408 of method 400, an input vector u may be converted into a set of corresponding voltages, and the voltages may be applied to common row lines of the first and second crossbar array portions. More specifically, each element u i of an input vector u is fed into a corresponding DAC to generate a corresponding voltage level proportional to u i .
[0040] In block 410 of method 400, a result of a vector-matrix multiplication between the input vector u and the target calculation matrix A' programmed into the first crossbar array section is calculated. More specifically, a first output signal from the first crossbar array portion may represent a first output vector c' = uA' representing a result of the desired matrix multiplication calculation between the input vector u and the calculation matrix A'. Similarly, a second output signal from the second crossbar array portion may represent a second output vector c'' = uA'' representing the result of a matrix multiplication operation between the input vector u and the encoder matrix A''. The second output vector c'' may contain redundancy symbols that allow correction of calculation errors in c'. The first and second output signals may be determined, i.e.the first and second output vectors representing the matrix multiplication output of the first and second crossbar array parts can be calculated by reading the currents collected at the column lines.
[0041] In block 412 of method 400, the current outputs, which represent the result of the matrix multiplication of the input vector u by the calculation and encoding matrices programmed into the first and second crossbar array sections, respectively, may be fed as input to the third crossbar array section, in which the decoding matrix values are programmed. Then, in block 414 of method 400, a measured syndrome vector may be obtained as output from the third crossbar array section, wherein the measured syndrome vector represents an output of the decoder matrix based on the received matrix multiplication result input.
[0042] In block 416 of method 400, the measured syndrome vector may be thresholded to obtain a thresholded syndrome vector. As previously mentioned, thresholding the measured syndrome vector may include converting values of the syndrome vector into binary values. More specifically, thresholding may include converting a value in the measured syndrome vector that is above a first threshold to a binary 1 and converting a value in the measured syndrome vector that is below a second threshold to a binary -1.
[0043] In block 418 of method 400, assuming that the syndrome vector indicates the presence of a calculation error that exceeds a threshold error value, the thresholded syndrome vector may be fed back to the third crossbar array portion to perform an inverse matrix multiplication operation to determine the location of the calculation error in the result of the matrix multiplication. The location may be determined in the manner previously described. Then, in block 420 of method 400, a calculation error value for correction may be determined from the off-threshold elements of the measured syndrome vector, as previously described.
[0044] Fig. Figure 5 shows a block diagram of an example computer system 500 in which various aspects of the technology described herein may be implemented. Computer system 500 includes a bus 502 or other communication mechanism for conveying information and one or more hardware processors 504 coupled to bus 502 for processing information. For example, hardware processor(s) 504 may be one or more general-purpose microprocessors.
[0045] Computer system 500 also includes main memory 506, such as random access memory (RAM), a cache, and / or other dynamic storage devices, connected to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When such instructions are stored in storage media accessible to processor 504, computer system 500 becomes a special-purpose machine adapted to perform the operations specified in the instructions.
[0046] The computer system 500 additionally includes a crossbar-based matrix multiplication hardware accelerator 508 that Fig.1 that implements analog ECC. In some aspects, the hardware accelerator 508 may be configured to execute instructions (i.e., programming or software code) stored in main memory 506, ROM 512, and / or storage 514. In one example implementation, the example hardware accelerator 508 may include multiple integrated circuits, which in turn may include ASICs, FPGAs, or other Very Large Scale Integrated Circuits (VLSIs). The integrated circuits of the example hardware accelerator 508 may be specifically optimized to perform a discrete subset of computer processing operations or to execute a discrete subset of computer-executable instructions in an expedited manner.For example, the hardware accelerator 508 may be configured or manufactured to perform analog crossbar-based vector-matrix multiplication and analog error detection and correction.
[0047] The circuit 100, which may be implemented in the accelerator 508, may include non-volatile memory constructed using technologies including, for example, resistive switching memory (i.e., a memristor), phase-change memory, magnetoresistive memory, ferroelectric memory, another resistive random access memory (Re-RAM), or combinations of these technologies. More generally, the circuit 100 may be implemented using technologies that allow the circuit 100 to retain its contents even when power is interrupted or otherwise removed. In this way, the data in the circuit 100 "persists," and the circuit 100 may function as what is known as "non-volatile memory."
[0048] Computer system 500 also includes a read-only memory (ROM) 512 or other static storage device connected to bus 502 for storing static information and instructions for processor 504. A storage device 514, such as a magnetic disk, an optical disk, or a USB flash drive, etc., is provided and connected to bus 502 for storing information and instructions.
[0049] Computer system 500 may be connected to a display 516, such as a liquid crystal display (LCD) (or a touchscreen), via bus 502 to display information to a computer user. An input device 518, which may include alphanumeric and other keys, is coupled to bus 502 to convey information and enable command selection to processor 504. Another type of user input device is cursor control 520, such as a mouse, trackball, or cursor direction keys, for conveying direction information and command selections to processor 504 and controlling cursor movement on display 516. In some aspects, the same direction information and command selections as provided by cursor control 520 may be implemented via receiving touches on a touchscreen without a cursor.
[0050] Computer system 500 may include a user interface module for implementing a graphical user interface, which may be stored on a mass storage device as executable software code executed by the computing device(s). This and other modules may include, for example, various components, such as software components, object-oriented software components, class components, and task components, including, but not limited to, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.
[0051] In general, the terms “component,” “engine,” “system,” “database,” “data store,” and the like, as used herein, may refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and be written in a programming language such as Java, C, or C++. A software component may be compiled and linked into an executable program, installed in a dynamic link library, or written in an interpreted programming language such as BASIC, Perl, or Python. It is understood that software components may be callable by other components or by themselves, and / or may be called in response to detected events or interrupts. Software components configured to run on computing devices may be embodied on a computer-readable medium, such as a hard disk.a compact disc, digital video disc, flash drive, magnetic disk, or other tangible medium, or as a digital download (and may be originally stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code may be stored partially or entirely in a memory of the executing computing device so that it can be executed by the computing device. Software instructions may be embedded in firmware, such as an EPROM. In addition, the hardware components may consist of connected logic units such as gates and flip-flops and / or programmable units such as programmable gate arrays or processors.
[0052] Computer system 500 may implement the techniques described herein using custom hard-wired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, makes or programs computer system 500 into a special-purpose machine. In one aspect, the techniques described herein are performed by computer system 500 in response to processor(s) 504 executing one or more sequences of one or more instructions contained in main memory 506. Such instructions may be read into main memory 506 from another storage medium, such as storage device 510. Execution of the instruction sequences contained in main memory 506 causes processor(s) 504 to perform the process steps described herein.In alternative aspects, hard-wired circuits may be used instead of or in combination with software instructions.
[0053] The term "non-transitory media" and similar terms such as machine-readable storage media, as used herein, refer to any media that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transferable media may include non-volatile media and / or volatile media. Examples of non-volatile media include optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of non-volatile media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or other magnetic data storage media, CD-ROMs, other optical data storage media, physical media with hole patterns, RAM, PROM and EPROM, FLASH EPROM, NVRAM, other memory chips or cartridges, and networked versions thereof.
[0054] Non-transitory media are distinct from transmission media but can be used in conjunction with them. Transmission media are involved in the transfer of information between non-transitory media. Examples of transmission media include coaxial cables, copper cables, and fiber optic cables, including the wires that make up bus 502. Transmission media can also take the form of sound or light waves, such as those generated in data communications via radio and infrared.
[0055] Computer system 500 also includes a communications interface 522 connected to bus 502. Communications interface 522 establishes a bidirectional data communications connection to one or more network connections connected to one or more local area networks. For example, communications interface 522 may be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for establishing a data communications connection to a corresponding type of telephone line. As another example, communications interface 522 may be a LAN card for establishing a data communications connection to a compatible LAN (or a WAN component for communicating with a WAN). Wireless connections may also be implemented.In each of these implementations, the communication interface 522 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams containing various types of information.
[0056] A network connection typically enables data communication across one or more networks to other data devices. For example, a network connection may establish a connection across a local area network to a host computer or data devices operated by an Internet service provider (ISP). The ISP, in turn, provides data communication services across the worldwide packet data communications network, now commonly referred to as the "Internet." Both the local area network and the Internet use electrical, electromagnetic, or optical signals that carry digital data streams. The signals across the various networks and the signals on the network connection and across the communications interface 522 that carry the digital data to and from the computer system 500 are examples of transmission media.
[0057] Computer system 500 can send messages and receive data, including program code, over the network(s), the network connection, and the communications interface 522. In the Internet example, a server could transmit requested code for an application program over the Internet, the ISP, the local area network, and the communications interface 522. The received code can be executed by processor 504 upon receipt and / or stored in storage device 510 or other non-volatile memory for later execution.
[0058] Each of the processes, methods, and algorithms described in the preceding sections may be embodied in, and fully or partially automated by, code components executed by one or more computer systems or computer processors comprising computer hardware. The one or more computer systems or computer processors may also operate to support the performance of the corresponding operations in a cloud computing environment or as software as a service (SaaS). The processes and algorithms may be partially or fully implemented in application-specific circuitry. The various features and methods described above may be used independently or combined in various ways.Various combinations and subcombinations are intended to be within the scope of this disclosure, and certain method or process blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular sequence, and the associated blocks or states may be performed in other suitable sequences, in parallel, or otherwise. Blocks or states may be added to or removed from the disclosed example aspects. The execution of certain operations or processes may be distributed among computer systems or computer processors located not only on a single machine, but distributed across a number of machines.
[0059] A circuit may be implemented in any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms may be implemented to form a circuit. In implementation, the various circuits described herein may be implemented as discrete circuits, or the described functions and features may be distributed, in part or in whole, among one or more circuits. Even though various features or functional elements are individually described or claimed as separate circuits, those features and functions may be shared by one or more common circuits, and such description is not intended to assume or imply that separate circuits are required to implement those features or functions.If a circuit is implemented in whole or in part with software, that software may be implemented to operate with a computer or processing system capable of performing the functionality described with respect to it, such as computer system 500.
Claims
[1] Analog error correction circuit (100) comprising: one or more non-volatile resistive analog memory crossbar arrays comprising: a first crossbar array portion including a circuit (102) programmed with values of a multiplication matrix (A'); a second crossbar array part for correcting calculation errors, which includes a circuit (104) programmed with values of an encoder matrix (A''), the values of the encoder matrix (A'') depending on the values of the multiplication matrix (A'); and a third crossbar array part for detecting calculation errors, which contains a circuit (114) programmed with values of a decoder matrix, wherein the first crossbar array part and the second crossbar array part are configured to output a first output (c') and a second output (c''), respectively, representing a result (c) of a matrix multiplication operation with an input vector (u), wherein the first and second outputs (c', c'') are provided as inputs for the third crossbar array part, and wherein a third output of the third crossbar array part indicates whether a calculation error is present in the first output (c') of the first crossbar array part, wherein the third output of the third crossbar array part represents a measured syndrome vector (212), and wherein the measured syndrome vector (212) comprises two off-threshold values indicating the presence of the calculation error in the first output (c') of the first crossbar array part, wherein the measured syndrome vector (212) is thresholded to obtain a thresholded syndrome vector (302), and wherein the thresholded syndrome vector (302) is fed back to the third crossbar array part to determine the location of the calculation error. [2] The analog error correction circuit (100) of claim 1, wherein the third output of the third crossbar array portion is used in combination with the values of the decoder matrix to determine a location of the calculation error present in the first output (c') of the first crossbar array portion. [3] The analog error correction circuit (100) of claim 1, wherein a fourth output (306) of the third crossbar array portion represents a result of an inverse matrix multiplication operation between the thresholded syndrome vector (302) and the decoder matrix. [4] The analog error correction circuit (100) of claim 3, wherein the fourth output (306) is an output vector, and wherein a particular element of the output vector has a unique value indicating a location of the calculation error in the first output (c') of the first crossbar array portion. [5] The analog error correction circuit (100) of claim 4, wherein each column of the decoder matrix is unique, and wherein the particular element of the output vector is a result of the inverse matrix multiplication operation between a particular column of the decoder matrix and the thresholded syndrome vector (302). [6] The analog error correction circuit (100) of claim 5, wherein the location of the calculation error is in a particular column of the first crossbar array portion producing a corresponding output provided as an input to a particular column of the third crossbar array portion programmed with values of the particular column of the decoder matrix. [7] The analog error correction circuit (100) of claim 1, wherein a value of the calculation error is an average absolute value of the two off-threshold values in the measured syndrome vector (212). [8] The analog error correction circuit (100) of claim 1, wherein a conductance difference between respective adjacent non-volatile resistive devices in the third crossbar array portion represents a corresponding value of the decoder matrix. [9] An analog error detection and correction method implemented in analog in-memory crossbars, the method comprising: Programming a first crossbar array part of one or more non-volatile resistive analog memory crossbar arrays with values of a multiplication matrix (A'); Programming a second crossbar array portion of the one or more non-volatile resistive analog memory crossbar arrays with values of an encoder matrix (A'') for correcting calculation errors, wherein the values of the encoder matrix (A'') depend on the values of the multiplication matrix (A'); and Programming a third crossbar array part of the one or more non-volatile resistive analog memory crossbar arrays with values of a decoder matrix for detecting calculation errors, Conversion of an input vector (u) into a series of corresponding voltages; Applying the set of corresponding voltages to the first crossbar array part and the second crossbar array part to obtain a first output (c') and a second output (c''), respectively, representing a result (c) of a matrix multiplication operation with the input vector (u); and Providing the first and second outputs (c', c'') as input to the third crossbar array part to obtain a third output of the third crossbar array part indicating whether a calculation error is present in the first output (c') of the first crossbar array part, wherein the third output of the third crossbar array part represents a measured syndrome vector (212) and wherein the measured syndrome vector (212) comprises two off-threshold values indicating the presence of the calculation error in the first output (c') of the first crossbar array part, wherein the measured syndrome vector (212) is thresholded to obtain a thresholded syndrome vector (302), and wherein the thresholded syndrome vector (302) is fed back to the third crossbar array part to determine the location of the calculation error. [10] A method according to claim 11, wherein the third output of the third crossbar array part is used in combination with the values of the decoder matrix to determine the location of the calculation error present in the first output (c') of the first crossbar array part. [11] The method of claim 9, further comprising: Obtaining a fourth output (306) of the third crossbar array portion representing a result of an inverse matrix multiplication operation between the thresholded syndrome vector (302) and the decoder matrix. [12] The method of claim 11, wherein the fourth output (306) is an output vector, and wherein a particular element of the output vector has a unique value indicating a location of the calculation error in the first output (c') of the first crossbar array portion. [13] The method of claim 12, wherein each column of the decoder matrix is unique, and wherein the particular element of the output vector is a result of the inverse matrix multiplication between a particular column of the decoder matrix and the thresholded syndrome vector (302). [14] The method of claim 13, wherein the location of the calculation error is in a particular column of the first crossbar array part producing a corresponding output provided as input to a particular column of the third crossbar array part programmed with values of the particular column of the decoder matrix. [15] The method of claim 9, wherein a conductance difference between respective adjacent non-volatile resistive devices in the third crossbar array portion represents a corresponding value of the decoder matrix. [16] A non-transitory computer-readable medium storing machine-executable instructions that, in response to execution by a processor, cause a method to be performed, the method comprising: Programming a first crossbar array part of one or more non-volatile resistive analog memory crossbar arrays with values of a multiplication matrix (A'); Programming a second crossbar array portion of the one or more non-volatile resistive analog memory crossbar arrays with values of an encoder matrix (A") for correcting calculation errors, wherein the values of the encoder matrix (A") depend on the values of the multiplication matrix (A'); and Programming a third crossbar array part of the one or more non-volatile resistive analog memory crossbar arrays with values of a decoder matrix for detecting calculation errors, Applying a set of voltages representative of an input vector (u) to the first crossbar array part and the second crossbar array part to obtain a first output (c') and a second output (c''), respectively, representing a result (c) of a matrix multiplication operation with the input vector (u); and Providing the first and second outputs (c', c'') as input signals to the third crossbar array part to obtain a third output of the third crossbar array part indicating whether a calculation error is present in the first output (c') of the first crossbar array part, wherein the third output of the third crossbar array part represents a measured syndrome vector (212), and wherein the measured syndrome vector (212) comprises two off-threshold values indicating the presence of the calculation error in the first output (c') of the first crossbar array part, wherein the measured syndrome vector (212) is thresholded to obtain a thresholded syndrome vector (302), and wherein the thresholded syndrome vector (302) is fed back to the third crossbar array part to determine the location of the calculation error.
Citation Information
Patent Citations
Precise data tuning method and apparatus for analog neural memory in an artificial neural network
US20210209456A1