Non-volatile storage device performing multiply-accumulate operations
By designing memory cell arrays and calculation output circuits in nonvolatile memory devices, the problem of inefficient MAC computing in neural networks is solved, and efficient MAC computing performance is achieved.
Patent Information
- Application Number
- CN202010327413.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-04-13
- Filing Date
- 2020-04-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2040-04-23
AI Technical Summary
The prior art is difficult to efficiently perform a large number of multiplication accumulation (MAC) operations in neural networks, resulting in inefficient computing.
Using non-volatile memory devices, including a memory cell array and a calculation output circuit, each storage element in the memory cell array stores weights and controls it according to the input signal, the calculation output circuit generates an internal product signal of the input vector and the weight vector.
It realizes efficient execution of MAC operations, improving the efficiency and performance of neural network computing.
Smart Images

Figure CN112447228B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to Korean Patent Application No. 10 - 2019 - 0109898, filed on September 5, 2019, and Korean Patent Application No. 10 - 2020 - 0044466, filed on April 13, 2020, the entire contents of which are incorporated herein by reference. Technical field
[0003] Various embodiments relate to non - volatile memory devices that perform multiplication and accumulation (MAC) operations. Background art
[0004] Neural networks are widely used in artificial intelligence applications, such as the technologies used in image recognition and autonomous vehicles.
[0005] In an example, a neural network includes an input layer, an output layer, and one or more inner layers between the input layer and the output layer.
[0006] Each of the output layer, the input layer, and the inner layers includes one or more neurons. Neurons included in adjacent layers are connected in various ways through synapses. For example, a synapse points from a neuron in a given layer to a neuron in the next layer. Alternatively or additionally, a synapse points from a neuron in an upper layer to a neuron in a given layer.
[0007] Each neuron stores a value. The value of a neuron included in the input layer is determined according to an input signal (e.g., an image to be recognized). The value of a neuron included in the inner layers and the output layer is based on the neurons and synapses included in the corresponding upper layer. For example, the value of a neuron in each inner layer is based on the value of a neuron in the upper layer in the neural network.
[0008] Each synapse has a weight. The weight of each synapse is based on the training operation of the neural network.
[0009] After the neural network is trained, the neural network can be used to perform inference operations. In the inference operation, the value of a neuron in the input layer is set based on an input, and the value of a neuron in the next layer (e.g., inner layers and output layer) is set based on the value of a neuron in the input layer and the weights of the trained synapses connecting the layers. The value of a neuron in the output layer represents the result of the inference operation.
[0010] For example, in an inference operation, in which image recognition is performed by a neural network after the neural network has been trained, values of neurons in an input layer are set based on an input image, a plurality of operations are performed in an inner layer based on the values of the neurons in the input layer, and a result of the image recognition is output from the inner layer in an output layer.
[0011] In such an inference operation, a large number of MAC operations must be performed by neurons in a convolutional neural network. As a result, a semiconductor device capable of efficiently performing a large number of MAC operations is desired. Summary of the Invention
[0012] According to an embodiment of the present disclosure, a non-volatile storage device may include: a memory cell array including a plurality of non-volatile memory elements and a plurality of bit lines coupled to the plurality of non-volatile memory elements, the plurality of non-volatile memory elements being configured to store a plurality of weights and being controlled respectively according to a plurality of input signals; and a calculation output circuit configured to generate a calculation signal corresponding to an inner product between an input vector and a weight vector, the input vector corresponding to the plurality of input signals, and the weight vector corresponding to the plurality of weights. Brief Description of the Drawings
[0013] The drawings (wherein like reference numerals refer to the same or functionally similar elements in each separate view) and the following detailed description are incorporated in and form a part of the specification, and are used to further illustrate embodiments including various features and to explain various principles and advantageous aspects of these embodiments.
[0014] Figure 1 A flash memory device according to an embodiment of the present disclosure is shown.
[0015] Figure 2 A NAND string according to an embodiment of the present disclosure is shown.
[0016] Figure 3A and Figure 3B A calculation operation using flash memory cells according to an embodiment of the present disclosure is shown.
[0017] Figure 4 An operation of an input circuit according to an embodiment of the present disclosure is shown.
[0018] Figure 5 A first input circuit according to an embodiment of the present disclosure is shown.
[0019] Figure 6 An output circuit according to an embodiment of the present disclosure is shown.
[0020] Figure 7Shows a computing output circuit according to an embodiment of the present disclosure.
[0021] Figure 8 Shows a timing diagram illustrating the operation of a conversion circuit according to an embodiment of the present disclosure.
[0022] Figure 9 Shows a timing diagram illustrating the computing operation of a flash memory device according to an embodiment of the present disclosure. Detailed Description
[0023] Various embodiments will be described below with reference to the accompanying drawings. The embodiments are provided for illustrative purposes, and other embodiments that are not explicitly shown or described are also possible. In addition, the embodiments of the present disclosure that will be described in detail below can be modified.
[0024] In the following disclosure, a non-volatile storage device is disclosed using a flash memory device as an example, but the type of non-volatile storage device that can be used need not be limited to a flash memory device.
[0025] Figure 1 Is a block diagram showing a flash memory device 1 according to an embodiment of the present disclosure.
[0026] The flash memory device 1 includes a flash memory cell array 100, an input circuit 200, an output circuit 300, a command decoder 400, and a calibration circuit 500.
[0027] The command decoder 400 controls read operations, program operations, and erase operations in the same manner as those performed in a command decoder included in a conventional flash memory device.
[0028] In this embodiment, the command decoder 400 also additionally performs control operations required for computing operations.
[0029] The flash memory device according to this embodiment has a storage operation mode and a computing operation mode.
[0030] In the storage operation mode, the operations of a normal flash memory device are performed. In the computing operation mode, MAC operations are performed.
[0031] The command decoder 400 can output a mode signal MODE to distinguish between the storage operation mode indicated by the logical value "0" and the computing operation mode indicated by the logical value "1".
[0032] The flash memory cell array 100 may be referred to as a storage cell array.
[0033] The flash memory cell array 100 includes a plurality of NAND strings, a plurality of bit lines, and a plurality of source lines.
[0034] Figure 2 Shows according to an embodiment of the present disclosure, such as may be included inFigure 1 in the NAND string 110 of the flash memory cell array 100.
[0035] The NAND string 110 includes a plurality of flash memory cells F1, F2, ..., Fn, and Fc, and the plurality of flash memory cells are connected between a corresponding bit line BL and a corresponding source line SL.
[0036] Hereinafter, the NAND string may be referred to as a cell string, and the flash memory cell may be referred to as a storage cell or a flash storage cell.
[0037] The NAND string 110 includes: a bit line selection switch N1 that couples the flash memory cell F1 to the bit line BL; and a source line selection switch N2 that couples the flash memory cell Fc to the source line SL.
[0038] In this embodiment, the flash memory cell Fc is used for calibration operations and may be referred to as a calibration cell.
[0039] In this embodiment, the bit line selection switch N1 and the source line selection switch N2 are NMOS transistors.
[0040] The plurality of flash memory cells F1, F2, ..., Fm and Fc may be floating gate flash memory cells or charge trapping flash memory cells.
[0041] Each flash memory cell stores a weight, and the step of storing the weight is performed during the programming operation of the flash memory device 1.
[0042] In the illustrated embodiment, each flash memory cell stores one bit of weight. In another embodiment, each flash memory cell may store multiple bits of weight.
[0043] Figure 3A is a diagram showing a calculation operation using a flash memory cell.
[0044] The flash memory cell has a low threshold voltage VTH,L or a high threshold voltage VTH,H depending on whether charge is injected into the floating gate or into the charge trapping region.
[0045] In this embodiment, when the threshold voltage is low (such as when the threshold voltage is the low threshold voltage VTH,L), the weight stored in the flash memory cell corresponds to the logical value "1", and when the threshold voltage is high (such as when the threshold voltage is the high threshold voltage VTH,H), the weight stored in the flash memory cell corresponds to the logical value "0".
[0046] In this case, when an input voltage VIN is applied to the gate of the flash memory cell, the drain-source voltage VDS of the flash memory cell changes according to the threshold voltage and the input voltage VIN.
[0047] At this time, the input voltage VIN can be set to the low input voltage VIN,L or the high input voltage VIN,H.
[0048] In this embodiment, the low input voltage VIN,L corresponds to the logical value "1", while the high input voltage VIN,H corresponds to the logical value "0".
[0049] As Figure 3A shown, the low input voltage VIN,L is set to be lower than the high input voltage VIN,H and higher than the high threshold voltage VTH,H.
[0050] For example, the low threshold voltage VTH,L can be distributed between 0V and 1V, while the high threshold voltage VTH,H can be distributed between 4V and 5V. Similarly, the low input voltage VIN,L can be 7V, and the high input voltage VIN,H can be 11V.
[0051] Therefore, regardless of the level of the input voltage VIN, the flash memory cell is turned on, but the magnitude of the on-resistance of the flash memory cell can vary according to the threshold voltage of the flash memory cell and the magnitude of the input voltage VIN.
[0052] Figure 3B A table showing various values of the drain-source voltage VDS of the flash memory cell according to the input voltage VIN and the threshold voltage VTH is shown. This table assumes a constant current is provided through the flash memory cell.
[0053] In the table, the product signal IWP corresponds to the product of the input voltage VIN and the threshold voltage VTH. In this embodiment, the input voltage VIN and the threshold voltage VTH are respectively 1-bit signals (i.e., each carrying 1-bit of information), so the value of the product signal IWP is "0" or "1".
[0054] The drain-source voltage VDS can be divided into three cases.
[0055] The value "V1" (the maximum value) of the drain-source voltage VDS corresponds to the low input voltage VIN,L and the high threshold voltage VTH,H, which corresponds to the case where the on-resistance of the flash memory cell is large due to the small difference between the input voltage VIN and the threshold voltage VTH.
[0056] The value "V2" (the intermediate value) of the drain-source voltage VDS corresponds to the low input voltage VIN,L and the low threshold voltage VTH,L or the high input voltage VIN,H and the high threshold voltage VTH,H, which corresponds to the case where the on-resistance of the flash memory cell is medium.
[0057] The value of the drain-source voltage VDS, "V3" (the minimum value), corresponds to the high input voltage VIN,H and the low threshold voltage VTH,L, which corresponds to the case where the on-resistance of the flash memory cell is minimized due to the largest difference between the input voltage VIN and the threshold voltage VTH.
[0058] In the present embodiment, when the drain-source voltage VDS is V2 or V3, the product signal IWP generated by the flash memory cell corresponds to "0". Since V2 and V3 correspond to the same product signal but have different levels, a calibration operation for adjusting the level of the drain-source voltage VDS is performed in the present embodiment. The calibration operation will be described in detail below.
[0059] In this embodiment, the NAND string 110 includes a plurality of flash memory cells connected in series. In this case, during the calculation operation, the maximum drain-source voltage (i.e., V1) of each flash memory cell is preferably set to be very small compared to the magnitude of the input voltage. The maximum drain-source voltage of the flash memory cell can be controlled, for example, by the design of the flash memory, by controlling the value of the high threshold voltage VTH,H, by controlling the value of the low input voltage VIN,L, by controlling the magnitude of the current flowing through the flash memory cell, or by a combination thereof.
[0060] Since the maximum drain-source voltage (i.e., V1) of each flash memory cell is set to be very small compared to the magnitude of the input voltage during the calculation operation, the voltage difference between the input voltage and the threshold voltage of a specific flash memory cell is substantially not affected by the drain-source voltage of the flash memory cell located below the specific flash memory cell.
[0061] In the calculation operation mode, a constant current is supplied to the NAND string. If the magnitude of the current is I and the maximum resistance of the flash memory cell is R1, the following conditions are preferred.
[0062]
[0063] For example, when the drain-source voltage of all flash memory cells becomes V1, the bit line voltage becomes the maximum value, and it is desired to set this maximum value to be much smaller than the difference between the low input voltage VIN,L and the high threshold voltage VTH,H. In one embodiment, the maximum value of the bit line voltage can be set to be less than 100 mV.
[0064] In the present embodiment, the bit line selection switch N1 is controlled by the bit line selection signal BSL, and the source line selection switch N2 is controlled by the source line selection signal CSL.
[0065] In this embodiment, the bit line select signal BSL and the source line select signal CSL may be provided by the input circuit 200, but the configuration for providing the bit line select signal BSL and the source line select signal CSL may be variously changed.
[0066] The input circuit 200 provides input signals X1, X2, ..., Xn and Xc to the flash memory cell array 100 according to the mode signal MODE.
[0067] Figure 4 is a block diagram of the input circuit 200 according to an embodiment of the present disclosure.
[0068] In this embodiment, the input circuit 200 includes a first input circuit 210 used in the calculation operation mode and a second input circuit 220 used in the storage operation mode.
[0069] The second input circuit 220 provides read voltages or pass voltages as word line signals PX1, PX2, ..., PXn and PXc to the flash memory cell array 100 according to the input signals X1, X2, ..., Xn and Xc.
[0070] The input signal Xc as the (n + 1)th input signal may be expressed as Xn+1, and the word line signal PXc as the signal corresponding to the (n + 1)th input signal may be expressed as PXn+1.
[0071] During the execution of the calculation operation mode, programming operations and erasing operations may be controlled.
[0072] Since this operation corresponds to a normal storage operation, the second input circuit 220 may be used to perform it.
[0073] That is, the second input circuit 220 may also be used for the storage operations required during the calculation operation mode.
[0074] Techniques for providing word line signals according to input signals in the storage operation mode are well known in the art, and thus their detailed descriptions will be omitted.
[0075] The first input circuit 210 converts the input signals X1, X2, ..., Xn into pulse input signals PX1, PX2, ..., PXn in the calculation operation mode.
[0076] In the calculation operation mode, the input signal Xc is used for calibration and may be referred to as the calibration signal Xc.
[0077] In the calculation operation mode, the input signals X1, X2, ..., Xn may each be provided as signals of corresponding multiple bits.
[0078] In this embodiment, the pulse input signals PX1, PX2, ..., PXn and PXc are pulse signals respectively having pulse widths corresponding to the values of the corresponding input signals X1, X2, ..., Xn and Xc.
[0079] In the calculation operation mode, the pulse input signal PXc can be referred to as the pulse calibration signal PXc.
[0080] In this embodiment, the input signals X1, X2, ..., Xn are signals input from the outside, and the input signal Xc is provided by the calibration circuit 500.
[0081] The calibration circuit 500 can perform a calibration operation using the weight information of the flash memory cells included in the NAND string and the values of the input signals X1, X2, ..., Xn. The operation of the calibration circuit 500 will be described in detail below.
[0082] Figure 5 is a block diagram showing a first input circuit 210 according to an embodiment of the present disclosure.
[0083] The first input circuit 210 includes a conversion circuit 211 and a delay circuit 212 that delays the input signals X1, X2, ..., Xn and provides the delayed input signals to the conversion circuit 211.
[0084] The delay circuit 212 delays the input of the input signals X1, X2, ..., Xn into the conversion circuit 211 until the value to be used for controlling the calibration signal Xc is determined. In one embodiment, the input signals X1, X2, ..., Xn are provided to the calibration circuit 500 without being delayed by the delay circuit 212.
[0085] As described above, in this embodiment, the pulse input signals PX1, PX2, ..., PXn and PXc are pulse signals respectively having pulse widths corresponding to the values of the input signals X1, X2, ..., Xn and Xc.
[0086] Figure 8 is a timing diagram showing the operation of the conversion circuit 211 in the calculation operation mode according to an embodiment of the present disclosure.
[0087] In Figure 8 X1 is "1111" (i.e., 15), X2 is "1000" (i.e., 8), X3 is "0100" (i.e., 4), and X4 is "0010" (i.e., 2).
[0088] When the period of the clock signal CLK input to the conversion circuit 211 is T, PX1 is a low-level pulse with a width of 15T, PX2 is a low-level pulse with a width of 8T, PX3 is a low-level pulse with a width of 4T, and PX4 is a low-level pulse with a width of 2T. In one embodiment, during the calculation operation, the low levels of the pulse input signals PX1, PX2,..., PXn and PXc are equal to the low input voltage VIN,L, and the high levels of the pulse input signals PX1, PX2,..., PXn and PXc during the calculation operation are equal to the high input voltage VIN,H.
[0089] Return reference Figure 1 , the output circuit 300 is connected to the bit line BL of the flash memory cell array 100 to output a data signal VOUT in the storage operation mode and output a calculation signal VMAC in the calculation operation mode.
[0090] Figure 6 is a block diagram showing the output circuit 300 according to an embodiment of the present disclosure.
[0091] The output circuit 300 includes a first switch 301, a second switch 302, a calculation output circuit 310, and a data output circuit 320.
[0092] In this embodiment, the first switch 301 is turned on in response to the mode signal MODE being at a high level, and the second switch 302 is turned on in response to the mode signal MODE being at a low level.
[0093] The calculation output circuit 310 outputs a calculation signal VMAC based on the bit line voltage VBL output from the bit line BL.
[0094] The data output circuit 320 outputs a data signal VOUT based on the bit line voltage VBL.
[0095] Since the configuration and operation of the data output circuit 320 are basically the same as those in a conventional flash memory device, its detailed description will be omitted.
[0096] Figure 7 is a circuit diagram showing the calculation output circuit 310 according to an embodiment of the present disclosure.
[0097] The calculation output circuit 310 includes a first current source 311 that provides a constant current I to the NAND string 110 through the bit line BL and a second current source 312 controlled by the bit line voltage VBL.
[0098] The calculation output circuit 310 further includes: a capacitor 313 that is charged by a calculation current IMAC provided from a second current source 312 and outputs a calculation signal VMAC; and a reset switch 314 that discharges the capacitor 313 according to a reset signal RESET.
[0099] The calculation output circuit 310 may further include a sampling switch 315 that couples the second current source 312 and the capacitor 313 according to a sampling clock signal SCLK.
[0100] In this embodiment, the sampling clock signal SCLK has the same frequency as the clock signal CLK input to the input circuit 200.
[0101] The sampling clock signal SCLK may have a predetermined phase difference from the clock signal CLK.
[0102] The second current source 312 includes: an operational amplifier 3121 that amplifies the voltage difference between a bit line voltage VBL and a feedback voltage VF; a PMOS transistor 3122 that includes a gate receiving the output voltage of the operational amplifier 3121 and a source outputting a calculation current IMAC; and a resistor 3213 that is connected between the source of the PMOS transistor 3122 and a power supply voltage VDD.
[0103] The bit line voltage VBL is input to the positive input terminal of the operational amplifier 3121, and the feedback voltage VF, which is the source voltage of the PMOS transistor 3122, is fed back to the negative input terminal of the operational amplifier 3121. Accordingly, the second current source 312 generates a calculation current IMAC that increases as the bit line voltage VBL decreases and decreases as the bit line voltage VBL increases.
[0104] The calculation output circuit 310 may further include an analog-to-digital converter (not shown) that converts the calculation signal VMAC into a digital signal.
[0105] The calculation output circuit 310 may further include a circuit that adjusts the level of the calculation signal VMAC, such as by adding or subtracting an offset from the calculation signal VMAC, amplifying the calculation signal VMAC, or both.
[0106] Figure 9 is a timing diagram showing the calculation operation of a flash memory device according to an embodiment of the present disclosure.
[0107] A calculation operation is performed during a calculation period from T0 to Tr. Prior to T0, a calibration operation is performed as described herein to determine a calibration value C for controlling a pulse calibration signal PXc.
[0108] When the period of the sampling clock SCLK is T and the number of bits of the input signal is m, the calculation period should be at least equal to (2 m -1)T. In this embodiment, it is assumed that the duration of the calculation period is (2 m -1)T.
[0109] exist Figure 9 , the pulse input signal PX1 is a pulse signal having a low level between T0 and T1, the pulse input signal PX2 is a pulse signal having a low level between T0 and T2, and the pulse input signal PX3 is a pulse signal having a low level between T0 and T3. The duration of the low level of the pulse input signals PX1, PX2, PX3, ... is determined according to the input signals X1, X2, X3, ..., respectively. The low level of the pulse calibration signal PXc starts from T0 and has a duration corresponding to the calibration value C.
[0110] As described above with reference to FIG. 3, given the Figure 7 Under the condition of a constant current provided by the first current source 311 of the calculation output circuit 310, the drain-source voltage of each flash memory cell is determined to be one of V1, V2 and V3 according to the value of the pulse input signal supplied to the flash memory cell and the weight value stored in the flash memory cell (corresponding to the threshold voltage of the flash memory cell), and the bit line voltage VBL is determined by the sum of the drain-source voltages of the flash memory cells.
[0111] In the present embodiment, the sampling clock signal SCLK used by the conversion circuit 211 is a signal having the same frequency as the clock signal CLK and having a different phase from the clock signal CLK.
[0112] exist Figure 9 In the embodiment, the width of the low level interval of the pulse input signal is determined as a multiple of the period of the clock signal CLK or the sampling clock signal SCLK according to the corresponding input value.
[0113] The sampling clock signal SCLK is generated by delaying the clock signal CLK so that the high level interval of the sampling clock signal SCLK does not overlap with the transition of the pulse input signal. As a result, during the low level interval of each pulse input signal, the sampling clock signal SCLK has one or more high level intervals.
[0114] In such Figure 7 In the illustrated embodiment using the sampling switch 315 , the capacitor 313 is charged by the calculation current IMAC generated by the second current source 312 only during the interval when the sampling clock signal SCLK is at a high level.
[0115] Therefore, in Figure 9In this case, the computed signal VMAC rises during each interval in which the sampling clock signal SCLK is at a high level, and maintains a constant level during each interval in which the sampling clock signal SCLK is at a low level.
[0116] An interval in which the sampling clock signal SCLK is at a high level may be referred to as a sampling interval.
[0117] In Figure 9 this case, for each sampling period, the bit line voltage VBL may have different values according to the number of pulse input signals having a low level during the sampling period and the weights of the flash memory cells. As a result, the shape of the computed signal VMAC may vary according to the shape of the pulse input signals and the weights stored in each flash memory cell.
[0118] After time Tr, the reset signal is activated in synchronization with the sampling clock signal SCLK, so that the capacitor 313 is discharged to initialize the computed signal VMAC for the next computation.
[0119] Hereinafter, the operation of the calibration circuit 500 will be described.
[0120] As described above, when the drain-source voltage VDS of the flash memory cell is V2 or V3, the product signal IWP is considered to be "0".
[0121] In an embodiment, when the product signal IWP has a value of "0", a calibration operation is performed to adjust the drain-source voltage to be actually equal to a predetermined voltage.
[0122] In this embodiment, when the product signal IWP has a value of "0", a calibration operation is performed to adjust the drain-source voltage to be actually V2. That is, a calibration operation is performed for the case where the drain-source voltage becomes V3 to adjust the drain-source voltage to be actually V2.
[0123] The threshold voltage of the calibration unit Fc is set to have a high threshold voltage VTH,H.
[0124] After receiving the input signals X1, X2, ..., Xn and before performing the computation operation, the calibration operation is performed by the calibration circuit 500. After the calibration operation, the pulse calibration signal PXc is controlled to have a low level VIN,L during C cycles during the computation operation, and a high level VIN,H during the rest of the computation operation, where C is a calibration value generated by the calibration operation.
[0125] For a calibration operation, the number of cases corresponding to combinations of an input voltage VIN and a threshold voltage VTH is counted for pulse input signals PX1, PX2, ..., PXn obtained from input signals X1, X2, ..., Xn. In one embodiment, the number of times the drain-source voltage across the flash memory cell will be equal to V3 during a calculation operation is determined, and this number is used to determine a calibration value for the calculation operation.
[0126] Next, as shown in Table 1, for flash memory cells having a high threshold voltage VTH,H, the number of intervals during which the pulse input voltage VIN is at a low level is denoted as N1, for flash memory cells having a low threshold voltage VTH,L, the number of intervals during which the pulse input voltage VIN is at a low level is denoted as N2, for flash memory cells having a high threshold voltage VTH,H, the number of intervals during which the pulse input voltage VIN is at a high level is denoted as N3, and for flash memory cells having a low threshold voltage VTH,L, the number of intervals during which the pulse input voltage VIN is at a high level is denoted as N4. Since the number N4 corresponds to the number of times the drain-source voltage is V3 instead of the target value V2 when IWP = 0, an error corresponding to N4×(V2 - V3) will occur in the output unless calibration is used.
[0127] Next, the number of cases corresponding to combinations of an input voltage VIN and a threshold voltage VTH is determined for a pulse calibration signal PXc obtained from a calibration signal Xc. Specifically, N4 is determined, which corresponds to the total number of times the drain-source voltage is V3 during a calculation operation because when the corresponding flash memory cell F i has a low threshold voltage VTH,L and the pulse input signal PX i is a high input voltage VIN,H, where i = 1…n.
[0128] The calibration circuit 500 uses the weights W1...W stored in the flash memory cell during a calculation operation n and the values of the input signals x1...x n to calculate N4. In one embodiment, the calibration circuit sums the one's complement of the input signals x i for which the corresponding weight W i is 0. Thus,
[0129]
[0130] where p is the number of bits in the input signal and n is the number of input signals.
[0131] In one embodiment, before performing the calibration operation, the weights W1..W nThe values are stored in the registers in the calibration circuit 500. For example, the weights W1..W n The values can be stored in the calibration circuit 500 when they are programmed into the flash memory cells.
[0132] The threshold voltage of the calibration cell Fc is fixed to the high threshold voltage VTH,H. Therefore, when the pulsed calibration signal PXc is at the low input voltage value VIN,L (corresponding to logic 1), the drain-source voltage across the calibration cell Fc will be the highest drain-source voltage V1, and when the pulsed calibration signal PXc is at the high input voltage value VIN,H (corresponding to logic 0), the drain-source voltage across the calibration cell Fc will be the intermediate drain-source voltage V2.
[0133] Therefore, when the value of the calibration signal Xc is C and m is the number of bits of the calibration signal Xc, the number of intervals with a low level in the pulsed calibration signal PXc corresponding to the calibration signal Xc is C and the number of intervals at a high level is 2 m -1 - C.
[0134]
[0135] The total bit line voltage VBL generated during one calculation period in the absence of calibration can be expressed by Equation 3 below, which corresponds to the result of the MAC operation.
[0136] In the following equations, the drain-source voltages of the bit line selection switches and the drain-source voltages of the source line selection switches are ignored.
[0137] VBL = N1·V1 + (N2 + N3)·V2 + N4×V3 [Equation 3]
[0138] When the calibration operation is performed, i.e., when considering the operation of the flash memory cell Fc of the NAND string, the bit line voltage VBL can be given by Equation 4.
[0139] VBL = (N1 + C)×V1 + (N2 + N3 + 2 m -1 - C)×V2 + N4×V3 [Equation 4]
[0140] When the calibration operation is performed, the calibration cell Fc should not have a negative impact on the calculation result, but should compensate for the difference between the drain-source voltage V2 and the drain-source voltage V3.
[0141] Therefore, by inserting the calibration value C given in Equation 6 below into Equation 4 above, the bit line voltage generated after performing the calibration operation and the calculation operation is given in Equation 5.
[0142] VBL = N1×V1 + (N2 + N3 + N4 + 2m -1) × V2 [Equation 5]
[0143] The value of Equation 4 representing the calibration result and the value of Equation 5 should be the same as each other. And thus, since the error to be corrected is equal to N4·(V2 - V3), and the bit line voltage difference between the high pulse calibration signal PXc and the low pulse calibration signal PXc is V1 - V2, when N4·(V2 - V3) = C·(V1 - V2), the error is compensated by the pulse calibration signal PXc, and thus the calibration value C of the calibration signal Xc can be determined as shown in the following Equation 6.
[0144]
[0145] The value (V2 - V3) / (V1 - V2) can be a constant value determined by the design of the flash memory cells and the device. Thus, in one embodiment, once the calibration circuit 500 has determined N4, the calibration value C can be determined, for example, by multiplying by a fixed value or using a look-up table.
[0146] Although various embodiments have been described for illustrative purposes, it will be apparent to those skilled in the art that various changes and modifications can be made to the described embodiments without departing from the spirit and scope of the present disclosure as defined by the appended claims.
Claims
1. A non-volatile memory device, comprising: A memory cell array, comprising: A plurality of non-volatile memory elements, configured to store a plurality of weights and respectively controlled according to a plurality of input signals provided during a storage operation mode, and Bit lines, coupled to the plurality of non-volatile memory elements; and A calculation output circuit, configured to generate a calculation signal corresponding to an inner product between an input vector and a weight vector, the input vector corresponding to the plurality of input signals provided during a calculation operation mode, the weight vector corresponding to the plurality of weights, and An input circuit, configured to generate a plurality of pulse input signals respectively corresponding to the plurality of input signals, Wherein, the calculation operation mode is performed after storing the plurality of weights in the plurality of memory elements; Wherein, the plurality of pulse input signals are provided to the plurality of non-volatile memory elements, and Wherein, each of the plurality of pulse input signals is a pulse signal having a pulse width corresponding to the value of the corresponding input signal.
2. The non-volatile memory device according to claim 1, wherein, The memory cell array includes cell strings, and the cell strings include the plurality of non-volatile memory elements connected in series.
3. The non-volatile memory device according to claim 2, wherein, Each of the plurality of non-volatile memory elements stores a corresponding weight among the plurality of weights, and includes a gate controlled according to a corresponding input signal among the plurality of input signals and a drain coupled to the source of an adjacent non-volatile memory element.
4. The non-volatile memory device according to claim 2, wherein, The memory cell array further includes: a bit line selection switch, which couples the cell string to the bit line according to a bit line selection signal; and a source line selection switch, which couples the cell string to a source line according to a source line selection signal.
5. The non-volatile storage device according to claim 1, wherein, The calculation output circuit includes a first current source configured to provide a constant current to the bit line, and wherein, the signal at the bit line is a voltage signal induced by the constant current.
6. The non-volatile memory device according to claim 5, wherein, The calculation output circuit further includes a second current source, configured to generate a calculation current according to the voltage of the bit line.
7. The non-volatile memory device according to claim 6, wherein, The calculation output circuit further includes a capacitor charged by the calculation current.
8. The non-volatile memory device according to claim 7, wherein, The calculation output circuit further includes a sampling switch, configured to provide the voltage of the bit line to the second current source according to a sampling clock.
9. The non-volatile memory device according to claim 8, wherein, The calculation output circuit further includes a reset switch, which is used to discharge the capacitor according to a reset signal.
10. The non-volatile memory device according to claim 6, wherein, The second current source includes: An operational amplifier, configured to amplify the difference between the voltage of the bit line and a feedback voltage; A transistor, including a gate receiving the output voltage of the operational amplifier, a source, and a drain; and A resistor, coupled between a power supply voltage and one of the source or the drain of the transistor, Wherein, the calculation current is provided from one of the source and the drain of the transistor, and Wherein, the feedback voltage is provided from the other of the source and the drain of the transistor.
11. The non-volatile memory device according to claim 1, wherein, The non-volatile memory device is a NAND flash memory device.
12. A non-volatile memory device, comprising: A memory cell array, comprising: A plurality of non-volatile memory elements configured to store a plurality of weights and respectively controlled according to a plurality of input signals provided during a storage operation mode, bit lines coupled to the plurality of non-volatile memory elements; and a non-volatile calibration memory element having a predetermined weight and controlled according to a calibration signal; and a calculation output circuit configured to generate a calculation signal corresponding to an inner product between an input vector and a weight vector, the input vector corresponding to the plurality of input signals provided during a calculation operation mode, the weight vector corresponding to the plurality of weights; and a calibration circuit configured to generate a calibration signal according to the plurality of input signals and the plurality of weights; wherein, after the plurality of weights are stored in the plurality of memory elements, a calculation operation mode is executed; wherein the voltage of the bit lines is adjusted according to the calibration signal and the predetermined weight.
13. The non-volatile memory device according to claim 12, further comprising an input circuit configured to: convert the plurality of input signals and the calibration signal into a plurality of pulse input signals and a pulse calibration signal, and provide the plurality of pulse input signals and the pulse calibration signal to the plurality of non-volatile memory elements and the non-volatile calibration memory element, Among them, each of the plurality of pulse input signals having a pulse width corresponding to the value of the corresponding input signal, and the pulse calibration signal having a pulse width corresponding to the value of the calibration signal.
Citation Information
Patent Citations
Folding circuit and nonvolatile memory devices
CN107305784A
Nonvolatile memory device, operation method thereof, and memory system
CN110047543A