Planar level compression of soft bit data in non-volatile memory using data latches

CN115809019BActive Publication Date: 2026-08-21SANDISK TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211102454.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-05-12
Filing Date
2022-09-09
Publication Date
2026-08-21
Estimated Expiration
2042-09-09

Smart Images

  • Figure CN115809019B_ABST
    Figure CN115809019B_ABST
Patent Text Reader

Abstract

For non-volatile memories that use hard bit and soft bit data in error correction operations, to reduce the amount of soft bit data that needs to be transferred from the memory to the controller and improve memory system performance, the soft bit data can be compressed prior to transfer. After the soft bit data is read and stored into internal data latches associated with sense amplifiers, the soft bit data is compressed within these internal data latches. The compressed soft bit data can then be transferred to transfer data latches of a cache buffer, where the compressed soft bit data can be merged and transferred out through an input-output interface. In the input-output interface, the compressed data can be reorganized, if needed, to place it in logical user data order.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Statement

[0002] This application is a continuation-in-part of U.S. Patent Application No. 17 / 666,657, filed February 8, 2022, entitled “Use of Data Latches for Compression of Soft Bit Data in Non-Volatile Memories,” which is a continuation-in-part of U.S. Patent Application No. 17 / 557,236, filed December 21, 2021, entitled “Efficient sensing of Soft Bit Data for Non-Volatile Memory,” which in turn claims priority to U.S. Provisional Patent Application No. 63 / 244,951, filed September 16, 2021, entitled “Plane Level Vertical Compression Scheme,” all of which are hereby incorporated herein by reference in their entirety. Background Technology

[0003] This disclosure relates to non-volatile storage devices.

[0004] Semiconductor memories are widely used in a variety of electronic devices, such as cellular phones, digital cameras, personal digital assistants, medical electronics, mobile computing devices, servers, solid-state drives, non-mobile computing devices, and other devices. Semiconductor memories can include non-volatile or volatile memories. Non-volatile memories allow information to be stored and retained even when not connected to a power source (e.g., a battery). An example of a non-volatile memory is flash memory (e.g., NAND flash memory and NOR flash memory).

[0005] Users of non-volatile memory can program (e.g., write) data into the non-volatile memory and then read that data back. For example, a digital camera can take a photo and store it in non-volatile memory. The user of the digital camera can then view the photo by having the camera read it from the non-volatile memory. Because users often rely on the data they store, it is important for users of non-volatile memory to be able to reliably store data so that it can be successfully read back. Attached Figure Description

[0006] Elements with similar numbers refer to common components in different drawings.

[0007] Figure 1It is a block diagram depicting one implementation of a storage system.

[0008] Figure 2A This is a block diagram of one implementation scheme for a memory die.

[0009] Figure 2B This is a block diagram of one implementation scheme for an integrated memory component.

[0010] Figure 3 A circuit for sensing data from a non-volatile memory is depicted.

[0011] Figure 4 This is a perspective view of one implementation of a monolithic three-dimensional memory architecture.

[0012] Figures 5A to 5F An example depicting the threshold voltage distribution is shown.

[0013] Figure 6 This is a flowchart describing one implementation of the process for programming non-volatile memory.

[0014] Figure 7 The diagram shows the overlap of the distributions of two adjacent data states and a set of reads that can be used to determine the data state of a cell and the reliability of such reads.

[0015] Figure 8 The concepts of hard bits and soft bits are shown.

[0016] Figure 9A and Figure 9B The read levels for calculating the hard and soft bit values ​​of the next page data are shown in the three-bit data implementation per memory cell.

[0017] Figure 10 The assignment of hard and soft bit values ​​and readout levels used in an implementation for effective soft sensing is shown.

[0018] Figure 11 The diagram illustrates how the coding in Table 2 is used in a three-bit data implementation per memory cell to apply an effective soft-sensing mode to the next page of data.

[0019] Figure 12 It shows the corresponding Figure 11 The implementation scheme shown is a sensing operation for the next page data reading operation in an effective soft sensing reading operation of the reading point.

[0020] Figure 13 An implementation scheme of a sensing amplifier circuit that can be used to determine the hard and soft bit values ​​of a memory cell is shown.

[0021] Figure 14 This is a flowchart of an implementation scheme for effective soft sensing operation.

[0022] Figure 15 It is a block diagram of an implementation scheme for some control circuit elements of a memory device, including a soft-bit compression element.

[0023] Figure 16 , Figure 17A and Figure 17B Further details are provided regarding implementation schemes for data latches that can be used in soft bit data compression processes.

[0024] Figure 18 and Figure 19 Two implementations for compressing raw soft bit data from one set of internal data latches to another set of data latches are shown.

[0025] Figure 20 This illustrates the rearrangement of compressed soft bit data within the internal data latch.

[0026] Figure 21 This illustrates transmission within a transmission latch to compress data.

[0027] Figure 22 This is a diagram illustrating how compressed data bits are rearranged into a logical order.

[0028] Figure 23 This is a block diagram of an alternative implementation of some control circuit elements in the control circuit elements of a memory device including a soft-bit compression element.

[0029] Figure 24A and Figure 24B Is using Figure 23 The implementation scheme reorganizes compressed data bits into a logical order, as illustrated in the diagram of an alternative implementation scheme.

[0030] Figure 25 This is a flowchart of an implementation scheme for performing data compression using a data latch associated with a sense amplifier of a non-volatile memory device.

[0031] Figure 26 An implementation scheme is shown for moving compressed data from an internal data latch to a buffer cache memory without first rearranging the compressed data within the internal data latch.

[0032] Figure 27 An implementation scheme for multiplexing data from a data transfer latch from a cache buffer to a global data bus is shown.

[0033] Figure 28 This is a flowchart of an additional implementation scheme for performing data compression within a data latch associated with a sense amplifier of a non-volatile memory device. Detailed Implementation

[0034] In some memory systems, error correction methods sometimes use "soft bit" data. Soft bit data provides information about the reliability of a standard or "hard bit" data value used to distinguish between data states. For example, when the data value is based on the threshold voltage of a memory cell, a hard bit read will determine whether the threshold voltage of the memory cell is higher or lower than the data read value in order to distinguish between stored data states. For memory cells with threshold voltages slightly higher or lower than this reference value, such a hard bit may be incorrect because the memory cell actually means being in a different data state. To determine memory cells with threshold voltages close to the hard bit read level, and therefore with hard bit values ​​of lower reliability, a pair of additional reads offset slightly above and slightly below the hard bit read level can be performed to generate soft bit values ​​for the hard bit values. The use of soft bits can be a powerful tool for extracting the data contents of memory cells, but because it requires additional reads to obtain the soft bit data that then needs to be transmitted to error correction circuitry, it is generally only used when the data cannot be accurately determined from the hard bit values ​​alone.

[0035] The following proposes an efficient soft-sensing readout mode that requires fewer reads to generate soft-bit data and generates less soft-bit data, thereby reducing the performance and power consumption losses typically associated with using soft-bit data. This allows the efficient soft-sensing mode to be used as the default readout mode. Compared to a typical hard-bit / soft-bit arrangement, the hard-bit readout point offset ensures that the hard bit value of one data state of the memory cell is reliable, but the hard bits of the other data state include a larger number of unreliable hard bit values. Performing a single soft-bit read for the less reliable hard bit value, but without providing reliability information for the more reliable hard bit value, reduces the number of reads and the amount of data obtained. To further improve performance, both hard-bit sensing and soft-bit sensing can be combined into a single sensing, such as by pre-charging the node of the sensing amplifier and performing a single discharge through the selected memory cell, but sensing the level twice for the single discharge at the node, once for the hard bit value and once for the soft bit value.

[0036] To further reduce the amount of data that needs to be transferred from memory to the controller and improve memory system performance, soft-bit data can be compressed before transfer. After the soft-bit data is read and stored in internal data latches associated with the sense amplifier, it is compressed within these internal data latches. The compressed soft-bit data can then be transferred to the transfer data latch of the cache buffer, where it can be merged and transmitted out through the input-output interface. At the input-output interface, the compressed data can be reorganized to place it in the logical user data order if needed.

[0037] Figure 1 The components of the storage system 100 depicted are electronic circuits. The storage system 100 includes a memory controller 120 connected to a non-volatile memory 130 and a local high-speed volatile memory 140 (e.g., DRAM). The memory controller 120 uses the local high-speed volatile memory 140 to perform certain functions. For example, the local high-speed volatile memory 140 stores logic in a physical address translation table (“L2P table”).

[0038] The memory controller 120 includes a host interface 152 that connects to and communicates with the host 102. In one embodiment, the host interface 152 implements NVM Express (NVMe) via PCI Express (PCIe). Other interfaces, such as SCSI, SATA, etc., may also be used. The host interface 152 is also connected to a network on-chip (NOC) 154. An NOC is a communication subsystem on an integrated circuit. NOCs can span synchronous and asynchronous clock domains or use non-clocked asynchronous logic. NOC technology applies network theory and methods to on-chip communication and provides significant improvements compared to conventional bus and cross-switch interconnects. Compared to other designs, NOCs improve the scalability of system-on-chip (SoC) and the power efficiency of complex SoCs. The wires and links in a NOC are shared by many signals. Because all links in a NOC can operate simultaneously on different data packets, a high degree of parallelism is achieved. Therefore, as the complexity of integrated subsystems increases, NOCs offer enhanced performance (such as throughput) and scalability compared to previous communication architectures (e.g., dedicated point-to-point signal lines, shared buses, or segmented buses with bridges). In other embodiments, NOC 154 may be replaced by a bus. Processor 156, ECC engine 158, memory interface 160, and DRAM controller 164 are connected to and communicate with NOC 154. DRAM controller 164 is used to operate and communicate with local high-speed volatile memory 140 (e.g., DRAM). In other embodiments, local high-speed volatile memory 140 may be SRAM or another type of volatile memory.

[0039] ECC engine 158 performs error correction services. For example, ECC engine 158 performs data encoding and decoding according to implemented ECC technology. In one embodiment, ECC engine 158 is a software-programmable electronic circuit. For example, ECC engine 158 may be a programmable processor. In other embodiments, ECC engine 158 is a custom-designed dedicated hardware circuit without any software. In yet another embodiment, the functionality of ECC engine 158 is implemented by processor 156.

[0040] Processor 156 performs various controller memory operations, such as programming, erasing, reading, and memory management processes. In one embodiment, processor 156 is programmed by firmware. In other embodiments, processor 156 is a custom-designed dedicated hardware circuit without any software. Processor 156 also implements a translation module, either as a software / firmware process or as dedicated hardware circuitry. In many systems, non-volatile memory is addressed inward to the memory system using physical addresses associated with one or more memory dies. However, the host system will use logical addresses to address various memory locations. This allows the host to assign data to consecutive logical addresses while the memory system is idle to store data between the locations of one or more memory dies as desired. To implement such a system, memory controller 120 (e.g., a translation module) performs address translation between logical addresses used by the host and physical addresses used by the memory dies. An exemplary embodiment maintains a table that identifies the current translation between logical and physical addresses (i.e., the L2P table described above). Entries in the L2P table may include a logical address and an identifier of the corresponding physical address. Although logical address to physical address tables (or L2P tables) include the word "table," they do not have to be tables in the literal sense. Instead, logical address to physical address tables (or L2P tables) can be any type of data structure. In some examples, the memory space of the storage system is so large that local memory 140 cannot hold all the L2P tables. In this case, the entire set of L2P tables is stored in memory die 130, and a subset of the L2P tables is cached (the L2P cache) in local high-speed volatile memory 140.

[0041] Memory interface 160 communicates with non-volatile memory 130. In one embodiment, the memory interface provides a switching mode interface. Other interfaces may also be used. In some example implementations, memory interface 160 (or another part of controller 120) implements a scheduler and buffer for transferring data to and receiving data from one or more memory dies.

[0042] In one embodiment, the non-volatile memory 130 includes one or more memory dies. Figure 2A This is a functional block diagram of one embodiment of a memory die 200 including non-volatile memory 130. Each of one or more memory dies of non-volatile memory 130 can be implemented as Figure 2A The memory die 200. Figure 2AThe components depicted are circuits. Memory die 200 includes a memory array 202, which may include non-volatile memory cells, as described in more detail below. The array terminal lines of memory array 202 include individual word line layers organized into rows and individual bit line layers organized into columns. However, other orientations may also be implemented. Memory die 200 includes row control circuitry 220, the outputs of which 208 are connected to the corresponding word lines of memory array 202. Row control circuitry 220 receives a set of M row address signals and one or more various control signals from system control logic circuitry 260, and typically includes circuitry such as row decoder 222, array terminal driver 224, and block select circuitry 226 for both read and write (programming) operations. Row control circuitry 220 may also include read / write circuitry. Memory die 200 also includes column control circuitry 210, which includes a sense amplifier 230, the inputs / outputs of which 206 are connected to the corresponding bit lines of memory array 202. Although only a single block is shown for array 202, the memory die may include multiple arrays that can be accessed individually. Column control circuitry 210 receives a set of N column address signals and one or more various control signals from system control logic unit 260, and typically includes circuitry such as column decoder 212, array terminal receiver or driver circuitry 214, block selection circuitry 216, and read / write circuitry and I / O multiplexers.

[0043] System control logic unit 260 receives data and commands from memory controller 120 and provides output data and status to the host. In some embodiments, system control logic unit 260 (which includes one or more circuits) includes a state machine 262 that provides die-level control for memory operations. In one embodiment, state machine 262 is programmable by software. In other embodiments, state machine 262 does not use software and is implemented entirely in hardware (e.g., electronic circuitry). In yet another embodiment, state machine 262 is replaced by a microcontroller or microprocessor, which is located on or outside the memory chip. System control logic unit 260 may also include a power control module 264 that controls the power and voltage supplied to rows and columns of memory structure 202 during memory operations and may include charge pump and regulator circuitry for generating regulated voltages. System control logic unit 260 includes a storage device 266 (e.g., RAM, registers, latches, etc.) that can be used to store parameters for operating memory array 202.

[0044] Commands and data are transmitted between memory controller 120 and memory die 200 via memory controller interface 268 (also referred to as the "communication interface"). Memory controller interface 268 is an electrical interface used for communicating with memory controller 120. Examples of memory controller interface 268 include a switching mode interface and an Open NAND Flash Interface (ONFI). Other I / O interfaces may also be used.

[0045] In some embodiments, all components of memory die 200 (including system control logic 360) may be formed as part of a single die. In other embodiments, some or all of the system control logic 260 may be formed on different dies.

[0046] In one embodiment, memory structure 202 includes a three-dimensional memory array of non-volatile memory cells, wherein multiple memory stages are formed over a single substrate such as a wafer. The memory structure may include any type of non-volatile memory integrally formed in one or more physical stages of memory cells having active regions disposed over a silicon (or other type of) substrate. In one example, the non-volatile memory cells include vertical NAND strings with charge trapping layers.

[0047] In another embodiment, memory structure 202 includes a two-dimensional memory array of non-volatile memory cells. In one example, the non-volatile memory cells are NAND flash memory cells utilizing floating gates. Other types of memory cells (e.g., NOR flash memory) may also be used.

[0048] The exact type of memory array architecture or memory cell included in memory structure 202 is not limited to the examples described above. Many different types of memory array architectures or memory technologies can be used to form memory structure 202. Implementing the novel embodiments claimed herein does not require a specific non-volatile memory technology. Other examples of suitable technologies for memory cells of memory structure 202 include ReRAM (Resistive Random Access Memory), magnetoresistive memory (e.g., MRAM, spin-torque MRAM, spin-orbit torque MRAM), FeRAM, phase-change memory (e.g., PCM), and so on. Examples of suitable technologies for memory cell architectures of memory structure 202 include two-dimensional arrays, three-dimensional arrays, cross-point arrays, stacked two-dimensional arrays, vertical bitline arrays, and so on.

[0049] One example of a ReRAM crosspoint memory includes reversible resistive switching elements arranged in a crosspoint array accessed by X-rays and Y-rays (e.g., word lines and bit lines). In another embodiment, the memory cell may include a conductive bridge memory element. A conductive bridge memory element may also be referred to as a programmable metallized cell. Based on the physical repositioning of ions within a solid electrolyte, the conductive bridge memory element can be used as a state-changing element. In some cases, the conductive bridge memory element may include two solid metal electrodes, one relatively inert (e.g., tungsten) and the other electrochemically active (e.g., silver or copper), with a thin film of solid electrolyte between the two electrodes. As temperature increases, ion mobility also increases, leading to a decrease in the programming threshold of the conductive bridge memory cell. Therefore, the conductive bridge memory element can have a wide range of programming thresholds across the entire temperature range.

[0050] Another example is magnetoresistive random access memory (MRAM), which stores data using magnetic storage elements. These elements are formed from two ferromagnetic layers separated by a thin insulating layer, each of which can remain magnetized. One of these layers is a permanent magnet set to a specific polarity; the magnetization of the other layer can be changed to match the magnetization of the memory by an external magnetic field. The memory device is constructed from a grid of such memory cells. In one implementation for programming, each memory cell is located between a pair of write lines arranged at right angles to each other, parallel to the cell, one above and one below. When current passes through them, an induced magnetic field is generated. MRAM-based memory implementations will be discussed in more detail below.

[0051] Phase-change memories (PCMs) utilize the unique properties of chalcogenide glasses. One embodiment uses a GeTe-Sb₂Te₃ superlattice to achieve a non-thermal phase transition by changing the coordination state of germanium atoms using only a laser pulse (or a light pulse from another source). Therefore, the programming dose is the laser pulse. Memory cells can be suppressed by preventing them from receiving light. In other PCM embodiments, memory cells are programmed by current pulses. It should be noted that the use of "pulse" in this document does not require a rectangular pulse, but includes (continuous or discontinuous) vibrations or pulse trains of sound, current, voltage, light, or other waves. These memory elements within the individual selectable memory cells or bits may include additional series elements as selectors, such as bidirectional threshold switches or metallic insulator substrates.

[0052] Those skilled in the art will recognize that the techniques described herein are not limited to a single particular memory structure, memory configuration, or material composition, but encompass many related memory structures within the technical essence and scope as described herein and as understood by those skilled in the art.

[0053] Figure 2A The components can be grouped into two parts: (1) memory structure 202 and (2) peripheral circuitry, which includes Figure 2A The diagram depicts all components except for memory structure 202. A crucial characteristic of memory circuitry is its capacity, which can be increased by increasing the area of ​​the memory die allocated to the memory system 100 for the specific purpose of memory structure 202; however, this reduces the area of ​​the memory die available for peripheral circuitry. This can impose significant limitations on these peripheral circuit components. For example, the need to mount sense amplifier circuitry within the available area can be a major constraint on sense amplifier design architecture. Regarding system control logic component 260, the reduced available area may limit the available functions that can be implemented on the chip. Therefore, a fundamental trade-off must be made between the amount of dedicated area for memory structure 202 and the amount of dedicated area for peripheral circuitry in the design of the memory die for memory system 100.

[0054] Another area where memory structure 202 often conflicts with peripheral circuitry is in the processing involved in forming these areas, as these areas typically involve different processing techniques and trade-offs when implementing different techniques on a single die. For example, when memory structure 202 is NAND flash memory, it is an NMOS structure, while the peripheral circuitry is typically CMOS-based. For instance, elements such as sense amplifier circuitry, charge pumps, logic elements in state machines, and other peripheral circuitry in system control logic unit 260 typically employ PMOS devices. The processing operations used to manufacture CMOS dies will differ in many ways from those optimized for NMOS flash NAND memory or other memory cell technologies.

[0055] To mitigate these limitations, the implementation scheme described below can... Figure 2AThe components are separated onto individually formed dies, and then these dies are bonded together. More specifically, the memory structure 202 can be formed on a single die (referred to as the memory die), and some or all of the peripheral circuitry elements (including one or more control circuits) can be formed on separate dies (referred to as the control die). For example, the memory die can be formed solely of memory elements, such as flash NAND memory, MRAM memory, PCM memory, ReRAM memory, or other memory cell arrays of other memory types. Some or all of the peripheral circuitry (even including elements such as decoders and sense amplifiers) can then be moved to separate control dies. This allows each die in the memory die to be optimized individually according to its technology. For example, a NAND memory die can be optimized for an NMOS-based memory array structure without worrying about CMOS elements now moved to a control die that can be optimized for CMOS processing. This provides more space for peripheral elements, and additional capabilities that might not be easily combined can now be incorporated if peripheral elements were confined to the edges of the same die housing the memory cell array. Two dies can then be bonded together in a bonded multi-die memory circuit, with an array on one die connected to peripheral components on the other die. For example, while the following will focus on a bonded memory circuit with one memory die and one control die, other embodiments may use more dies, such as two memory dies and one control die.

[0056] Figure 2B It shows Figure 2A An alternative arrangement of the arrangement can be implemented using wafer-to-wafer bonding to provide bonded die pairs. Figure 2B A functional block diagram of one embodiment of integrated memory component 207 is depicted. One or more integrated memory components 207 may be used to implement non-volatile memory 130 of memory system 100. Integrated memory component 207 includes two types of semiconductor dies (or more simply, "dies"). Memory die 201 includes memory structure 202. Memory structure 202 includes non-volatile memory cells. Control die 211 includes control circuitry 260, 210, and 220 (as described above). In some embodiments, control die 211 is configured to be connected to memory structure 202 within memory die 201. In some embodiments, memory die 201 and control die 211 are coupled together.

[0057] Figure 2B An example of peripheral circuitry is shown, including control circuitry formed in the peripheral circuitry or control die 311, which is coupled to the memory structure 202 formed in the memory die 201. General components are similar to... Figure 2AThe system control logic unit 260, row control circuitry 220, and column control circuitry 210 are located in control die 211. In some embodiments, all or part of the column control circuitry 210 and all or part of the row control circuitry 220 are located on memory die 201. In some embodiments, some circuitry in the system control logic unit 260 is located on memory die 201.

[0058] System control logic unit 260, row control circuitry 220, and column control circuitry 210 can be formed using conventional processes (e.g., CMOS processes), making it possible to add elements and functions more commonly found on memory controller 120, such as ECC, with few or no additional process steps (i.e., the same process steps used to manufacture controller 120 can also be used to manufacture system control logic unit 260, row control circuitry 220, and column control circuitry 210). Therefore, while removing such circuitry from a die (e.g., memory 2 die 201) reduces the number of steps required to manufacture such a die, adding such circuitry to a die (e.g., control die 311) may not require many additional process steps. Because some or all of the control circuitry 260, 210, and 220 are implemented using CMOS technology, control die 211 may also be referred to as a CMOS die.

[0059] Figure 2B A column control circuit 210, including a sense amplifier 230, is shown on a control die 211. This column control circuit is coupled to a memory structure 202 on a memory die 201 via an electrical path 206. For example, the electrical path 206 can provide electrical connections between the column decoder 212, the driver circuit 214, the block selector 216, and the bit lines of the memory structure 202. The electrical path can extend from the column control circuit 210 in the control die 211 through pads on the control die 211 that bond to corresponding pads on the memory die 201 that are connected to the bit lines of the memory structure 202. Each bit line of the memory structure 202 may have a corresponding electrical path in the electrical path 206 connected to the column control circuit 210, including a pair of bonded pads. Similarly, a row control circuit 220 (including a row decoder 222, an array driver 224, and a block selector 226) is coupled to the memory structure 202 via an electrical path 208. Each electrical path in electrical path 208 may correspond to a word line, a dummy word line, or a select gate line. Additional electrical paths may also be provided between the control die 211 and the memory die 201.

[0060] For the purposes of this document, the phrase "control circuitry" or "one or more control circuits" may include all or part of the memory controller 120, state machine 262, system control logic unit 260, all or part of row control circuitry 220, all or part of column control circuitry 210, microcontroller, microprocessor, and / or other similar functional circuitry or any combination thereof. Control circuitry may consist only of hardware or a combination of hardware and software (including firmware). For example, a controller programmed by firmware to perform the functions described herein is an example of control circuitry. Control circuitry may include a processor, FGA, ASIC, integrated circuit, or other types of circuitry. In some embodiments, more than one control die 211 and more than one memory die 201 are present in the integrated memory assembly 207. In some embodiments, the integrated memory assembly 207 includes a stack of multiple control dies 211 and multiple memory dies 201.

[0061] Figure 3 This is a block diagram depicting one embodiment of a column control circuit 210, which is divided into a plurality of sense amplifiers 230 and a common section referred to as management circuitry 302. In one embodiment, each sense amplifier 230 is connected to a corresponding bit line, which in turn is connected to one or more NAND strings. In one exemplary embodiment, each bit line is connected to six NAND strings, with one NAND string per sub-block. Management circuitry 302 is connected to a group of a plurality (e.g., four, eight, etc.) of sense amplifiers 230. Each of the sense amplifiers 230 in the group communicates with the associated management circuitry via a data bus 304.

[0062] Each sense amplifier 230 operates to provide voltage to bit lines (see BL0, BL1, BL2, BL3) during programming, verification, erasing, and read operations. The sense amplifiers are also used to sense the conditions (e.g., data state) of memory cells in a NAND string connected to the respective sense amplifier.

[0063] Each sense amplifier 230 includes a selector 306 or switch connected to a transistor 308 (e.g., an NMOS). Based on the voltage at the control gate 310 and drain 312 of transistor 308, the transistor can operate as a transmission gate or as a bit line clamp. When the voltage at the control gate is sufficiently higher than the voltage at the drain, the transistor operates as a transmission gate to pass the voltage at the drain to the bit line (BL) at the source 314 of the transistor. For example, a programming-suppression voltage such as 1V to 2V can be passed when pre-charging and suppressing an unselected NAND string. Alternatively, a programming-enable voltage such as 0V can be passed to allow programming in a selected NAND string. The selector 306 can pass a supply voltage Vdd (e.g., 3V to 4V) to the control gate of transistor 308 to make it operate as a transmission gate.

[0064] When the voltage at the control gate is lower than the voltage at the drain, transistor 308 operates as a source follower to set or clamp the bit line voltage at Vcg - Vth, where Vcg is the voltage at the control gate 310 and Vth (e.g., 0.7V) is the threshold voltage of transistor 308. This assumes the source line is at 0V. If Vcelsrc is non-zero, the bit line voltage is clamped at Vcg - Vcelsrc - Vth. Therefore, this transistor is sometimes referred to as a bit line clamp (BLC) transistor, and the voltage Vcg at the control gate 310 is referred to as the bit line clamp voltage Vblc. This mode can be used during sensing operations, such as read and verification operations. Thus, the bit line voltage is set by transistor 308 based on the voltage output by selector 306. For example, selector 306 can pass Vsense + Vth (e.g., 1.5V) to the control gate of transistor 308 to provide Vsense (e.g., 0.8V) on the bit line. Vbl selector 316 can pass a relatively high voltage, such as Vdd, to drain 312, which is higher than the control gate voltage on transistor 308 to provide source follower mode during sensing operation. Vbl refers to the bit line voltage.

[0065] Vbl selector 316 can transmit one of a plurality of voltage signals. For example, the Vbl selector can transmit a program-suppress voltage signal that increases from an initial voltage (e.g., 0V) to a program-suppress voltage (e.g., the voltage Vbl_inh for the corresponding bit line of an unselected NAND string during a programming cycle). Vbl selector 316 can also transmit a programming-enable voltage signal, such as 0V, for the corresponding bit line of a selected NAND string during a programming cycle.

[0066] In one approach, the selector 306 for each sensing circuit can be controlled separately from the selectors of other sensing circuits. The Vbl selector 316 for each sensing circuit can also be controlled separately from the Vbl selectors of other sensing circuits.

[0067] During sensing, sensing node 318 is charged to an initial voltage Vsense_init, such as 3V. The sensing node is then passed to a bitline via transistor 308, and the amount of decay of the sensing node is used to determine whether the memory cell is in a conductive or non-conductive state. The amount of decay of the sensing node also indicates whether the current Icell in the memory cell exceeds a reference current Iref. A larger decay corresponds to a larger current. If Icell ≤ Iref, the memory cell is in a non-conductive state, and if Icell > Iref, the memory cell is in a conductive state.

[0068] Specifically, the comparator circuit 320 determines the attenuation amount by comparing the sense node voltage with the trip voltage during sensing. If the sense node voltage attenuates below the trip voltage Vtrip, the memory cell is in a conductive state and its Vth is equal to or lower than the verification voltage. If the sense node voltage does not attenuate below Vtrip, the memory cell is in a non-conductive state and its Vth is higher than the verification voltage. For example, the comparator circuit 320 sets the sense node latch 322 to 0 or 1 based on whether the memory cell is in a conductive or non-conductive state. For example, in a program-verify test, 0 can indicate failure, and 1 can indicate success. The bits in the sense node latch can be read during a status bit scan operation of a scan operation or toggled from 0 to 1 during a fill operation. The bits in the sense node latch 322 can also be used in a lock scan to determine whether the bit line voltage is set to an inhibit level or a programming level in the next programming cycle.

[0069] The management circuitry 302 includes a processor 330, four sets of exemplary data latches 340, 342, 344 and 346, and an I / O interface 332 coupled between the sets of data latches and the data bus 334. Figure 3 Four exemplary sets of data latches 340, 342, 344, and 346 are shown; however, in other embodiments, more or fewer than four sets may be implemented. In one embodiment, each sense amplifier 230 has one set of latches. A set of three data latches may be provided for each sense circuit, for example, including individual latches ADL, BDL, CDL, and XDL. In some cases, different numbers of data latches may be used. In a three-bit per memory cell embodiment, ADL stores bits for next page data, BDL stores bits for intermediate page data, CDL stores bits for previous page data, and XDL acts as an interface latch for storing / latching data from the memory controller.

[0070] Processor 330 performs calculations, such as determining data stored in sensed memory cells and storing the determined data in the set of data latches. Each set of data latches 340-346 is used to store data bits determined by processor 330 during a read operation and data bits imported from data bus 334 during a programming operation, these data bits representing write data to be programmed into memory. I / O interface 332 provides an interface between data latches 340-346 and data bus 334.

[0071] During a read operation, the system operates under the control of state machine 262, which controls the supply of different control gate voltages to the addressed memory cell. As it progresses through various predefined control gate voltages corresponding to different memory states supported by the memory, a sensing circuit can trip at one of these voltages, and the corresponding output is provided to the processor 330 from the sensing amplifier via data bus 304. The processor 330 then determines the resulting memory state by considering the tripping event of the sensing circuit and information about the control gate voltages applied via input line 348 from the state machine. It then calculates the binary code of the memory state and stores the resulting data bits in data latches 340-346.

[0072] Some specific implementations may include multiple processors 330. In one implementation, each processor 330 will include output lines (not depicted) such that each output line is connected by a line or connection. A line or connection or line can be provided by connecting multiple lines together at a node, where each line carries a high or low input signal from the corresponding processor, and the node's output is high if any of the input signals is high. In some implementations, the output lines are inverted before being connected to the line or line. This configuration allows for rapid determination of when the programming process is complete during programming verification testing, as the state machine receiving the line or line can determine when all programmed bits have reached the desired level. For example, when each bit reaches its desired level, a logic zero for that bit is sent to the line or line (or data one is inverted). When all bits output data 0 (or data one is inverted), the state machine knows to terminate the programming process. Since each processor communicates with eight sensing circuits, the state machine needs to read the line or line eight times, or logic can be added to the processor 330 to accumulate the results of the relevant bit lines, so that the state machine only needs to read the line or line once. Similarly, by correctly selecting the logic level, the global state machine can detect when the first bit changes its state and adjust the algorithm accordingly.

[0073] During the programming or verification operation of a memory cell, the data to be programmed (write data) is stored from the data bus 334 into the data latch group 340-346. During reprogramming, the corresponding set of data latches for the memory cell can store data indicating when the memory cell can be reprogrammed based on the programming pulse magnitude value.

[0074] Under the control of state machine 262, the programming operation applies a series of programming voltage pulses to the control gate of the addressed memory cell. The amplitude of each voltage pulse can be incrementally increased by one step from the previous programming pulse during the process, a process known as incremental step pulse programming. Each programming voltage is followed by a verification operation to determine whether the memory cell has been programmed to the desired memory state. In some cases, processor 330 monitors the read-back memory state relative to the desired memory state. When both are consistent, processor 330 sets the bit line to a programming-inhibited mode, such as by updating its latch. This prevents further programming of the memory cell coupled to the bit line, even if additional programming pulses are applied to its control gate.

[0075] Figure 4 This is a perspective view as part of an exemplary embodiment of a monolithic three-dimensional memory array / structure that may include memory structure 202, which includes a plurality of non-volatile memory cells arranged as vertical NAND strings. For example, Figure 4 A portion 400 of a memory block is shown. The depicted structure includes a set of bit lines BL, which lie above a stack 401 of alternating dielectric and conductive layers. For illustrative purposes, one of the dielectric layers is labeled D, and one of the conductive layers (also referred to as a word line layer) is labeled W. The number of alternating dielectric and conductive layers can vary based on specific implementation requirements. As will be explained below, in one embodiment, the alternating dielectric and conductive layers are divided into six (or a different number of) regions (e.g., sub-blocks) by isolation regions IR. Figure 4 An isolation region IR separating two sub-blocks is shown. The source line layer SL lies beneath alternating dielectric and word line layers. Memory vias are formed within the stack of alternating dielectric and conductive layers. For example, a memory via is labeled MH. Note that in... Figure 4 In the diagram, the dielectric layers are depicted as a perspective view, allowing the reader to see the memory holes located within the stack of alternating dielectric and conductive layers. In one embodiment, NAND strings are formed by filling the memory holes with a material including a charge-trapping material to form vertical columns of memory cells. Each memory cell can store one or more data bits. Further details of a three-dimensional monolithic memory array including memory structure 202 are provided below.

[0076] The memory system discussed above can be erased, programmed, and read. At the end of a successful programming process, the threshold voltage of the memory cell should, where appropriate, be within one or more distributions of the threshold voltages of the memory cells used for programming or within the distribution of the threshold voltages of the erased memory cells. Figure 5A This is a graph of threshold voltage versus the number of memory cells, illustrating an exemplary threshold voltage distribution of the memory array when each memory cell stores one bit of data per memory cell. A memory cell that stores one bit of data per memory cell is called a single-level cell (“SLC”). The data stored in an SLC memory cell is called SLC data; therefore, SLC data comprises one bit per memory cell. Data stored as one bit per memory cell is SLC data. Figure 5A Two threshold voltage distributions are shown: E and P. Threshold voltage distribution E corresponds to the erase data state. Threshold voltage distribution P corresponds to the program data state. Therefore, memory cells with a threshold voltage in threshold voltage distribution E are in the erase data state (e.g., they are erased). Therefore, memory cells with a threshold voltage in threshold voltage distribution P are in the program data state (e.g., they are programmed). In one embodiment, erased memory cells store data "1", and programmed memory cells store data "0". Figure 5A The reference voltage Vr is described. By testing (e.g., performing one or more sensing operations) whether the threshold voltage of a given memory cell is higher or lower than Vr, the system can determine whether the memory cell is erased (state E) or programmed (state P). Figure 5A The verification reference voltage Vv is also described. In some implementations, when memory cells are programmed to data state P, the system will test whether these memory cells have a threshold voltage greater than or equal to Vv.

[0077] Figures 5B to 5F An exemplary threshold voltage distribution for a memory array is shown when each memory cell stores multiple bits of data per memory cell. A memory cell that stores multiple bits of data per memory cell is called a multi-level cell (“MLC”). The data stored in an MLC memory cell is called MLC data; therefore, MLC data comprises multiple bits per memory cell. Data stored as multiple bits per memory cell data is MLC data. Figure 5B In one exemplary implementation, each memory cell stores two bits of data. Other implementations may use other data capacities per memory cell (e.g., three, four, five, or six bits of data per memory cell).

[0078] Figure 5BA first threshold voltage distribution E for erasing memory cells is shown. Three threshold voltage distributions A, B, and C for programming memory cells are also depicted. In one embodiment, the threshold voltage in distribution E is negative, and the threshold voltages in distributions A, B, and C are positive. Figure 5B Each distinct threshold voltage distribution corresponds to a predetermined set of values ​​for a set of data bits. In one implementation, each of the two data bits stored in the memory cell resides in a different logical page, referred to as the next page (LP) and the previous page (UP). In other implementations, all data bits stored in the memory cell reside in a common logical page. The specific relationship between the data programmed into the memory cell and the threshold voltage level of that cell depends on the data encoding scheme adopted by the cell. Table 1 provides exemplary encoding schemes.

[0079] Table 1

[0080] E A B C LP 1 0 0 1 UP 1 1 0 0

[0081] In one implementation known as full-sequence programming, it is possible to use Figure 6 The process directly programs memory cells from an erased data state E to any of a programmed data state A, B, or C (discussed below). For example, a group of memory cells to be programmed can be erased first, leaving all memory cells in the group in an erased data state E. The programming process then directly programs the memory cells to data states A, B, and / or C. For example, while some memory cells are being programmed from data state E to data state A, other memory cells are being programmed from data state E to data state B and / or from data state E to data state C. Figure 5B The arrow indicates full-sequence programming. In some implementations, data states AC can overlap, where the memory controller 120 (or memory die 211) relies on error correction to identify the correct data being stored.

[0082] Figure 5C An exemplary threshold voltage distribution for memory cells is depicted, where each memory cell stores three bits of data per memory cell (another example of MLC data). Figure 5CEight threshold voltage distributions corresponding to eight data states are shown. The first threshold voltage distribution (data state) Er represents an erased memory cell. The other seven threshold voltage distributions (data states) AG represent programmed memory cells and are therefore also referred to as programmed states. Each threshold voltage distribution (data state) corresponds to a predetermined set of data bits. The specific relationship between the data programmed into the memory cell and the threshold voltage level of that cell depends on the data encoding scheme adopted by that cell. In one implementation, Gray code allocation is used to assign data values ​​to a range of threshold voltages such that if the memory's threshold voltage is erroneously shifted to its adjacent physical state, only one bit will be affected. Table 2 provides examples of encoding schemes for implementations where each of the three bits of data stored in the memory cell is in a different logical page, referred to as the next page (LP), middle page (MP), and previous page (UP).

[0083] Table 2

[0084] Er A B C D E F G UP 1 1 1 0 0 0 0 1 MP 1 1 0 0 1 1 0 0 LP 1 0 0 0 0 1 1 1

[0085] Figure 5C Seven read reference voltages, VrA, VrB, VrC, VrD, VrE, VrF, and VrG, are shown for reading data from memory cells. By testing (e.g., performing a sensing operation) whether the threshold voltage of a given memory cell is higher or lower than the seven read reference voltages, the system can determine the data state (i.e., A, B, C, D, ...) of the memory cell.

[0086] Figure 5C Seven verification reference voltages, VvA, VvB, VvC, VvD, VvE, VvF, and VvG, are also shown. In some embodiments, when memory cells are programmed to data state A, the system tests whether these memory cells have a threshold voltage greater than or equal to VvA. When memory cells are programmed to data state B, the system tests whether these memory cells have a threshold voltage greater than or equal to VvB. When memory cells are programmed to data state C, the system determines whether these memory cells have a threshold voltage greater than or equal to VvC. When memory cells are programmed to data state D, the system tests whether these memory cells have a threshold voltage greater than or equal to VvD. When memory cells are programmed to data state E, the system tests whether these memory cells have a threshold voltage greater than or equal to VvE. When memory cells are programmed to data state F, the system tests whether these memory cells have a threshold voltage greater than or equal to VvF. When memory cells are programmed to data state G, the system tests whether these memory cells have a threshold voltage greater than or equal to VvG. Figure 5CIt also shows Vev, which is the voltage level used to test whether a memory cell has been correctly erased.

[0087] In implementations utilizing full sequence programming, the following can be used: Figure 6 The process involves directly programming memory cells from an erased data state Er to any of the programmed data states AG (discussed below). For example, the group of memory cells to be programmed can be erased first, leaving all memory cells in the group in an erased data state Er. Then, the programming process is used to directly program the memory cells to data states A, B, C, D, E, F, and / or G. For example, while some memory cells are being programmed from data state ER to data state A, other memory cells are being programmed from data state ER to data state B and / or from data state ER to data state C, and so on. Figure 5C The arrow indicates full-sequence programming. In some embodiments, the data state AG may overlap, where the control die 211 and / or memory controller 120 rely on error correction to identify the correct data being stored. It should be noted that in some embodiments, the system may use a multi-pass programming process known in the art instead of full-sequence programming.

[0088] Generally, during verification and read operations, the selected word line is connected to a voltage (an example of a reference signal), the level of which is specific to each read operation (see, for example, [reference]). Figure 5C The read comparison levels VrA, VrB, VrC, VrD, VrE, VrF, and VrG) or verification operations (e.g., see [link to relevant documentation]). Figure 5C The verification target levels (VvA, VvB, VvC, VvD, VvE, VvF, and VvG) are specified to determine whether the threshold voltage of the relevant memory cell has been reached. After the word line voltage is applied, the conduction current of the memory cell is measured to determine whether the memory cell is turned on (conducted current) in response to the voltage applied to the word line. If the conduction current is measured to be greater than a certain value, then it is assumed that the memory cell is turned on and the voltage applied to the word line is greater than the threshold voltage of the memory cell. If the conduction current is not measured to be greater than a certain value, then it is assumed that the memory cell is not turned on and the voltage applied to the word line is not greater than the threshold voltage of the memory cell. During the read or verification process, unselected memory cells are provided with one or more read pass voltages (also known as bypass voltages) at their control gate, causing these memory cells to operate as transmit gates (e.g., conducting current regardless of whether these memory cells are being programmed or erased).

[0089] There are many methods to measure the conduction current of a memory cell during a read or verification operation. In one example, the conduction current of the memory cell is measured as the rate at which the memory cell discharges or charges a dedicated capacitor in a sense amplifier. In another example, the conduction current of a selected memory cell allows (or does not allow) the NAND string of the memory cell to discharge to the corresponding bit line. The voltage on the bit line is measured after a certain period of time to see if it has discharged. It should be noted that the techniques described herein can be used in conjunction with various methods known in the art for verification / reading. Other read and verification techniques known in the art can also be used.

[0090] Figure 5D The threshold voltage distribution is depicted when each memory cell stores four bits of data (another example of MLC data). Figure 5D The diagram illustrates that some overlap may exist between the threshold voltage distributions (data states) S0-S15. This overlap can occur due to factors such as memory cell charge loss (and thus a drop in threshold voltage). Programming interference can unintentionally increase the threshold voltage of a memory cell. Similarly, read interference can unintentionally increase the threshold voltage of a memory cell. Over time, the position of the threshold voltage distribution can change. Such changes can increase the bit error rate, thereby increasing decoding time or even making decoding impossible. Changing the read reference voltage can help mitigate such effects. Using ECC during the read process can correct errors and ambiguities. Note that in some implementations, the threshold voltage distributions of a group of memory cells storing four bits of data per memory cell do not overlap and are separated from each other; for example, as... Figure 5E What is depicted. Figure 5D The threshold voltage distribution will include reading the reference voltage and verifying the reference voltage, as discussed above.

[0091] When each memory cell uses four bits, the memory can be programmed using the full-sequence programming discussed above or the multi-pass programming process known in the art. Figure 5D Each threshold voltage distribution (data state) corresponds to a predetermined set of values ​​for a set of data bits. The specific relationship between the data programmed into the memory cell and the threshold voltage level of that cell depends on the data encoding scheme adopted by that cell. Table 3 provides examples of encoding schemes in which each of the four bits of data stored in the memory cell resides in different logical pages, referred to as the next page (LP), middle page (MP), previous page (UP), and top page (TP).

[0092] Table 3

[0093] S0 S1 S2 S3 S4 S5 S6 S7 S8 S9 S10 S11 S12 S13 S14 S15 TP 1 1 1 1 1 0 0 0 0 0 1 1 0 0 0 1 UP 1 1 0 0 0 0 0 0 1 1 1 1 1 1 0 0 MP 1 1 1 0 0 0 0 1 1 0 0 0 0 1 1 1 LP 1 0 0 0 1 1 0 0 0 0 0 1 1 1 1 1

[0094] Figure 5FThe threshold voltage distribution is depicted when each memory cell stores five bits of data (another example of MLC data). In one exemplary implementation, when the memory cell stores five bits of data, the data is stored in any of thirty-two data states (e.g., S0-S31).

[0095] Figure 6 This is a flowchart describing one implementation of the process for programming memory cells. For the purposes of this document, the terms program and programming are synonymous with write and writing. In one exemplary implementation, the memory array 202 is executed using one or more control circuits discussed above (e.g., system control logic unit 260, column control circuit 210, row control circuit 220). Figure 6 The process. In one exemplary implementation, Figure 6 The process is performed by integrated memory component 207 using one or more control circuits of control die 211 (e.g., system control logic unit 260, column control circuit 210, row control circuit 220) to program memory cells on memory die 201. This process includes multiple loops, each including a programming phase and a verification phase. Figure 6 The process is used to achieve full-sequence programming as well as other programming schemes including multi-pass programming. When implementing multi-pass programming, Figure 6 The process is used to implement any / every pass of the multi-pass programming process.

[0096] Typically, during programming operations (via selected data word lines), the programming voltage applied to the control gate is applied as a series of programming pulses (e.g., voltage pulses). Between the programming pulses is a set of verification pulses (e.g., voltage pulses) to perform verification. In many implementations, the amplitude of the programming pulses increases by a predetermined step size with each successive pulse. Figure 6In step 602, the programming voltage signal (Vpgm) is initialized to an initial amplitude (e.g., approximately 12V to 16V or another suitable level), and the programming counter PC maintained by state machine 262 is initialized to 1. In one embodiment, a set of memory cells selected for programming (referred to herein as the selected memory cells) are programmed simultaneously and all connected to the same word line (the selected word line). There may be other memory cells not selected for programming (unselected memory cells) also connected to the selected word line. That is, the selected word line will also be connected to memory cells that should be disabled for programming. Furthermore, when the memory cells reach their expected target data state, they will be disabled for further programming. These NAND strings (e.g., unselected NAND strings) boost their channels to disable programming; these strings include the memory cells to be disabled for programming connected to the selected word line. When the channel has a boosted voltage, the voltage difference between the channel and the word line is insufficient to induce programming. To assist in boosting, in step 604, the control die precharges the channels of the NAND strings that include the memory cells connected to the selected word lines that will be disabled for programming. In step 606, a NAND string including a memory cell connected to the selected word line to be disabled for programming is boosted to disable programming. Such a NAND string is referred to herein as an "unselected NAND string". In one embodiment, the unselected word line receives one or more boost voltages (e.g., about 7 to 11 volts) (also referred to as pass voltages) to perform a boost scheme. A programming disable voltage is applied to the bit line coupled to the unselected NAND string.

[0097] In step 608, a programming voltage pulse of the programming voltage signal Vpgm is applied to the selected word line (the word line selected for programming). If a memory cell on the NAND string should be programmed, the corresponding bit line is biased at the programming enable voltage. In step 608, programming pulses are simultaneously applied to all memory cells connected to the selected word line, such that all memory cells connected to the selected word line are programmed simultaneously (unless they are disabled for programming). That is, they are programmed at the same time or during an overlap period (both are considered simultaneous). In this way, all memory cells connected to the selected word line will have their threshold voltage changes simultaneously, unless they are disabled for programming.

[0098] In step 610, the memory cell that has undergone programming verification and reached its target state is locked and cannot be further programmed by the control die. Step 610 includes performing programming verification by sensing at one or more verification reference levels. In one embodiment, the verification process is performed by testing whether the threshold voltage of the selected memory cell for programming has reached the appropriate verification reference voltage. In step 610, after the memory cell has been verified (by the test of Vt) that the memory cell has reached its target state, the memory cell can be locked.

[0099] If, in step 612, it is determined that all memory cells have reached their target threshold voltage (pass), the programming process is complete and successful because all selected memory cells have been programmed and verified to their target state. In step 614, a "pass" status is reported. Otherwise, if it is determined in 612 that not all memory cells have reached their target threshold voltage (failure), the programming process continues to step 616.

[0100] In step 616, the number of memory cells that have not yet reached their respective target threshold voltage distributions is counted. That is, the number of memory cells that have failed to reach their target state so far is counted. This counting can be performed by state machine 262, memory controller 120, or another circuit. In one embodiment, there is a total count that reflects the total number of currently programmed memory cells for which the last verification step failed. In another embodiment, a separate count is maintained for each data state.

[0101] In step 618, it is determined whether the count from step 616 is less than or equal to a predetermined limit. In one embodiment, the predetermined limit is the number of bits that can be corrected by error correction codes (ECC) during the page read process of a memory cell. If the number of failed cells is less than or equal to the predetermined limit, the programming process can stop and a "pass" status is reported in step 614. In this case, enough memory cells have been correctly programmed so that ECC can be used during the read process to correct the remaining few memory cells that have not yet been fully programmed. In some embodiments, the predetermined limit used in step 618 is lower than the number of bits that can be corrected by error correction codes (ECC) during the read process to allow for future / additional errors. The predetermined limit can be a fraction (proportional or non-proportional) of the number of bits that can be corrected by ECC during the page read process of a memory cell when programming fewer than all memory cells of a page, or when comparing counts of only one data state (or fewer than all states). In some embodiments, the limit is not predetermined. Instead, it varies based on the number of errors already counted for the page, the number of program erase cycles performed, or other criteria.

[0102] If the number of failed memory cells is not less than a predetermined limit, the programming process continues at step 620 and the programming counter PC is checked against the programming limit value (PL). Examples of programming limit values ​​include 6, 12, 16, 19, 20, and 30; however, other values ​​can be used. If the programming counter PC is not less than the programming limit value PL, the programming process is considered to have failed and a "failure" status is reported in step 624. If the programming counter PC is less than the programming limit value PL, the process continues at step 626, during which the programming counter PC is incremented by 1, and the programming voltage signal Vpgm is stepped to the next amplitude. For example, the next pulse will have an amplitude ΔVpgm larger than the previous pulse (e.g., a step size of 0.1 volts to 1.0 volts). After step 626, the process loops back to step 604, and another programming pulse is applied (to the selected word line) to cause execution... Figure 6 Another iteration of the programming process (steps 604 to 626).

[0103] In one implementation, memory cells are erased before programming, and erasure is the process of changing the threshold voltage of one or more memory cells from a programming data state to an erase data state. For example, changing the threshold voltage of one or more memory cells from... Figure 5A The state changes from P to E. Figure 5B The state changes from A / B / C to E. Figure 5C The state AG changes to state Er or from Figure 5DThe states S1-S15 change to state S0.

[0104] One technique for erasing memory cells in some memory devices is to bias a p-well (or other type) substrate to a high voltage to charge the NAND channel. When the NAND channel is at a high voltage, an erase enable voltage (e.g., a low voltage) is applied to the control gate of the memory cell to erase the non-volatile memory element (memory cell). In this document, this is referred to as p-well erase.

[0105] Another method for erasing memory cells is to generate a gate-induced drain leakage (GIDL) current to charge the NAND string channel. An erase enable voltage is applied to the control gate of the memory cell while maintaining the NAND string channel potential to erase the memory cell. In this paper, this is referred to as GIDL erase. Both p-well erase and GIDL erase can be used to reduce the threshold voltage (Vt) of the memory cell.

[0106] In one embodiment, a GIDL current is generated by inducing a drain-to-gate voltage at a select transistor (e.g., SGD and / or SGS). The drain-to-gate voltage of the transistor that generates the GIDL current is referred to herein as the GIDL voltage. A GIDL current is generated when the drain voltage of the select transistor is significantly higher than the control gate voltage of the select transistor. The GIDL current is a result of carrier generation, i.e., electron-hole pair generation due to band-band tunneling and / or trap-assisted generation. In one embodiment, the GIDL current can cause one type of carrier (e.g., holes) to move predominantly into the NAND channel, thereby raising the channel potential. Another type of carrier (e.g., electrons) is extracted from the channel by an electric field along the direction of the bit line or along the direction of the source line. During erasure, holes can tunnel from the channel to the charge storage region of the memory cell and recombine with electrons therein to lower the threshold voltage of the memory cell.

[0107] A GIDL current can be generated at either end of the NAND string. A first GIDL voltage can be generated between the two terminals of a select transistor (e.g., a drain-side select transistor) connected to or near the bit line to generate a first GIDL current. A second GIDL voltage can be generated between the two terminals of a select transistor (e.g., a source-side select transistor) connected to or near the source line to generate a second GIDL current. An erase based on the GIDL current at only one end of the NAND string is called a single-sided GIDL erase. An erase based on the GIDL current at both ends of the NAND string is called a double-sided GIDL erase.

[0108] In some implementations, the control die or memory die performs the ECC decoding process (see ECC engine). Error correction is used to help correct errors that may occur when storing data. During the programming process, the ECC engine encodes the data to add ECC information. For example, the ECC engine is used to create codewords. In one implementation, data is programmed on a page-by-page basis. Error correction is used in conjunction with the programming of data pages because errors can occur during programming or reading, and errors can occur when storing data (e.g., due to electronic drift, data retention problems, or other phenomena). Many error correction coding schemes are well known in the art. These conventional error correction codes (ECCs) are particularly useful in high-capacity memories, including flash memory (and other non-volatile) memories, because such coding schemes can have a significant impact on manufacturing yield and device reliability, making devices with a small number of unprogrammable or defective cells usable. Of course, there is a trade-off between yield savings and the cost of providing additional memory cells to store code bits (i.e., encoding "rate"). Therefore, some ECC codes are better suited to flash memory devices than others. Generally, ECC codes for flash memory devices tend to have a higher encoding rate (i.e., a lower code bit / data bit ratio) than codes used in data communication applications (which can have encoding rates as low as half). Well-known examples of ECC codes commonly used with flash memory storage devices include Reed-Solomon codes, other BCH codes, Hamming codes, etc. Sometimes, error correction codes used with flash memory storage devices are "systematic" because the data portion of the final codeword does not change from the actual data being encoded, with code or parity bits appended to the data bits to form the complete codeword. In other implementations, the actual data is altered.

[0109] Specific parameters for a given error-correcting code include the type of code, the size of the block from which the actual data from which the codeword is derived, and the total length of the encoded codeword. For example, a typical BCH code applied to 512 bytes (4096 bits) of data can correct up to four error bits if at least 60 ECC or parity bits are used. Reed-Solomon codes are a subset of BCH codes and are also commonly used for error correction. For example, a typical Reed-Solomon code can correct up to four errors in a 512-byte data sector using approximately 72 ECC bits. In the case of flash memory, error-correcting codes provide significant improvements in manufacturing yield and the reliability of flash memory over time.

[0110] In some implementations, the controller receives host data, also known as information bits, which will be stored in a memory structure. Information bits are represented by a matrix i = [1 0] (note that two bits are used for illustrative purposes only, and many implementations have codewords longer than two bits). An error-correcting coding process (such as any process mentioned above or below) is implemented where parity bits are added to the information bits to provide data represented by a matrix or codeword v = [1 0 1 0], indicating that two parity bits have been appended to the data bits. Other techniques can be used to map input data to output data in a more complex manner. For example, low-density parity-check (LDPC) codes, also known as Gallager codes, can be used. Further details on LDPC codes can be found in RG Gallager's "Low-density parity-check codes," *IRE Trans. Inform. Theory*, vol. IT-8, pp. 21-28, Jan. 1962; and D. MacKay's *Information Theory, Inference and Learning Algorithms*, Cambridge University Press, 2003, chapter 47. In practice, such LDPC codes are typically applied to multiple pages encoded across multiple memory elements, but they do not necessarily need to be applied across multiple pages. Data bits can be mapped to logical pages and stored in memory structure 326 by programming one or more memory cells to one or more programming states corresponding to v.

[0111] In one possible implementation, an iterative probabilistic decoding process is used, which corresponds to the error-correcting decoding of the encoding implemented in controller 120. Further details regarding iterative probabilistic decoding can be found in the aforementioned D. MacKay text. Iterative probabilistic decoding attempts to decode the codeword by assigning an initial probability metric to each bit in the codeword. The probability metric indicates the reliability of each bit, that is, the probability that the bit is not erroneous. In one approach, the probability metric is the log-likelihood ratio (LLR) obtained from an LLR table. The LLR value is a measure of the known reliability of the values ​​of the various binary bits read from the storage element.

[0112] The bit LLR is given by the following formula:

[0113]

[0114] Where P(v = 0 | Y) is the probability that a bit is 0 given a read of Y, and P(v = 1 | Y) is the probability that a bit is 1 given a read of Y. Therefore, an LLR > 0 indicates that the bit is more likely to be 0 than 1, while an LLR < 0 indicates that the bit is more likely to be 1 than 0, to satisfy one or more parity checks of the error-correcting code. Furthermore, a larger magnitude indicates a greater probability or reliability. Therefore, a bit with LLR = 63 is more likely to be 0 than a bit with LLR = 5, and a bit with LLR = -63 is more likely to be 1 than a bit with LLR = -5. An LLR = 0 indicates that the bit can also be either 0 or 1.

[0115] An LLR value can be provided for each bit position in a codeword. Furthermore, the LLR table can specify multiple reads, allowing a larger LLR value to be used when bit values ​​are consistent across different codewords.

[0116] The controller receives codeword Y1 and LLR and iterates in successive iterations, where the controller determines whether the parity (equation) of the error encoding process has been satisfied. If all parity checks have been satisfied, the decoding process has converged and the codeword has been corrected. If one or more parity checks have not been satisfied, the decoder adjusts the LLR of one or more bits that are inconsistent with the parity check, and then reapplies the parity check or the next check in the process to determine whether it has been satisfied. For example, the magnitude and / or polarity of the LLR can be adjusted. If the parity check in question is still not satisfied, the LLR can be adjusted again in another iteration. In some, but not all, adjusting the LLR may result in bit flipping (e.g., from 0 to 1 or from 1 to 0). In one implementation, once the parity check in question is satisfied, another parity check is applied to the codeword, if applicable. In other cases, the process moves to the next parity check and later loops back to the failed check. The process continues to attempt to satisfy all parity checks. Thus, the decoding process of Y1 is completed to obtain decoding information including parity bit v and decoding information bit i.

[0117] Figure 7 The standard read flow combining ECC correction and read error handling is illustrated. Step 701 involves reading data stored in a memory cell to determine a "hard bit" (HB), where the hard bit value corresponds to the value used. Figures 5A to 5C The standard reading of the value Vri is used to distinguish different states (if they are as follows). Figures 5A to 5C(A well-defined, separated distribution as in the example). Step 703 uses ECC technology to determine whether the read data is correctable, and if so, the read process is completed at step 705. When the hard bit data becomes uncorrectable by ECC in step 703, a read error handling flow can be invoked at step 707, which may involve various read types to recover the read data. Depending on the implementation, examples of read types that can be used to recover the data content are: “CFh read” 711, which is a reread of the hard bits but allows the bias level, such as the voltage of the selected word line, to have a longer settling time; “soft bit” read 713, which provides information about the reliability of the hard bit value; “BES read” 715, which attempts to offset the hard bit read level to extract the data; and “DLA read” 717, which takes into account the influence of adjacent word lines on the word line selected for reading. One or more of these can be combined in various sequences or combinations to attempt and extract the data content in the event that the basic ECC process fails. For any of the implementation schemes, performance is typically severely degraded once the read error handling process 707 is invoked as step 703. The following considers techniques for using soft-bit data while mitigating its impact on memory performance. Figure 8 Consider the uses of soft positions in more detail.

[0118] Figure 8 It can be used to illustrate the concepts of hard bits and soft bits. Figure 8 The diagram illustrates the overlap of distributions of two adjacent data states and a set of read values ​​that can be used to determine the data state of a cell and the reliability of such reads. The corresponding hard and soft bits for a specific encoding of the value are shown in the table below. The read value VH is the initial data state value or hard read value used to determine the hard bit (HB) value, and corresponds to... Figure 5A , Figure 5B or Figure 5C The value Vri is used to distinguish different states (if they are as follows). Figures 5A to 5C (A well-defined, separated distribution like that in [the original text]). Additional read levels, VS+ with a margin slightly above VH and VS- with a margin slightly below VH, are "soft read" values ​​and can be used to provide "soft bit" (SB) values. Soft bit values ​​give information about the quality or reliability of the initial data state or hard bit data, as soft bit data provides information about the extent to which the distribution has been expanded. Some implementations of ECC codes (such as Low-Density Parity-Check (LDPC) codes) can use both hard and soft bit data to increase their capabilities. Although... Figure 8Only one pair of soft-bit readouts is shown, but other implementations can use additional readouts with margin to generate more soft-bit values ​​for a given hard bit if higher resolution is desired. More generally, a hard bit corresponds to a hypothetical data value based on sensing operations, and soft information (which can be a single binary soft bit, multiple soft bits, or a decimal / minute value) indicates the reliability or confidence level of the hard-bit value. When used in ECC methods that utilize soft information, the soft information can be considered as the probability that the corresponding hard-bit value is correct.

[0119] During a read operation, if VH is below the memory cell threshold, the memory cell will be non-conductive, and the read data value (HB) will be read as "0". If the memory cell is below the threshold, the memory cell will be non-conductive. Figure 8 If the data is within the central region of either distribution, then reads of VS+ and VS- will provide the same result; if these reads differ, the threshold voltage of the memory cell lies between these values ​​and may originate from the tail region of either the upper or lower distribution, making the HB data unreliable. If the data is considered reliable, reading at both levels and performing an XOR NOT operation on the result gives an SB value of "1"; if unreliable, an SB value of "0" is given.

[0120] For example, when both SB+ and SB- read "0", then:

[0121] SB = (SB+) × NOR (SB-)

[0122] = "0" x NOR "0"

[0123] =1,

[0124] SB=1 and the HB read value will be considered reliable. During soft bit decoding in ECC, this will result in memory cells in the upper distribution having HB="0" and SB="1" to indicate a reliable correct bit (RCB), while memory cells with a threshold voltage between SB+ and SB- will result in SB="0" to indicate that the HB value is unreliable.

[0125] Figure 9A and Figure 9B The read levels for calculating the hard and soft bit values ​​of the next page data in a three-bit data implementation per memory cell using the encoding in Table 2 above are shown respectively, wherein the soft bit values ​​1 and 0 indicate that the hard bit value is reliable and unreliable, respectively. Figure 9A This shows the threshold voltage distribution of memory cells in each 3-bit unit, similar to... Figure 5CAs shown, the distribution is not well-defined and exhibits some overlap. This overlap can arise from several causes, such as charge leakage or interference, where an operation on one word line or bit line affects the data state stored in a nearby memory cell. Furthermore, in actual write operations, the distribution typically does not behave as expected. Figure 5C The definition is not ideal because writing to memory cells with such accuracy is detrimental to performance, as a large number of fine-grained programming steps and some cells will be difficult or overly fast to program. Therefore, programming algorithms typically allow for a degree of overlap, relying on ECC to accurately extract user data content.

[0126] The read points used to distinguish the data values ​​for the next page are represented by vertical dashed lines between Er and A states, and between D and E states, and the corresponding hard bit values ​​written below. Due to overlapping distributions, multiple memory cells storing Er or E data will be incorrectly read as HB = 0, and multiple memory cells storing A or D data will be incorrectly read as HB = 1. For example, the optimal read value can be determined as part of the device characterization and stored as the fuse value of the control circuit. In some implementations, the control circuit can offset these values ​​to improve its accuracy as part of a standard read operation or as part of the read error handling flow 707 of the BES read 715.

[0127] To handle higher error rates, a more robust ECC can be used. However, this requires storing more parity bits, reducing the proportion of memory cells available for user data and thus effectively decreasing memory capacity. Furthermore, performance is impacted by the increased computation involved in encoding / decoding codewords and writing and reading additional ECC data. Additionally, more ECC data needs to be transferred to and from the ECC circuitry via the data bus structure.

[0128] Figure 9B It shows that it can be used to determine the corresponding Figure 9A The soft bit value and read point of the next page hard bit value. As shown, the soft bit value is determined to either side of the basic hard bit read value based on a pair of reads. For example, these soft bit read values ​​may be based on an offset from the hard bit read value, and may be symmetrical or asymmetrical, and are stored as fuse values ​​in a register that is determined as part of the device characterization. In other implementations, they may be determined or updated dynamically. Although using soft bits at step 713 may be quite efficient in retrieving data content that cannot be retrieved in step 703, it comes with a performance penalty because it requires being invoked in response to an ECC failure at step 703, uses two additional reads for each hard bit read, requires the soft bit data to be transferred out after the additional reads, and requires additional calculations.

[0129] To improve this situation, an implementation scheme of "Effective Soft Sensing Mode" is introduced below. In this sensing mode, hard and soft reads can be combined into a sequence using two sensing levels to sense time efficiency. By using Effective Soft Sensing Read as the default mode, additional soft bit information can be provided for ECC correction, while triggering a read error handling process. Since only two sensing operations are used to generate hard and soft bit data, this technique avoids the threefold increase in sensing time caused by standard hard and soft reads. Furthermore, by merging hard and soft bit sensing into a single sequence, most of the additional overhead involved in read sequence operations (e.g., enabling charge pumps, ramping word lines, etc.) can be avoided. Figure 10 The use of the effective soft sensing mode is shown.

[0130] Figure 10 The assignment of hard and soft bit values, along with the read level, is shown in an implementation for effective soft sensing. Figure 10 Similar to Figure 8 The diagram illustrates a distribution of memory cells Vth that again overlaps with two data states in the central region. A hard read is performed again, but instead of attempting to place it at or near the center of the overlapping region at an optimized point to distinguish the two states, in this embodiment, the hard read is offset to the lower Vth side, such that any memory cell read at or below VH is reliably in the lower data state (shown here as "1", as in the exemplary diagram). Figure 8 (in Chinese). It also assigns a soft bit value of "0", which is related to... Figure 8 Compared to the previous implementation, the SB=0 value now indicates a reliable HB value. If a memory cell reads above VH, its hard bit value corresponds to a higher Vth data state with HB=0. Figure 10 The implementation plan is different from Figure 8 Two soft-bit reads are performed, but only a single soft-bit read is executed with the VS value offset to the high Vth side. If a memory cell's Vth is found to be higher than VS, it is assigned an HB value of HB=0 and is considered reliable (HS=0). For memory cells with a Vth found between VH and VS, the memory cell is assigned HB=0 but is considered unreliable (SB=1). Note that in Figure 10In this implementation, only one of the two states is checked for the soft bit data, such that only the HB=0 state can have any SB value, while the HB=1 memory cell will always have SB=0. In other words, the soft bit data is determined only on one side of the overlapping distribution (here, the lower side, for HB=0), and not on the other side (here, the higher side, for HB=1). In this implementation, a single VS read is performed on the left side of the VH read (higher Vth), but in other implementations, this arrangement can be reversed.

[0131] although Figure 10 The total amount of data generated in the implementation plan is less than Figure 8 The total amount of data, but Figure 10 An effective soft-sensing mode will generally be sufficient to extract user data content without excluding further read error handling. Because in Figure 10 The determination involves only two readings, so the sensing time is short, and can be further reduced by performing the two readings as a single sensing operation, such as relative to... Figure 12 As stated above. Additionally, less data is transmitted to one or more ECC engines: in Figure 8 In the middle, four combinations of (HB,SB) data are generated, while Figure 10 Of these, only three combinations exist, or 25% less data is lost. The increased error tolerance provided by effective soft sensing also improves write performance because the data does not need to be precisely programmed, allowing for relaxed programming tolerances.

[0132] Figure 11 The diagram illustrates how the coding in Table 2 is used in a three-bit data implementation per memory cell to apply an effective soft-sensing mode to the next page of data. Figure 11 Similar to Figure 9A and Figure 9BHowever, the HB and SB values ​​are combined into a single graph, and a single SB read level is used for a given HB read level for effective soft sensing, instead of a pair of SB reads for a given HB. For example, observe the difference between Er and A states. For a left-hand read, the left memory cell is reliably "1" for the next page value, where (HB,SB) = (1,0). Note again that in this encoding, SB = 0 indicates a reliable HB value, and SB = 1 indicates an unreliable HB value. For a right-hand read of Er and A, the right memory cell indicates a memory cell with a reliable next page value of "0" or (HB,SB) = (0,0). Memory cells with Vth between the left and right read levels are assigned a next page hard value of 0 but are considered unreliable, such that (HB,SB) = (0,1). Similarly, for reads that distinguish between states D and E, the memory cell to the left of the left read is reliably "0" ((HB,SB)=(0,0)), the memory cell above the right read is a reliable next page "1" data ((HB,SB)=(1,0)), and the memory cell in between is assigned an unreliable next page value of "1" ((HB,SB)=(1,1)).

[0133] Figure 12 It shows the corresponding Figure 11 The illustration shows an implementation scheme for the sensing operation of the read point in an effective soft-sensing read operation for the next page data read operation. At the top, Figure 12 This diagram illustrates the relationship between the control gate read voltage VCGRV waveform that can be applied to the word line of a selected memory cell and the effective soft sensing time of the 3 bits of next-page data per memory cell, where the vertical dashed line corresponds to... Figure 11 The four readings are also marked with dashed lines (but the determined order is different, as will be explained below). Below the waveform, it is shown how these readings using the waveform at the top correspond to the Vth values ​​of the D and E state distributions.

[0134] To improve read time performance Figure 12The implementation uses a "reverse order" read mode, but other implementations can use a standard sequence. In a standard read sequence, the read voltage applied to the selected memory cell starts at a lower limit and gradually increases. In reverse order read mode, the control gate read voltage (VCGRV) applied to the selected word line initially ramps up to a high value and then the read is performed from a higher Vth state to a lower Vth state. In this example of a next-page read, reads distinguishing between D and E states are performed before reads distinguishing between A state and Er state. Therefore, after the initial ramp, the VCGRV voltage drops to the read level of E state (ER) and then drops to the read level of A state (AR). This sequence reduces the time required for the significant additional overhead involved in read sequence operations (e.g., enabling charge pumps, ramping up word lines, etc.).

[0135] For each read voltage level, two sensing operations are performed to generate a hard bit and a soft value, allowing for a faster sensing time than using a single read voltage. (Reference) Figure 12 The D and E state distributions at the bottom, the dashed lines for the HB and SB boundaries, both represent relatively close Vth values, but the SB boundary is offset to the right with a higher Vth value. Therefore, in embodiments where sensing is based on discharging voltage through selected memory cells, if a read voltage ER is chosen, the HB and SB Vth values ​​are conducted to some extent, but by different amounts. The HB boundary corresponds to a lower Vth value because the memory cell at the SB boundary will be more conductive, thus discharging faster, and can be determined using a shorter sensing interval. The slower-discharging SB boundary is sensed using the same control gate voltage, but with a longer sensing time.

[0136] Figure 13 An implementation scheme of a sensing amplifier circuit that can be used to determine the hard and soft bit values ​​of a memory cell is shown. Figure 13 The sensing amplifier circuit can correspond to Figure 2A or Figure 2B The sensing amplifier 230, and as included in Figure 3 In the structure. In Figure 13 In one implementation, the state of the memory cell is determined by pre-charging the sensing line or node SEN 1305 to a predetermined level, connecting the sensing node to the bit line of the bias-selected memory cell, and determining the degree to which the node SEN 1305 discharges within the sensing interval. Many variations are possible depending on the implementation, however... Figure 13The implementation scheme illustrates some typical components. Node SEN 1305 can be precharged to level VHLB via switch SPC 1323, where many MOSFET switches are notated here using the same names and corresponding control signals as transistors, where various control signals can be controlled by processor 330, state machine 262, and / or... Figure 2A , Figure 2B and Figure 3 Other control elements are provided for the implementation scheme. Node SEN1305 can be connected to the selected memory cell via switch XXL 1319 to node SCOM 1307 along bit line BL 1309, and then connected after possible intermediate elements to bit line selection switch BLS 1327 corresponding to the decoding and selection circuitry of the memory device. SEN node 1305 is connected to the local data bus LBUS 1301 via switch BLQ 1313, which can then be connected to data DBUS 1303 via switch DSW 1311. Switch LPC 1321 can be precharged to level VLPC, where the values ​​of VHLB and VLPC depend on the details of the implementation scheme and specific implementation.

[0137] In the sensing operation, the selected memory cell is biased by setting its corresponding selected word line to the read voltage level as described above. In a NAND array implementation, the selected gate of the NAND string of the selected word line and the unselected word line are also biased to ON. Once the array is biased, the selected memory cell conducts a level based on the relationship between the applied read voltage and the threshold voltage of the memory cell. Capacitor 1325 can be used to store charge on SEN node 1305, wherein during pre-charging, level CLK (and the lower plate of capacitor 1325) can be set to a low voltage (e.g., ground or VSS) such that the voltage on SEN node 1305 references this low voltage. The pre-charge SEN node 1305 of the selected memory is connected to the corresponding bit line 1309 via XXL 1319 and BLS 1327 to the selected bit line and is allowed to discharge to a level depending on the threshold voltage of the memory cell within the sensing interval relative to the voltage level applied to the control gate of the selected memory cell. At the end of the sensing interval, XXL 1319 can be turned off to trap the resulting charge on SEN 1305. At this time, the CLK level can be slightly increased, similarly increasing the voltage on SEN 1305 to account for the voltage drop across intermediate components (such as XXL 1319) in the discharge path. Therefore, the voltage level on SEN 1305 that controls the degree to which transistor 1317 is turned on will reflect the data state of the selected memory cell relative to the applied read voltage. Local data LBUS 1301 is also pre-charged so that LBUS will discharge to the CLK node as determined by the voltage level on SEN 1305 during the continuous gating interval when turn-on transistor STB 1315 is on. At the end of the gating interval, STB 1315 is turned off to set the sensed value on LBUS, and the result can be latched into one of the latches, such as... Figure 3 As shown.

[0138] Now return to the reference. Figure 12 After biasing the selected memory cell to the ER voltage level and other array biases (select gate, unselected word line, etc.) as needed, the precharged SEN node 1305 is discharged against the interval ER between the dashed lines: if the level on SEN is high enough to discharge LBUS 1301 when STB1315 is selected, the Vth of the memory cell is lower than HB; otherwise, it is higher than HB. After discharging the additional interval ER+, STB 1315 is re-selected: if LBUS 1301 is now discharged, the Vth of the memory cell is between HB and SB; otherwise, it is higher than SB. This process is then repeated with the VCGRV value at the AR level to determine the HB and SB values ​​used to distinguish between the A state and the erase state.

[0139] Therefore, in relation to Figure 12 Under the illustrated implementation, for each VCGRV level, the left sensing result is used to generate HB data, and the right sensing result is combined with the left sensing result to generate SB data. To optimize the performance of the two senses (left / right), Figure 12 The implementation scheme uses "sensing time modulation" for Vth separation without word line voltage level changes.

[0140] In contrast to the effective soft-sensing read level control and parameters, similar to typical implementations of read parameters, these can be determined as part of the device characterization process and stored as register values ​​(such as control data parameters set to fuse values ​​in memory device 266), dynamically determined, or some combination thereof. In one set of implementations, the hard-bit and soft-bit read levels for effective soft sensing can be referenced to standard hard read values. Even when using the effective soft-sensing read process as the default read operation, memory devices typically have a standard read (i.e., hard-bit only) read mode option, making... Figures 5A to 5C The standard read value will be available as a read option. For example, return to see Figure 11 The read levels associated with distinguishing between the D and E state distributions can be referenced to the effective soft-sensing level relative to the normal HB read trim value, represented as a heavier dashed line at the apex of the D and E state distributions. The effective soft-sensing read levels for left reads (effective soft-sensing hard bit, decremented) and right reads (effective soft-sensing soft bit, incremented) can be specified relative to the normal HB read levels. This allows for the reuse of the group feature register to generate effective soft-sensing left / right offsets, and in one set of implementations, a common setting can be used across all planes, with separate settings for each state.

[0141] Figure 14 This is a high-level flowchart of an implementation scheme for effective soft sensing operations. (See the above section regarding...) Figures 1 to 4 The memory system and relative to Figure 12 The process is described within the context of the described embodiment. The process begins at step 1401 to perform a first sensing operation on a plurality of memory cells to determine a hard bit value that distinguishes between two data states in the data states of the memory cells. In an effective soft-sensing embodiment, both the hard bit read in step 1401 and the soft bit read in step 1403 can be responded to a single read command. For example, see [link to previous section] Figure 1 The host 102 and / or non-volatile memory controller 120 can issue valid soft-sensing commands to one or more of the memories 130. Then, the system control logic unit 260 ( Figure 2A and Figure 2B Performing sensing operations, such as reading the next page data in the example above, to determine both the hard bit value and the soft bit value of the memory cell, such as... Figure 11 As shown.

[0142] To perform the hard bit determination in step 1401, in the above embodiment, the memory array is biased for read operations, and the sensing nodes of one or more corresponding sense amplifiers are pre-charged. More specifically, for the embodiment used as an example herein, the control gate of the selected memory cell is biased by the read voltage through its corresponding word line for distinguishing between data states, and other array elements (e.g., selected gates and unselected word lines of NAND strings) are biased as needed based on the memory architecture. When using, such as Figure 13 When using a sensing amplifier such as a sensing amplifier, in the case of determining the data state while discharging the sensing node SEN 1305, the sensing node SEN 1305 is pre-charged and connected to the bit line of the selected memory cell in the first sensing interval ( Figure 12 Discharge within the ER(HB) boundary region to determine the hard potential value.

[0143] As relative to Figure 11 As shown in the implementation, the hard bit determination is offset to a lower Vth value, so memory cells sensed below this value are reliably within that value, while memory cells sensed above this value include both reliable and unreliable hard bit values. In an implementation using more conventional sequential sensing, hard bit sensing for the hard bit is performed first, followed by soft bit sensing for differentiation between Er and A states, and then both hard and soft bit sensing for differentiation between D and E states, each of which involves different biases and sensing node pre-charges for each sensing operation. In relation to... Figure 12 In the reverse sequence sensing operation shown, hard and soft bit values ​​are first determined for the D and E states, and then hard and soft bit values ​​are determined for the Er and A states. Although Figure 14 The process typically involves hard bit determination (step 1401) preceding soft bit determination (step 1403), but in some implementations, the order can be reversed. Additionally, Figure 14 The process involves only a single hard bit and a single soft bit determination, which is insufficient in many cases (such as in...). Figure 12 (In the middle), multiple hard / soft bit pairs will be identified.

[0144] At step 1403, a second sensing operation is performed to determine soft bits. During effective soft sensing, this is reliability information determined only for memory cells that have a first hard bit value but not a second hard bit value. For example, in Figure 11 In the implementation scheme, when the hard bit boundary shifts downward, the soft bit value is used only for the higher of the hard bit values. In relation to... Figure 12In the described implementation, the second sensing operation is based on the longer discharge time of the pre-charged sensing node SEN 1305. If the read involves distinguishing between a pair of states (as in the binary memory cell implementation), only one hard-bit / soft-bit pair is determined. In the case of a multilevel memory cell, as described above... Figure 11 and Figure 12 In the example, additional hard and soft bit pairs are determined, where the next page sensing operation is similar to steps 1401 and 1403 used for Er / A state determination in determining the hard and soft bit pairs. Once the hard and soft bit data values ​​are determined, they can be used to perform ECC operations at step 1405. This can be done on the non-volatile memory controller 120 in ECC engine 158, on control die 211, or some combination thereof.

[0145] While the use of effective soft sensing reduces the amount of soft data determined and thus the amount of soft data to be transferred to the ECC engine compared to a standard hard-soft arrangement, there is still a significant increase in data compared to using only hard data. To reduce the amount of data that needs to be transferred from the memory die to the ECC engine, the soft data can be compressed in memory before being transferred to the non-volatile memory controller via the bus structure. The following discussion presents techniques for compressing soft data. These techniques can be applied to both effective soft sensing and standard soft sensing, but the following discussion will primarily use examples of effective soft sensing implementations.

[0146] More specifically, the exemplary implementations presented below will be primarily based on the above description relative to... Figures 10 to 14 The described effective soft-sensing mode. As mentioned above, when using soft bit data, the effective soft-sensing mode reduces performance degradation, making it practically the default read mode, where one page of hard bit data and one page of soft bit data are output in a read sequence. These pages of soft bit data and hard bit data are then transmitted to the error correction engine to extract the data content of that page of user data. In some implementations, it is possible to... Figure 2B Control die 211 or Figure 2A Some or all of the ECC operations are performed on the memory die 200, but typically the ECC operations are performed on the ECC engine 158 on the non-volatile memory controller 120, thus requiring the read hard and soft bit data to be transferred to the controller 120 via an external data bus structure through interface 269. For example, in a 3D NAND memory implementation, data from a single page of a single plane could be 16KB of user data plus corresponding parity bits and redundant data for defective memory locations. Therefore, without compression, in addition to 16+ kilobytes of hard data per plane, 16+ kilobytes of soft data per plane will also be transferred.

[0147] To maintain memory performance, soft-bit data can be compressed on memory die 200 or control die 211 before transmission. For example, if a compression factor N is used, the amount of soft-bit data transmitted is reduced by 1 / N; therefore, the choice of compression factor is a trade-off between the speed and amount of soft-bit data available to the ECC engine. Various compression techniques can be used with different compression factors. For example, a compression factor of N=4 can be achieved by performing a logical AND operation on the soft-bit data in a group of four soft bits. While this will not indicate the individual reliability of the corresponding hard bit values, it will indicate that at least one of the four hard bit values ​​in a group should be considered unreliable.

[0148] Figure 15 This is a block diagram illustrating an implementation of some control circuitry elements for a memory device, including soft-bit compression elements. The example shown is for a four-plane memory device, and most of the shown elements can be repeated for each plane; however, other implementations may use fewer or more planes. Depending on the implementation, these one or more control circuits may be as follows: Figure 2B Such control circuitry is located on a control die 211 bonded to one or more memory dies 201. In other embodiments, the one or more control circuits may be located on the memory die 200 containing the memory array 202, such as on the periphery of the memory die 200 or on a substrate formed beneath the aforementioned 3D NAND memory structure.

[0149] exist Figure 15 To simplify the diagram, only the common blocks of plane 3 1501-3 are labeled. However, it should be understood that each of the common blocks in planes 0 1501-0, 1 1501-1, 2 1501-2, and 3 1501-3 includes corresponding common blocks 1505, 1507, and 1509. These blocks correspond to... Figure 2A and Figure 2B The text describes the elements of the row control circuit 220, column control circuit, and system control logic unit 260, but more specifically how these elements are physically arranged in some embodiments. On either side of each plane are row decoders 1503-L and 1503-R, which decode the connections of the array to the plane's word lines and select lines and can correspond to... Figure 2A and Figure 2B The row decoder 222 and other components of the row control circuit 220. The column control circuit 1509 can correspond to Figure 2A and Figure 2B The column control circuit 210. Above and below the column control circuit of column 1509 is a set of sense amplifiers 1505 (including internal data latches) and cache buffers 1507. See again. Figure 3The sense amplifier circuit 1505's internal data latch can correspond to ADL, BDL, and CDL data latches, and the cache buffer 1507 can correspond to the transfer data latch XDL. Although not labeled, another plane includes similar elements. Conversely, other planes include arrows indicating the data flow between the plane's memory cells and I / O interfaces, where similar transfers may also occur in planes 3 1501-3, but are not shown so that block markings can be indicated.

[0150] Figure 15 The one or more control circuits presented also include input-output or I / O circuitry (including I / O pads 1517) and a data path (DP) block 1515 that performs (multi-bit) serial-to-parallel conversion for inbound write data and parallel-to-(multi-bit) serial conversion for outbound read data. The DP block 1515 is connected to byte-wide (in this example) I / O pads 1517 to transfer data to and from the non-volatile memory controller 120 via an external data bus. Figure 15 In the block diagram, DP block 1515 and IO pad 1517 are located at plane 1 1501-1. However, these components can be placed on any of these planes or distributed between them, but positioning these components on one of the central planes (plane 1 1501-1 or plane 2 1501-2) reduces wiring. The global data bus GDB1511 within the memory device spans these planes, allowing data to be transferred to and from individual planes and DP block 1515. Figure 15 The vertical arrows illustrate the data flow between the upper part of the sense amplifier block 1505 and the IO pads 1517, where these arrows are not shown on planes 3 1501-3 to allow for block marking. During a read, data pages from the memory array of the plane are sensed by the sense amplifier 1505 and stored in the corresponding internal data latch, then shifted to the cache buffer 1507 of the transfer latch, and continue to be decoded by the control circuitry of column 1509 to reach the global data bus 1511. The hard-bit data then continues from the global data bus 1511 through the DP block 1515 to be arranged into (byte-wide) serial data for outward transmission through the IO pads 1517. When writing data, the data flow can be reversed along the path used by the hard-bit data.

[0151] Regarding the corresponding soft bit data, after determining the soft bit data (for both valid and conventional soft sensing operations), the soft bit data is compressed before being transferred from the memory device to the ECC engine. The implementation for compressing soft bit data presented in the following discussion performs compression within the SA / internal data latch 1505 and the transfer latch of the cache buffer 1507. After compression, the compressed soft bit data can travel along the same path as the hard bit data from the cache buffer 1507 to the I / O pad 1517. Since the compression process can affect the logical address allocation of the soft bit data, the DP block 1515 may include mapping logic to ensure that the compressed soft bit data is properly allocated. Figure 16 , Figure 17A and Figure 17B Further details are provided regarding implementation schemes for data latches that can be used in soft bit data compression processes.

[0152] Figure 16 yes Figure 15 The SA / internal data latch 1505 and the cache buffer 1507 transfer latch and Figure 3 A schematic diagram illustrating the correspondence between data latch groups 340, 342, 344, and 346. Internal data latches associated with the sense amplifier of the SA / internal DL may include latches ADL, BDL, and CDL, as well as the sense amplifier data latch (SDL) and possibly other data latches, depending on the implementation. The cache memory includes the transfer data latch XDL and also includes additional latches, such as the DTCT latch for temporary data storage and operations such as bit scan operations. Internal data latches are connected along the local data bus LBUS, where internal data latches are connected to the transfer data latches along the data bus DBUS. This is relative to... Figure 17A Shown in more detail.

[0153] Figure 17A This is a schematic diagram of the structure of one implementation of a data latch. Figure 17AThe example is an implementation with 3 bits per cell, where each sense amplifier (SA) has a set of associated data latches forming a "layer," which includes a sense amplifier data latch (SDL), data latches for 3-bit data states (ADL, BDL, CDL), and auxiliary data latches (TDL) for implementing, for example, fast write operations. Within each of these data latch stacks, data can be transferred along the local bus LBUS between the sense amplifier and its associated group of latches. In some implementations, each sense amplifier in the sense amplifier and the corresponding internal group of data latches associated with a bit line in the layer can be grouped together for corresponding bit line "columns" and formed on the memory die, within the spacing of the memory cell array along the periphery of the memory cell array. The example discussed here uses an implementation where 16 bit lines are formed into a column so that 16-bit words are physically positioned together in the array. An example of the memory array could have 1000 such columns, corresponding to 16K bit lines. In one implementation topology, each sense amplifier and its associated layer's data latch group are connected along the DBUS internal bus structure, allowing data to be transferred between each latch in that layer and its corresponding XDL along the DBUS internal bus structure. In the implementation described below, the XDL transfer latch can transfer data to and from the I / O interface, but other data latches in that layer (e.g., ADLs) are not arranged to transfer data directly to or from the I / O interface and must be mediated by the transfer data latch XDL.

[0154] Figure 17B Show Figure 17A The implementation plan of the group of columns. Figure 17B Will Figure 17A The structure is repeated 16 times, with only the internal data latches of layer 0 for each group shown. Each DBUS connects to a set of 16 XDLs. Each horizontal row (as shown in the figure) connects to one XBUS line in the XBUS line, such that the lowest row of XDLs (or "XDL layers") connects to the XBUS. <0> The next line or next level XDL connects to XBUS. <1> Similarly, the XDL layer 15 is connected to the XBUS. <15> . Figure 17B An arrangement of DTCT latches is also shown, wherein one DTCT latch corresponds to each sense amplifier layer / DBUS value, and this DTCT latch is also connected to one of the XDL layer / XBUS values. In this arrangement, each DTCT latch is connected to the DBUS for the i-th value. and XBUS This connects the leftmost DTCT to the DBUS. <0> and XBUS <0> The next DTCT connects to DBUS <1> and XBUS <1> And so on, until the rightmost DTCT is connected to the DBUS. <15> and XBUS <15> The compression of soft-bit data values ​​within the data latch structure described below will... Figure 17B The implementation is presented in the context of the present invention, but other implementations may be used. Furthermore, although discussed in the context of compressed soft bit data and more specifically in the context of effective soft sensing implementations, these compression techniques can be applied to compress other data stored in memory devices.

[0155] In the vertical compression scheme, data is compressed within a word unit, where soft bits (or other data) are stored in an internal data latch, compressed, and then written to the XDL latch, or alternatively, written back to the internal data latch. Then, during stream output, the compressed data is reordered within the mapping logic of DP block 1515 to place the compressed data in logical user column order. Figure 18 The first step in the process of an exemplary implementation is shown.

[0156] Figure 18 This illustrates the compression of raw soft-bit data from one set of internal data latches to another. In this example, the raw soft-bit data is stored in an ADL latch and compressed with a compression factor N=4, then stored in a BDL latch; however, other implementations may use other combinations of internal data latches. Figure 18 The top consists of a 16×16 table with vertically arranged sense amplifier (SA) layers and horizontally arranged XDL layers, as shown below. Figure 17B As shown, the entries correspond to soft-bit data from valid soft-sensing operations, and the squares with "0" values ​​are highlighted by dotted lines. In the vertical compression scheme, for each XDL layer (i.e., Figure 18 In the N=4 compression, the data is compressed from 16 values ​​to 4 values, allowing the compressed data to be stored only in the four sense amplifier layers of the BDL latch. SA layers 0-3 store the compressed data, and dummy entries (here, "1") are input to the other SA layers of the BDL latch. More specifically, in Figure 18 In the implementation scheme, the soft bit data of each XDL layer is grouped into 4 SA layer groups according to the following compression algorithm, and the values ​​are ANDed:

[0157] SA layer[3:0]:BDL[0]=&ADL[3:0];

[0158] SA layer[7:4]:BDL[1]=&ADL[7:4];

[0159] SA layer[11:8]:BDL[2]=&ADL[11:8];and

[0160] SA layer[15:12]:BDL[3]=&ADL[15:12],

[0161] The & represents the logical AND of ADL entries: for example, &ADL[3:0] = ADL[3] AND ADL[2] AND ADL[1] AND ADL[0].

[0162] Figure 18 The lower part shows the entries of compressed valid soft-sensing data values ​​in the BDL latch, where the compressed soft bit data is in the bottom four rows of SA layers 0-3, as highlighted by the dotted lines, and the other rows are filled with dummy "1" values. For example, looking at the first column of XDL layer 0 values, the SA layer 0 value is "0" to reflect the presence of "0" in ADL[3:0] at SA layer 1, while SA layers 1, 2, and 3 are all "1" because there are no other "0" values ​​in ADL[7:4], ADL[11:8], or ADL[15:12]. In the example of XDL layer 15 (rightmost column), the "0" soft bit values ​​in SA layers 0 and 13 of ADL are reflected in the compressed value of BDL[3:0] = (0110). Note that under this N=4 compression algorithm, the compressed value "0" indicates that at least one of the four soft bit values ​​is "0", making all four corresponding hard bit values ​​considered unreliable. Although compressed soft-bit data does not offer the same level of resolution as uncompressed soft-bit data, it is often still sufficient to significantly aid in decoding the data. It should also be noted that, for the purposes of discussion, Figure 18 The example of soft data at the top has a relatively high number of "0" values.

[0163] Figure 19 The N=4 compression from the original ADL<15:0> data to BDL<3:0> is shown again, but this is in an implementation using a position-based compression algorithm. The original soft bits are compressed again and stored in SA layers 0-3 of the BDL latch, where dummy "1" values ​​are entered in the other SA layers. To illustrate the position-based algorithm, exemplary uncompressed soft bit data in the ADL latch has "0" along the anti-diagonal of the entry, where the number of SA layers is the same as the number of XDL layers. The BDL<3:0> values ​​are 4-bit values ​​indicating the SA layer with a "0" soft bit value. For example, in XDL layer 2, "0" is the 4-bit binary value (0010) = 2 corresponding to SA layer 2. For other cases, such as no "0" soft bit value or more than one "0", one of these values ​​(such as (1111)) can also be used.

[0164] Figure 20 This illustrates how compressed soft-bit data is moved from the BDL SA layer <3:0> to different SA layers based on XDL layer information. In this implementation, Figure 18 The compressed soft-bit data values ​​of the XDL layers are shifted upwards by a certain amount to different SA layers, which is done cyclically. XDL layers 0, 4, 8, and 12 are not shifted; XDL layers 1, 5, 9, and 13 are shifted upwards by 4 SA layers; XDL layers 2, 6, 10, and 14 are shifted upwards by 8 SA layers; and XDL layers 3, 7, 11, and 15 are shifted upwards by 12 SA layers. In some implementations, regarding... Figure 18 or Figure 19 The compression steps shown can be combined with Figure 20 The rearrangement and combination. Other latches can be used to store the XDL layer information of this movement, where the data can be created for either the XBUS side or the DBUS side. Figure 20 In the implementation scheme, XDL layer information is stored as a 2-bit value in the CDL and TDL latches, (00) indicating that the compressed data has not been moved, (01) indicating that it has been moved up 4 SA layers, (10) indicating that it has been moved up 8 SA layers, and (11) indicating that it has been moved up 12 SA layers. Then, the compressed data is moved out of the internal data latch and into the XDL transfer latch, which can be a standard direct move from BDL to XDL, so that the data will be as follows: Figure 20 It was rearranged as before, but now it's in the XDL latch, which is... Figure 21 It is shown at the top.

[0165] Figure 21 This illustrates transmission within a transmission latch to compress data. See again. Figure 17B XDL-to-XDL transfers can be performed using temporary storage in the DTCT latch via the XBUS structure. For example... Figure 21 As shown, compressed data in each of XDL layers 1, 2, and 3 can be transferred to XDL layer 0; compressed data in each of XDL layers 4, 5, 6, and 7 can be transferred to XDL layer 1, and so on, for the other XDL layers... Figure 21 The grouping and arrows indicate this. At the end of the process, the compressed data is merged or compressed into XDL layers 0-3, where previous data can remain in other layers until it needs to be rewritten at a later time. The purpose of this compression is to combine compressed data into a limited number of XDL layers to reduce data output time (in the example with compression factor N=4, outputting from layer 16 to layer 4).

[0166] See again Figure 15 Once the compressed soft data has been loaded into the transfer latch of cache buffer 1507, such as Figure 21 As shown, it can be transmitted to the global data bus 1511 via the control circuit of column 1509, and then to the data path block 1515, so that it can be transmitted out via the IO pad 1517. However, relative to Figures 18 to 21 The data movement shown can affect the position of data within an external data latch relative to its local user address. To address this issue, the mapping logic in data path block 1515 can be used to reorganize the data bits of the compressed soft data into a continuous logical column order during data output.

[0167] Figure 22 This is a diagram illustrating the rearrangement of compressed data bits into a logical order. For example... Figure 21 As shown at the bottom, the compressed data is formed as a 16-bit word in the XDL layer, which is transferred from the transfer latch of cache buffer 1507 to DP block 1515; however, due to the... Figures 18 to 21 The operation shown, as part of the data output process, involves reorganizing data bits to move them in logical user column address order. In an exemplary implementation, where... Figure 21 The diagram shows how to shift 4-bit groups of compressed data to form a 16-bit word in the XDL layer, using the mapping logic in DP block 1515 to rearrange these 4-bit units. Figure 22 One implementation is shown where data is received from a global data bus 1511 in an 8×16-bit word-parallel format, where word W0 has bits 0-15, W1 has bits 16-31, and so on, up to W7 with bits 112-127, as shown on the left. These words W0-W7 are divided into groups of 4 bits to form words W0'-W7', where the first 4 bits of each of W0-W3 go into W0', the next 4 bits of each of W0-W1 go into W2', the third 4 bits of each of W0-W1 go into W4', and the last 4 bits of each of W0-W1 go into W6'. Words W4-W7 are similarly rearranged into W1', W3', W5', and W7'. These words W0'-W7' can then be output on IO pads 1517 in a (byte-width) serial format with this rearranged word order.

[0168] Compared to Figure 2A and Figure 2B The DP block 1515, including mapping logic, may be part of the interface 268 under the control of the system control logic unit 260. According to the implementation, additional latches or FIFOs and multiplexing circuitry may be included in the DP block 1515 to facilitate the reorganization of compressed data into logical user column address order.

[0169] Figure 23 This is a block diagram of an alternative implementation of some control circuit elements in the control circuit elements of a memory device including a soft-bit compression element. Figure 23 The implementation plan is repeated Figure 15 The elements, which have a similar number (i.e., cache buffer 1505 is now 2305, global data bus GDB1511 is now 2311, etc.), now also include the mapping logic within the control circuitry of column 2309, as a supplement to or replacement of the mapping logic in DP block 1515. Figure 15 or Figure 23 In the specific implementation of the control circuit, it may be easier or more efficient to reorder all or part of the compressed data bits into a logical order in the column control circuit of column 2309 rather than in DP block 1515 or 2315.

[0170] For example, due to constraints on timing, available circuit area, or circuit topology, it may be difficult to completely reorder the compressed data bits into user logic order using DP block 2315. To completely reorder the compressed data bits, in one implementation, the compressed data bits are partially reordered within the mapping logic of the control circuitry of column 2309, then transmitted to the global data bus 2311, transmitted to DP block 2315, further reordered into full user logic order within the mapping logic of DP block 2315, and then transmitted away from the device via IO pad 2317. Under this arrangement, Figures 18 to 21 The process of compressing and reading soft bits or other data shown can be the same until... Figure 21 The data is merged or compressed at the end. However, the compressed data at this moment is not transmitted via the global data bus 2311 at this time, and only as... Figure 22 The remodeling process is not performed in DP block 2315 as shown, but can be performed as per the description of... Figure 24A and Figure 24B The reorganization is performed as shown.

[0171] Figure 24A and Figure 24B Is using Figure 23 The implementation scheme reorganizes compressed data bits into a logical order, as illustrated in the diagram of an alternative implementation scheme. Figure 24A and Figure 24B The steps shown can begin as follows: Figure 21 The bottom shows the compressed data bits in the XDL latch of the cache buffer 2307. Figure 21 The diagram illustrates data compressed into three XDL layers, such as W0<3:0> for word zero, while other XDL layers remain open for other compressed data words. In this N=4 example, three other data compressions could be stored. For example, if reading a page (e.g., a word line) corresponds to M such words, then (in the N=4 example) the word line could be divided into four partitions of M / 4 words, such that W0 is in the first XDL layer <3:0>, W(M / 4) in XDL layer <7:4>, W(2M / 4) in XDL layer <11:8>, and W(3M / 4) in XDL layer <15:12>, with the next word in each of these partitions grouped with W1, and so on. This is in Figure 24A The top of the table is shown for an example of (M / 4) = 288. Therefore, as can be seen in each line, different bits in a compressed word can include bits from a word that, before compression, comes from widely separated logical words, such that W1 in, for example, the XDL latch of cache buffer 2307, can have compressed data bits from logical words W1, W289, W577, and W864. To classify this, as... Figure 24A As shown at the bottom, when compressed and condensed data bits are transferred from the XBU of cache buffer 2307 to the IOBUS of the control circuitry of column 2309, a first step for refactoring can be performed in one implementation.

[0172] exist Figure 24A In the example, the I / O bus within the control circuitry of column 2309 has two 16-bit wide buses, IOBUS A and IOBUS B. In this first phase of reorganization, the compressed and condensed data words are rearranged by the mapping logic of the control circuitry of column 2309 such that the logical word addresses of a given partition are grouped together. The paired I / O bus values ​​0-8 for IOBUS A and IOBUS B illustrate one implementation for placing the compressed data of the first 64 logical words. In this example, the compressed bits for logical words W0, W16, W32, and W48 are on IOBUS0A, with the logical word value incremented by 1 on IOBUS1A, and so on, with the IOBUS B value incremented by 8 logical words. Details regarding the I / O bus width and number of buses are specific to the implementation and depend on the degree of parallelism used at different stages of the data transmission process.

[0173] By using the multiplexing circuit within the control circuitry of column 2309, compressed data bits can be further processed from... Figure 24A The bottom layout has been rearranged as follows Figure 24B The words in the top arrangement. In one implementation, compressed data bits are transmitted to DP block 2315 via global data bus 2311 in the form of partial reordering, where the multiplexing circuitry within the mapping logic can then perform reordering to conform to the order of sequential logical addresses, such as Figure 24B As shown in the lower part. Within DP block 2315, the multiplexing circuitry within the mapping logic can then complete the reorganization of the compressed data, similar to the... Figure 22 The described process, and then the reorganized compressed data can be transmitted through IO pad 2317.

[0174] Figure 25 This is a flowchart of an implementation scheme for performing data compression within a data latch associated with a sense amplifier of a non-volatile memory device. Starting at step 2501, a read operation is performed using sense amplifier circuitry on multiple memory cells, wherein in step 2503 the result of the read operation is stored in an internal data latch corresponding to the sense amplifier. (Reference) Figure 15 These steps are performed in SA / internal DL 1505, where internal data latches (ADL, BDL, ...) can be arranged to... Figure 17A and Figure 17B In the layered structure shown. Regarding... Figure 2A and Figure 2B The reading process is performed using the control circuits in row control circuit 220, column control circuit 210, and system control logic unit 260, as described above. In the main example here, the data is about... Figures 10 to 14 The soft bit data determined in the effective soft sensing operation. More generally, the compression process can be applied to, for example, regarding... Figures 8 to 9B The described typical type of soft bit data or other data. In any case, in the example above, at the end of step 2503, the soft bit data or other read data is stored in the ADL latch.

[0175] The compression process begins at step 2505, where data is compressed within an internal data latch group, such as... Figure 18 or Figure 19 As shown, the data in ADL<15:0> is compressed and stored in BDL<3:0> of each XDL layer. See again. Figure 3 This process and subsequent data latch operations can be controlled by processor 330 and system logic unit 260. Once compressed, at step 2507, the data can be... Figure 20 The internal data latch shown is rearranged, and then at step 2509, it is copied from the internal data latch to the external data latch XDL of the cache buffer 1507. Depending on the implementation, various variations are possible, such as combining steps 2505 and 2507, using different latches, different compression algorithms, or different rearrangements of the compressed data. Figures 18 to 22 This is just one example. In any case, soft bit data or other data can be compressed within the internal data latch before being transmitted compressed data to the cache buffer 1507.

[0176] Once the compressed data is stored in the transfer latch of the cache buffer 1507, the compressed data can be merged at step 2511, such as in... Figure 21 During compression, the compressed data bits are transmitted more efficiently via the global data bus 1511 to the input-output interface elements and I / O pads 1517 of the DP block 1515. At step 2513, the compressed data bits can be reordered into user logic order, and then transmitted via the external data bus at step 2515. (As mentioned above...) Figure 15 and Figure 22 As described, reorganization can be performed in the mapping logic of DP block 1515, in the mapping logic of the control circuit in column 2309, or in a combination thereof, as per [reference to...]. Figure 23 , Figure 24A and Figure 24B In the case of soft bit data, this may include sending data to the ECC engine (e.g., Figure 1 The transmission of 158). Figure 22 An example of reorganizing compressed data by a specific rearrangement for a previous step in an exemplary implementation is shown, but in other cases, the rearrangement may be different or may not be performed before transmission, but may be reorganized later if necessary.

[0177] Return to Figure 23 Several alternative implementations for data compression operations are possible, using a combination of the mapping logic in the control circuit of column 2309 and the mapping logic within DP block 2315. (Relative to...) Figure 25 The implementation shown below does not perform the rearrangement of compressed data using an internal latch as described in step 2507 or the compression of data within a transfer latch as described in step 2511. Instead, it moves multiple copies of the compressed data to a buffer cache memory and performs different reorganizations on the compressed data.

[0178] In this alternative implementation, the process again begins with data stored in an internal data latch, which, in the exemplary implementation, could again be the result of an effective soft-sensing operation. This could again be Figure 18 or Figure 19 The data at the top of the example. Then, the data is compressed using a compression algorithm in the SA layer direction, again as shown. Figure 18 or Figure 19 As shown at the bottom, as discussed above regarding these two figures or other compression algorithms. This implementation, after vertical compression within the data latch, is similar to... Figures 18 to 25 The difference is not described. The data is not rearranged within the internal data latch 2305 before being moved to the transfer latch of the buffer cache memory 2307 (as described). Figure 20 , Figure 25 (As shown in step 2507), instead, it is moved to the transfer latch of the buffer cache memory 2307, as per the description of step 2507. Figure 26 As shown.

[0179] Figure 26 An implementation scheme is shown for moving compressed data from an internal data latch to a buffer cache memory without first rearranging the compressed data within the internal data latch. At the top, Figure 26 The compressed data in the internal data latch 1505 is shown. In this example, the compressed data in the latch's BDL group is... Figure 18 The bottom is the same, but this can also be like... Figure 19 The bottom or according to other compression algorithms. For example, from... Figures 20 to 21 The top of the transformation ( Figure 25 In step 2509), the compressed data is transferred from the internal data latch 2305 (BDL in this example) to the transfer latch of the buffer cache memory 2307. In this case, the BDL SA layer <3:0> is moved to the XDL SA layer <3:0> and also copied to the XDL SA layers <7:4>, <11:8>, and <15:12>, such that the compressed data in the BDL SA layer <3:0> is copied four times in this example with a compression factor N=4. These four copies are... Figure 26 The bottom left side is enclosed in parentheses. In an exemplary embodiment, four copies can be transmitted simultaneously, but they can also be transmitted sequentially. By copying multiple copies of the compressed data into the transmission latch of the buffer cache memory 2307, skipping... Figure 21 and Figure 25 Step 2511 involves compressing the data within the transfer latch, and the data is transmitted via the global data bus to each partition of the control circuitry in column 2309 and to DP block 2315. In some implementations, (e.g., Figure 18 or Figure 19 Compression and Figure 26 The transmission can be combined into one step.

[0180] During the transmission of compressed data from the XDL latch of cache register 2307 to the global data bus 2311 and DP block 2315 via the control circuitry of column 2309, different portions of different copies of the transmitted compressed data are used for different portions of the data placed on the global data bus 2311. In an exemplary embodiment, the intermediate bus in Figure 17B The XBUS is connected to the global data bus GDB 2311. Figure 27 An implementation scheme is shown.

[0181] Figure 27 An implementation scheme for multiplexing data from a data transfer latch from a cache buffer onto the global data bus is shown. For example... Figure 27 As shown, the intermediate bus IOBUS is connected to Figure 17B The XBUS is connected to the global data bus 2711. Multiplexing circuit 2799, part of the mapping logic in the control circuitry of column 2309, can multiplex data from the XDL transfer latches of different XBUS cache buffers 2307 (XBUS<3:0>, XBUS<7:4>, XBUS<11:8>, XBUS<15:12>) onto the intermediate bus IOBUS, and then transfer it to the global data bus 2711 without additional mapping logic. Using multiplexing circuit 2799, a first copy of the compressed data in the XDL latches is transferred to IOBUS for <3:0>, a second copy for <7:4>, a third copy for <11:8>, and a fourth copy for <15:12>. Then, the compressed data on the IOBUS can be transferred to the global data bus 2711 in a loop, and then the data is transferred to the DP block 2315 via the global data bus 2311.

[0182] Once received within DP block 2315, the compressed data can be reorganized into contiguous logical column addresses, similar to... Figure 22 and Figure 25 As described in step 2513. This can again be implemented via multiplexing and logic circuitry to place the compressed data in a FIFO, as represented in the mapped logic block of DP block 2315, to place the compressed data in the FIFO for transmission to the IO pad 2317. When transmitting the compressed data from the FIFO to the IO pad 2317, a process similar to... Figure 24B The final restructuring.

[0183] Figure 28 This is a flowchart of an additional implementation scheme for performing data compression within a data latch associated with a sense amplifier of a non-volatile memory device. Relative to... Figure 25 The process can be described as above. Figure 25 Steps 2501, 2503, and 2505 are implemented respectively to achieve steps 2801, 2803, and 2805. At step 2807, multiple copies of the compressed data from the internal data latch 2305 are transferred to the transfer latch of the buffer cache memory 2307, thereby filling as... Figure 26 The example shown illustrates the transfer latch. The primary example here uses a compression factor of N=4, where N=4 copies are transferred to the XDL latch. Similarly, if N=2 compression is used, two copies will be transferred, and if N=8 is used, eight copies will be transferred, and so on for other values, where the number of latches will be a multiple of the compression factor. In some implementations, steps 2805 and 2807 may be combined into a single step.

[0184] In step 2809, compressed data in the transfer latches (XDL latches) of the buffer cache memory 2307 is loaded onto the global data bus 2311 and transferred to DP block 2315. In one embodiment, by using the multiplexing circuit 2309 in the control circuitry of column 2309, a first copy of the compressed data in the XDL latches is transferred to the global data bus 2311 for <3:0>, a second copy of the compressed data in the XDL latches is transferred to the global data bus 2311 for <7:4>, a third copy of the compressed data in the XDL latches is transferred to the global data bus 2311 for <11:8>, and a fourth copy of the compressed data in the XDL latches is transferred to the global data bus 2311 for <15:12>. In DP block 2315, where at step 2811, the multiplexing circuitry within the mapping logic can then perform reordering to conform to the order of sequential logical addresses, such as... Figure 24B The lower part and Figure 25 As shown in step 2513, the reorganized compressed data is loaded into the FIFO in the output circuit of DP block 2315. In step 2813, the data from the FIFO is transferred to the IO pads 2317 of the input-output circuit, similar to... Figure 25 Step 2515, where the final reorganization is similar to Figure 24B .

[0185] According to a first aspect, a non-volatile memory device includes control circuitry configured to be connected to a plurality of bit lines, each bit line being connected to a corresponding plurality of memory cells. The control circuitry includes: a plurality of sense amplifiers, each sense amplifier configured to read data from a memory cell connected to one or more corresponding bit lines; a plurality of internal data latch groups, each internal data latch group configured to store data associated with a corresponding sense amplifier; a cache buffer including a plurality of transfer data latch groups, each transfer data latch group corresponding to one of the internal data latch groups; an input-output interface configured to provide data to an external data bus; and an internal data bus configured to transfer data from the cache buffer to the input-output interface. The control circuit is configured to: perform a read operation on a plurality of memory cells by each of the sense amplifiers; store the result of the read operation performed by each of the sense amplifiers in a corresponding internal data latch group; compress the result of the read operation performed by each of the sense amplifiers in the corresponding internal data latch group; transfer multiple copies of the compressed result of the read operation from each of the internal data latch groups to a corresponding transmission data latch group; transfer multiple copies of the compressed result of the read operation from the transmission data latch to an input-output interface via an internal data bus; and transfer the compressed result of the read operation to the external data bus via the input-output interface.

[0186] In another aspect, the method includes: performing a read operation on a plurality of memory cells by each of a plurality of sense amplifiers; storing the result of the read operation performed by each of the sense amplifiers in a corresponding internal data latch group; performing a data compression operation on the result of the read operation performed by each of the sense amplifiers in the corresponding internal data latch group; transferring a plurality of copies of the compressed result of the read operation from each of the internal data latch groups to a corresponding transfer data latch group; and transferring the plurality of copies of the compressed result of the read operation from the transfer data latch to an input-output interface via an internal data bus.

[0187] The additional aspect includes a non-volatile memory device comprising: a plurality of bit lines, each bit line connected to a corresponding plurality of non-volatile memory cells; a plurality of sense amplifier circuits, each sense amplifier circuit configured to read data from memory connected to one or more corresponding bit lines; a plurality of internal data latch groups, each internal data latch group configured to store data associated with a corresponding sense amplifier among the sense amplifiers; a cache buffer including a plurality of transfer data latch groups, each transfer data latch group corresponding to one of the internal data latch groups; an input-output interface configured to provide data to an external data bus; an internal data bus configured to transfer data from the transfer data latches to the input-output interface; and one or more control circuits connected to the sense amplifier circuits, the internal data latch groups, the cache buffer, and the input-output interface. One or more control circuits are configured to: perform a read operation on a plurality of corresponding memory cells by each of a plurality of sense amplifiers; store the result of the read operation performed by each of the sense amplifiers in a corresponding internal data latch group; compress the result of the read operation performed by each of the sense amplifiers in the corresponding internal data latch group; transfer a plurality of copies of the compressed result of the read operation from each of the internal data latch groups to a corresponding transmission data latch group; and transfer a plurality of copies of the compressed result of the read operation from the transmission data latch group to an input-output interface via an internal data bus.

[0188] For the purposes of this document, references to "implementation scheme," "one implementation scheme," "some implementation schemes," or "another implementation scheme" in the specification may be used to describe different implementation schemes or the same implementation scheme.

[0189] For the purposes of this document, a connection may be a direct connection or an indirect connection (e.g., via one or more other components). In some cases, when a component is referred to as being connected or coupled to another component, the component may be directly connected to the other component or indirectly connected to the other component via an intermediary component. When a component is referred to as being directly connected to another component, there is no intermediary component between the two components. If two devices are directly or indirectly connected, the two devices are "communicating," enabling them to exchange electronic signals with each other.

[0190] For the purposes of this document, the term "based on" may be understood as "at least partially based on".

[0191] For the purposes of this document, the use of numerical terms such as “first” object, “second” object, and “third” object without additional context may not imply an ordering of objects, but may be used for identification purposes to distinguish different objects.

[0192] For the purposes of this document, the term "group" of objects may refer to a "group" of one or more objects.

[0193] The detailed description above has been provided for purposes of illustration and description. It is not intended to be exhaustive or to limit the precise forms disclosed in the invention. Many modifications and variations are possible based on the teachings above. The described embodiments were chosen to best explain the principles of the proposed technology and its practical application, thereby enabling others skilled in the art to best utilize it in various embodiments and various modifications suitable for the specific intended use. The scope of the invention is intended to be defined by the appended claims.

Claims

1. A non-volatile memory device, comprising: A control circuit, configured to be connected to a plurality of bit lines, each bit line being connected to a corresponding plurality of memory cells, the control circuit comprising: Multiple sense amplifiers, each of which is configured to read data from the memory cell connected to one or more corresponding bit lines; Multiple internal data latch groups, each configured to store data associated with a corresponding sense amplifier among the sense amplifiers; A cache buffer, the cache buffer comprising a plurality of transfer data latch groups, each transfer data latch group corresponding to one of the internal data latch groups; Input-output interface, configured to provide data to an external data bus; and An internal data bus, configured to transfer data from the cache buffer to the input-output interface. The control circuit is configured as follows: Each of the sensing amplifiers performs a read operation on multiple memory cells; The result of the read operation performed by each of the sensing amplifiers is stored in the corresponding internal data latch group; Within the corresponding internal data latch group, the result of the read operation performed by each of the sense amplifiers is compressed; Multiple copies of the compressed result of the read operation are transferred from each of the internal data latch groups to the corresponding transport data latch group. The multiple copies of the compressed result of the read operation are transferred from the transfer data latch to the input-output interface via the internal data bus; and The compressed result of the read operation is transmitted to the external data bus via the input-output interface.

2. The non-volatile memory device according to claim 1, wherein the control circuit is formed on a control die, and the non-volatile memory device further comprises: The memory die includes the plurality of bit lines and the corresponding plurality of non-volatile memory cells, and the memory die is separately formed and coupled to the control die.

3. The non-volatile memory device of claim 1, wherein the result of the read operation is a soft bit data value.

4. The non-volatile memory device according to claim 3, wherein the control circuit is further configured to: Each of the sensing amplifiers performs a hard bit read operation on the plurality of memory cells to determine a hard bit value for each of the plurality of memory cells, the hard bit value indicating whether the memory cell is reliably in a first data state or unreliably in a second data state, wherein each soft bit data value corresponds to one of the hard bit values, and indicates a reliability value for memory cells determined to be in the second data state, but not for memory cells determined to be in the first data state.

5. The non-volatile memory device according to claim 1, wherein: During the transfer of the plurality of copies of the compressed result of the read operation from each of the internal data latch groups to the corresponding transfer data latch group, the control circuit is further configured to simultaneously transfer the plurality of copies of the compressed result of the read operation from each of the internal data latch groups to the corresponding transfer data latch group. and During the transfer of multiple copies of the compressed result of the read operation from the transfer data latch to the input-output interface via the internal data bus, the control circuitry is further configured to multiplex the multiple copies of the compressed result of the read operation from the transfer data latch onto the internal data bus.

6. The non-volatile memory device according to claim 1, wherein the control circuit is further configured to: Before transmitting the compressed result of the read operation to the external data bus, the bits of the compressed result of the read operation are reordered.

7. The non-volatile memory device according to claim 1, wherein the control circuit is further configured to: The compressed result of the read operation is received in parallel format at the input-output interface; and The compressed result of the read operation is converted before being transmitted to the external data bus.

8. The non-volatile memory device of claim 1, wherein, in order to compress the result of the read operation performed by each of the sense amplifiers in the corresponding internal data latch group, the control circuit is further configured to: A logical combination is performed on multiple results of the readout operation performed by each of the sensing amplifiers.

9. The non-volatile memory device according to claim 1, further comprising: A non-volatile memory cell array, comprising the plurality of bit lines and the plurality of memory cells corresponding to each bit line, wherein the non-volatile memory cell array is formed according to a three-dimensional NAND architecture.

10. A method comprising: Each of the multiple sense amplifiers performs a read operation on multiple memory cells; The result of the read operation performed by each of the sensing amplifiers is stored in the corresponding internal data latch group; Within the corresponding internal data latch group, a data compression operation is performed on the result of the read operation performed by each of the sensing amplifiers; Multiple copies of the compressed result of the read operation are transferred from each of the internal data latch groups to the corresponding transport data latch group. as well as The multiple copies of the compressed result of the read operation are transferred from the transfer data latch to the input-output interface via the internal data bus.

11. The method of claim 10, wherein the plurality of copies of the compressed result of the read operation are simultaneously transferred from each of the internal data latch groups to the corresponding transport data latch group, and The transfer of multiple copies of the compressed result of the read operation from the transmission data latch to the input-output interface via the internal data bus includes: The multiple copies of the compressed result of the read operation are multiplexed from the transmission data latch onto the internal data bus.

12. The method of claim 10, further comprising: The bits of the compressed result of the read operation in the input-output interface are reordered.

13. The method of claim 12, further comprising: The compressed result of the read operation is received in parallel format at the input-output interface; and Before transmitting the compressed result of the read operation to the external data bus, the compressed result of the read operation is converted into a serial format.

14. The method of claim 10, wherein the result of the read operation is a soft bit data value.

15. The method of claim 14, further comprising: Each of the sensing amplifiers performs a hard bit read operation on the plurality of memory cells to determine a hard bit value for each of the plurality of memory cells, the hard bit value indicating whether the memory cell is reliably in a first data state or unreliably in a second data state, wherein each soft bit data value corresponds to one of the hard bit values, and indicates a reliability value for memory cells determined to be in the second data state, but not for memory cells determined to be in the first data state.

16. The method of claim 15, further comprising: The hard bit value and the corresponding soft bit data value are transmitted from the input-output interface to the error correction code engine.

17. A non-volatile memory device, comprising: Multiple bit lines, each bit line is connected to a corresponding set of multiple non-volatile memory cells; Multiple sense amplifier circuits, each of which is configured to read data from a memory connected to one or more corresponding bit lines; Multiple internal data latch groups, each configured to store data associated with a corresponding sense amplifier among the sense amplifiers; A cache buffer, the cache buffer comprising a plurality of transfer data latch groups, each transfer data latch group corresponding to one of the internal data latch groups; An input-output interface configured to provide data to an external data bus; An internal data bus configured to transfer data from the transfer data latch to the input-output interface; and One or more control circuits, connected to the sense amplifier circuit, the internal data latch group, the cache buffer, and the input-output interface, are configured to: Each of the plurality of sensing amplifiers performs a read operation on a plurality of corresponding memory cells; The result of the read operation performed by each of the sensing amplifiers is stored in the corresponding internal data latch group; Within the corresponding internal data latch group, the result of the read operation performed by each of the sense amplifiers is compressed; Multiple copies of the compressed result of the read operation are transferred from each of the internal data latch groups to the corresponding transport data latch group. as well as The multiple copies of the compressed result of the read operation are transferred from the transfer data latch group to the input-output interface via the internal data bus.

18. The non-volatile memory device of claim 17, wherein the one or more control circuits are further configured to: During the transfer of the plurality of copies of the compressed result of the read operation from each of the internal data latch groups to the corresponding transport data latch group, the plurality of copies of the compressed result of the read operation are simultaneously transferred from each of the internal data latch groups to the corresponding transport data latch group; and During the transfer of multiple copies of the compressed result of the read operation from the transfer data latch to the input-output interface via the internal data bus, the multiple copies of the compressed result of the read operation are multiplexed from the transfer data latch onto the internal data bus.

19. The non-volatile memory device of claim 17, wherein the result of the read operation is a soft bit data value.

20. The non-volatile memory device of claim 17, wherein the non-volatile memory device comprises: A control die, on which one or more control circuits are formed; and The memory die includes the plurality of bit lines and the corresponding plurality of non-volatile memory cells, and the memory die is separately formed and coupled to the control die.

Citation Information

Patent Citations

  • Rewritable multibit non-volatile memory with soft decode optimization

    CN105679364A

  • Soft bit techniques for reading a data storage device

    CN108140407A