Storage System and Method for Direct Quad-Level Cell (QLC) Programming
Through the two-stage programming solution and Gray code mapping, the problem of large write buffers and imbalanced bit error rate in QLC memory is solved, and efficient and stable four-level unit programming is achieved.
Patent Information
- Application Number
- CN202110376176.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-29
- Filing Date
- 2021-04-08
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-04-08
AI Technical Summary
When programming multi-level unit memory, the prior art has problems such as large write buffer requirements, low programming efficiency, imbalance of bit error rate and instability, especially in the four-level unit (QLC) programming.
Using a two-stage programming scheme, three pages are first programmed through direct three-level units (TLC), and then the fourth page is programmed using unbalanced grey code mapping to reduce the write buffer requirement, and balance the bit error rate and read time by adjusting ECC redundancy.
It realizes the reduction of write buffer requirements in QLC memory, improves programming efficiency, ensures the balance of bit error rate and read time, and improves the stability and performance of the memory.
Smart Images

Figure CN113936725B_ABST
Abstract
Description
BACKGROUND OF THE INVENTION
[0001] When writing data to a non-volatile memory having a multi-level cell (MLC) configuration, this process is typically achieved by storing each bit of a cell for all cells in a full word line in the memory in a random access memory (RAM) in a memory controller, and then performing a multi-stage programming process for injecting charge into each multi-bit cell to achieve the desired programmed state of the cell. As part of this multi-step programming process, and for each of the multiple programming steps, the memory in the controller can store a copy of all data bits to be programmed in the cell and process error correction code (ECC) bits of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0002] Figure 1A is a block diagram of a non-volatile storage system of an embodiment.
[0003] Figure 1B is a block diagram showing a storage module of an embodiment.
[0004] Figure 1C is a block diagram showing a hierarchical storage system of an embodiment.
[0005] Figure 2A is a block diagram showing components of a controller of a non-volatile storage system shown in Figure 1A in accordance with an embodiment.
[0006] Figure 2B is a block diagram showing components of a non-volatile storage system shown in Figure 1A in accordance with an embodiment.
[0007] Figure 3 is a diagram of a 2-3-2-8 Gray code mapping used with an embodiment.
[0008] Figure 4 is a diagram showing a two-stage programming technique of an embodiment.
[0009] Figure 5 is a flowchart of a method of an embodiment for direct quad-level cell (QLC) programming.
[0010] Figure 6 is an illustration of a codeword of an embodiment.
[0011] Figure 7 is a flowchart of a method of another embodiment for direct quad-level cell (QLC) programming.
[0012] Figure 8 is an illustration of a codeword of another embodiment.
[0013] Figure 9 A flowchart of a method of another embodiment for direct quad-level cell (QLC) programming.
[0014] Figure 10 An illustration of a codeword of another embodiment. Detailed Description
[0015] By way of introduction, the embodiments below relate to a storage system and method for direct quad-level cell (QLC) programming. In one embodiment, a controller of the storage system is configured to generate codewords for lower, middle, and upper data pages; program the codewords for the lower, middle, and upper data pages in a memory of the storage system using a three-level cell programming operation; verify the programming of the codewords for the lower, middle, and upper data pages in the memory; generate a codeword for a top data page; and program the codeword for the top data page in the memory. Other embodiments are possible, and each of the embodiments may be used alone or in combination with one another. Accordingly, various embodiments will now be described with reference to the accompanying drawings.
[0016] Turning now to the drawings, a storage system suitable for aspects of implementing these embodiments is shown in Figures 1A to 1C . Figure 1A A block diagram of a non-volatile storage system 100 (sometimes referred to herein as a storage device or simply as a device) illustrating an embodiment of the subject matter described herein. Referring to Figure 1A , the non-volatile storage system 100 includes a controller 102 and non-volatile memory, which may consist of one or more non-volatile memory dies 104. As used herein, the term "die" refers to a collection of non-volatile memory cells formed on a single semiconductor substrate and associated circuitry for managing physical operations of those non-volatile memory cells. The controller 102 interfaces with a host system and transmits command sequences for read, program, and erase operations to the non-volatile memory die 104.
[0017] The controller 102 (which may be a non-volatile memory controller such as a flash memory, resistive random access memory (ReRAM), phase change memory (PCM), or magnetoresistive random access memory (MRAM) controller) may take the following forms: a processing circuit, a microprocessor, or a processor and a computer-readable medium storing computer-readable program code (such as firmware), the computer-readable program code being executable by, for example, a (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. The controller 102 may be configured with hardware and / or firmware to perform the various functions described below and shown in the flowcharts. Also, some components shown as being within the controller may also be stored outside the controller and other components may be used. Additionally, the phrase "operatively communicate with" may mean communicate directly with or communicate indirectly (wired or wirelessly) with via one or more components that may or may not be shown or described herein.
[0018] As used herein, a non-volatile memory controller is a device that manages data stored on a non-volatile memory and communicates with a host such as a computer or an electronic device. The non-volatile memory controller may have various functions in addition to the specific functionality described herein. For example, the non-volatile memory controller may format the non-volatile memory to ensure that the memory operates properly, map out bad non-volatile memory cells, and allocate spare cells to replace future failing cells. A portion of the spare cells may be used to store firmware to operate the non-volatile memory controller and implement other features. In operation, when the host needs to read data from or write data to the non-volatile memory, the host may communicate with the non-volatile memory controller. If the host provides a logical address at which to read / write data, the non-volatile memory controller may convert the logical address received from the host into a physical address in the non-volatile memory. (Alternatively, the host may provide a physical address). The non-volatile memory controller may also perform various memory management functions such as, but not limited to, wear leveling (distributing writes to avoid wearing out a particular memory block that would otherwise be repeatedly written to) and garbage collection (after a block is full, moving only valid data pages to a new block so that the full block can be erased and reused). Also, the structure of the "means" recited in the claims may include some or all of the structure of the controller described herein, the structure being programmed or fabricated to cause the controller to operate to perform the recited function.
[0019] The non-volatile memory die 104 may include any suitable non-volatile memory medium, including resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), phase change memory (PCM), NAND flash memory cells, and / or NOR flash memory cells. The memory cells may be in the form of solid-state (e.g., flash) memory cells and may be one-time programmable, few-time programmable, or multiple-time programmable. The memory cells may also be single-level cells (SLCs), multi-level cells (MLCs), three-level cells (TLCs), or other memory cell level technologies known today or developed in the future. Also, the memory cells may be fabricated in two-dimensional or three-dimensional fashions.
[0020] The interface between the controller 102 and the non-volatile memory die 104 may be any suitable flash memory interface, such as Toggle Mode 200, 400, or 800. In one embodiment, the storage system 100 may be a card-based system, such as a Secure Digital (SD) or micro Secure Digital (micro SD) card. In an alternative embodiment, the storage system 100 may be part of an embedded storage system.
[0021] Although in the Figure 1A example shown in Figure 1B and Figure 1C the non-volatile storage system 100 (sometimes referred to herein as the storage module) includes a single channel between the controller 102 and the non-volatile memory die 104, the subject matter described herein is not limited to having a single memory channel. For example, in some storage system architectures, such as
[0022] Figure 1B the storage system architectures shown in
[0023] Figure 1C a storage module 200 is shown that includes multiple non-volatile storage systems 100. Thus, the storage module 200 may include a storage controller 202 that interfaces with a host and with a storage system 204 that includes multiple non-volatile storage systems 100. The interface between the storage controller 202 and the non-volatile storage system 100 may be a bus interface, such as a Serial Advanced Technology Attachment (SATA), a Peripheral Component Interconnect Express (PCIe) interface, or a Double Data Rate (DDR) interface. In one embodiment, the storage module 200 may be a solid-state drive (SSD) or a non-volatile dual in-line memory module (NVDIMM), such as found in a server PC or a portable computing device, such as a laptop computer and a tablet computer.is a block diagram showing a hierarchical storage system. The hierarchical storage system 250 includes a plurality of storage controllers 202, each of which controls a corresponding storage system 204. The host system 252 can access the memory within the storage system via a bus interface. In one embodiment, the bus interface can be a Non-Volatile Memory Express (NVMe) interface or an Ethernet Fibre Channel (FCoE) interface. In one embodiment, Figure 1C the system shown in can be a rack-mounted mass storage system accessible by multiple host computers, such as those found in a data center or other locations that require mass storage devices.
[0024] Figure 2A is a block diagram showing the components of the controller 102 in more detail. The controller 102 includes a front-end module 108 that interfaces with the host, a back-end module 110 that interfaces with one or more non-volatile memory dies 104, and various other modules that perform functions that will be described in detail below. The modules can take, for example, the form of an encapsulated functional hardware unit designed to be used with other components, a portion of program code (such as software or firmware) executable by a (micro)processor or processing circuitry that typically performs a specific function among related functions, or a self-contained hardware or software component that interfaces with a larger system. The controller 102 may sometimes be referred to herein as a NAND controller or a flash controller, but it should be understood that the controller 102 can be used with any suitable memory technology, and some examples of memory technologies are provided below.
[0025] Referring again to the modules of the controller 102, the buffer manager / bus controller 114 manages the buffers in the random access memory (RAM) 116 and controls the internal bus arbitration of the controller 102. The read-only memory (ROM) 118 stores the system boot code. Although shown as being located separately from the controller 102 in Figure 2A one embodiment, in other embodiments, one or both of the RAM 116 and the ROM 118 may be located within the controller. In still other embodiments, portions of the RAM and the ROM may be located within the controller 102 and outside the controller.
[0026] The front-end module 108 includes a host interface 120 that provides an electrical interface to the host or the next-level storage controller and a physical layer interface (PHY) 122. The choice of the type of the host interface 120 can depend on the type of memory being used. Examples of the host interface 120 include, but are not limited to, SATA, SATA Express, Serial Attached SCSI (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 generally facilitates the transfer of data, control signals, and timing signals.
[0027] The backend module 110 includes an Error Correction Code (ECC) engine 124 that encodes data bytes received from a host and decodes and corrects errors in data bytes read from the non-volatile memory. A command sequencer 126 generates command sequences to be transmitted to the non-volatile memory die 104, such as programming and erase command sequences. An independent Redundant Array of Independent Drives (RAID) module 128 manages the generation of RAID parity and the recovery of failed data. The RAID parity can be used as an additional level of integrity protection for writing data to the memory device 104. In some cases, the RAID module 128 can be part of the ECC engine 124. A memory interface 130 provides the command sequences to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104. In one embodiment, the memory interface 130 can be a Double Data Rate (DDR) interface, such as a toggle mode 200, 400, or 800 interface. A flash control layer 132 controls the overall operation of the backend module 110.
[0028] The storage system 100 also includes other discrete components 140, such as an external electrical interface, external RAM, resistors, capacitors, or other components that can interface with the controller 102. In an alternative embodiment, one or more of the physical layer interface 122, the RAID module 128, the media management layer 138, and the buffer management / bus controller 114 are optional components that are not necessarily in the controller 102.
[0029] Figure 2B is a block diagram that more particularly illustrates the components of the non-volatile memory die 104. The non-volatile memory die 104 includes a peripheral circuit 141 and a non-volatile memory array 142. The non-volatile memory array 142 includes non-volatile memory cells for storing data. The non-volatile memory cells can be any suitable non-volatile memory cells, including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in two-dimensional and / or three-dimensional configurations. The non-volatile memory die 104 further includes a data cache 156 that caches data. The peripheral circuit 141 includes a state machine 152 that provides status information to the controller 102.
[0030] Returning again to Figure 2A, the flash control layer 132 (which will be referred to herein as the flash translation layer (FTL) or more generally as the "media management layer" when the memory may not be flash memory) processes flash errors and interfaces with the host. Specifically, the FTL, which can be an algorithm in the firmware, is responsible for internal memory management and translates writes from the host into writes to the memory 104. Because the memory 104 may have limited durability, may only be written in multi-page form, and / or may not be written to unless the memory 104 is erased as a block, an FTL may be required. The FTL is aware of these potential limitations of the memory 104, which may not be visible to the host. Thus, the FTL attempts to translate writes from the host into writes in the memory 104.
[0031] The FTL can include a logical-to-physical address (L2P) mapping (sometimes referred to herein as a table or data structure) and an allocated cache memory. In this way, the FTL translates a logical block address ("LBA") from the host into a physical address in the memory 104. The FTL can include other features such as, but not limited to, power-fail recovery (to enable recovery of the FTL's data structures in the event of a sudden power loss) and wear leveling (to make the wear on the storage blocks more even to prevent some blocks from wearing out excessively, which would lead to a greater probability of failure).
[0032] As mentioned above, when writing data to a non-volatile memory having a multi-level cell (MLC) configuration, this process is typically achieved by: storing each of the multiple bits of a cell in a random access memory (RAM) in a memory controller for all cells in a complete word line in the memory, and then performing a multi-stage programming process for injecting charge into each multi-bit cell to achieve the desired programmed state of the cell. Generally, the multi-stage programming involves an initial programming portion (i.e., the "fuzzy" programming step) of the states using a widened voltage distribution, and a subsequent final programming (i.e., the "fine" programming step) of all states using a tight voltage distribution. As part of this multi-step programming process, and for each of the multiple programming steps, the memory in the controller can store a copy of all the data bits to be programmed in the cell and process error correction code (ECC) bits for the data. It is well known that the fuzzy-fine programming scheme is used to program multi-level cell memories.
[0033] To improve memory cost efficiency by increasing memory density, four-level cells (QLCs) that store four bits per cell to provide 16 data states can be used. When designing a QLC programming scheme, several factors need to be considered. One consideration is to support a smaller write buffer. Conventional fuzzy-fine programming schemes may require a large write buffer (e.g., about 1.5 MB per die), especially when the number of memory plane / string / page sizes grows (as expected in the development of successive generations of memory), but this may be excessive. A programming scheme that allows some pages (e.g., 1 or 2 or 3 pages) to be programmed during a first pass such that reliable reading of the pages can be achieved without applying ECC correction before the remaining pages are added in subsequent passes can achieve a significant reduction in the write buffer (because the pages programmed in the first pass do not need to be stored in the controller write buffer for the next pass as long as they can be reliably read from the memory array).
[0034] An example of such a programming scheme is MLC fine programming, where two pages ("MLC") are programmed during a first "fuzzy" pass and the remaining two pages are added in a subsequent "fine" pass. The main drawback of such MLC fine programming is that in order to have sufficient tolerance between the 4 MLC states to enable internal reading within the memory array without ECC correction (i.e., IDL reading), unbalanced state encoding may be required. That is, the resulting state encoding may be a Gray code with a significantly different number of 1 / 0 transitions per page. For example, a 2-3-5-5 or 1-2-6-6 Gray code may be required instead of a balanced (as balanced as possible) state encoding such as a 3-4-4-4 encoding.
[0035] Unbalanced state encoding results in an unbalanced bit error rate (BER) per page, which means that more ECC redundancy is needed to achieve the same reliability (because ECC needs to cope with the worst page). This in turn reduces memory cost efficiency because more overhead needs to be allocated for ECC. Another alternative is to encode the data programmed during the "fuzzy" step using a very simple ECC (e.g., a simple XOR page) such that "fuzzy" page readback is possible inside the memory die, provided that the decoding of the ECC is based on logic with a low enough complexity. Such an encoded fuzzy-fine programming scheme may require a relatively small write buffer and can be achieved using a balanced Gray state encoding (i.e., a 3-4-4-4 state encoding). However, this depends on the memory quality to ensure that the simple ECC applied to the fuzzy data can provide sufficient reliability, which also incurs a performance penalty for fuzzy decoding and fuzzy parity transfer.
[0036] As mentioned, another important consideration is to use as balanced a state encoding as possible. Unbalanced encodings introduce unbalanced bit error rates (BERs) across pages and are sensitive to jitter in the state positions. This can significantly reduce the BER distribution and widen the voltage (Vt) tolerance. Also, unbalanced encodings can produce unbalanced read times (tR). In terms of balancing BER and tR, a 3-4-4-4 encoding scheme (where each number indicates the number of transitions) may be preferred. However, it does not allow MLC fine programming because there is no tolerance for IDL reads or resulting near WL interference (NWI) being poor. 2-3-5-5 or 1-2-6-6 encodings may allow MLC fine programming but may result in poor BER and tR balance. Another condition is robustness against unexpected shutdown (UGSD). For UGSD, only the 1-2-6-6 encoding seems robust enough (even though it is still under discussion), but may not be feasible in terms of BER balance due to high imbalance.
[0037] The following embodiments present a QLC programming scheme that meets the above conditions while avoiding the above problems. Generally, these embodiments recognize that it may be necessary to use a direct QLC programming scheme in order to reduce the required write buffer and avoid the need to go through single-level cells (SLCs). Additionally, these embodiments recognize that it may be necessary to use balanced Gray state encoding in order to have a balanced bit error rate (BER) and balanced read time (tR) across different pages, both in the presence of robustness against unexpected shutdown.
[0038] Generally, through these embodiments, the controller 102 of the storage system 100 transfers three pages during a three-level cell (TLC) programming phase and then transfers one additional page using a fine phase, where the first three pages are read internally (e.g., via IDL reads). Thus, through these embodiments, a two-phase method can be used to program QLC memory cells. The first phase is a direct TLC programming phase where three pages are programmed in the memory. The second phase programs the QLC (one additional page) using an unbalanced mapping (e.g., 2-3-2-8). The unbalanced mapping compensates by using a different redundancy for the three-level cell (TLC) pages than for the additional QLC page added on top. This two-phase programming technique can require a smaller write buffer, provide high performance, and have extremely low NWI. However, the system and ECC may need to change according to the different ECC redundancies (or different data payloads per page).
[0039] The following paragraphs provide an example implementation of this embodiment. It should be understood that these are only examples and other implementations may be used.
[0040] Turning to the drawings, Figure 3It is a schematic diagram of a 2-3-2-8 Gray code mapping used with embodiments. This stage encoding schematic shows 16 states (S0 to S15) in each of the lower (L), middle (M), upper (U), and top (P) pages. "2-3-2-8" refers to the number of transitions in each page. Thus, the lower and upper pages have two transitions, the middle page has three transitions, and the top page has eight transitions.
[0041] Figure 4 It is a schematic diagram showing a two-stage programming technique of an embodiment. As Figure 4 shown, in the first programming stage, the lower, middle, and upper pages are programmed using a direct 2-3-2 three-level cell (TLC) programming technique with a relatively small step programming voltage (dVPGM). In the second programming stage, the top page is programmed based on internal soft reading. With this programming scheme, the top page has a high BER due to its eight transitions, which is compensated by its high ECC redundancy. Higher top page ECC redundancy can be achieved by storing less data (e.g., 12 KB instead of 16 KB) on the top page, thus providing more space for its extended parity check. Alternatively, the extra ECC redundancy of the top page can "overflow" to other pages, i.e., parts of other lower / middle / upper pages (which require less ECC redundancy) can be allocated for storing the top page ECC redundancy. Soft read errors can be further minimized by squeezing the top page state transitions and leveraging its high ECC redundancy. The described embodiments have minimal write buffer requirements (direct TLC programming + direct top page programming). And, if even / odd TLC states are stored, then these embodiments can provide power loss immunity, which may require storing one SLC page together with the exclusive OR (XOR) of the lower, middle, and upper pages.
[0042] In one example, each word line uses a reduced payload of 60 KB: 16 KB in the lower page, 16 KB in the middle page, 16 KB in the upper page, and 12 KB in the top page. To compensate for the reduced amount of payload per WL (60 KB, different from the conventional QLC payload of 4 x 16 KB = 64 KB) and maintain the same memory density (i.e., the same memory cost efficiency), a reduced WL size can be used. For example, the reduced word line size can be: 16 KB + 1408 B = 17784 B, and the density is 60 KB / 17784 B = 3.4548 information bits / cell. This can maintain the same density as an exemplary conventional QLC memory, where a 64 KB payload can be stored in a WL size of 18976 B, with a roughly equal density of 64 KB / 18976 B = 3.4536 information bits / cell. In this example, assuming the ECC codeword payload is 4 KB plus some metadata, such as a 32 B firmware (FW) header, the ECC redundancy per 4 KB of the lower / upper / middle (four codewords per page) is (17784 - 4 x 4 KB data - 4 x 32 B FW header) / 4 = 318 B = 7.15%, which can correct the BER by approximately 1.05%. Since the top page is expected to have a higher BER in the case of unbalanced state encoding (e.g., Figure 3 the 2-3-2-8 encoding shown above), more ECC redundancy needs to be allocated to the top page. This is achieved by storing a smaller payload of only 12 KB on the top page. Thus, the ECC redundancy per 4 KB of the top (three codewords per page) is: (17784 - 3 x 4 KB data - 3 x 32 B FW header) / 3 = 1800 B = 30.36%, which can correct the BER by approximately 6.3%.
[0043] This implementation is further described in detail in Figure 5 flowchart 500 of Figure 5As shown, in this embodiment, a QLC 2-3-2-2 Gray code mapping is provided (act 510). Next, the controller 102 encodes the first three pages (lower, middle, and upper pages) with nominal ECC redundancy and programs them in a direct TLC manner using a balanced 2-3-2 Gray code (act 520). Then, the storage system 100 reads (verifies) the TLC programming phase (without error correction) on the memory chips 104, thus reducing the need to store all four pages in the controller's write buffer (act 530). In cases where pages in the ambiguous state cannot be reliably read without ECC correction (i.e., the BER is non-negligible), a simple temporary encoding can be applied to the ambiguous pages, e.g., by additionally storing the XOR page of the lower, middle, and upper pages in the controller write buffer and using it to perform a simple decoding of the page in the ambiguous page. Alternatively, a low-complexity ECC decoder (e.g., a bit-flip LDPC decoder) can be implemented within the memory die, provided that the BER of the pages expected to be in the ambiguous state is very low. Such low-complexity decoders can also be implemented in the CMOS wafer bonded to the memory die (also known as a CMOS bonded array - CbA). Finally, the controller 102 encodes the QLC top page with increased ECC redundancy to compensate for its higher BER. The increased top-page ECC redundancy can be achieved by storing less data (a smaller data payload) in the top page. The top-page programming triggers the QLC distribution using a 2-3-2-8 Gray mapping (act 540).
[0044] Figure 6 is an illustration of the codewords produced by the method of this embodiment. As Figure 6 shown, the top page has fewer codewords than the other three pages, and the extra space in the word line is used for extra parity bits.
[0045] There are many alternatives that can be used with these embodiments. For example, in the case where the balance between the top page and the lower, middle, and upper pages becomes sub-optimal (e.g., too much ECC is spent on the top page and insufficient ECC is spent on the lower, middle, and upper pages), some of the top page can be allocated for additional lower, middle, and upper page parity. For example, beyond 17784 - 3x4 KB - 3x32 B = 5400B of parity available on the top page, we can allocate: 4128B for three codewords for the top page. This results in 1376B ECC / 4KB = 25% ECC redundancy, thus yielding an error correction capability of approximately 5%. As additional parity for the lower, middle, and upper pages, 1272B results in an additional 4x318B of parity bits per page. The additional parity bits for the lower, middle, and upper pages can be XORed, and the result can be stored in the top page. In the case of a single page (lower, middle, or upper) failure, we can recover an additional 318B of ECC per codeword, thus doubling the redundancy to approximately 14.3%. This provides an error correction capability of approximately 2.5%.
[0046] Compared to other embodiments described above, this embodiment may be more complex because an additional write buffer may be required to store 1272B for each fuzzy programming until fine programming is complete. For example, an additional 1272B x 4 pages x 6 strings is approximately 30KB. Also, uncorrectable error events for the lower, middle, and upper pages may require a more complex recovery process to read additional parity from the top page and may require an ECC design change.
[0047] Returning to the drawings, Figure 7 is a flowchart 700 of the method of this embodiment. As Figure 7 shown, in this embodiment, the controller 102 calculates the first and second parity bits for the first three pages (lower, middle, and upper) (action 710). Next, the controller 102 programs the first three pages with the first parity bit using a direct three-level cell (TLC) programming technique with a balanced 2-3-2 Gray code mapping (action 720). Next, the TLC programming phase is verified on the memory chip 104 without using error correction and without storing the data in the write buffer of the controller (action 730). Then, the top page is programmed with the XOR signature of the lower payload, the first parity, and the second parity of the first three pages, thus triggering a QLC distribution with a 2-3-3-8 Gray code mapping (action 740). Figure 8 is an illustration of the codewords produced by the method of this embodiment.
[0048] In yet another alternative embodiment, there is 64 KB per word line, with 16 KB for the lower page, 16 KB for the middle page, 16 KB for the upper page, and 16 KB for the top page. Assume, for example, a QLC word line of size 18976B. This provides a density of 64 KB / 18976B = 3.4536 information bits / cell (roughly the same as the previous example). The ECC redundancy can be allocated according to the number of transitions per page. The lower / upper pages have two transitions per page. Thus, the ECC redundancy is 432B = ~9.5%, which provides a correction capability of approximately 1.55%. The middle page has three transitions per page, and there is 448B = ~9.8% of ECC redundancy, which provides a correction capability of approximately 1.6%. The top page has eight transitions per page. There is 1152B = ~21.8% of ECC redundancy, divided into two levels: parity-1 of 616B and parity-2 of 536B. The correction capability using parity-1 is approximately 4.3%. The correction capability using full parity (parity-1 + parity-2) is approximately 2.2%.
[0049] Returning to the drawings, Figure 9 is a flow chart 900 of the method of this embodiment. As Figure 9 shown, in this embodiment, a QLC 2-3-2-8 Gray code is provided (action 910). Next, the controller 102 encodes each page with an amount of ECC parity proportional to the number of transitions in the page (action 920). In this embodiment, the ECC parity of the top page that exceeds the amount of the ECC column in the word line is split into two parts: parity-1 and parity-2, such that parity-1 is placed within the page, and parity-2 will be stored as part of the other three pages (lower, middle, and upper pages). Next, the controller 102 programs the three pages (lower, middle, and upper) in a direct TLC manner using a balanced 2-3-2 Gray mapping (action 930). Then, the TLC programming phase is verified on the memory chip 104 without error correction and without storing all the pages in the memory buffer of the controller (action 940). Next, the controller 102 programs the top page, thereby inducing a QLC distribution using a 2-3-2-8 mapping (action 950). Figure 10It is a diagram of a codeword generated by the method of this embodiment. For previous examples, in cases where a page in a blurred state cannot be reliably read without performing ECC correction (i.e., the BER is non-negligible), a simple temporary encoding can be applied to the blurred page. For example, by additionally storing the XOR page of the lower, middle, and upper pages in the controller write buffer and using it to perform simple decoding of the page in the blurred page. Alternatively, a low-complexity ECC decoder (e.g., a bit-flip LDPC decoder) can be implemented within the memory die, provided that the BER of the pages expected to be in a blurred state is very low. Such low-complexity decoders can also be implemented in a CMOS wafer bonded to the memory die (also known as a CMOS bonded array - CbA).
[0050] Finally, as mentioned above, any suitable type of memory can be used. Semiconductor memory devices include: volatile memory devices, such as dynamic random access memory (“DRAM”) or static random access memory (“SRAM”) devices; non-volatile memory devices, such as resistive random access memory (“ReRAM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory (which can also be considered a subset of EEPROM), ferroelectric random access memory (“FRAM”), and magnetoresistive random access memory (“MRAM”); and other semiconductor elements capable of storing information. Each type of memory device can have a different configuration. For example, flash memory devices can be configured in a NAND or NOR configuration.
[0051] Memory devices can be formed from passive and / or active elements in any combination. By way of non-limiting example, passive semiconductor memory elements include ReRAM device elements, which in some embodiments include resistivity-switching memory elements, such as antifuses, phase change materials, etc., and optionally include steering elements, such as diodes, etc. Additionally, by way of non-limiting example, active semiconductor memory elements include EEPROM and flash memory device elements, which in some embodiments include elements containing charge storage regions, such as floating gates, conductive nanoparticles, or charge storage dielectric materials.
[0052] Multiple memory elements can be configured such that they are connected in series or such that each element can be accessed individually. By way of non-limiting example, a flash memory device (NAND memory) in a NAND configuration typically contains memory elements connected in series. A NAND memory array can be configured such that the array consists of multiple memory strings, where a string consists of multiple memory elements that share a single bit line and are accessed as a group. Alternatively, the memory elements can be configured such that each element can be accessed individually, such as in a NOR memory array. NAND and NOR memory configurations are examples, and the memory elements can be configured in other ways.
[0053] Semiconductor memory elements located within and / or above a substrate can be configured in two-dimensional or three-dimensional forms, such as a two-dimensional memory structure or a three-dimensional memory structure.
[0054] In a two-dimensional memory structure, the semiconductor memory elements are arranged in a single plane or a single memory device level. Generally, in a two-dimensional memory structure, the memory elements are arranged in a plane that extends substantially parallel to the main surface of the substrate that supports the memory elements (e.g., in the x-z direction plane). The substrate can be a wafer of a layer on or in which the memory elements are formed, or can be a carrier substrate attached to the memory elements after the memory elements are formed. By way of non-limiting example, the substrate can comprise a semiconductor such as silicon.
[0055] The memory elements can be arranged in an ordered array such as multiple rows and / or columns in a single memory device level. However, the memory elements can be arranged in a non-regular or non-orthogonal configuration. Each memory element can have two or more electrodes or contact lines, such as bit lines and word lines.
[0056] A three-dimensional memory array is arranged such that the memory elements occupy multiple planes or multiple memory device levels, thereby forming a three-dimensional (i.e., in the x, y, and z directions, where the y direction is substantially perpendicular to the main surface of the substrate, and the x and z directions are substantially parallel to the main surface of the substrate) structure.
[0057] By way of non-limiting example, a three-dimensional memory structure can be arranged vertically as a stack of multiple two-dimensional memory device levels. As another non-limiting example, a three-dimensional memory array can be arranged as multiple vertical columns (e.g., columns that extend substantially perpendicular to the main surface of the substrate (i.e., in the y direction)), where each column has multiple memory elements in each column. The columns can be arranged in a two-dimensional configuration (e.g., in the x-z plane), thereby resulting in a three-dimensional arrangement of memory elements having elements on multiple vertically stacked memory planes. Other configurations of memory elements in three-dimensional form can also constitute a three-dimensional memory array.
[0058] By way of non-limiting example, in a three-dimensional NAND memory array, memory elements may be coupled together to form NAND strings within a single horizontal (e.g., x-z) memory device tier. Alternatively, memory elements may be coupled together to form vertical NAND strings that traverse multiple horizontal memory device tiers. Other three-dimensional configurations are conceivable, where some NAND strings contain memory elements within a single memory tier, while other strings contain memory elements spanning multiple memory tiers. The three-dimensional memory array may also be designed in a NOR configuration and in a ReRAM configuration.
[0059] Typically, in a monolithic three-dimensional memory array, one or more memory device tiers are formed above a single substrate. Optionally, the monolithic three-dimensional memory array may also have one or more memory layers at least partially within the single substrate. By way of non-limiting example, the substrate may comprise a semiconductor such as silicon. In a monolithic three-dimensional array, the layers that make up each memory device tier of the array are typically formed on the layers of the underlying memory device tier of the array. However, the layers of adjacent memory device tiers of the monolithic three-dimensional memory array may be shared, or there may be intervening layers between the memory device tiers.
[0060] Moreover, two-dimensional arrays may be formed separately and then packaged together to form a non-monolithic memory device having multiple memory layers. For example, a non-monolithic stacked memory may be constructed by forming memory tiers on separate substrates and then stacking the memory tiers on top of each other. The substrates may be thinned or removed from the memory device tiers prior to stacking, but since the memory device tiers are initially formed above separate substrates, the resulting memory array is not a monolithic three-dimensional memory array. Additionally, multiple two-dimensional memory arrays or three-dimensional memory arrays (monolithic or non-monolithic) may be formed on separate chips and then packaged together to form a stacked chip memory device.
[0061] Associated circuitry is typically required to operate and communicate with the memory elements. By way of non-limiting example, a memory device may have circuitry for controlling and driving the memory elements to perform functions such as programming and reading. This associated circuitry may be located on the same substrate as the memory elements and / or on a separate substrate. For example, a controller for memory read and write operations may be located on a separate controller chip and / or on the same substrate as the memory elements.
[0062] Those skilled in the art will recognize that the present invention is not limited to the two-dimensional and three-dimensional structures described, but encompasses all relevant memory structures as described herein and as understood by those skilled in the art within the spirit and scope of the present invention.
[0063] The foregoing detailed description is to be understood as illustrative of selected forms of the invention, and not as limiting of the invention. Only the appended claims, including all equivalents, are intended to define the scope of the claimed invention. Finally, it should be noted that any aspect of any of the embodiments described herein may be used alone or in combination with one another.
Claims
1. A storage system, comprising: A memory; And A controller configured to: Generate codewords for the lower page, middle page, and upper page of data; Program the codewords for the lower page, middle page, and upper page of the data in the memory using a three-level cell programming operation; Read the programming of the codewords for the lower page, middle page, and upper page of the data in the memory; Generate a codeword for the top page of data; And Program the codeword for the top page of the data in the memory using the lower page, middle page, and upper page of the read data.
2. The storage system according to claim 1, wherein the codewords for the lower page, middle page, and upper page of the data are programmed using balanced 2-3-2 Gray code mapping.
3. The storage system according to claim 2, wherein programming the codeword for the top page of the data in the memory using 2-3-2-8 Gray code mapping induces a four-level cell (QLC) distribution.
4. The storage system according to claim 1, wherein the read data is retrieved without using error correction.
5. The storage system according to claim 1, wherein the lower page, middle page, and upper page are read from the memory array without retrieving them in a write buffer in the controller.
6. The storage system according to claim 1, wherein the codewords are generated by encoding the lower page, middle page, and upper page of the data using error correction code parity bits.
7. The storage system according to claim 6, wherein the codeword for the top page of the data is generated by encoding the top page of the data using more error correction code parity bits than the lower page, middle page, and upper page of the data.
8. The storage system according to claim 1, wherein fuzzy data is retrieved using a low-complexity ECC decoder on the memory die or on a CMOS die (CbA) bonded to the memory die.
9. A method in a storage system including a memory and a controller, comprising: Calculating first and second parity bits for the lower page, middle page, and upper page of data; Programming the lower page, middle page, and upper page of the data using the first parity bit using a three-level cell programming operation; Reading the programming of the lower page, middle page, and upper page of the data in the memory; And Programming the top page of data in the memory using the lower page, middle page, and upper page of the read data.
10. The method according to claim 9, wherein the lower page, middle page, and upper page of the data are programmed using balanced 2-3-2 Gray code mapping.
11. The method according to claim 10, wherein programming the top page of the data in the memory using 2-3-2-8 Gray code mapping induces a four-level cell (QLC) distribution.
12. The method according to claim 9, wherein the top page of the data is programmed using the exclusive OR signature of the first parity bit and the second parity bit.
13. The method according to claim 9, wherein the lower page, the middle page, and the upper page are read from the memory array without retrieving them in a write buffer in the controller.
14. The method according to claim 9, wherein the top page includes a payload that is less than the payloads of the lower page, the middle page, and the upper page.
15. A storage system, comprising: a memory; means for encoding each of a lower, a middle, an upper, and a top page of data with an amount of error correction code parity bits proportional to the number of transitions in each page; means for programming the lower page, the middle page, and the upper page of the data in the memory using a three-level cell programming operation; means for reading the programming of the lower page, the middle page, and the upper page of the data in the memory; and means for programming the top page of the data in the memory using the lower page, the middle page, and the upper page of the read data.
16. The storage system according to claim 15, wherein the lower page, the middle page, and the upper page of the data are programmed using a balanced 2-3-2 Gray code mapping.
17. The storage system according to claim 16, wherein programming the top page of the data in the memory using a 2-3-2-8 Gray code mapping results in a four-level cell (QLC) distribution.
18. The storage system according to claim 15, wherein a number of error correction code parity bits in the top page that exceed the number of error correction code columns in a word line are divided into a first part and a second part, wherein the first part is placed within the top page, and wherein the second part is stored as part of the lower page, the middle page, and the upper page.
19. The storage system according to claim 15, further comprising means for storing the same data on the top page but overflowing its ECC redundancy to other pages.
20. The storage system according to claim 15, wherein the lower page, the middle page, and the upper page are read from the memory array without retrieving them in a write buffer in a controller of the storage system.
Citation Information
Patent Citations
Systems and methods of storing data
CN107357678A
Memory system and operation method thereof
CN110874191A