Implementing fault tolerant page banding on low density memory systems

By adopting fault-tolerant striping and word line separation technology of specified size in low-density memory systems, the problem of easy loss and excessive loss of fault-tolerant striping data is solved, and higher memory system reliability and efficiency are achieved.

CN119938391APending Publication Date: 2025-05-06MICRON TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510015202.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2020-10-23
Filing Date
2020-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In low-density memory systems, data of fault-tolerant stripes are easily lost due to failures, and the prior art tends to lead to excessive provisioning losses when implementing fault-tolerant solutions.

Method used

By storing host data in a fault-tolerant stripe of a specified size, each fault-tolerant stripe contains multiple data pages and a redundant metadata page, using word-line separation to ensure independence between pages, dynamically adjusting the fault-tolerant stripe size to provide the necessary word-line separation.

Benefits of technology

It effectively reduces the risk of data loss caused by failures, and limits the excessive provision loss caused by fault tolerance solutions, improving the reliability and efficiency of the memory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938391A_ABST
    Figure CN119938391A_ABST
Patent Text Reader

Abstract

The invention relates to implementing fault tolerant page banding on low density memory systems. An example memory subsystem includes a memory device; and a processing device operatively coupled with the memory device. The processing device is configured to: receive a first host data item; storing the first host data item in a first page of a first logic unit of a memory device, wherein the first page is associated with a fault tolerant stripe; receiving a second host data item; storing the second item of host data in a second page of the first logical unit of the memory device, where the second page is associated with the fault tolerant stripe, and where the second page is separated from the first page by one or more word lines including dummy word lines that do not store host data; and storing redundant metadata associated with the fault tolerant stripe in a third page of a second logic unit of the memory device.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Information about divisional applications

[0002] This application is a divisional application of the invention patent application with application date of December 30, 2020, application number "202011606424.4", and invention name "Implementing Fault-Tolerant Page Striping on Low-Density Memory Systems". Technical Field

[0003] The present disclosure relates generally to memory systems, and more particularly to implementing fault-tolerant page striping on low-density memory systems. Background Art

[0004] A memory subsystem may include one or more memory devices that store data. The memory devices may be, for example, non-volatile memory devices and volatile memory devices. Generally speaking, a host system may utilize a memory subsystem to store data at the memory devices and retrieve data from the memory devices. Summary of the invention

[0005] One aspect of the present application is directed to a system, the system comprising: a memory device; and a processing device operatively coupled to the memory device to perform operations comprising: receiving a first host data item; storing the first host data item in a first page of a first logical unit of the memory device, wherein the first page is associated with a fault-tolerant stripe; receiving a second host data item; storing the second host data item in a second page of the first logical unit of the memory device, wherein the second page is associated with the fault-tolerant stripe, and wherein the second page is separated from the first page by at least a predetermined number of word lines; storing redundant metadata associated with the fault-tolerant stripe in a page of the second logical unit of the memory device.

[0006] Another aspect of the present application is directed to a method, the method comprising: receiving, by a processing device of a memory subsystem controller, a first host data item; storing the first host data item in a first page of a first logical unit of a memory device, wherein the first page is associated with a fault-tolerant stripe; receiving a second host data item; storing the second host data item in a second page of the first logical unit of the memory device, wherein the second page is associated with the fault-tolerant stripe, and wherein the second page is separated from the first page by at least one dummy word line that does not store host data; and storing redundant metadata associated with the fault-tolerant stripe in a page of the second logical unit of the memory device.

[0007] Yet another aspect of the present application is directed to a non-transitory computer-readable storage medium comprising executable instructions which, when executed by a processing device, cause the processing device to perform operations, the operations comprising: receiving a first host data item; storing the first host data item in a first page of a first logical unit of a memory device, wherein the first page is associated with a fault-tolerant stripe; receiving a second host data item; storing the second host data item in a second page of the first logical unit of the memory device, wherein the second page is associated with the fault-tolerant stripe, and wherein the second page is separated from the first page by at least a predetermined number of word lines; and storing redundant metadata associated with the fault-tolerant stripe in a page of the second logical unit of the memory device. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The present disclosure will be more fully understood from the detailed description given below and the accompanying drawings of various embodiments of the present disclosure.

[0009] Figure 1 An example computing system including a memory subsystem according to some embodiments of the present disclosure is shown.

[0010] Figure 2 An example layout of a memory device according to an embodiment of the present disclosure is schematically shown.

[0011] Figure 3 An example fault-tolerant layout of a memory device according to an embodiment of the present disclosure is schematically shown.

[0012] Figure 4 Another example redundancy scheme of a memory device according to an embodiment of the present disclosure is schematically shown.

[0013] Figure 5 Another example redundancy scheme of a memory device according to an embodiment of the present disclosure is schematically shown.

[0014] Figure 6 Another example redundancy scheme of a memory device according to an embodiment of the present disclosure is schematically shown.

[0015] Figure 7 Another example redundancy scheme of a memory device according to an embodiment of the present disclosure is schematically shown.

[0016] Figure 8 is a flow chart of an example method of implementing fault-tolerant page striping by a memory subsystem controller operating according to some embodiments of the present disclosure.

[0017] Fig. 9 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION

[0018] Aspects of the present disclosure are directed to storing data for a fault-tolerant stripe at locations based on the boundaries of a memory device. The memory subsystem may be a storage device, a memory module, or a mixture of a storage device and a memory module. Figure 1 Examples of storage devices and memory modules are described. In general, a host system may utilize a memory subsystem that includes one or more components (eg, memory devices). The host system may provide data for storage at the memory subsystem and may request retrieval of data from the memory subsystem.

[0019] The memory device may be a non-volatile memory device, which is a package of one or more dies. The dies in the package may be assigned to one or more channels for communicating with the memory subsystem controller. The non-volatile memory device includes cells (i.e., electronic circuits storing information) grouped into pages to store data bits. One example of a non-volatile memory device is a NAND memory device. Figure 1 Other examples of non-volatile memory devices are described.

[0020] Various memory subsystems may implement fault-tolerant redundancy schemes, such as Redundant Arrays of Independent NAND (RAIN), for error checking and correction. The fault-tolerant redundancy schemes may store host data in groups of pages (referred to herein as fault-tolerant stripes), such that each stripe contains redundant metadata pages (e.g., parity pages), thus enabling reconstruction of the data in the event that one of the pages of the stripe fails.

[0021] The memory device may include a plurality of memory cell arrays grouped by word lines. A failure of the memory device at a particular word line may result in at least partial loss of data stored at the word line. In addition, a defect that results in a failure of a particular word line may further trigger failures of other word lines close to the word line. Thus, a defect may result in the loss of multiple data pages of a fault-tolerant stripe at different locations (e.g., at different word lines). If multiple data pages of the same fault-tolerant stripe are located near a word line, too many host data elements may be lost at the same time, thus making it impossible to reconstruct the lost host data elements based on available redundant metadata. Accordingly, in the event of a failure of the memory device, storing data pages of a fault-tolerant stripe at adjacent word lines may result in the loss of data of the fault-tolerant stripe.

[0022] Aspects of the present disclosure address the above and other deficiencies by storing host data in fault-tolerant stripes of a specified size, which may be predetermined in order to limit over-provisioning losses due to implementation of a fault-tolerant scheme. Each fault-tolerant stripe may include multiple pages of data, with the last page of the fault-tolerant stripe being dedicated to storing redundant metadata that may be used for error detection and correction. This fault-tolerant stripe may be formed using multiple pages with sequential page numbering from each plane of each logical unit, provided that pages residing within the same plane have sufficient wordline separation (e.g., separated by at least a predetermined number of wordlines that include a level separation wordline). The fault-tolerant stripe size may be dynamically adjusted based on the layout of the memory device in order to provide the necessary wordline separation, as explained in more detail below herein.

[0023] Figure 1 An example computing system 100 is shown that includes a memory subsystem 110 according to some embodiments of the present disclosure. Memory subsystem 110 may include media such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., memory device 130), or a combination of such media.

[0024] The memory subsystem 110 may be a storage device, a memory module, or a mixture of storage devices and memory modules. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).

[0025] The computing system 100 can be a computing device, such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (e.g., an airplane, drone, train, car, or other transportation vehicle), an Internet of Things (IoT)-enabled device, an embedded computer (e.g., a computer included in a vehicle, industrial equipment, or a networked commercial device), or such computing device that includes a memory and a processing device (e.g., a processor).

[0026] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to memory subsystems 110 of different types. Figure 1An example of a host system 120 coupled to one memory subsystem 110 is shown. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without intervening components), whether wired or wireless, including connections such as electrical connections, optical connections, magnetic connections, etc.

[0027] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110, for example, to write data to the memory subsystem 110 and read data from the memory subsystem 110.

[0028] The host system 120 may be coupled to the memory subsystem 110 via a physical host interface. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, a Fibre Channel, a Serial Attached SCSI (SAS), a Double Data Rate (DDR) memory bus, a Small Computer System Interface (SCSI), a Dual In-line Memory Module (DIMM) interface (e.g., a DIMM socket interface supporting Double Data Rate (DDR)), an Open NAND Flash Interface (ONFI), Double Data Rate (DDR), Low Power Double Data Rate (LPDDR), etc. The physical host interface may be used to transmit data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 through a PCIe interface, the host system 120 may further utilize an NVM Express (NVMe) interface to access components (e.g., the memory device 130). The physical host interface may provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120 . Figure 1 Memory subsystem 110 is shown as an example. In general, host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0029] The memory devices 130, 140 may include any combination of different types of non-volatile memory devices and / or volatile memory devices. Volatile memory devices (e.g., memory device 140) may be, but are not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).

[0030] Some examples of non-volatile memory devices (e.g., memory device 130) include NAND-type flash memory and write-in-place memory, such as a three-dimensional cross-point ("3D cross-point") memory device, which is a cross-point array of non-volatile memory cells. The cross-point array of non-volatile memory can be combined with a stackable cross-grid data access array to perform bit storage based on changes in body resistance. In addition, in contrast to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, where non-volatile memory cells can be programmed without pre-erasing the non-volatile memory cells. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).

[0031] Each of the memory devices 130 may include one or more memory cell arrays. One type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), three-level cells (TLC), and four-level cells (QLC), may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more memory cell arrays, such as SLC, MLC, TLC, QLC, or any combination of these. In some embodiments, a particular memory device may include an SLC portion of a memory cell, and an MLC portion, a TLC portion, or a QLC portion. The memory cells of the memory device 130 may be grouped into pages, which may refer to a logical unit of a memory device for storing data. For some types of memory (e.g., NAND), pages may be grouped to form blocks. "Block" herein will refer to a collection of adjacent or non-adjacent memory pages. An example of a "block" is an "erasable block," which is the smallest erasable unit of memory, and a "page" is the smallest writable unit of memory. Each page includes a collection of memory cells.

[0032] Although nonvolatile memory components such as a 3D cross-point nonvolatile memory cell array and NAND-type flash memory (e.g., 2D NAND, 3D NAND) are described, the memory device 130 may be based on any other type of nonvolatile memory, such as read-only memory (ROM), phase-change memory (PCM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridging RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), "NOR" (NOR) flash memory, and electrically erasable programmable read-only memory (EEPROM).

[0033] The memory subsystem controller 115 (for simplicity, controller 115) can communicate with the memory device 130 to perform operations, such as reading data, writing data, or erasing data at the memory device 130, as well as other such operations. The memory subsystem controller 115 can include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware can include digital circuitry with dedicated (i.e., hard-coded) logic to perform the operations described herein. The memory subsystem controller 115 can be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or other suitable processor.

[0034] The memory subsystem controller 115 may include a processor 117 (e.g., a processing device) configured to execute instructions stored in a local memory 119. In the example shown, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.

[0035] In some embodiments, local memory 119 may include memory registers that store memory pointers, fetched data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1 The example memory subsystem 110 in FIG. 1 has been shown as including a memory subsystem controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115, but may rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).

[0036] In general, the memory subsystem controller 115 may receive commands or operations from the host system 120, and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations, such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block addresses (LBA), name space) and physical addresses (e.g., physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include a host interface circuit system to communicate with the host system 120 via a physical host interface. The host interface circuit system may convert commands received from the host system into command instructions to access the memory device 130, and convert responses associated with the memory device 130 into information for the host system 120.

[0037] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.

[0038] In some embodiments, the memory device 130 includes a local media controller 135 that operates in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller (e.g., the memory subsystem controller 115) may manage the memory device 130 externally (e.g., perform media management operations on the memory device 130). In some embodiments, the memory device 130 is a managed memory device, which is a raw memory device combined with a local controller (e.g., the local controller 135) to perform media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.

[0039] The memory subsystem 110 includes a fault tolerance manager 113 that manages storage host data in a fault tolerant manner. In some embodiments, the memory subsystem controller 115 includes at least a portion of the fault tolerance manager 113. For example, the memory subsystem controller 115 may include a processor 117 (processing device) configured to execute instructions stored in a local memory 119 for performing the operations described herein. In some embodiments, the fault tolerance manager 113 is part of the host system 110, an application, or an operating system.

[0040] The fault tolerance manager 113 may be used to implement a fault tolerance layout for storing host data at the memory device 130. As host data is arriving, the fault tolerance manager 113 may program the pages of the memory device to form a fault tolerance stripe of a specified size, which may be predetermined in order to limit over-provisioning losses due to implementing the fault tolerance scheme. Each fault tolerance stripe may include multiple pages of data, with the last page of the fault tolerance stripe being dedicated to storing redundant metadata that may be used for error detection and correction. This fault tolerance stripe may be formed using multiple pages with sequential page numbers for each plane of each logical unit of at least a subset of the logical units of the memory device, provided that pages residing in the same plane have sufficient word line separation (e.g., separated by at least a predetermined number of word lines including a plane separation word line).

[0041] As pages are programmed (i.e., as the contents of the pages are stored to the memory device), redundant metadata may be calculated by, for example, summing the contents of the pages by means of an exclusive disjunction (XOR) operation, and intermediate results of the XOR operation may be stored in memory until the fault-tolerant stripe is closed (i.e., until the metadata is written to the last page of the fault-tolerant stripe). Because the pages have sequential page numbering, no additional memory is required to store the intermediate XOR results, since the XOR results for page #i may continue to page #i+1. After the last data page of the fault-tolerant stripe is programmed, the XOR results are written to the memory device. The fault-tolerant stripe size may be dynamically adjusted based on the layout of the memory device in order to provide the necessary word line separation, as described in more detail herein below.

[0042] Figure 2 Schematically illustrates an example layout 200 of a memory device according to an embodiment of the present disclosure. As mentioned herein above and by Figure 2 Schematically shown, host data may be stored on a memory device, which may include multiple logical units (also referred to as "LUNs" or "die"). Each logical unit may include multiple blocks 210 residing on multiple planes. Each block may include a corresponding word line WL n -WL n+k Grouped multiple pages 210. Program and / or erase operations can be performed on two or more pages simultaneously, provided that each page is residing on a corresponding plane.

[0043] Multiple blocks may be logically combined to form a super block SB0-SB3, which includes at least one block from each plane of each logical unit. Programming operations with respect to the memory device may be performed by the super block, i.e., by writing host data to a page of one super block after writing host data to a page of another super block.

[0044] The memory subsystem controller may store host data in a fault-tolerant manner by sequentially writing host data to another page after writing the host data to one page so that the pages are grouped into fault-tolerant stripes. Each fault-tolerant stripe contains a specific number of data pages (i.e., pages storing host data) and redundant metadata pages storing metadata to be used for error detection and recovery. As mentioned herein above, the redundant metadata may be represented by parity metadata, so that each bit of the metadata page of the fault-tolerant stripe may be generated by performing a bitwise exclusive or (also referred to as "XOR") operation of the corresponding data bit pages of the fault-tolerant stripe. This redundancy scheme will provide fault tolerance in the event that no more than one page of a given fault-tolerant stripe is erroneous. The erroneous page may be reconstructed by performing a bitwise exclusive or of all remaining data pages and metadata pages.

[0045] Because the fault tolerance scheme described above allows no more than one erroneous page per fault-tolerant stripe, pages that share one or more adjacent word lines within a single plane of any given logical unit may not be present in a fault-tolerant stripe because the presence of one erroneous page on a given word line may indicate that other pages on the same word line are also erroneous. In other words, no more than one page from any given word line of any given plane of a logical unit may be present in a fault-tolerant stripe. Thus, in some implementations, a fault-tolerant stripe may include pages from each plane of each logical unit of a memory device, such that all but one of the pages of a fault-tolerant stripe are utilized to store host data, while the remaining pages are utilized to store redundant metadata.

[0046] Figure 3 Schematically illustrates an example fault-tolerant layout 300 of a memory device according to an embodiment of the present disclosure. Figure 3 In the illustrative example of , one fault-tolerant stripe is formed by pages (e.g., page #1) with a first number within its plane across planes and logical units of a memory device, such that all but one of the pages of the fault-tolerant stripe are used to store host data, while the last page P1 is used to store redundant metadata for the fault-tolerant stripe. Another fault-tolerant stripe is formed by pages (e.g., page #2) with a second number within its plane across planes and logical units of a memory device, such that all but one of the pages of the fault-tolerant stripe are used to store host data, while the last page P2 is used to store redundant metadata for the fault-tolerant stripe. Figure 3 The numbers of logical units, planes, pages, word lines, and fault-tolerant stripes in the illustrative examples of are chosen for illustrative purposes and are not limiting; other implementations may use various other numbers of logical units, planes, pages, and fault-tolerant stripes.

[0047] Because the fault-tolerant stripes described above contain one page for each plane of each logical unit, the fault-tolerant stripe size is effectively limited by the number of planes multiplied by the number of logical units in the memory device. This may cause efficiency issues in low-density devices with a relatively small number of logical units and / or planes, because dedicating one page of each fault-tolerant stripe to store redundant metadata will result in an over-provisioning penalty of the inverse product of the number of logical units and the number of planes (i.e., over-provisioning penalty = 1 / (number of LUNs * number of planes)). "Over-provisioning" herein will refer to the number of physical memory blocks allocated exceeding the logical capacity presented as available memory to the host.

[0048] Accordingly, one way to reduce over-provisioning penalties involves forming a fault-tolerant stripe that utilizes multiple pages per plane of each logical unit of at least a subset of the logical units of the memory device, rather than one page, as described by Figure 4 Schematically shown, Figure 4Schematically illustrates an example fault-tolerant layout 400 of a memory device according to an embodiment of the present disclosure. Figure 4 In an illustrative example of , a first fault-tolerant stripe is formed by pages (e.g., page #1) with a first predefined number within its plane and pages (e.g., page #5) with a second predefined number within its plane across planes and logical units of a memory device, such that all but one of the pages of the fault-tolerant stripe are used to store host data, while the last page P1 is used to store redundant metadata for the fault-tolerant stripe. A second fault-tolerant stripe is formed by pages (e.g., page #2) with numbers after the first predefined number within its plane and pages (e.g., page #6) with numbers after the second predefined number within its plane across planes and logical units of the memory device, such that all but one of the pages of the fault-tolerant stripe are used to store host data, while the last page P2 is used to store redundant metadata for the fault-tolerant stripe. A third fault-tolerant stripe is formed by pages #3 and #7 across planes and logical units of the memory device, such that all but one of the pages of the fault-tolerant stripe are used to store host data, while the last page P3 is used to store redundant metadata for the fault-tolerant stripe. A fourth fault-tolerant stripe is formed by pages #4 and #8 across planes and logical units of the memory device, such that all but one of the pages of the fault-tolerant stripe are utilized to store host data, while the last page P4 is utilized to store redundant metadata for the fault-tolerant stripe, and so on.

[0049] It is worth noting that a physical defect on a word line may affect not only pages residing on the same word line, but also pages on one or more adjacent word lines. Accordingly, each pair of pages from a given plane that participate in the same fault-tolerant stripe may be separated by at least one word line (e.g., page #1 and page #5, page #2 and page #6, etc.), as indicated by Figure 5 Schematically shown, Figure 5 An example fault-tolerant layout 500 of a memory device according to an embodiment of the present disclosure is schematically shown.

[0050] In various other examples and implementations, pages from a given plane that participate in the same fault tolerance stripe may be separated by at least a specified number of word lines (i.e., one or more word lines). The probability that two pages of the same plane are simultaneously faulty decreases as the number of word lines separating the two pages increases.

[0051] Figure 5 The numbers of logical units, planes, pages, word lines, and fault-tolerant stripes in the illustrative example of are chosen for illustrative purposes and are not limiting; other implementations may use various other numbers of logical units, planes, pages, and fault-tolerant stripes.

[0052] Due to physical limitations of various memory device implementations, pages in may only be programmed in ascending order of their corresponding page numbers within the corresponding planes. Accordingly, forming a fault-tolerant stripe that utilizes multiple pages from each plane of each logical unit of at least a subset of logical units may require keeping multiple fault-tolerant stripes open, and thus will further require additional memory (e.g., on a non-volatile memory device of the memory subsystem) to store intermediate metadata. In an illustrative example, when a page is programmed (i.e., when the contents of the page are stored to the memory device) the contents of the page may be summed, for example, by an exclusive or (XOR) operation, and the intermediate result of the XOR operation may be stored in memory until the fault-tolerant stripe is closed (i.e., until the metadata is written to the last page of the fault-tolerant stripe).

[0053] Accordingly, if Figure 6 Schematically shown, Figure 6 Schematically illustrating an example fault-tolerant layout 600 of a memory device according to an embodiment of the present disclosure, a fault-tolerant stripe may be formed using a plurality of pages with sequential page numbers from each plane of each logical unit of at least a subset of logical units of the memory device, provided that pages residing in the same plane have sufficient word line separation (e.g., separated by at least a predetermined number of word lines including a level separation word line). Level separation word lines refer to dummy word lines that do not contain data pages. Level separation word lines effectively split the logical units across all planes into an upper level located above the level separation word lines and a lower level located below the word separation word lines.

[0054] exist Figure 6 In the illustrative example of FIG. 1 , a first fault-tolerant stripe is formed by pages having a first predefined number (e.g., page #0) and pages having a second predefined number following the first predefined number (e.g., page #1) across planes and logical units of a memory device, such that all but one of the pages of the fault-tolerant stripe are used to store host data, while the last page P1 is used to store redundant metadata for the fault-tolerant stripe. Page #0 is separated by a level word line WL n=5 Separated from page #1.

[0055] In operation, as host data is arriving, a first predefined number (e.g., page #0) having all planes across all logical units is programmed, and then a second predefined number (e.g., page #1) having all planes across all logical units is programmed. As the pages are programmed (i.e., as the contents of the pages are stored to the memory device), redundant metadata may be calculated by, for example, summing the contents of the pages with an exclusive OR (XOR) operation, and intermediate results of the XOR operation may be stored in memory until the fault-tolerant stripe is closed (i.e., until the metadata is written to the last page of the fault-tolerant stripe). Thus, no additional memory is required to store intermediate XOR results, as the XOR result for page #0 may continue to page #1. After the last data page of the fault-tolerant stripe is programmed, the XOR result is written to the memory device ( Figure 6 Page P) in the illustrative example.

[0056] Figure 6 The numbers of logical units, planes, pages, word lines, and fault-tolerant stripes in the illustrative example of are chosen for illustrative purposes and are not limiting; other implementations may use various other numbers of logical units, planes, pages, and fault-tolerant stripes.

[0057] In some embodiments, the memory device layout may fail to provide the necessary word line separation for at least some sequentially numbered pages. Figure 7 In the illustrative example of Figure 7 Schematically illustrating an example fault-tolerant layout 600 of a memory device according to an embodiment of the present disclosure, a page with a first predetermined number (e.g., page #0) and a page with a second predetermined number subsequent to the first predetermined number (e.g., page #1) are separated by a single word line (a level separation word line), and the number of required separation word lines may be two. Accordingly, at least two fault-tolerant stripes of different sizes may be formed on the memory device: for a page residing adjacent to the level separation word line WL n+5 Word line WL n+4 and WL n+6 , "short" fault-tolerant stripes may be formed such that each fault-tolerant stripe will only include a single page from each plane of each logical unit of at least a subset of the logical units of the memory device (similar to Figure 4 The remaining “long” fault-tolerant stripe may be formed to include two or more pages with sequential page numbering from each plane of each logical unit, because pages residing in the same plane (residing adjacent to the level separation word line WL n+5 Word line WL n+4 and WL n+6 The pages above (except for the pages above) have sufficient word line separation.

[0058] Because the word line separation can increase as the page number increases, the fault-tolerant stripe size can be dynamically adjusted based on the layout of the memory device in order to provide the necessary word line separation. Figure 7 In the illustrative example of FIG. 1 , pages #0 and #1 are separated by a single word line (a level separation word line), while the number of required separation word lines may be three. Accordingly, at least two fault-tolerant stripes of different sizes may be formed on the memory device: for pages residing adjacent to the level separation word line WL n+5 Word line WL n+4 and WL n+6 , a "short" fault-tolerant stripe can be formed so that each fault-tolerant stripe will only contain a single page from each plane of each logical unit (similar to Figure 4 illustrative example). Thus, each of pages #0 through #7 will have its own fault-tolerant stripe.

[0059] On the contrary, from WL n+3 and WL n+7 The starting word line provides the necessary word line separation of at least three word lines. Accordingly, a "long" fault-tolerant stripe can be formed to include two or more pages with sequential page numbering from each plane of each logical unit, provided that the pages reside in WL n+3…n+1 and WL n+7…n+9 For example, Page #8 and Page #9 have sufficient word line separation and therefore can participate in a single fault-tolerant stripe.

[0060] Accordingly, as host data is arriving, pages are programmed to form a fault-tolerant stripe of a specified size, which may be predetermined in order to limit over-provisioning losses due to implementation of a fault-tolerant scheme. Each fault-tolerant stripe may include multiple data pages, with the last page of the fault-tolerant stripe being dedicated to storing redundant metadata that may be used for error detection and correction. This fault-tolerant stripe may be formed using multiple pages with sequential page numbers from each plane of each logical unit of at least a subset of logical units of the memory device, provided that pages residing in the same plane have sufficient word line separation (e.g., separated by at least a predetermined number of word lines including a plane separation word line).

[0061] As pages are programmed (i.e., as the contents of the pages are stored to the memory device), redundant metadata may be calculated by summing the contents of the pages, for example, by means of an exclusive-OR (XOR) operation, and intermediate results of the XOR operation may be stored in memory until the fault-tolerant stripe is closed (i.e., until the metadata is written to the last page of the fault-tolerant stripe). Because pages have sequential page numbering, no additional memory is required to store intermediate XOR results, since the XOR results for page #i may continue to page #i+1. After the last data page of the fault-tolerant stripe is programmed, the XOR results are written to the memory device. The fault-tolerant stripe size may be dynamically adjusted based on the layout of the memory device in order to provide the necessary word line separation.

[0062] Figure 8 is a flow chart of an example method 800 for implementing fault-tolerant page striping by a memory subsystem controller operating according to some embodiments of the present disclosure. The method 800 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 800 is performed by Figure 1 The fault-tolerant manager 113 of the embodiment of the present invention is executed. Although shown in a specific sequence or order, unless otherwise specified, the order of operations may be modified. Therefore, it should be understood that the described embodiments are only examples, and the described operations may be performed in a different order, and some operations may be performed in parallel. In addition, in some embodiments, one or more operations may be omitted. Therefore, not all of the operations described are required in every embodiment, and other process flows are possible.

[0063] At operation 810, a processing device of a memory subsystem controller receives an i-th host data item (ie, data to be stored on a memory device) from a host.

[0064] At operation 820, the processing device stores the i-th host data item in page #j of plane #k of LUN #m of the memory device. The page is associated with a currently opened fault-tolerant stripe.

[0065] At operation 830 , the processing device receives an (i+1)th host data item from the host.

[0066] At operation 840, the processing device stores the second host data item in page #j+1 of plane #k+1 of LUN #m of the memory device. The page is associated with the currently opened fault-tolerant stripe. Based on the layout of the memory device, page #j of plane #k+1 is separated from page #j+1 of plane #k by at least a predetermined number of word lines.

[0067] At operation 850, the processing device stores redundant metadata associated with the currently opened fault-tolerant stripe in page #(j+p) of plane #(k+q) of LUN #(m+r) of the storage device, thereby closing the fault-tolerant stripe.

[0068] Fig. 9 An example machine of computer system 900 is shown, within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, computer system 900 may correspond to a host system (e.g., Figure 1 1) a host system 120 that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 The memory subsystem 110 of the controller may be used to execute the operation of the controller (for example, execute the operating system to execute the corresponding Figure 1 In some embodiments, the machine may be connected (e.g., using a network) to other machines in a LAN, an intranet, an extranet, and / or the Internet. The machine may operate in the capacity of a server or a client machine in a client-server network environment, or in the capacity of a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment.

[0069] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, a digital or non-digital circuit system, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be performed by the machine. In addition, while a single machine is described, the term "machine" shall also be construed to include any collection of machines that individually or collectively execute one (or more) sets of instructions to perform any one or more of the methodologies discussed herein.

[0070] The example computer system 900 includes a processing device 902, a main memory 904 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), a static memory 906 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 918, which communicate with each other via a bus 930.

[0071] The processing device 902 represents one or more general processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device 902 can also be one or more special processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 902 is configured to execute instructions 926 for performing the operations and steps discussed herein. The computer system 900 may further include a network interface device 908 to communicate on a network 920.

[0072] The data storage system 918 may include a machine-readable storage medium 924 (also referred to as a computer-readable medium) on which is stored one or more sets of instructions 926 or software embodying any one or more of the methods or functions described herein. The instructions 926 may also reside, in whole or in part, within the main memory 904 and / or within the processing device 902 during execution thereof by the computer system 900, the main memory 904 and the processing device 902 also constituting machine-readable storage media. The machine-readable storage medium 924, the data storage system 918, and / or the main memory 904 may correspond to Figure 1 Memory subsystem 110.

[0073] In one embodiment, instructions 926 include instructions for implementing a method corresponding to a fault-tolerant stripe component (e.g., Figure 1 Although the machine-readable storage medium 924 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. Therefore, the term "machine-readable storage medium" should be considered to include, but not limited to, solid-state memory, optical media, and magnetic media.

[0074] Some portions of the previously detailed description have been presented with respect to algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing can most effectively communicate the substance of their work to other skilled in the art. Algorithms are here and generally considered to be self-consistent sequences of operations that produce desired results. Operations are operations that require physical manipulation of physical quantities. These quantities are usually, but not necessarily, in the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. Sometimes, primarily for general reasons, it has proven convenient to refer to these signals as bits, values, elements, symbols, characters, items, numbers, etc.

[0075] It should be borne in mind, however, that all of these and similar terms should be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may be directed to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within a computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage systems.

[0076] The present disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as, but not limited to, any type of disk (including floppy disks, optical disks, CD-ROMs, and magnetic optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0077] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used with programs according to the teachings herein, or it may prove convenient to construct more specialized devices to perform the methods. The structures of a variety of these systems will be presented as set forth in the description below. In addition, the present disclosure is not described with reference to any particular programming language. It should be appreciated that a variety of programming languages ​​may be used to implement the teachings of the present disclosure described herein.

[0078] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having stored thereon instructions that can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.

[0079] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments thereof. It should be apparent that various modifications may be made to the present disclosure without departing from the broader spirit and scope of embodiments of the present disclosure as set forth in the appended claims. Accordingly, the description and drawings should be viewed in an illustrative rather than a restrictive sense.

Claims

1. A system comprising: Memory device; as well as A processing device operatively coupled to the memory device to perform operations comprising: receiving host data items; storing the host data item in a first page of the memory device, wherein the first page is associated with a fault-tolerant stripe, and wherein the first page is separated from a second page of the fault-tolerant stripe by one or more word lines including a level separation word line, wherein the first page and the second page have sequential page numbers; and Redundant metadata associated with the fault-tolerant stripe is stored on the storage device.

2. The system of claim 1, wherein the one or more word lines include dummy word lines that do not store host data. 3 . The system of claim 1 , wherein the fault-tolerant stripe comprises a specified number of pages. 4 . The system of claim 1 , wherein the redundant metadata represents a bit-wise exclusive OR of a plurality of pages associated with the fault-tolerant stripe.

5. The system of claim 1, wherein the operations further comprise: A plurality of host data items are stored in a plurality of pages associated with a second fault-tolerant stripe, wherein a first size of the fault-tolerant stripe is different than a second size of the second fault-tolerant stripe. The system of claim 1 , wherein a size of the fault-tolerant stripe is determined based on a layout of the memory device.

7. A method comprising: receiving, by a processing device of a memory subsystem controller, a host data item; storing the host data item in a first page of a memory device, wherein the first page is associated with a fault-tolerant stripe, and wherein the first page is separated from a second page of the fault-tolerant stripe by one or more word lines including a level separation word line, wherein the first page and the second page have sequential page numbers; and Redundant metadata associated with the fault-tolerant stripe is stored on the storage device.

8. The method of claim 7, wherein the one or more word lines include dummy word lines that do not store host data.

9. The method of claim 7, wherein the fault-tolerant stripe comprises a specified number of pages.

10. The method of claim 7, wherein the redundant metadata represents a bit-wise exclusive OR of a plurality of pages associated with the fault-tolerant stripe.

11. The method according to claim 7, further comprising: A plurality of host data items are stored in a plurality of pages associated with a second fault-tolerant stripe, wherein a first size of the fault-tolerant stripe is different than a second size of the second fault-tolerant stripe.

12. The method of claim 7, determining a size of the fault-tolerant stripe based on a layout of the memory device.

13. A non-transitory computer-readable storage medium comprising executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising: receiving host data items; storing the host data item in a first page of a memory device, wherein the first page is associated with a fault-tolerant stripe, and wherein the second page is separated from a second page of the fault-tolerant stripe by one or more word lines including a level separation word line, wherein the first page and the second page have sequential page numbers; Redundant metadata associated with the fault-tolerant stripe is stored on the storage device.

14. The non-transitory computer-readable storage medium of claim 13, wherein the one or more word lines include a dummy word line that does not store host data.

15. The non-transitory computer-readable storage medium of claim 13, wherein the fault-tolerant stripe comprises a specified number of pages.

16. The non-transitory computer-readable storage medium of claim 13, wherein the redundant metadata represents a bit-wise exclusive OR of a plurality of pages associated with the fault-tolerant stripe.

17. The non-transitory computer-readable storage medium of claim 13, further comprising executable instructions that, when executed by a processing device, cause the processing device to perform operations comprising: A plurality of host data items are stored in a plurality of pages associated with a second fault-tolerant stripe, wherein a first size of the fault-tolerant stripe is different than a second size of the second fault-tolerant stripe.