Dual device data correction in memory devices using extended Reed-Solomon codewords
By associating multiple memory strips in a memory system and using extended Reed-Solomon codewords, data correction problems in the event of multiple memory strip failures are solved, improving system reliability and reducing energy consumption.
Patent Information
- Application Number
- CN202510073961.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-12-18
- Filing Date
- 2025-01-17
- Publication Date
- 2025-07-18
AI Technical Summary
In the event of more than one memory strip failure in existing memory systems, the traditional Reed-Solomon code cannot effectively correct errors, resulting in system unreliability and data loss.
Dual device data correction is achieved by associating multiple memory strips and using extended Reed-Solomon codewords, the lost data is recovered using the error correction element of the second memory strip.
Improves the reliability of the memory system, reduces data loss and read/write errors, and reduces power and compute storage consumption.
Smart Images

Figure CN120336074A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This patent application claims priority to U.S. Provisional Patent Application No. 63 / 622,495, filed on Jan. 18, 2024, titled “Double Device Data Correction in Memory Devices Using Enlarged Reed-Solomon Codewords” and assigned to its assignee. The disclosure of the prior application is considered part of this patent application and is incorporated herein by reference. Technical Field
[0003] This disclosure generally relates to memory devices, memory device operations, and, for example, double device data correction in memory devices using enlarged Reed-Solomon codewords. Background Art
[0004] Memory devices are widely used to store information in various electronic devices. A memory device includes memory cells. A memory cell is an electronic circuit capable of being programmed to a data state of two or more data states. For example, a memory cell can be programmed to a data state representing a single binary value (commonly represented by a binary “1” or a binary “0”). As another example, a memory cell can be programmed to a data state representing a fractional value (e.g., 0.5, 1.5, or the like). To store information, an electronic device can write to or program a group of memory cells. To access the stored information, the electronic device can read from or sense the stored state of the group of memory cells.
[0005] There are various types of memory devices, including random access memory (RAM), read only memory (ROM), dynamic RAM (DRAM), static RAM (SRAM), synchronous dynamic RAM (SDRAM), ferroelectric RAM (FeRAM), magnetic RAM (MRAM), resistive RAM (RRAM), holographic RAM (HRAM), flash memory (e.g., NAND memory and NOR memory), and others. A memory device can be volatile or non-volatile. A non-volatile memory (e.g., flash memory) can store data for a long time even without an external power source. A volatile memory (e.g., DRAM) loses stored data over time unless the volatile memory is refreshed by a power source. In some instances, a memory device can be associated with a Compute Express Link (CXL). For example, a memory device can be a CXL-compliant memory device and / or can include a CXL interface. Summary of the Invention
[0006] In one aspect, the present disclosure relates to a memory device that includes: one or more components configured to: associate a first memory strip with a second memory strip, wherein the first memory strip is associated with a first set of data storage elements and a first set of error correction elements, and wherein the second memory strip is associated with a second set of data storage elements and a second set of error correction elements; receive a first codeword associated with the first memory strip, wherein the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity information stored at the first set of error correction elements; identify a first error in the first set of data bits using the first codeword; correct the first error using the first codeword; receive a second codeword associated with the first memory strip and the second memory strip, wherein the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one error correction element of the first set of error correction elements, and wherein the second codeword includes a second set of error correction bits associated with parity information stored at the second set of error correction elements; identify a second error in the second set of data bits; and correct the second error using the second codeword.
[0007] On the other hand, the present disclosure relates to a method that includes: associating, by a memory device, a first memory strip with a second memory strip, where the first memory strip is associated with a first set of data storage elements and a first set of error correction elements, and where the second memory strip is associated with a second set of data storage elements and a second set of error correction elements; receiving, by the memory device, a first codeword associated with the first memory strip, where the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity check information stored at the first set of error correction elements; identifying, by the memory device, a first error in the first set of data bits using the first codeword; correcting, by the memory device, the first error using the first codeword; receiving, by the memory device, a second codeword associated with the first memory strip and the second memory strip, where the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one error correction element of the first set of error correction elements, and where the second codeword includes a second set of error correction bits associated with parity check information stored at the second set of error correction elements; identifying, by the memory device, a second error in the second set of data bits; and correcting, by the memory device, the second error using the second codeword.
[0008] In another aspect, the present disclosure relates to a memory system, comprising: a memory controller; and a plurality of encoder / decoder components associated with the memory controller, wherein the memory system is configured to: associate a first memory stripe with a second memory stripe by the memory controller, wherein the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and wherein the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receive, by a first encoder / decoder component of the plurality of encoder / decoder components, a first codeword associated with the first memory stripe, wherein the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity check information stored at the first set of error correction elements; identify, by the first encoder / decoder component, a first error in the first set of data bits using the first codeword; correct, by the first encoder / decoder component, the first error using the first codeword; receive, by a second encoder / decoder component, a second codeword associated with the first memory stripe and the second memory stripe, wherein the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one error correction element of the first set of error correction elements, and wherein the second codeword includes a second set of error correction bits associated with parity check information stored at the second set of error correction elements; identify, by the second encoder / decoder component, a second error in the second set of data bits; and correct, by the second encoder / decoder component, the second error using the second codeword. Description of the Drawings
[0009] Figure 1 is a diagram illustrating an example system capable of performing dual device data correction in a memory device using extended Reed-Solomon (RS) codewords.
[0010] Figures 2A to 2C is a diagram of an example associated with an error correction code.
[0011] Figures 3A to 3D is a diagram of an example of performing dual device data correction (DDDC) in a memory device using extended RS codewords.
[0012] Figure 4 is a flowchart of an example method associated with performing DDDC in a memory device using extended RS codewords. Detailed Description
[0013] Memory systems and / or devices may utilize error correction code (ECC) to identify and / or correct errors in data accessed from memory. For example, data may be striped across multiple memory dies (sometimes referred to herein as a memory stripe), where multiple dies are used to store data bits and / or parity bits. For example, a memory stripe may be associated with 10 dies (e.g., 10 dynamic random access memory (DRAM) dies), where 8 dies are used to store data bits and 2 dies are used to store parity bits. In some instances, the parity bits may store information such that in the event that an entire die fails, the information may be used in conjunction with ECC to correct the data (sometimes referred to as chip kill protection). For example, if an entire die of a DRAM stack fails, the stored parity bits may be encoded in such a way that the parity bits can be used to recover the data stored on the failed die.
[0014] In some instances, ECC may be associated with Reed-Solomon (RS) codes and / or the memory stripe may be associated with an RS chip kill protection scheme. For example, a memory system and / or device may utilize an 8-bit RS code, a 16-bit RS code, or a similar RS code to correct a number of bits corresponding to one failed die in a memory stripe, thereby providing chip kill protection in the event that an entire die of the memory stripe fails. However, if more than one data die of the memory stripe fails and / or contains errors, the ECC procedure implementing the RS code may be ineffective. Thus, if a first chip kill event occurs in conjunction with a memory stripe, the RS code is able to correct errors and / or retrieve lost data. However, if a second or subsequent chip kill event occurs in conjunction with the memory stripe, the RS code is unable to correct errors and / or retrieve lost data, resulting in uncorrectable errors. This can lead to an unreliable memory system, unrecoverable host data, read / write errors, and high power, computational, and storage consumption to move host data, rewrite host data, and / or recover host data.
[0015] Some embodiments described herein implement dual device data correction (DDDC) (e.g., correction of errors associated with two or more failed dies of a memory stripe) for certain memory systems (e.g., memory systems employing an RS-based error correction scheme). In some embodiments, a memory system may associate multiple memory stripes (e.g., two memory stripes) with each other, where each memory stripe includes a respective data storage element (e.g., a data die) and a respective error correction element (e.g., a parity die). In some embodiments, a memory controller, an encoder / decoder component, and / or another component of the memory system is capable of encoding and / or decoding an extended RS codeword after a first die failure, e.g., to correct errors associated with a second or subsequent die failure. For example, in some embodiments, the memory system may associate two memory stripes with each other and / or pair original RS codewords. Original (e.g., unextended) RS codewords may be used to correct the first die failure. Additionally, after the first die failure, recovered data may be written to the error correction element (e.g., the parity die) of the first memory stripe, and the error correction element of the second memory stripe may be used to store error correction bits of an extended codeword associated with both the first memory stripe and the second memory stripe. In this way, if another data storage element (e.g., a data die) of the first memory stripe and / or the second memory stripe fails, the memory system may use the error correction bits stored in the error correction element of the second memory stripe to recover lost data, thereby implementing DDDC at the memory system. This can result in increased reliability of the memory system, reduced data loss and / or read / write errors, and reduced power, computation, and storage consumption required to move, rewrite, and / or recover host data.
[0016] Figure 1 FIG. is a diagram illustrating an example system 100 capable of performing dual device data correction in a memory device using extended Reed-Solomon (RS) codewords. System 100 may include one or more devices, apparatuses, and / or components for performing the operations described herein. For example, system 100 may include a host system 105 and a memory system 110. Memory system 110 may include a memory system controller 115 and one or more memory devices 120, shown as memory devices 120-1 through 120-N (where N≥1). The memory devices may include local controllers 125 and one or more memory arrays 130. Host system 105 may communicate with memory system 110 (e.g., memory system controller 115 of memory system 110) via a host interface 140. Memory system controller 115 and memory devices 120 may communicate via respective memory interfaces 145, shown as memory interfaces 145-1 through 145-N (where N≥1).
[0017] System 100 can be any electronic device configured to store data in a memory. For example, system 100 can be a computer, a mobile phone, a wired or wireless communication device, a network device, a server, a device in a data center, a device in a cloud computing environment, a vehicle (such as a car or an airplane), and / or an Internet of Things (IoT) device. Host system 105 can include host processor 150. Host processor 150 can include one or more processors configured to execute instructions and store data in memory system 110. For example, host processor 150 can include a central processing unit (CPU), a graphics processing unit (GPU), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), and / or another type of processing component.
[0018] Memory system 110 can be any electronic device or apparatus configured to store data in a memory. For example, memory system 110 can be a hard disk, a solid state drive (SSD), a flash memory system (such as a NAND flash memory system or a NOR flash memory system), a universal serial bus (USB) drive, a memory card (such as a secure digital (SD) card), an auxiliary storage device, a non-volatile memory express (NVMe) device, an embedded multimedia card (eMMC) device, a dual in-line memory module (DIMM), and / or a random access memory (RAM) device, such as a dynamic RAM (DRAM) device or a static RAM (SRAM) device.
[0019] Memory system controller 115 can be any device configured to control the operation of memory system 110 and / or the operation of memory device 120. For example, memory system controller 115 can include control logic, a memory controller, a system controller, an ASIC, an FPGA, a processor, a microcontroller, and / or one or more processing components. In some embodiments, memory system controller 115 can communicate with host system 105 and can instruct one or more memory devices 120 regarding memory operations performed by the one or more memory devices 120 based on one or more instructions from host system 105. For example, memory system controller 115 can provide instructions regarding memory operations performed by local controller 125 in conjunction with corresponding memory device 120 to local controller 125.
[0020] Memory device 120 may include a local controller 125 and one or more memory arrays 130. In some embodiments, memory device 120 includes a single memory array 130. In some embodiments, each memory device 120 of memory system 110 may be implemented in a separate semiconductor package or on a separate die that includes the corresponding local controller 125 and corresponding memory array 130 of the memory device 120. Memory system 110 may include multiple memory devices 120.
[0021] The local controller 125 may be any device configured to control the memory operations of the memory device 120 in which the local controller 125 is included (e.g., and not control the memory operations of other memory devices 120). For example, the local controller 125 may include control logic, a memory controller, a system controller, an ASIC, an FPGA, a processor, a microcontroller, and / or one or more processing components. In some embodiments, the local controller 125 may communicate with the memory system controller 115 and may control operations performed on the memory array 130 coupled to the local controller 125 based on one or more instructions from the memory system controller 115. As an example, the memory system controller 115 may be an SSD controller, and the local controller 125 may be a NAND controller.
[0022] The memory array 130 may include an array of memory cells configured to store data. For example, the memory array 130 may include a non-volatile memory array (e.g., a NAND memory array or a NOR memory array) or a volatile memory array (e.g., a SRAM array or a DRAM array). In some embodiments, memory system 110 may include one or more volatile memory arrays 135. The volatile memory arrays 135 may include SRAM arrays and / or DRAM arrays and other examples. The one or more volatile memory arrays 135 may be included in the memory system controller 115, one or more memory devices 120, and / or both the memory system controller 115 and one or more memory devices 120. In some embodiments, memory system 110 may include both non-volatile memory capable of maintaining stored data after the memory system 110 is powered off and volatile memory (e.g., volatile memory arrays 135) that requires power to maintain stored data and loses stored data after the memory system 110 is powered off. For example, the volatile memory arrays 135 may cache data read from or written to non-volatile memory and / or may cache instructions executed by the controller of the memory system 110.
[0023] The host interface 140 enables communication between the host system 105 (e.g., the host processor 150) and the memory system 110 (e.g., the memory system controller 115). The host interface 140 may include, for example, a Small Computer System Interface (SCSI), a Serial Attached SCSI (SAS), a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, an NVMe interface, a USB interface, a Universal Flash Storage (UFS) interface, an eMMC interface, a Double Data Rate (DDR) interface, and / or a DIMM interface.
[0024] The memory interface 145 enables communication between the memory system 110 and the memory device 120. The memory interface 145 may include a non-volatile memory interface (e.g., for communicating with non-volatile memory), such as a NAND interface or a NOR interface. Additionally or alternatively, the memory interface 145 may include a volatile memory interface (e.g., for communicating with volatile memory), such as a DDR interface.
[0025] In some instances, the memory system 110 may be a Compute Express Link (CXL)-compliant memory system (sometimes referred to herein simply as a CXL memory system), and / or one or more of the memory devices 120 may be CXL-compliant memory devices (sometimes referred to herein simply as CXL memory devices). CXL is a high-speed CPU-to-device and CPU-to-memory interconnect designed to accelerate next-generation performance. CXL technology maintains memory coherence between the CPU memory space and the memory on attached devices, which allows resource sharing to achieve higher performance, reduced software stack complexity, and reduced total system cost. CXL is designed as an industry-open standard interface for high-speed communication. CXL technology builds on the PCIe infrastructure to provide advanced protocols, such as input / output (I / O) protocols, memory protocols, and coherence interfaces, in the fabric using the PCIe physical and electrical interfaces.
[0026] In some instances, the memory system 110 may include a PCIe / CXL interface (e.g., the host interface 140 may be associated with the PCIe / CXL interface), which may be a physical interface configured to connect a CXL memory system and / or a CXL memory device to a CXL-compliant host device. In such instances, the PCIe / CXL interface may comply with the CXL physical connection standard specification, ensuring broad compatibility and ease of integration into existing systems using the CXL protocol. Additionally or alternatively, the CXL memory system and / or the CXL memory device may be designed to efficiently interface with a computing system (e.g., the host system 105) by leveraging the CXL protocol. For example, the CXL memory system and / or the CXL memory device may be configured to utilize the high-speed, low-latency interconnect capabilities of CXL, such as for the purpose of adapting the CXL memory system and / or the CXL memory device to high-performance computing, data center applications, artificial intelligence (AI) applications, and / or similar applications.
[0027] The CXL memory system and / or the CXL memory device may include a CXL memory controller (e.g., the memory system controller 115 and / or the local controller 125), which may be configured to manage the data flow between a memory array (e.g., the volatile memory array 135 and / or the memory array 130) and the CXL interface (e.g., the PCIe / CXL interface, e.g., the host interface 140). In some instances, the CXL memory controller may be configured to handle one or more CXL protocol layers, such as: an I / O layer (e.g., a layer associated with the CXL.io protocol, which may be used for purposes such as device discovery, configuration, initialization, I / O virtualization, direct memory access (DMA) using non-uniform load-store semantics, and / or similar purposes); a cache coherence layer (e.g., a layer associated with the CXL.cache protocol, which may be used for purposes such as caching host memory using the modified, exclusive, shared, invalid (MESI) coherence protocol or similar purposes); or a memory protocol layer (e.g., a layer associated with the CXL.memory (sometimes referred to as CXL.mem) protocol, which may enable the CXL memory device to expose host-managed device memory (HDM) to permit the host device to manage and access memory similar to native DDR connected to the host); and other instances.
[0028] A CXL memory system and / or a CXL memory device may further include and / or be associated with one or more high-bandwidth memory modules (HBMMs) or similar memory arrays (such as volatile memory array 135 and / or memory array 130). For example, a CXL memory system and / or a CXL memory device may include multiple layers of DRAM (such as stacked and / or interconnected via advanced through-silicon via (TSV) technology) to maximize storage density and / or enhance data transfer speed between memory layers. Additionally or alternatively, a CXL memory system and / or a CXL memory device may include a power management unit, which may be configured to regulate the power consumption associated with the CXL memory system and / or the CXL memory device and / or may be configured to improve the energy efficiency of the CXL memory system and / or the CXL memory device. Additionally or alternatively, a CXL memory system and / or a CXL memory device may include additional components, such as one or more error correction code (ECC) engines, for example for the purpose of detecting and / or correcting data errors to ensure data integrity and / or improve the overall reliability of the CXL memory system and / or the CXL memory device.
[0029] Although the example memory system 110 described above includes a memory system controller 115, in some embodiments, the memory system 110 does not include a memory system controller 115. For example, an external controller (such as included in host system 105) and / or one or more local controllers 125 included in one or more corresponding memory devices 120 may perform the operations described herein as being performed by the memory system controller 115. Additionally, as used herein, "controller" may refer to the memory system controller 115, the local controller 125, or the external controller. In some embodiments, a set of operations described herein as being performed by a controller may be performed by a single controller. For example, the entire set of operations may be performed by a single memory system controller 115, a single local controller 125, or a single external controller. Alternatively, a set of operations described herein as being performed by a controller may be performed by more than one controller. For example, a first subset of the operations may be performed by the memory system controller 115, and a second subset of the operations may be performed by the local controller 125. Additionally, the term "memory device" may refer to the memory system 110 or the memory device 120, depending on the context.
[0030] A controller (such as memory system controller 115, local controller 125, or external controller) can control operations performed on a memory (such as memory array 130), for example, by executing one or more instructions. For example, memory system 110 and / or memory device 120 can store one or more instructions in the memory as firmware, and the controller can execute the one or more instructions. Additionally or alternatively, the controller can receive one or more instructions from host system 105 and / or memory system controller 115 and can execute the one or more instructions. In some embodiments, a non-transitory computer-readable medium (such as volatile memory and / or non-volatile memory) can store a set of instructions (such as one or more instructions or code) for the controller to execute. The controller can execute the set of instructions to perform one or more operations or methods described herein. In some embodiments, the execution of the set of instructions by the controller causes the controller, memory system 110, and / or memory device 120 to perform one or more operations or methods described herein. In some embodiments, hardwired circuitry is used instead of or in combination with one or more instructions to perform one or more operations or methods described herein. Additionally or alternatively, the controller can be configured to perform one or more operations or methods described herein. Instructions are sometimes referred to as "commands."
[0031] For example, a controller (such as memory system controller 115, local controller 125, or external controller) can transmit signals to and / or receive signals from a memory (such as one or more memory arrays 130) based on one or more instructions, such as transferring (e.g., writing or programming) data to all or a portion of the memory (such as one or more memory cells, pages, sub-blocks, blocks, or planes of the memory), transferring (e.g., reading) data from all or a portion of the memory, and / or refreshing all or a portion of the memory. Additionally or alternatively, the controller can be configured to control access to the memory and / or provide a translation layer between host system 105 and the memory (such as for mapping logical addresses to physical addresses of memory array 130). In some embodiments, the controller can translate host interface commands (such as commands received from host system 105) into memory interface commands (such as commands for performing operations on memory array 130).
[0032] In some embodiments, Figure 1One or more systems, devices, apparatuses, components, and / or controllers may be configured to: associate a first memory stripe with a second memory stripe, where the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and where the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receive a first codeword associated with the first memory stripe, where the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity information stored at the first set of error correction elements; identify a first error in the first set of data bits using the first codeword; correct the first error using the first codeword; receive a second codeword associated with the first memory stripe and the second memory stripe, where the second codeword includes a second set of data bits associated with data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one error correction element of the first set of error correction elements, and where the second codeword includes a second set of error correction bits associated with parity information stored at the second set of error correction elements; identify a second error in the second set of data bits; and correct the second error using the second codeword.
[0033] Figure 1 The number and arrangement of components shown in Figure 1 are for illustration only. In fact, there may be additional components, fewer components, different components, or different arrangements of components as compared to the components shown in Figure 1 Two or more components shown in Figure 1 may be implemented within a single component, or Figure 1 a single component shown in Figure 1 may be implemented as multiple distributed components. Additionally or alternatively,
[0034] Figures 2A to 2C is a diagram of an example associated with an error correction code. The operations described in conjunction with Figures 2A to 2C may be performed by the memory system 110 and / or one or more components of the memory system 110 (such as the memory system controller 115, one or more memory devices 120, one or more local controllers 125, and / or one or more ECC engines associated with the memory system 110 and / or one or more memory devices 120).
[0035] As Figure 2A shown, ECC may be used in conjunction with memory stripes 200 (sometimes referred to as data blocks, data frames, and / or similar terms), and the memory stripes 200 may correspond to those described above in conjunction with Figure 1The described volatile memory array 135. In some instances, a memory strip 200 may be associated with a memory channel (e.g., a data path between the memory and other components of the memory device such as a memory controller and / or a processor), where the "width" of the memory channel (e.g., measured in bits) refers to the number of bits that can be transferred in one operation and / or one memory cycle. For example, as described in more detail below, in some instances, the memory strip 200 may be associated with a 40-bit channel, and thus, the memory device associated with the memory strip 200 may be referred to as a 40-bit memory device. For example, the memory device may be a Double Data Rate 5 (DDR5) 40-bit memory device or a similar device.
[0036] The memory strip 200 may be associated with a plurality of memory dies for storing data bits and / or parity bits. In other words, in some instances, a plurality of data bits and / or parity bits may be striped across the plurality of dies associated with the memory strip 200. For example, Figure 2A The memory strip 200 shown in is associated with 10 dies (e.g., 10 DRAM dies) indexed as die 0 to die 9, where dies 0 to 7 are used to store data bits (and thus referred to as data dies, as indicated by reference numeral 202) and where dies 8 to 9 are used to store parity bits for error correction purposes (and thus referred to as parity dies, as indicated by reference numeral 204). As indicated by reference numeral 206, each die may be associated with 16 bit lines (BL), and / or as indicated by reference numeral 208, each die may be configured in a "times 4" (x4) configuration such that each die includes 4 input / output pins (sometimes referred to as DQ pins). In this regard, each die is capable of storing 64 bits (e.g., 8 bytes). In some instances, the memory strip may be associated with 64 bytes of data (corresponding to 8 data dies indicated by reference numeral 202, each capable of storing 8 bytes) and 16 bytes of parity information (corresponding to 2 parity dies indicated by reference numeral 204, each capable of storing 8 bytes). In other words, the data dies of the memory strip 200 may collectively store 512 data bits and / or the parity dies of the memory strip 200 may collectively store 128 parity bits, where each of the 128 parity bits is a function of the 512 data bits. In this way, a 16BL access to 64 bytes of data may include 10 dies in an x4 mode, where 8 dies provide 64 bytes of data (e.g., 8 bytes per die) and where 2 dies provide 16 bytes of redundancy (e.g., 8 bytes per die) for error correction purposes.
[0037] In addition, as indicated by reference numeral 210, the memory strip 200 may be associated with a 40-bit channel, where 32 bits may be associated with data bits (as indicated by reference numeral 212) and 8 bits may be associated with parity bits (as indicated by reference numeral 214). In some instances, a memory system (such as memory system 110) may be organized into channels and / or columns. For example, a memory system may include 4 columns and / or 4 x 40-bit channels. In this regard, Figure 2A the memory strip 200 shown in may be associated with data provided by access in a particular column or a particular channel.
[0038] In some instances, a parity die may store information such that, for example, in the event that an entire die fails, the information may be used in conjunction with ECC to correct data (such as chip kill protection). In other words, the error correction system associated with the memory strip 200 is capable of correcting errors caused by an entire die failure. For example, as indicated by reference numeral 216, in some events, an entire die of a DRAM stack may fail (such as die 3 failing in the depicted example). In such cases, the parity bits stored in the parity die may be encoded in such a way that the parity bits are available for recovering the data stored on the failed die.
[0039] More specifically, Figure 2B and 2C illustrate instances where the parity die is associated with an RS code and / or where the memory strip is associated with an RS chip kill protection scheme. As Figure 2B shown in and as indicated by reference numeral 218, the chip kill protection scheme may be obtained by using an RS code with 8-bit symbols. In such cases, the size of the symbol set (sometimes referred to as q) used in the RS coding scheme for the 40-bit memory strip 200 described above in conjunction with Figure 2A may be equal to 256 (e.g., 2 8 ), the length of the RS codeword 219 (sometimes referred to as n and / or the shortened codeword) may be 80 symbols, and the length of the data portion of the RS codeword (sometimes referred to as k) may be 64 symbols. In some instances, the RS code is capable of correcting up to t symbols, where t is equal to Thus, for the 8-bit symbol instance shown in Figure 2B , the RS code is capable of correcting up to (e.g., 8 bytes), which is equal to the amount of data stored on one die of the memory strip 200. In this regard, in the event that an entire die of the memory strip fails, the 8-bit RS code may be used to provide chip kill protection.
[0040] Similarly, as Figure 2CAs shown and indicated by reference numeral 220, the chip kill protection scheme may alternatively be obtained by using an RS code with 16-bit symbols. In such cases, the size of the symbol set (e.g., q) used in the RS coding scheme of the 40-bit memory stripe 200 described above in connection with Figure 2A may be equal to 65,536 (e.g., 2 16 ), the length of the RS codeword 221 (e.g., n) may be 40 symbols, and the length of the data portion of the RS codeword (e.g., k) may be 32 symbols. Thus, 16-bit symbol instances are capable of correcting up to 4 symbols (e.g., or 8 bytes), which is equal to the amount of data stored on one die. In this regard, in the event that an entire die of the memory stripe fails, the 16-bit RS code may also be used to provide chip kill protection.
[0041] In this way, certain ECC programs (e.g., ECC programs implementing RS codes, such as the program described above in connection with Figures 2A to 2C ) are only effective when a single data storage element (e.g., one data die) fails and / or contains an error. This is because for the above-described 40-bit memory instance, the parity bits stored on the parity die are only capable of correcting up to 8 bytes of data, which is equal to the amount of data stored on one die of the memory stripe 200. Thus, such ECC programs may become ineffective when more than one data storage element of the memory stripe contains an error and / or fails (e.g., when two or more data dies associated with the memory stripe fail). In other words, if a first chip kill event occurs in the memory system and / or in connection with the memory stripe 200, then the RS code is capable of correcting errors and / or retrieving lost data. However, if a second chip kill event occurs in the memory system and / or in connection with the memory stripe 200, then the RS code is unable to correct errors and / or retrieve lost data, resulting in uncorrectable errors. This can lead to an unreliable memory system, unrecoverable host data, read / write errors, and high power, computational, and storage consumption for moving, rewriting, and / or recovering host data.
[0042] Some embodiments described herein implement DDDC for certain memory systems (e.g., memory systems employing an RS-based error correction scheme). In some embodiments, a memory system may associate multiple memory stripes (e.g., two memory stripes) with each other, where each memory stripe includes a corresponding data storage element (e.g., a data die) and a corresponding error correction element (e.g., a parity die). In some embodiments, a memory controller, an encoder / decoder component of the memory system, and / or another component of the memory system are capable of encoding and / or decoding an extended codeword after a first die failure, e.g., for the purpose of correcting errors associated with a second or subsequent die failure. For example, in some embodiments, a memory system may associate two memory stripes with each other and / or pair original RS codewords. The original (e.g., unextended) RS codewords may be used to correct a first die failure, e.g., by implementing an error correction procedure similar to the procedure described above in connection with Figures 2A to 2C In addition, after a first die failure, the recovered data may be written to the error correction element (e.g., the parity die) of the first memory stripe, and the error correction element of the second memory stripe may be used to store error correction bits of an extended codeword associated with both the first memory stripe and the second memory stripe. In this way, if another data storage element (e.g., a data die) of the first memory stripe and / or the second memory stripe fails, then the memory system may use the error correction bits stored in the error correction element of the second memory stripe to recover the lost data, thereby implementing DDDC at the memory system. This may result in increased reliability of the memory system, reduced data loss and / or read / write errors, and reduced power, computation, and storage consumption required to move, rewrite, and / or recover host data.
[0043] As indicated above, Figures 2A to 2C is for illustration only. Other examples may differ from what is described in connection with Figures 2A to 2C Figure 13 is a diagram of example 300 of performing DDDC in a memory device using extended RS codewords. The operations described in connection with
[0044] Figures 3A to 3D Figures 3A to 3D may be performed by the memory system 110 and / or one or more components of the memory system 110 (e.g., the memory system controller 115, one or more memory devices 120, one or more local controllers 125, and / or one or more encoder / decoder components of the memory system 110, which are described in more detail below in connection with Figure 3D
[0045] In some embodiments, an ECC scheme (e.g., an ECC scheme associated with DDDC) may involve associating multiple memory stripes (e.g., multiple memory stripes 200 and / or similar memory stripes) with each other. For example, as shown in instance 300, a memory controller, an encoder / decoder component of a memory system, and / or a similar component of the memory system may associate a first memory stripe 302 with a second memory stripe 304. In some embodiments, each memory stripe 302, 304 may be associated with multiple data storage elements (e.g., data dates) and / or multiple error correction elements (e.g., parity dies) in a manner similar to that described above in connection with memory stripe 200. For example, each memory stripe 302, 304 may be associated with 8 data storage components and / or data dies (e.g., dies indexed 0 to 7 in instance 300) and / or 2 error correction components and / or parity dies (e.g., dies indexed 8 to 9 in instance 300).
[0046] In some embodiments, a memory controller, an encoder / decoder component of a memory system, and / or a similar component of the memory system may store an indication of the association between a first memory stripe 302 and a second memory stripe 304 in a dynamic storage component (e.g., an SRAM component and / or a similar dynamic storage component) associated with the memory system. In this regard, when a host device (e.g., host system 105) accesses one of the memory stripes 302, 304, the memory controller, the encoder / decoder component of the memory system, and / or a similar component of the memory system may access a codeword associated with the paired memory stripe (e.g., the first memory stripe 302 and the second memory stripe 304) to retrieve the data requested by the host, as described in more detail below.
[0047] In some embodiments, each memory stripe 302, 304 may be associated with an ECC scheme, such as an RS-based ECC scheme (e.g., the 8-bit RS-based ECC scheme described above in connection with Figure 2B or the RS-based ECC scheme described above in connection with Figure 2CThe described 16-bit RS-based ECC scheme and other examples). In this regard, the RS codewords associated with each memory stripe can include data bits stored at the data storage elements (e.g., dies 0 to 7) of the corresponding memory stripe and / or error correction bits stored at the error correction elements (e.g., dies 8 to 9) of the corresponding memory stripe. For example, the first RS codeword 303 can be associated with the first memory stripe 302, which can include data bits stored at dies 0 to 7 of the first memory stripe 302 and / or error correction bits (e.g., parity bits) stored at dies 8 to 9 of the first memory stripe 302. Similarly, the second RS codeword 305 can be associated with the second memory stripe 304, which can include data bits stored at dies 0 to 7 of the second memory stripe 304 and / or error correction bits (e.g., parity bits) stored at dies 8 to 9 of the second memory stripe 304.
[0048] In this regard, when data is retrieved in response to a read command received from a host device and / or for a similar purpose, in a manner similar to that described above in conjunction with Figures 2A to 2C As described, any errors detected in the first RS codeword 303 can be corrected using the parity information (e.g., error correction bits) included in the first RS codeword 303, and / or any errors detected in the second RS codeword 305 can be corrected using the parity information (e.g., error correction bits) included in the second RS codeword 305. For example, as Figure 3A shown and indicated by reference numeral 306, the dies associated with the first memory stripe 302 may fail, resulting in the encoder / decoder component detecting an error in the first RS codeword 303 during a read operation and / or a similar memory operation. In other words, the encoder / decoder component can receive the first RS codeword 303 associated with the first memory stripe 302 (e.g., in response to a read command received from a host device associated with the data stored at the first memory stripe 302), the encoder / decoder component can use the first RS codeword 303 to identify the first error in the first set of data bits (e.g., can detect an error cluster associated with die 2, indicating that die 2 has failed) and / or can correct the first error in a manner similar to that described above in conjunction with Figures 2A to 2C using the first RS codeword 303.
[0049] In some embodiments, the memory system may store correction and / or recovery data, e.g., for further error correction at the paired memory strips (sometimes referred to herein as DDDC, indicating that errors from two or more failed dies can be corrected), using one of the error correction elements of one of the paired memory strips 302, 304 (e.g., one of the parity dies). More specifically, as indicated by reference numeral 308, the memory controller, encoder / decoder component, and / or similar components of the memory system may store the recovery and / or correction data (e.g., data recovered using the parity bits of the first RS codeword 303 associated with the failed die) at the parity die (e.g., die 8) of the first memory strip 302. In other words, the memory controller, encoder / decoder component, and / or similar components of the memory system may replace the parity information stored at the first error correction element (e.g., die 8) associated with the first memory strip 302 with the first set of data associated with the error detected using the first RS codeword 303 (e.g., the error caused by the failed die (die 2)).
[0050] By replacing the parity information stored at the error correction element (e.g., die 8 of the first memory strip 302) with the data of the failed die (e.g., die 2), the memory system is able to detect and / or correct subsequent errors associated with the first memory strip 302 and / or the second memory strip 304. More specifically, after the data replacement described above in connection with reference numeral 308, the encoder / decoder component and / or similar components of the memory system may store a third RS codeword 310, which may include an RS payload that extends and / or spans the paired memory strips 302, 304 more than the first RS codeword 303 and / or the second RS codeword 305. More specifically, as Figure 3A shown, the third RS codeword 310 may be associated with both the first memory strip 302 and the second memory strip 304, such that the RS payload of the third RS codeword 310 includes the remaining (e.g., operable) data dies of the first memory strip 302 (e.g., Figure 3A dies 0 to 1 and 3 to 7 in the example shown), at least one parity die of the first memory strip 302 (e.g., the die used to store the data of the failed die, e.g., Figure 3A die 8 in the example shown) and the operable data dies of the second memory strip 304 (e.g., dies 0 to 7). In addition, the third RS codeword 310 may be associated with two parity dies (e.g., dies 8 to 9 of the second memory strip 304), and the parity dies may store parity information for error correction of the extended payload associated with the third RS codeword 310.
[0051] In this regard, for example, in the event that another die associated with the first memory stripe 302 or the second memory stripe 304 fails, the third RS codeword 310 can be used for subsequent error correction. For example, as Figure 3B shown in Figure 3B , an encoder / decoder component and / or a similar component of the memory system can receive the third RS codeword 310, which can include data bits associated with the remaining operable data dies of the first memory stripe 302 (e.g., dies 0 to 1 and 3 to 7), data bits associated with the parity die that is currently being used to store the recovery data for the failed die 2 (e.g., die 8), and data bits associated with the data dies of the second memory stripe 304 (e.g., dies 0 to 7), collectively referred to herein as the extended RS payload. The third RS codeword 310 can also include parity bits associated with the extended RS payload (e.g., dies 0 to 1 and 3 to 8 of the first memory stripe 302 and dies 0 to 7 of the second memory stripe 304), which can be associated with the parity dies of the second memory stripe 304 (e.g., dies 8 to 9), as described above in connection with reference numeral 312. In this way, if a subsequent error occurs in the extended RS payload, the encoder / decoder component can use the new parity information to identify and / or correct the error.
[0052] More specifically and as indicated by reference numeral 314, another die associated with the first memory stripe 302 may fail, resulting in the encoder / decoder component and / or a similar component of the memory system detecting an error in the third RS codeword 310. In other words, the encoder / decoder component can receive the third RS codeword 310 associated with the extended RS payload, and the encoder / decoder component can use the third RS codeword 310 to identify a second error in the extended RS payload (e.g., an error cluster associated with die 6 of the first memory stripe 302 can be detected, indicating that die 6 has failed) and / or can use the third RS codeword 310 to correct the second error.
[0053] In a manner similar to that described above in connection with reference numeral 308, in some embodiments, the memory system may store correction and / or recovery data from a second failed die (e.g., die 6 of the first memory strip 302) using one of the error correction elements of one of the paired memory strips 302, 304 (e.g., one of the parity dies), e.g., for the purpose of further error correction at the paired memory strips (e.g., for the purpose of identifying and / or correcting a third error and / or a third failed die). More specifically, as indicated by reference numeral 316, a memory controller, an encoder / decoder component, and / or a similar component of the memory system may store second recovery and / or correction data (e.g., data associated with the second failed die) at a parity die (e.g., die 9) of the first memory strip 302. In other words, a memory controller, an encoder / decoder component, and / or a similar component of the memory system may replace the parity information stored at a second error correction element (e.g., die 9) associated with the first memory strip 302 with a second set of data associated with an error detected using the third RS codeword 310 (e.g., an error caused by the failed die (die 6)).
[0054] By replacing the parity information stored at an error correction element (e.g., die 9 of the first memory strip 302) with data from a failed die (e.g., die 6), the memory system is able to detect and / or correct subsequent errors associated with the first memory strip 302 and / or the second memory strip 304. More specifically, after the data replacement described above in connection with reference numeral 316, an encoder / decoder component and / or a similar component of the memory system may store a fourth RS codeword, which may include an RS payload that extends and / or spans the paired memory strips 302, 304 more than the first RS codeword 303 and / or the second RS codeword 305. For example, the fourth RS codeword may be associated with both the first memory strip 302 and the second memory strip 304 such that the RS payload of the third RS codeword 310 includes the remaining (e.g., operational) data dies of the first memory strip 302 (e.g., dies 0 to 1, 3 to 5, and 7 in the example shown in Figure 3B and at least one parity die of the first memory strip 302 (e.g., the die used to store the data of the failed die, e.g., Figure 3BThe die (e.g., die 8 to 9) in the example shown in FIG. and the operational data die (e.g., die 0 to 7) of the second memory strip 304. In addition, the fourth RS codeword can be associated with two parity dies (e.g., die 8 to 9 of the second memory strip 304), and the parity dies can store parity information for error correction of the extended payload associated with the fourth RS codeword. In this way, the encoder / decoder component and / or similar components of the memory system can use the fourth RS codeword to detect and / or correct subsequent errors (e.g., the third failed die).
[0055] Although the embodiments described above in conjunction with Figure 3B show the second failed die on the same memory strip (e.g., the first memory strip 302) as the first failed die, in some other embodiments, the third RS codeword 310 (e.g., the extended RS codeword) can be used for the purpose of identifying and / or correcting errors on the second memory strip 304 (e.g., identifying and / or correcting errors caused by the failed die in the second memory strip 304). More specifically, as Figure 3C shown in FIG. and indicated by reference numeral 322, after the RS payload and / or RS codeword extension, the die (e.g., die 0) associated with the second memory strip 304 may fail, resulting in the encoder / decoder component and / or similar components of the memory system detecting an error in the third RS codeword 310. In other words, the encoder / decoder component can receive the third RS codeword 310 associated with the extended RS payload, and the encoder / decoder component can use the third RS codeword 310 to identify the second error in the extended RS payload (e.g., can detect an error cluster associated with die 0 of the second memory strip 304, indicating that die 0 has failed) and / or can use the third RS codeword 310 to correct the second error in a manner similar to that described above in conjunction with Figure 3B described.
[0056] In addition, in a manner similar to that described above in conjunction with Figure 3B reference numeral 316, in some embodiments, the memory system can use one of the error correction elements of one of the paired memory strips 302, 304 (e.g., one of the parity dies) to store data from the second failed die (e.g., Figure 3CCorrection and / or recovery data for die 0 of the second memory stripe 304 in the example shown, e.g., for the purpose of further error correction at the paired memory stripe (e.g., for the purpose of identifying and / or correcting a third error and / or a third failed die). More specifically, as indicated by reference numeral 324, a memory controller, an encoder / decoder component, and / or a similar component of the memory system may store the second recovery and / or correction data (e.g., data associated with the second failed die) at the parity die (e.g., die 9) of the first memory stripe 302. In other words, a memory controller, an encoder / decoder component, and / or a similar component of the memory system may replace the parity information stored at the second error correction element (e.g., die 9) associated with the first memory stripe 302 with a second set of data associated with an error detected using the third RS codeword 310 (e.g., an error caused by the failed die (die 0 of the second memory stripe 304)). By replacing the parity information stored at the error correction element (e.g., die 9 of the first memory stripe 302) with data from the failed die (e.g., die 0 of the second memory stripe 304), the memory system is able to detect and / or correct subsequent errors associated with the first memory stripe 302 and / or the second memory stripe 304 in a manner similar to that described above in conjunction with Figure 3B subsequent errors associated with the first memory stripe 302 and / or the second memory stripe 304 (e.g., using a fourth RS codeword that will include the RS payload of dies 0 to 1 and 3 to 9 of the first memory stripe 302 and dies 1 to 7 of the second memory stripe 304 and will include the parity information stored at dies 8 to 9 of the second memory stripe 304).
[0057] Figure 3D shows that can be used to identify and / or correct errors associated with a memory stripe (e.g., as described above in conjunction with Figures 3A to 3CAn example encoder / decoder component (sometimes referred to herein as an encoder / decoder block) of a memory system (e.g., memory system 110) associated with multiple data device errors of the described paired memory strips 302, 304. As described above, each memory strip 302, 304 can be associated with a 40-bit channel and / or can be paired with another memory strip (e.g., logically associated with another memory strip, where an indication of the association is stored in the dynamic storage structure (e.g., SRAM) of the memory system). Thus, before a first error is detected in the first memory strip 302 or the second memory strip 304 (e.g., before a first chip kill event occurs), each memory strip 302, 304 can be associated with a separate encoder / decoder component and / or a separate RS codeword. For example, the first memory strip 302 can be associated with the first RS encoder / decoder component 330, and / or the second memory strip 304 can be associated with the second RS encoder / decoder component 332. In this way, the first RS encoder / decoder component 330 can identify and / or correct errors associated with the first memory strip 302, for example, by using the parity information stored in the first RS codeword 303, and / or the second RS encoder / decoder component 332 can identify and / or correct errors associated with the second memory strip 304, for example, by using the parity information stored in the second RS codeword 305.
[0058] After a first error correction procedure (e.g., a correction associated with a first failed die in either the first memory strip 302 or the second memory strip 304), the memory system can use an extended RS payload that includes data from both the first memory strip 302 and the second memory strip 304 and / or can use an extended RS codeword (e.g., the third RS codeword 310) for the purpose of correcting subsequent errors in the first memory strip 302 and / or the second memory strip 304, as described above in connection with Figures 3A to 3C Description. In this regard, after the first error correction (described above in connection with reference numeral 306) and / or after the first data replacement (described above in connection with reference numeral 308), the paired memory strips 302, 304 can be associated with a third RS encoder / decoder component 334, which can be an encoder / decoder component capable of encoding and / or decoding the extended RS codeword (e.g., the third RS codeword 310) and / or handling the extended RS payload. For example, the third RS encoder / decoder component 334 can identify and correct a second failed die, a third failed die, or the like, as described above in connection with Figures 3B to 3C Description.
[0059] In this regard, after the first chip kill event, accessing the affected stripes (e.g., the first memory stripe 302 and the second memory stripe 304 in instance 300) can be affected by performance degradation because the memory system (e.g., the third RS encoder / decoder component 334 of the memory system) needs to access two 40-bit channels at once, as Figure 3D shown. In this way, the reliability improvement benefits of the embodiments described herein can come at the cost of performance degradation. Additionally, although the first RS encoder / decoder component 330, the second RS encoder / decoder component, and the third RS encoder / decoder component 334 are shown as separate components for ease of description, in some other embodiments, the RS encoder / decoder components can share certain components and / or logic. For example, in some embodiments, the third RS encoder / decoder component 334 (e.g., an encoder / decoder component capable of encoding and / or decoding extended RS codewords) can share logic with the first RS encoder / decoder component 330 and / or the second RS encoder / decoder component 332.
[0060] As indicated above, Figures 3A to 3D is for illustration only. Other examples can be different from what is described with respect to Figures 3A to 3D the description.
[0061] Figure 4 is a flowchart of an example method 400 associated with performing dual device data correction in a memory device using extended RS codewords. In some embodiments, a memory system (e.g., memory system 110) can execute or can be configured to execute method 400. In some embodiments, another device or a group of devices separate from or including the memory system 110 (e.g., memory device 120) can execute or can be configured to execute method 400. Additionally or alternatively, one or more components of the memory system 110 and / or the memory device 120 (e.g., memory system controller 115, local controller 125, first RS encoder / decoder component 330, second RS encoder / decoder component 332, and / or third RS encoder / decoder component 334) can execute or can be configured to execute method 400. Thus, the means for performing method 400 can include the memory system and / or the memory device and / or one or more components of the memory system and / or the memory device. Additionally or alternatively, a non-transitory computer-readable medium can store one or more instructions that, when executed by the memory system and / or the memory device (e.g., the memory system controller 115 of the memory system 110 and / or the local controller 125 of the memory device 120), cause the memory system and / or the memory device to execute method 400.
[0062] As Figure 4As shown in, method 400 may include associating a first memory stripe (e.g., first memory stripe 302) with a second memory stripe (e.g., second memory stripe), where the first memory stripe is associated with a first set of data storage elements (e.g., data dies (dies 0 to 7)) and a first set of error correction elements (e.g., parity dies (dies 8 to 9)), and where the second memory stripe is associated with a second set of data storage elements (e.g., data dies (dies 0 to 7)) and a second set of error correction elements (e.g., parity dies (dies 8 to 9)) (block 410). As Figure 4 Further shown in, method 400 may include receiving a first codeword (e.g., first RS codeword 303) associated with the first memory stripe, where the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity check information stored at the first set of error correction elements (block 420). As Figure 4 Further shown in, method 400 may include using the first codeword to identify a first error in the first set of data bits (e.g., the first faulty die described above in connection with reference numeral 306) (block 430). As Figure 4 Further shown in, method 400 may include using the first codeword to correct the first error (block 440). As Figure 4 Further shown in, method 400 may include receiving a second codeword (e.g., third RS codeword 310) associated with the first memory stripe and the second memory stripe, where the second codeword includes a second set of data bits associated with data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one of the error correction elements in the first set of error correction elements (e.g., replacement data stored at die 8 of the first memory stripe 302, as described above in connection with reference numeral 308), and where the second codeword includes a second set of error correction bits associated with parity check information stored at the second set of error correction elements (e.g., novel parity check information stored at dies 8 to 9 of the second memory stripe 304, as described above in connection with reference numeral 312) (block 450). As Figure 4 Further shown in, method 400 may include identifying a second error in the second set of data bits (e.g., the second faulty die, as described above in connection with reference numerals 314 and 322) (block 460). As Figure 4 Further shown in, method 400 may include using the second codeword to correct the second error (block 470).
[0063] Method 400 may include additional aspects, such as any single aspect or any combination of aspects described below and / or described in connection with one or more other methods or operations described elsewhere herein.
[0064] In a first aspect, a first codeword is a first RS codeword associated with a first RS payload, where a second codeword is associated with a second RS codeword associated with a second RS payload, and where the first RS payload is greater than the second RS payload (e.g., the second RS payload is the extended RS payload described above in connection with Figures 3A to 3D ).
[0065] In a second aspect, which may be separate or in combination with the first aspect, method 400 includes replacing, by a memory device, parity information at a first error correction element stored in a first set of error correction elements with a first set of data associated with a first error.
[0066] In a third aspect, which may be separate or in combination with one or more of the first and second aspects, method 400 includes replacing, by a memory device, parity information at a second error correction element (e.g., die 9 of the first memory strip 302, as described above in connection with reference numerals 316 and 324) stored in the first set of error correction elements with a second set of data associated with a second error.
[0067] In a fourth aspect, which may be separate or in combination with one or more of the first to third aspects, method 400 includes determining, by a memory device, parity information stored at a second set of error correction elements based on replacing parity information stored at the first error correction element with the first set of data.
[0068] In a fifth aspect, which may be separate or in combination with one or more of the first to fourth aspects, method 400 includes: receiving, by a memory device, a third codeword (e.g., the fourth RS codeword described above in connection with Figure 3B and 3C ) associated with a first memory strip and a second memory strip, where the third codeword includes a third set of data bits associated with data stored at a first set of data storage elements (e.g., the remaining operational data dies of the first memory strip 302), data stored at a second set of data storage elements (e.g., the remaining operational data dies of the second memory strip 304), and data stored at the first set of error correction elements (e.g., dies 8 to 9 of the first memory strip, after the second data replacement described above in connection with reference numerals 316 and 324), and where the third codeword includes a third set of error correction bits associated with parity information stored at a second set of error correction elements (e.g., dies 8 to 9 of the second memory strip 304); identifying, by the memory device, a third error in the third set of data bits; and correcting, by the memory device, the third error using the third codeword.
[0069] In a sixth aspect, either alone or in combination with one or more of the first through fifth aspects, method 400 includes storing an indication of the association between a first memory stripe and a second memory stripe through a memory device and in a dynamic storage component such as an SRAM.
[0070] Although Figure 4 illustrative blocks of method 400 are shown, in some embodiments, method 400 may include additional blocks, fewer blocks, different blocks, or a different arrangement of blocks compared to the blocks depicted in Figure 4 In addition or alternatively, two or more of the blocks of method 400 may be performed in parallel. Method 400 is an example of one method that may be performed by one or more of the devices described herein. The one or more devices may perform or may be configured to perform one or more other methods based on the operations described herein.
[0071] In some embodiments, a memory device includes one or more components configured to: associate a first memory stripe with a second memory stripe, where the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and where the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receive a first codeword associated with the first memory stripe, where the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity check information stored at the first set of error correction elements; identify a first error in the first set of data bits using the first codeword; correct the first error using the first codeword; receive a second codeword associated with the first memory stripe and the second memory stripe, where the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one of the error correction elements of the first set of error correction elements, and where the second codeword includes a second set of error correction bits associated with parity check information stored at the second set of error correction elements; identify a second error in the second set of data bits; and correct the second error using the second codeword.
[0072] In some embodiments, a method includes: associating, by a memory device, a first memory stripe with a second memory stripe, wherein the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and wherein the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receiving, by the memory device, a first codeword associated with the first memory stripe, wherein the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity information stored at the first set of error correction elements; identifying, by the memory device, a first error in the first set of data bits using the first codeword; correcting, by the memory device, the first error using the first codeword; receiving, by the memory device, a second codeword associated with the first memory stripe and the second memory stripe, wherein the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one of the error correction elements in the first set of error correction elements, and wherein the second codeword includes a second set of error correction bits associated with parity information stored at the second set of error correction elements; identifying, by the memory device, a second error in the second set of data bits; and correcting, by the memory device, the second error using the second codeword.
[0073] In some embodiments, a memory system includes: a memory controller; and a plurality of encoder / decoder components associated with the memory controller, wherein the memory system is configured to: associate a first memory stripe with a second memory stripe by the memory controller, wherein the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and wherein the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receive, by a first encoder / decoder component of the plurality of encoder / decoder components, a first codeword associated with the first memory stripe, wherein the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity information stored at the first set of error correction elements; identify, by the first encoder / decoder component, a first error in the first set of data bits using the first codeword; correct, by the first encoder / decoder component, the first error using the first codeword; receive, by a second encoder / decoder component, a second codeword associated with the first memory stripe and the second memory stripe, wherein the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one error correction element of the first set of error correction elements, and wherein the second codeword includes a second set of error correction bits associated with parity information stored at the second set of error correction elements; identify, by the second encoder / decoder component, a second error in the second set of data bits; and correct, by the second encoder / decoder component, the second error using the second codeword.
[0074] The foregoing disclosure provides illustration and description, but is not intended to be exhaustive or to limit the embodiments to the precise forms disclosed. Modifications and variations may be made in light of the above disclosure or may be acquired from practice of the embodiments described herein.
[0075] Even though particular combinations of features are recited in the claims and / or disclosed in the specification, these combinations are not intended to limit the disclosure of the embodiments described herein. Many of these features may be combined in ways not explicitly recited in the claims and / or not explicitly disclosed in the specification. For example, the present disclosure includes each dependent claim in a group of claims in combination with every other individual claim in the group of claims and every combination of multiple claims in the group of claims. As used herein, a phrase referring to “at least one” of a list of items refers to any combination of the items, including a single member. As an example, “at least one of a, b, or c” is intended to cover a, b, c, a + b, a + c, b + c, and a + b + c, as well as any combination with multiple identical elements (e.g., a + a, a + a + a, a + a + b, a + a + c, a + b + b, a + c + c, b + b, b + b + b, b + b + c, c + c, and c + c + c or any other ordering of a, b, and c).
[0076] When a “component” or “one or more components” (or another element, such as a “controller” or “one or more controllers”) is described or claimed (within a single claim or across multiple claims) as performing multiple operations or being configured to perform multiple operations, this language is intended to broadly cover a variety of architectures and environments. For example, unless otherwise explicitly claimed (e.g., by using “a first component” and “a second component” or other language that differentiates components in the claims), this language is intended to cover a single component performing or being configured to perform all operations, a group of components jointly performing or being configured to perform all operations, a first component performing or being configured to perform a first operation and a second component performing or being configured to perform a second operation, or any combination of components performing or being configured to perform the operations. For example, when a claim has the form “one or more components are configured to: perform X; perform Y; and perform Z,” the claim should be interpreted to mean “one or more components are configured to perform X; one or more (possibly different) components are configured to perform Y; and one or more (also possibly different) components are configured to perform Z.”
[0077] None of the components, acts, or instructions used herein should be construed as critical or essential unless expressly described as such. Also, as used herein, the article "a" is intended to include one or more items and may be used interchangeably with "one or more." Further, as used herein, the article "the" is intended to include one or more items referred to in conjunction with the article "the" and may be used interchangeably with "the one or more." When only one item is desired, the phrases "only one," "single," or similar language are used. Also, as used herein, the term "having" or the like is intended to be an open-ended term with respect to the component it modifies (e.g., a component "having" A may also have B). Further, unless expressly stated otherwise, the phrase "based on" is intended to mean "at least partially based on." As used herein, the term "multiple" may be replaced with "a plurality of," and vice versa. Also, as used herein, unless expressly stated otherwise (e.g., if used in combination with "either... of" or "only one of..."), the term "or" when used in a series is intended to be inclusive and may be used interchangeably with "and / or."
Claims
1. A memory device, comprising: One or more components configured to: Associate a first memory stripe with a second memory stripe, Wherein the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and Wherein the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; Receive a first codeword associated with the first memory stripe, Wherein the first codeword includes a first set of data bits associated with data stored in the first set of data storage elements and a first set of error correction bits associated with parity information stored in the first set of error correction elements; Identify a first error in the first set of data bits using the first codeword; Correct the first error using the first codeword; Receive a second codeword associated with the first memory stripe and the second memory stripe, Wherein the second codeword includes a second set of data bits associated with the data stored in the first set of data storage elements, the data stored in the second set of data storage elements, and data stored in at least one error correction element of the first set of error correction elements, and Wherein the second codeword includes a second set of error correction bits associated with parity information stored in the second set of error correction elements; Identify a second error in the second set of data bits; and Correct the second error using the second codeword.
2. The memory device according to claim 1, wherein the first codeword is a first RS codeword associated with a first Reed - Solomon RS payload, Wherein the second codeword is associated with a second RS codeword associated with a second RS payload, and Wherein the first RS payload is greater than the second RS payload.
3. The memory device according to claim 1, wherein the one or more components are further configured to replace the parity information stored at a first error correction element of the first set of error correction elements with a first set of data associated with the first error.
4. The memory device according to claim 3, wherein the one or more components are further configured to replace the parity information stored at a second error correction element of the first set of error correction elements with a second set of data associated with the second error.
5. The memory device according to claim 3, wherein the one or more components are further configured to determine the parity information stored at the second set of error correction elements based on replacing the parity information stored at the first error correction element with the first set of data.
6. The memory device according to claim 1, wherein the one or more components are further configured to: Receive a third codeword associated with the first memory stripe and the second memory stripe, Wherein the third codeword includes a third set of data bits associated with the data stored in the first set of data storage elements, the data stored in the second set of data storage elements, and data stored in the first set of error correction elements, and wherein the third codeword includes a third set of error correction bits associated with parity information stored at the second set of error correction elements; identifying a third error in the third set of data bits; and correcting the third error using the third codeword.
7. The memory device according to claim 1, wherein the one or more components are further configured to store an indication of the association between the first memory stripe and the second memory stripe in a dynamic storage component.
8. A method, comprising: associating, by a memory device, a first memory stripe with a second memory stripe, wherein the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and wherein the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receiving, by the memory device, a first codeword associated with the first memory stripe, wherein the first codeword includes a first set of data bits associated with data stored at the first set of data storage elements and a first set of error correction bits associated with parity information stored at the first set of error correction elements; identifying, by the memory device, a first error in the first set of data bits using the first codeword; correcting, by the memory device, the first error using the first codeword; receiving, by the memory device, a second codeword associated with the first memory stripe and the second memory stripe, wherein the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, data stored at the second set of data storage elements, and data stored at at least one of the error correction elements in the first set of error correction elements, and wherein the second codeword includes a second set of error correction bits associated with parity information stored at the second set of error correction elements; identifying, by the memory device, a second error in the second set of data bits; and correcting, by the memory device, the second error using the second codeword.
9. The method according to claim 8, wherein the first codeword is a first RS codeword associated with a first Reed-Solomon RS payload, wherein the second codeword is associated with a second RS codeword associated with a second RS payload, and wherein the first RS payload is greater than the second RS payload.
10. The method according to claim 8, further comprising replacing, by the memory device, parity information stored at a first error correction element in the first set of error correction elements with a first set of data associated with the first error.
11. The method according to claim 10, further comprising replacing, by the memory device, parity information stored at a second error correction element in the first set of error correction elements with a second set of data associated with the second error.
12. The method according to claim 10, further comprising determining, by the memory device, parity information stored at the second set of error correction elements based on replacing the parity information stored at the first error correction element with the first set of data.
13. The method according to claim 8, further comprising: receiving, by the memory device, a third codeword associated with the first memory stripe and the second memory stripe, wherein the third codeword includes a third set of data bits associated with the data stored at the first set of data storage elements, the data stored at the second set of data storage elements, and the data stored at the first set of error correction elements, and wherein the third codeword includes a third set of error correction bits associated with parity information stored at the second set of error correction elements; identifying, by the memory device, a third error in the third set of data bits; and correcting, by the memory device, the third error using the third codeword.
14. The method according to claim 8, further comprising storing, by the memory device and in a dynamic storage component, an indication of the association between the first memory stripe and the second memory stripe.
15. A memory system, comprising: a memory controller; and a plurality of encoder / decoder components associated with the memory controller, wherein the memory system is configured to: associate, by the memory controller, a first memory stripe with a second memory stripe, wherein the first memory stripe is associated with a first set of data storage elements and a first set of error correction elements, and wherein the second memory stripe is associated with a second set of data storage elements and a second set of error correction elements; receive, by a first encoder / decoder component of the plurality of encoder / decoder components, a first codeword associated with the first memory stripe, wherein the first codeword includes a first set of data bits associated with the data stored at the first set of data storage elements and a first set of error correction bits associated with parity information stored at the first set of error correction elements; identify, by the first encoder / decoder component, a first error in the first set of data bits using the first codeword; correct, by the first encoder / decoder component, the first error using the first codeword; receive, by a second encoder / decoder component, a second codeword associated with the first memory stripe and the second memory stripe, wherein the second codeword includes a second set of data bits associated with the data stored at the first set of data storage elements, the data stored at the second set of data storage elements, and the data stored at at least one error correction element of the first set of error correction elements, and wherein the second codeword includes a second set of error correction bits associated with parity information stored at the second set of error correction elements; identify, by the second encoder / decoder component, a second error in the second set of data bits; and The second error is corrected by the second encoder / decoder component using the second codeword.
16. The memory system of claim 15, wherein the first codeword is a first Reed-Solomon (RS) codeword associated with a first RS payload, wherein the second codeword is associated with a second RS codeword associated with a second RS payload, and wherein the first RS payload is greater than the second RS payload.
17. The memory system of claim 15, wherein the memory system is further configured to replace parity information at a first error correction element stored in the first set of error correction elements with a first set of data associated with the first error by the first encoder / decoder component.
18. The memory system of claim 17, wherein the memory system is further configured to replace parity information at a second error correction element stored in the first set of error correction elements with a second set of data associated with the second error by the second encoder / decoder component.
19. The memory system of claim 17, wherein the memory system is further configured to determine the parity information stored at the second set of error correction elements by the second encoder / decoder component based on replacing the parity information stored at the first error correction element with the first set of data.
20. The memory system of claim 15, wherein the memory system is further configured to: receive, by the second encoder / decoder component, a third codeword associated with the first memory stripe and the second memory stripe, wherein the third codeword includes a third set of data bits associated with the data stored at the first set of data storage elements, the data stored at the second set of data storage elements, and the data stored at the first set of error correction elements, and wherein the third codeword includes a third set of error correction bits associated with the parity information stored at the second set of error correction elements; identify, by the second encoder / decoder component, a third error in the third set of data bits; and correct, by the second encoder / decoder component, the third error using the third codeword.