Data protection and recovery
By introducing additional error correction and detection capabilities into the RAID solution, the limitations of the RAID solution in data recovery and protection are resolved, efficient data protection and recovery of the storage system is achieved, and single point failures are avoided.
Patent Information
- Application Number
- CN202480011782.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-16
- Filing Date
- 2024-02-26
- Publication Date
- 2025-09-16
AI Technical Summary
Existing RAID solutions have limitations in error correction and data recovery, and are unable to effectively protect and recover data in damaged chips, causing the memory system to become a single point of failure.
Combined with RAID solutions, it provides additional error correction and detection capabilities, ensuring data integrity by performing error correction and detection after RAID operations, and enabling data recovery through parity data.
The data protection capability of the RAID solution is improved, preventing the storage system from becoming a single point of failure due to chip damage, and ensuring data reliability and availability.
Smart Images

Figure CN120660065A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates generally to semiconductor memories and methods, and more particularly to apparatus, systems, and methods related to providing data protection and recovery schemes. Background Art
[0002] Memory devices are typically provided as internal semiconductor integrated circuits in computers or other electronic systems. There are many different types of memory, including volatile and non-volatile memory. Volatile memory may require power to maintain its data (e.g., host data, error data, etc.), and includes random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), synchronous dynamic random access memory (SDRAM), and thyristor random access memory (TRAM), among others. Non-volatile memory can provide persistent data by retaining stored data when not powered, and may include NAND flash memory, NOR flash memory, ferroelectric random access memory (FeRAM), and resistance variable memory, such as phase change random access memory (PCRAM), resistive random access memory (RRAM), and magnetoresistive random access memory (MRAM), such as spin torque transfer random access memory (STT RAM), among others.
[0003] A memory device may be coupled to a host (e.g., a host computing device) to store data, commands, and / or instructions for use by the host during operation of a computer or electronic system. For example, during operation of a computing or other electronic system, data, commands, and / or instructions may be transferred between the host and the memory device(s). A controller may be used to manage the transfer of data, commands, and / or instructions between the host and the memory devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figure 1 is a functional block diagram of a computing system including a memory controller according to several embodiments of the present disclosure.
[0005] Figure 2A is a functional block diagram of a memory controller with a redundant array of independent disks (RAID) encoder / decoder, alternatively referred to as a redundant array of independent devices (RAID) encoder / decoder, according to several embodiments of the present disclosure.
[0006] Figure 2B is another functional block diagram of a memory controller with a RAID encoder / decoder according to several embodiments of the present disclosure.
[0007] Figure 3is a block diagram schematically illustrating data subsets corresponding to RAID channels and error correction operations performed on the data subsets according to several embodiments of the present disclosure.
[0008] Figures 4A to 4B is a flow chart illustrating a locked RAID process for data protection of a subset of one or more user data blocks (UDBs) corresponding to a cache line, according to several embodiments of the present disclosure. DETAILED DESCRIPTION
[0009] Systems, devices, and methods related to providing protection and recovery schemes are described. Data protection and recovery schemes are often important aspects of the RAS (reliability, availability, and serviceability) associated with memory systems. Such schemes can provide "chip kill," wherein the memory system can function correctly even if a constituent chip (e.g., a memory die) is damaged; thereby avoiding a situation where one of the chips becomes a single point of failure (SPOF) for the memory system. Typically, chip kill capability is provided through various error correction code (ECC) schemes, such as "Redundant Array of Independent Disks" or "Redundant Array of Independent Devices" (RAID) schemes, which allow data recovery from a damaged chip by reading a subset of data from the other constituent chips of the RAID scheme.
[0010] RAID operations performed using any subset of a RAID scheme that contains errors may further result in errors in the recovered subset of damaged chips. Therefore, if errors already exist on data read from multiple component chips of a RAID scheme, the RAID scheme alone may not be able to protect data across all component chips.
[0011] Embodiments are directed to providing additional error correction and / or detection capabilities in conjunction with RAID scheme operations. For example, error correction capabilities may be provided after a RAID operation to correct any residual errors (e.g., bit errors) that existed on the subset used for the RAID operation, as well as those errors that propagated to the recovered subset. Furthermore, error detection capabilities may be provided after the error correction capabilities to ensure that errors are corrected by the error correction capabilities. Thus, the additional error correction and / or detection capabilities provided by the present disclosure may supplement the limitations of RAID schemes.
[0012] As used herein, the singular forms "a," "an," and "the" include both singular and plural referents, unless the context clearly dictates otherwise. Furthermore, the word "may" is used throughout this application in a permissive sense (i.e., having the potential to, being able to), rather than in a mandatory sense (i.e., having to). The term "include" and its derivatives mean "including, but not limited to." The term "coupled" means connected directly or indirectly. It should be understood that data can be transmitted, read, transferred, received, or exchanged via electronic signals (e.g., current, voltage, etc.).
[0013] The figures herein follow a numbering convention in which the first digit or digits correspond to the figure number of the accompanying drawing and the remaining digits identify an element or component in the accompanying drawing. Similar elements or components between different figures may be identified by using similar numerals. For example, 110 may refer to Figure 1 2 and similar elements may be referenced as 210 in FIG. 2. Similar elements within the figures may be referenced with a hyphen and an additional number or letter. For example, see Figure 1 1, 102-2, 102-M in FIG. 102-1, 102-2, 102-M. Such similar elements may generally be referenced without hyphens and additional numbers or letters. For example, elements 102-1, 102-2, 102-M may be collectively referred to as element 102. As used herein, particularly with respect to the designators "M" and "N" in the accompanying drawings, an indication may include several specific features so designated. As will be appreciated, elements shown in various embodiments herein may be added, exchanged, and / or eliminated to provide several additional embodiments of the present disclosure. In addition, as will be appreciated, the proportions and relative scales of the elements provided in the figures are intended to illustrate certain embodiments of the present disclosure and should not be interpreted in a limiting sense.
[0014] Figure 1 1 is a functional block diagram of a computing system 101 (alternatively referred to as a "memory system") including a memory controller 100 according to several embodiments of the present disclosure. The memory controller 100 may include a front-end portion 104, a central controller portion 110, and a back-end portion 119. The computing system 101 may include a host 103 coupled to the memory controller 100 and memory devices 126-1, ..., 126-N.
[0015] Front-end portion 104 includes interfaces and interface management circuitry for coupling memory controller 100 to host 103 via input / output (I / O) pathways 102-1, 102-2, ..., 102-M, as well as circuitry for managing I / O pathways 102. There may be any number of I / O pathways 102, such as eight, sixteen, or another number of I / O pathways 102. In some embodiments, I / O pathways 102 may be configured as a single port.
[0016] In some embodiments, the memory controller 100 may be a Compute Express Link (CXL)-compatible memory controller. The host interface (e.g., front-end portion 104) may be managed using the CXL protocol and coupled to the host 103 via an interface configured for the Peripheral Component Interconnect Express (PCIe) protocol. CXL is a high-speed central processing unit (CPU)-to-device and CPU-to-memory interconnect designed to accelerate next-generation data center performance. CXL technology maintains memory coherency between the CPU memory space and the memory on the attached device, which allows resource sharing for higher performance, reduced software stack complexity, and lower overall system cost. CXL is designed as an industry-standard interface for high-speed communication, as accelerators are increasingly used to supplement CPUs to support emerging applications such as artificial intelligence and machine learning. CXL technology is built on top of the PCIe infrastructure, leveraging the PCIe physical and electrical interfaces to provide advanced protocols in areas such as input / output (I / O) protocols, memory protocols (e.g., initially allowing hosts and accelerators to share memory), and coherent interfaces. As an example, the interface of the front end 104 may be a PCIe 5.0 or 6.0 interface coupled to the I / O lanes 102. In some embodiments, the memory controller 100 may receive access requests involving the memory device 126 via the PCIe 5.0 or 6.0 interface according to the CXL protocol.
[0017] The central controller portion 110 may include and / or be referred to as data management circuitry. The central controller portion 110 may control the execution of memory operations in response to requests received from the host 103. Examples of memory operations include a read operation to read data from the memory device 126 or a write operation to write data to the memory device 126.
[0018] The central controller portion 110 may generate error detection information and / or data recovery information based on data received from the host 103. The central controller portion 110 may perform error detection operations and / or data recovery operations on data received from the host 103 or from the memory device 126. An example of an error detection operation is a cyclic redundancy check (CRC) operation. A CRC may be referred to as an algebraic error detector. A CRC may include the use of a check value generated by performing an algebraic calculation using the data to be protected. The CRC may detect unexpected changes in the data by comparing a check value associated with the data storage with a check value calculated based on the data. An error correction operation (alternatively referred to as an error correction code (ECC) operation) may be performed to correct a certain number of bit errors and / or detect a certain number of bit errors that may not have been corrected using an ECC operation. The error correction information used to perform the ECC operation may be parity data (alternatively referred to as “ECC bits” or “ECC data”) generated by comparing (e.g., performing an XOR operation on) at least a portion of the rows (e.g., bit patterns) of an encoding matrix (alternatively referred to as a parity check matrix), the rows respectively corresponding to bits of user data (e.g., data received from the host 103) having specific values.
[0019] A data recovery operation (alternatively referred to as a "RAID operation") may be a chip-hunting operation, even if the constituent chips (e.g., memory dies, such as Figure 2A It can also protect the memory system from damage to the memory die 227 (and / or memory die 227 illustrated in FIG2B ), which can avoid a situation where one of the chips acts as a single point of failure (SPOF) for the memory system. Typically, chip hunting capability is provided by various ECC schemes, including "Redundant Array of Independent Disks" (RAID) schemes, which allow data recovery from a damaged chip by reading all the constituent chips of the memory system.
[0020] Chip hunting may involve parity data (e.g., RAID parity) that is specifically designed for data recovery of damaged chips. RAID parity data can be obtained by comparing each subset of user data (e.g., Figure 3 , 331-8, and 331-10) (e.g., performing an XOR operation on these subsets). User data that shares the same RAID parity data may be said to be grouped together. A RAID operation is an example of a "data recovery operation" and may also be referred to as such. When an error correction operation (e.g., an ECC operation) is performed to correct (e.g., flip) one or more bits of a subset indicated as having an error, the data recovery operation reconstructs and recovers the subset using the other subsets (without flipping one or more bits of the subset).
[0021] Back-end portion 119 may include a media controller and a physical (PHY) layer that couples memory controller 100 to memory devices 126. As used herein, the term "PHY layer" generally refers to the physical layer in the Open Systems Interconnection (OSI) model of computing systems. The PHY layer may be the first (e.g., lowest) layer of the OSI model and may be used to transmit data over a physical data transmission medium. In some embodiments, the physical data transmission medium may include channels 125-1, ..., 125-N. Channels 125 may include various types of data buses, such as an eight-pin data bus (e.g., a data input / output (DQ) bus) and a single-pin data mask inversion (DMI) bus, among other possible busses.
[0022] Memory device 126 can be a variety of memory devices. For example, the memory device can include an array of RAM, ROM, DRAM, SDRAM, PCRAM, RRAM, and flash memory cells, among others. In embodiments where memory device 126 includes persistent or non-volatile memory, memory device 126 can be a flash memory device, such as a NAND or NOR flash memory device. However, embodiments are not limited thereto, and memory device 126 can include an array of other non-volatile memory cells, such as non-volatile random access memory cells (e.g., non-volatile RAM (NVRAM), ReRAM, ferroelectric RAM (FeRAM), MRAM, PCRAM), "emerging" memory cells (e.g., ferroelectric RAM cells including ferroelectric capacitors that can exhibit hysteresis characteristics, memory devices having resistive, phase change, or similar memory cells, etc.), or combinations thereof.
[0023] As an example, an FeRAM device (e.g., memory device 126 includes an array of FeRAM cells) may include a ferroelectric capacitor and may perform bit storage based on the voltage or charge applied thereto. In such an example, relatively small and relatively large voltages allow the ferroelectric RAM device to exhibit properties similar to those of ordinary dielectric materials (e.g., dielectric materials with a relatively high dielectric constant), but at various voltages between such relatively small and relatively large voltages, the ferroelectric RAM device may exhibit polarization reversal that produces nonlinear dielectric behavior.
[0024] In another example, the memory device 126 may be a dynamic random access memory (DRAM) device (e.g., a memory device 126 including a DRAM cell array) operating according to a protocol such as low-power double data rate (LPDDRx), which may be referred to herein as an LPDDRx DRAM device, LPDDRx memory, etc. The "x" in LPDDRx refers to any one of the generations of the protocol (e.g., LPDDR5). In at least one embodiment, at least one of the memory devices 126-1 is operated as an LPDDRx DRAM device with low-power features enabled, and at least one of the memory devices 126-N is operated as an LPDDRx DRAM device with at least one low-power feature disabled. In some embodiments, although the memory device 126 is an LPDDRx memory device, the memory device 126 does not include circuitry configured to provide low-power functionality for the memory device 126, such as a dynamic voltage frequency scaling core (DVFSC), a subthreshold current reduction circuit (SCRC), or other low-power functionality-providing circuitry. Providing the LPDDRx memory device 126 without such circuitry can advantageously reduce the cost, size, and / or complexity of the LPDDRx memory device 126. For example, the LPDDRx memory device 126 with reduced low-power functionality-providing circuitry can be used in applications other than mobile applications (e.g., if the memory is not intended for use in mobile applications, some or all low-power functionality can be sacrificed to reduce the cost of producing the memory).
[0025] Data may be communicated between the backend portion 119 and the memory device 126 primarily in the form of memory transfer blocks (MTBs) comprising a number of user data blocks (UDBs). As used herein, the term "MTB" refers to a group of UDBs grouped with the same parity data block (PDB) (e.g., sharing the same PDB); thus, for each read or write command, they are transferred together from a cache (e.g., cache 212) and / or the memory device 126. For example, a group of UDBs of the same MTB may be transferred to / from the memory device 126 (e.g., written to / read from the memory device 126) via the channel 125 within a predefined burst length (e.g., a 16-bit or 32-bit BL) with which the memory controller 100 operates.
[0026] A burst is a series of data transfers over multiple cycles (e.g., ticks). As used herein, the term "tick" refers to a clock cycle increment during which an amount of data equal to the width of the memory bus can be transferred. For example, a 32-bit burst length may consist of 32 data transfer ticks, while a 16-bit burst length may consist of 16 data transfer ticks. Although embodiments are not limited thereto, the bus width corresponding to the size of each tick may be 8 (e.g., alternatively referred to as "x8").
[0027] As used herein, the term "PDB" refers to a block of data containing parity data (e.g., RAID parity) used for chip hunting (e.g., RAID) operations on UDBs grouped with the PDB. As further described herein, the MTB may be in plaintext or ciphertext form, depending on whether the MTB has been registered in the memory controller 100 (e.g., Figure 2A and 2B is encrypted at the security encoder 217-1 described in .
[0028] As used herein, the term "UDB" refers to a block of data containing host data (e.g., received from host 103 and alternatively referred to as "user data"). While a UDB may correspond to the size of a host read and / or write request, an MTB may be the unit of read and / or write access to a memory device. In addition to MTBs, PDBs may also be transferred between back-end portion 119 and memory device 126. The host data or parity data of a single UDB or PDB may correspond to multiple codewords (e.g., 64 codewords).
[0029] In addition to the UDBs, other "extra" data bits (e.g., data other than the data corresponding to the UDBs and alternatively referred to as "auxiliary data") may also be transferred between the backend portion 119 and the memory device 126. The extra data may include data used to correct and / or detect errors in the UDBs and / or authenticate and / or check the data integrity of the UDBs and / or metadata, but the embodiment is not limited thereto. Further details of the extra bits are illustrated and described in conjunction with Figures 2-3.
[0030] In some embodiments, some (e.g., one or more) memory devices 126 may be dedicated to PDBs or auxiliary data. For example, a memory device configured to store UDBs may be different from a memory device (e.g., one or more memory devices) configured to store PDBs or auxiliary data.
[0031] In some embodiments, the memory controller 100 may include a management unit 105 for initializing, configuring, and / or monitoring the characteristics of the memory controller 100. The management unit 105 may include an I / O bus for managing out-of-band data and / or commands, a management unit controller for executing instructions associated with initializing, configuring, and / or monitoring the characteristics of the memory controller, and a management unit memory for storing data associated with initializing, configuring, and / or monitoring the characteristics of the memory controller 100. As used herein, the term "out-of-band" generally refers to a transmission medium that is different from the primary transmission medium of a network. For example, out-of-band data and / or commands may be data and / or commands transmitted to a network using a transmission medium that is different from the transmission medium used to transmit data within the network.
[0032] Figure 2A is a functional block diagram of a memory controller 200 for cache line data protection according to several embodiments of the present disclosure. Figure 2A The memory controller 200, central controller portion 210, backend portion 219, and memory device 226 illustrated in FIG. Figure 1 Memory controller 100, central controller portion 210, back-end portion 119, and memory device 126 are illustrated in FIG.
[0033] The central controller portion 210 includes a front-end CRC ("FCRC") encoder 211-1 (e.g., paired with an FCRC decoder 211-2) for generating error detection information (e.g., alternatively referred to as an end-to-end CRC (e2e CRC)) based on data received as part of a write command (e.g., received from the host 103) (e.g., a UDB in "plaintext") before writing the data to the cache 212. The error detection information generated at the FCRC encoder 211-1 can be a check value, such as CRC data. Read and write commands of the CXL memory system can be the size of a UDB, e.g., 64 bytes. Therefore, the data received at the FCRC encoder 211-1 can correspond to a UDB.
[0034] Central controller portion 210 includes a cache 212 to store data (e.g., user data), error detection information, error correction information, and / or metadata associated with the performance of memory operations. An example of cache 212 is a 32-way set-associative cache that includes multiple cache lines. Although host read and write commands may be the size of a UDB (e.g., 64 bytes), the cache line size may be larger than the size of a UDB (e.g., equal to the size of a multiple of the UDBs). For example, the cache line size may correspond to the size of two UDBs (where each UDB is a 64-byte block), such as 128 bytes.
[0035] These UDBs stored in each cache line (e.g., alternatively referred to as "UDBs corresponding to a cache line") may be the unit of data transfer for the data path between the cache 212 and the memory device 226. For example, even if a host read / write command is the size of a UDB (e.g., 64 bytes), the UDBs corresponding to a cache line may be collectively transferred between the cache 212 and the memory device 226 as blocks (e.g., by Figure 2A ). Thus, the UDB corresponding to a cache line may be in Figure 2A 2 and are collectively encrypted / decrypted at the various encoders / decoders described in and located between cache 212 and memory device 226. Thus, the UDBs corresponding to cache lines may correspond to the same RAID stripe, as further described below. As used herein, the term "RAID stripe" refers to data including RAID parity data (e.g., Figure 3 ) and data used to generate RAID parity data (eg, subsets 331 - 1 , . . . , 331 - 8 and 331 - 10).
[0036] The data (e.g., UDB) stored in the cache 212 (e.g., its corresponding cache lines) may be further transferred to other components of the central controller portion 210 (e.g., security encoder 217-1 and / or authenticity / integrity check encoder 218-1, which is shown as "authentication encoder" 218-1) (e.g., as part of a cache write policy, such as cache writeback and / or cache writethrough) to ultimately be stored in the memory device 226 for access from a host (e.g., Figure 1 When the data received by the host 103 described in has not yet been written to the memory device 226, the cache 212 and the memory device 226 are synchronized.
[0037] Using cache 212 to store data associated with read or write operations can increase the speed and / or efficiency of accessing data because cache 212 can prefetch data in the event of a cache miss and store the data in multiple 64-byte blocks. Data can be read from cache 212 rather than searching a separate memory device in the event of a cache miss. Prefetched data can be accessed using less time and energy than would be used if the memory system had to search for the data before accessing it.
[0038] The central controller portion 210 further includes a security encoder 217-1 (e.g., paired with a security decoder 217-2) to encrypt data (e.g., a UDB corresponding to a cache line) before transmitting the data to the ECC encoder 216-1 (to be written to the memory device 226). Although not limited in this regard, the security encoder / decoder pair 217 may operate using an AES encryption / decryption algorithm. Unencrypted data (e.g., plaintext) may be converted to ciphertext via encryption by the security encoder 217-1. The central controller portion 210 further includes an authenticity / integrity check encoder 218-1 to generate authentication data based on data received from the cache 212. Although not limited in this regard, the authentication data generated by the authenticity / integrity check encoder 218-1 may be a MAC, such as a KECCAK MAC (KMAC) (e.g., a SHA-3-256 MAC).
[0039] In some embodiments, the TEE may be based on Trusted Execution Environment (TEE) data (alternatively referred to as a "TEE flag"), a Host Physical Address (HPA) (e.g., Figure 1 The MAC generated at the authenticity / integrity check encoder 218-1 is calculated based on the memory address associated with the host read / write transaction used / identified by the host 103 as described in ), the security key identifier (ID) associated with the physical address (of the memory device 226) to be accessed for executing the host write command.
[0040] The security encoder 217-1 and the authenticity / integrity check encoder 218-1 can operate in parallel. For example, data stored in the cache 212 and in plain text can be input (e.g., transmitted) to both the security encoder 217-1 and the authenticity / integrity check encoder 218-1. In some embodiments, a security key ID (along with the data in plain text) can be further input to the security encoder 217-1. In addition, in some embodiments, a security key ID, a TEE flag, and an HPA associated with a host write command (along with the data in plain text) can be further input to the authenticity / integrity check encoder 218-1.
[0041] The central controller portion 210 includes a CRC encoder 213-1 (e.g., paired with a CRC decoder 213-2) to collectively generate error detection information based on one or more UDBs corresponding to a cache line and transmitted from the security encoder 217-1. The CRC encoder 213-1 is alternatively referred to as a CRC medium (CRCm). The data transmitted and input to the CRC encoder 213-1 may be in ciphertext form because the data was previously encrypted at the security encoder 217-1. The error detection information generated by the error detection information generator 213-1 may be a check value, such as CRC data (alternatively referred to as "error detection data"). The CRC encoder 213-1 and the CRC decoder 213-2 can operate on data of a size equal to or greater than the cache line size.
[0042] like Figure 2A , the central controller portion 210 may include an ECC encoder 216-1. The ECC encoder 216-1 may be configured to collectively generate ECC data (alternatively referred to as "error correction data") based on one or more UDBs corresponding to a cache line and error detection information generated at the corresponding CRC encoder 213-1. The ECC data may later be used at the ECC decoder 216-2 to correct multiple bit errors on a subset (corresponding to the UDBs of the cache line) to be written to each memory device 226 and / or die, respectively. In some embodiments, the ECC data may include parity data.
[0043] The central controller portion 210 includes a RAID encoder 214-1 (e.g., paired with a RAID decoder 214-2) to generate and / or update RAID parity data (e.g., PDBs) based at least in part on data received from the CRC encoder 213-1 (e.g., one or more UDBs corresponding to a cache line, error correction information generated at a corresponding ECC encoder 216-1, and error detection information generated at the CRC encoder 213-1). The data transmitted from the CRC encoder 213-1 to the RAID encoder 214-1 may be in ciphertext form because the data is encrypted at the security encoder 217-1.
[0044] In some embodiments, the RAID encoder 214-1 may update the PDB to conform to a new UDB received from the host as part of a write command. To update the PDB, the old UDB (to be replaced by the new UDB) and the old PDB (having the same stripe as the old UDB) may be read (e.g., transferred to the RAID encoder 214-1) and compared (e.g., XORed) with the new UDB, and the result of the comparison (e.g., XORed) may be further compared (e.g., XORed) with the old PDB (to be updated) to generate a new (e.g., updated) PDB.
[0045] "Extra" data bits (alternatively referred to as "helper data") may be transferred (along with the UDBs) to the back end portion 219 to ultimately be transferred and written to the memory device 226. The "extra" bits may include RAID parity data (e.g., in the form of PDBs) generated at the RAID 214-1, CRC data generated at the FCRC encoder 211-1 and / or the CRC encoder 213-1, ECC data generated at the ECC encoder 216-1, and / or authentication data (e.g., MAC data) associated with the UDBs and metadata and / or TEE data generated at the authenticity / integrity check encoder 218-1.
[0046] In one embodiment, auxiliary data (including at least CRC data, ECC data, and metadata) may be written to a memory die 227 (alternatively referred to as a "memory unit") that is different from those memory dies 227 to which one or more UDBs and / or PDBs are to be written. In a different embodiment, the auxiliary data may be written to the same (e.g., one or more) memory dies 227 as the UDBs. Furthermore, the PDBs may be written to a memory device 226 or memory die 227 that is different / dedicated from those memory devices 226 and / or memory dies 227 to which the UDBs or auxiliary data are to be written.
[0047] Compared to other approaches, error correction information generated collectively on the data corresponding to a RAID stripe (e.g., to correct "mirrored" bit errors) can provide benefits such as a smaller number of bits used for error correction information. As used herein, the term "mirrored bit errors" refers to those bit errors on a recovered subset that were caused by bit errors already present in an input subset to the RAID process. As used herein, the term "input subset" refers to a subset that is input to a RAID process and used as part of a RAID operation to recover another subset.
[0048] For example, consider an example where each subset is provided with error correction capability that corrects a single bit error to cover a single "mirrored" bit error (caused by an input subset having a single bit error). In this example, if there are 10 subsets corresponding to the RAID stripe, then error correction information (e.g., having 8 bits) needs to be generated for each subset (each having 32 bytes), which can total 80 bits (e.g., 8 bits / subset * 10 subsets) and generate 10 different codewords, with the error correction information generated for each subset operating on the codeword. In contrast, the error correction information that can collectively correct two bit errors across the 10 subsets (of a single codeword) (thus being able to correct both the existing single bit error and the single "mirrored" bit error) can be only 22 bits, which is less than 80 bits.
[0049] like Figure 2A , the memory controller 200 may include a back-end portion 219 coupled to the central controller portion 210. The back-end portion 219 may include media controllers 221-1, ..., 221-N. The back-end portion 219 may include PHY memory interfaces 224-1, ..., 224-N. Each physical interface 224 is configured to couple to a respective memory device 226.
[0050] The media controllers 221-1, ..., 221-N can be used to drive the channels 225-1, ..., 225-N at substantially the same time. In at least one embodiment, each of the media controllers 221 can receive the same command and address and drive the channels 225 at substantially the same time. By using the same command and address, each of the media controllers 221 can utilize the channels 225 to perform the same memory operation on the same memory cell.
[0051] As used herein, the term "substantially" means that a characteristic need not be absolute, but close enough so that the advantage of the characteristic is achieved. For example, "substantially simultaneously" is not limited to operations performed absolutely simultaneously and can include timing that is intended to occur at the same time but may not be precisely simultaneous due to manufacturing limitations. For example, due to the read / write latency that may be exhibited by various interfaces (e.g., LPDDR5 and PCIe), media controllers utilized "substantially simultaneously" may not start or end at exactly the same time. For example, memory controllers may be utilized such that they write data to a memory device simultaneously, regardless of whether one of the media controllers starts or ends before the other.
[0052] The channels 225 may include several separate data protection channels (alternatively referred to as RAS (reliability, availability, and serviceability) channels), which may each include several memory devices (e.g., dies) 226 that are accessed together in association with a particular data protection scheme (e.g., RAID, LPCK, etc.). The data protection channels may include RAID (e.g., locked RAID) channels. In a "locked" RAID process, all subsets corresponding to a RAID stripe are accessed collectively together, regardless of whether the corresponding RAID process is triggered. For example, the subsets may be accessed collectively together even in response to only a host read request for accessing a portion (e.g., one) of the subsets, which makes the RAID process readily available without incurring additional / separate access to other subsets. As used herein, the term "RAID channel" refers to one or more channels (e.g., each in a RAID array) that are accessed together for RAID access. Figure 12 and 2). In other words, a RAID channel can be an access unit for transmitting a single RAID stripe. For example, channel 225 can be organized into several RAID channels, where each RAID channel includes a specific number of channels 225.
[0053] The PHY memory interface 224 may be an LPDDRx memory interface. In some embodiments, each of the PHY memory interfaces 224 may include data and DMI pins. For example, each PHY memory interface 224 may include twenty data pins (DQ pins) and five DMI pins. The media controller 221 may be configured to exchange data with the corresponding memory device 226 via the data pins. Instead of exchanging this information via the data pins, the media controller 221 may be configured to exchange error correction information (e.g., ECC data), error detection information, and / or metadata via the DMI pins. By setting a mode register, the DMI pins can provide multiple functions, such as data masking, data bus inversion, and parity checking for read operations. The DMI bus uses bidirectional signaling. In some examples, each transmitted data byte has a corresponding signal sent via the DMI pins for selecting the data. In at least one embodiment, data can be exchanged simultaneously with error correction information and / or error detection information. For example, UDBs can be exchanged (transmitted or received) via the data pins while additional bits are exchanged via the DMI pins. Such embodiments reduce overhead that would otherwise be used to communicate error correction information, error detection information, and / or metadata on a data input / output (eg, also referred to in the art as "DQ") bus.
[0054] Backend portion 219 can couple PHY memory interfaces 224-1, ..., 224-N to respective memory devices 226-1, ..., 226-N. Each memory device 226 includes at least one memory cell array. In some embodiments, memory devices 226 may be different types of memory. Media controller 221 may be configured to control at least two different types of memory. For example, memory device 226-1 may be an LPDDRx memory operating according to a first protocol, and memory device 226-N may be an LPDDRx memory operating according to a second protocol different from the first protocol. In this example, first media controller 221-1 may be configured to control a first subset of memory devices 226-1 according to the first protocol, and second media controller 221-N may be configured to control a second subset of memory devices 226-N according to the second protocol.
[0055] like Figure 2AAs illustrated in FIG, each memory device 226 may include one or more memory dies. For example, memory device 226-1 includes memory dies 227-1-1, ..., 227-1-X, while memory device 226-N includes memory dies 227-1-1, ..., 227-1-X. Although embodiments are not limited thereto, at least two memory devices 226 may include different numbers of memory dies.
[0056] Data stored in the memory device 226 (corresponding to a UDB of a cache line) may be transferred to the back end portion 219 to be ultimately transferred and written to the cache 212 and / or transferred to a host (eg, Figure 1 ). In some embodiments, data is transferred in response to a read command to access a subset of the data (e.g., one UDB) and / or to synchronize the cache 212 and the memory device 226 to clear "dirty" data in the cache 212.
[0057] In addition to the UDB, other "extra" data bits (alternatively referred to as "helper data") may also be transmitted to the backend portion 219. The "extra" bits may include CRC data generated by the FCRC encoders 211-1 and / or 213-1, ECC data generated by the ECC encoder 216-1, and authentication data associated with the UDB and metadata and / or TEE data generated by the authenticity / integrity check encoder 218-1. As described herein, the UDB transmitted to the backend portion 219 may be in ciphertext form.
[0058] The UDB corresponding to the cache line may be further transmitted to the CRC decoder 213-2 along with at least the error detection information previously generated at the CRC encoder 213-1 (e.g., from the back end portion 219). At the CRC decoder 213-2, an error detection operation may be performed to detect any errors in the UDB (e.g., and / or auxiliary data, such as ECC data, metadata, etc.) using the error detection information (e.g., CRC data). The CRC decoder 213-2 may operate on the data in conjunction with the RAID decoder 214-2 to provide checksum and recovery correction. More specifically, the CRC decoder 213-2 may detect errors in the data (e.g., received from the corresponding ECC decoder 216-2) and, in response, the RAID decoder 214-2 may recover the data.
[0059] When a RAID process is triggered, the RAID operations performed on the data (e.g., one or more UDBs corresponding to a cache line and auxiliary data including at least ECC data, metadata, etc.) can recover, for example, a subset of the data transferred from one (e.g., failed) memory die. Since all subsets are collectively input (e.g., transferred) to a CRC decoder (e.g., Figure 2A 2) and are collectively checked for one or more errors (alternatively referred to as "locked RAID"), the CRC check performed at the CRC decoder may not indicate which subset has the one or more errors. Thus, a triggered RAID process involves multiple RAID operations that may be performed separately and independently on each subset to correct the one subset that does have errors.
[0060] After performing the RAID operations, the results corresponding to the RAID operations (e.g., each including a UDB corresponding to a cache line and a corresponding recovered subset, alternatively referred to as a "resulting RAID stripe") can be input to the ECC decoder 216-2. For example, the ECC decoder 216-2 can perform multiple error correction operations on the results to correct any residual errors on the subset corresponding to the plurality of memory devices 226 and / or dies.
[0061] At the ECC decoder 216-2, error correction operations may be performed on the data to correct errors up to a certain number (number of errors) and / or to detect errors exceeding a certain number without correcting those errors. For example, the ECC decoder 216-2 may use the error correction information to correct multiple bit errors on different subsets. Although embodiments are not limited in this regard, the error correction information may provide error correction capabilities that correct multiples of 2 bit errors. For example, the error correction information may provide error correction capabilities that correct 2 bit errors on one or more subsets. x (x is a positive integer) bit errors. In a specific example, ECC decoder 216-2 can use error correction information (e.g., ECC data) to correct two bit errors without detecting more than two bit errors, which is called double error correction (DEC) operation. In this example, ECC decoder 216-2 can provide error correction capability to correct each single bit error in two subsets.
[0062] The bit errors that can be corrected at ECC decoder 216-2 may include a certain number of bit errors further caused by the corresponding RAID operation. For example, a RAID operation performed using one or more subsets with one or more bit errors may further cause (e.g., mirror) one or more corresponding bit errors in the recovered subset. In this example, the result input to ECC decoder 216-2 may include the bit errors present in the subset used for the RAID operation and the bit errors in the recovered subset.
[0063] The results from ECC decoder 216-2 (e.g., each including one or more UDBs corresponding to a cache line that may have corrected one or more bit errors) may be further input to CRC decoder 215, which provides the same functionality as CRC decoder 213-2, but performs error detection operations (e.g., CRC checks) on the results. RAID decoder 214-2 may determine one of the results that is not indicated as including a bit error and transmit that result to security decoder 217-2. If none of the results is indicated as including a bit error, central controller 210 may send a request to a host (e.g., Figure 1 The host 103) described in notifies the poisoned data.
[0064] As described above, the output from the RAID decoder 214-2 (e.g., a UDB corresponding to a cache line) may be further transmitted to the security decoder 217-2 and the authenticity / integrity check decoder 218-2 (e.g., the UDB corresponding to the cache line) along with at least the authentication data previously generated at the authenticity / integrity check encoder 218-1. Figure 2A 2). At security decoder 217-2, the data may be decrypted (e.g., converted from ciphertext back to plaintext as originally received from the host). Security decoder 217-2 may use AES decryption to decrypt the data.
[0065] At the authenticity / integrity check decoder 218-2, the authentication data (e.g., MAC data) previously generated at the authenticity / integrity check encoder 218-1 can be used to authenticate the data decrypted at the security decoder 217-2 (and / or check the data integrity). In some embodiments, the authenticity / integrity check decoder 218-2 can calculate a MAC based on the TEE data, the HPA, and the security key ID associated with the physical address to be accessed for executing the host read command. The MAC calculated during the read operation can be compared with the MAC transmitted from the memory device 226 (corresponding to the location of its physical address). If the calculated MAC matches the transmitted MAC, the UDB is written to the cache 212 (and further transmitted to the host, if necessary). If the calculated MAC does not match the transmitted MAC, the host is notified of the mismatch (and / or poisoning).
[0066] The data (eg, a UDB corresponding to a cache line) verified (and / or checked for data integrity) at the authenticity / integrity check decoder 218-2 may be transferred and written to the cache 212. In some embodiments, for example, in response to a request from a host (eg, Figure 1In response to a read command received from the host 103 (as described in
[15] ), data may be further transferred from cache 212 to FCRC decoder 211-2. As described herein, host read and write commands for the CXL memory system may be of the size of a UDB, such as 64 bytes. For example, data may be requested by the host at the granularity of a UDB. In this example, even if the data transferred from memory device 226 is a plurality of UDBs (corresponding to cache lines), the data may be transferred from cache 212 to the host at the granularity of a UDB. At FCRC decoder 211-2, the data (e.g., the UDB requested by the host) may be checked for any errors (CRC check) using the CRC data previously generated by FCRC encoder 211-1. The data decrypted by FCRC decoder 211-2 may be further transferred to the host.
[0067] Figure 2B is another functional block diagram of a memory controller 200 for cache line data protection according to several embodiments of the present disclosure. Figure 2B The memory controller 200, central controller portion 210, backend portion 219, and memory device 226 illustrated in FIG. Figure 1 1. The memory controller 100, the central controller portion 110, the back end portion 119, and the memory device 126 illustrated in FIG.
[0068] The memory controller 200 may include a central controller portion 210 and a back-end portion 219. The central controller portion 210 may include a front-end CRC ("FCRC") encoder 211-1-1 paired with an FCRC decoder 211-2 and an FCRC encoder 211-2-1 paired with an FCRC decoder 211-2-1, a cache memory 212 coupled between the paired CRC encoder / decoder 211-1 and the CRC encoder / decoder 211-2, a security encoder 217-1 paired with a security decoder 217-2, an authenticity / integrity check decoder 218-2 ( Figure 2B 2) paired with an authenticity / integrity check encoder 218-1 (shown as an "authentication decoder" 218-2). Figure 2B 2), a CRC encoder 213-1 paired with a CRC decoder 213-2, a RAID encoder 214-1 paired with a RAID decoder 214-2, and an ECC encoder 216-1 paired with an ECC decoder 216-2. A pair of security encoders / decoders 217, a pair of authenticity / integrity check encoders / decoders 218, a pair of CRC encoders / decoders 213, a pair of RAID 214, and a pair of ECC encoders / decoders 216 may be similar to Figure 2AA pair of security encoder / decoders 217, a pair of authenticity / integrity check encoder / decoders 218, a pair of CRC encoder / decoders 213, a pair of RAID encoder / decoders 214, and a pair of ECC encoder / decoders 216 are described in FIG. Figure 2B , 226-N and PHY memory interfaces 224-1, . . . , 224-N configured to couple to memory devices 226-1, . . . , 226-N via channels 225-1, . . . , 225-N.
[0069] Figure 2B Similar to Figure 2A , except that it includes additional circuitry to use the CRC data to check for any errors on the UDB without transmitting / storing the CRC to the memory device 226. For example, Figure 2B , the FCRC decoder 211-1-2 coupled between the cache 212 and the security encoder 217-1 (and / or the authenticity / integrity check encoder 218-1) may be configured to use the error detection information (e.g., CRC data) generated at the FCRC encoder 211-1-1 to check for any errors on the UDBs stored in the cache 212. The FCRC encoder 211-2-1 coupled between the cache 212 and the security decoder 217-2 (and / or the authenticity / integrity check decoder 218-2) may be configured to generate the error detection information (e.g., CRC data) generated at the FCRC encoder 211-1-1 before the UDBs are to be transmitted to the host (e.g., Figure 1 Error detection information (eg, CRC data) is generated on a UDB of the host 103 as described in FIG. The error detection information generated at the FCRC encoder 211 - 2 - 1 may be used at the FCRC decoder 211 - 2 - 2 to check for any errors on the UDB transferred from the cache 212 .
[0070] In some embodiments, only the CRC encoder / decoder pair 211-1 and 211-2 may be used to check for errors on data stored in the cache. Therefore, the error detection information (e.g., CRC data) used at the pair 211-1 and 211-2 may not be transferred and written to the memory device 226.
[0071] Figure 3 is a block diagram schematically illustrating data subsets corresponding to RAID channels and error correction operations performed on the data subsets according to several embodiments of the present disclosure. Figure 3Ten data subsets (alternatively referred to simply as "subsets") 331-1, ..., 331-10 corresponding to RAID (e.g., locked RAID) channels are illustrated such that they are transferred as a unit from a memory device (e.g., Figure 2A and 2B Memory device 226) and / or die read as described in . Figure 3 In the embodiment described in , the memory devices (eg, Figure 1 2 and 126 and / or 226). As used herein, the term "memory rank" generally refers to a plurality of memory chips (e.g., Figure 2A and / or memory die 227 illustrated in 2B).
[0072] Although the embodiment is not limited thereto, each subset of RAID channels may be a unit of a RAID protection scheme. For example, a RAID process of a RAID protection scheme may recover data corresponding to a single subset when triggered.
[0073] Furthermore, the ten data subsets may correspond to data transferred from the ten memory dies, respectively. Although embodiments are not limited thereto, subsets 331-1, ..., 331-8 may correspond to one or more UDBs and subset 331-9 may correspond to a PDB. Furthermore, subset 331-10 may correspond to a memory device containing error correction information (e.g., Figure 2A and / or the corresponding ECC encoder 216-1 described in 2B), error detection information (e.g., at Figure 2A and / or auxiliary data generated at the CRC encoder 213-1 described in 2B), metadata and / or TEE data.
[0074] Figure 3 The example where subset 331-7 contains one or more bit errors (eg, a single bit error) and a RAID operation is performed (eg, Figure 3 331 - 1 ) to recover the subset 331 - 1. Although not limited thereto, a RAID operation may be performed using one or more UDBs corresponding to the subsets 331 - 2, ..., 331 - 8, a PDB corresponding to the subset 331 - 9, and auxiliary data corresponding to the subset 331 - 10.
[0075] If subsets 331-2, ..., 331-8, and 331-10 contain a certain number of bit errors, the same number of bit errors may be propagated (e.g., mirrored) to the recovered subset. For example, if there is a single bit error in subsets 331-2, ..., 331-8, and 331-10, the recovered subset may also include the single bit error. Similarly, if there are two bit errors in subsets 331-2, ..., 331-8, and 331-10, the recovered subset may also include the two bit errors. Thus, in Figure 3 In the example illustrated in , one or more bit errors of subset 331 - 7 may be propagated to recovered subset 331 - 1 due to the RAID operation.
[0076] like Figure 3 As described in , error correction operations (e.g., Figure 3 331-9) can correct bit errors not only in subset 331-7, but also on subset 331-1. Consider an example where a single bit error already existed in subset 331-7, which would propagate the single bit error in recovered subset 331-1. In this example, there would be two bit errors on the resulting RAID stripe (e.g., subsets 331-1, ..., 331-9), and an error correction operation with the ability to correct at least two bit errors on subsets 331-1, ..., 331-9 can correct those two bit errors.
[0077] Figures 4A to 4B 4 is a flow chart illustrating a process for locking a RAID array corresponding to a subset of memory cells according to several embodiments of the present disclosure. Figure 4A As illustrated in FIG. 4 , at 452, a method for accessing data from one or more memory devices (eg, corresponding to Figure 2A and 2B A read command for data (eg, UDB) of one or more memory devices 226) described in FIG.
[0078] like Figure 4A As illustrated in FIG. 4 , at 454, the memory dies (eg, corresponding to Figure 1 、 2A and one or more memory devices 126 and / or 226 illustrated in FIG. 2B) read a subset corresponding to a RAID channel (e.g., Figure 3, 331-10 illustrated in FIG. ). A subset may include UDBs requested by the host as well as other data corresponding to the same RAID stripe (e.g., other UDBs, error correction information, error detection information, etc.). As described herein, one or more subsets including PDBs and auxiliary data may be read from a different memory die than those from which the subsets including UDBs are read.
[0079] like Figure 4A As illustrated in FIG, at 456, a RAID process is triggered, followed by error correction / detection operations. For example, at 456-1, a RAID operation is performed on a corresponding subset read from each memory die storing one or more UDBs and auxiliary data (e.g., error correction information, error detection information, etc.). For example, at 456-1-1, ..., 456-1-9, RAID operations may be performed to recover subsets read from nine memory dies (e.g., subsets 331-1, ..., 331-8, and 331-10), respectively.
[0080] like Figure 4A As illustrated in FIG. 456-2, respective error correction operations may be further performed on each RAID stripe after respective RAID operations 456-1-1, ..., 456-1-9 (e.g., at ECC decoder 216-2 and / or as shown in FIG. Figure 2A and 2B . , 456-2-9, error correction operations may be performed to perform error correction on the resulting RAID stripes passed from 456-1-1, . . . , 456-1-9, respectively. In some embodiments, each error correction operation performed at 456-2 may correct bit errors on multiple subsets (e.g., corresponding to multiple memory dies), such as two bit errors (e.g., on two different subsets, respectively), but embodiments are not limited thereto.
[0081] like Figure 4A As illustrated in FIG, at 456-3, each resulting (eg, error-corrected) RAID stripe may be further checked for remaining errors (eg, Figure 2A and 2B . . , 456-3-9, CRC checks may be performed on the resulting RAID stripes passed from 456-2-1, . . . , 456-2-9, respectively.
[0082] As used herein, the term "resulting RAID stripe" may refer to data corresponding to the results of each process described in conjunction with 456 of FIG. 4. For example, the resulting RAID stripe (e.g., the result) of each step 456-1 may be a portion of a RAID stripe that includes one or more subsets originally input to the RAID operation and a recovered subset. For example, the resulting RAID stripe (e.g., the result) of each step 456-2 may be a portion of a RAID stripe that includes one or more subsets originally input to the RAID operation and a recovered subset that has been error-corrected (e.g., using ECC data at ECC decoder 216-2).
[0083] In 456-3 and 458 (e.g. Figure 4B After performing the CRC check at (described in ), it is determined whether each of the resulting RAID stripes has passed the corresponding CRC check. If so, the flowchart 450 proceeds to 458-1 and 458-2, and any of the resulting RAID stripes can be provided as output to, for example, Figure 2A and 2B The security decoder 214-2 described in Figure 2A and 2B If not, the flowchart proceeds to 460-1 and 460-2 (as shown in FIG. Figure 4B ), and only those resulting RAID stripes that have passed the corresponding CRC check performed at 456-3 (e.g., one of the resulting RAID stripes passed from 456-3) may be provided as output to security decoder 214-2. If none of the resulting RAID stripes have passed the corresponding CRC check, then flowchart 450 proceeds to 462 (as Figure 4B ), and can be sent to a host (e.g., Figure 1 The host 103 described in ) notifies that data (e.g., a RAID stripe) is poisoned.
[0084] RAID operations (e.g., performed at 456-1) along with error correction / error detection operations (e.g., performed at 456-2 and 456-3, respectively) can be performed in various ways. In one example, each set of RAID operations 456-1-1, ..., 456-1-9, error correction operations 456-2-1, ..., 456-2-9, and / or error detection operations 456-3-1, ..., 456-3-9 can be performed in parallel.
[0085] Although specific embodiments have been illustrated and described herein, it will be understood by those skilled in the art that arrangements calculated to achieve the same results may replace the specific embodiments shown. The present disclosure is intended to cover adaptations or modifications of one or more embodiments of the present disclosure. It should be understood that the above description has been made in an illustrative manner, not a restrictive manner. Combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art upon reviewing the above description. The scope of one or more embodiments of the present disclosure includes other applications in which the above structures and processes are used. Therefore, the scope of one or more embodiments of the present disclosure should be determined with reference to the appended claims, together with the full scope of equivalents to which such claims are entitled.
[0086] In the foregoing detailed description, various features are grouped together in a single embodiment for the purpose of streamlining the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the disclosed embodiments of the disclosure necessarily utilize more features than are expressly recited in each claim. Rather, as reflected in the appended claims, the inventive subject matter lies in less than all features of a single disclosed embodiment. The appended claims are therefore hereby incorporated into the detailed description, with each claim standing on its own as a separate embodiment.
Claims
1. A method comprising: receiving a read command for accessing a first user data block (UDB) stored in one or more memory cells in a first memory cell group; In response to receiving the read command: performing a data recovery operation on one or more subsets of a plurality of subsets corresponding to the first UDB using a parity data block including data recovery information and read from one or more second memory cells, wherein each subset of the plurality of subsets corresponds to a respective memory cell in the first group of memory cells; To correct the remaining one or more bit errors in at least two of the plurality of subsets, an error correction operation is performed on the first UDB using auxiliary data including error correction information and read from one or more third memory cells.
2. The method according to claim 1, wherein: The auxiliary data further includes error detection information; and The method further includes, before performing the data recovery operation, performing an error detection operation on the first UDB using the error detection information to indicate whether the first UDB includes the one or more bit errors.
3. The method according to any one of claims 1 to 2, further comprising: performing a corresponding data recovery operation on each subset of the plurality of subsets; and A respective error correction operation is performed on each of the plurality of first results of the plurality of respective data recovery operations. The method of claim 3 , further comprising performing the plurality of corresponding data recovery operations in parallel.
5. The method according to claim 3, wherein: The auxiliary data further includes error detection information; and The method further includes performing a respective error detection operation on each of a plurality of second results of the plurality of respective error correction operations using the error detection information. 6 . The method of claim 5 , further comprising providing, in response to the read command, one of the plurality of second results indicated as not including the one or more bit errors. 7 . The method of claim 5 , further comprising providing, in response to the read command, a notification indicating that none of the plurality of second results is indicated as not including the one or more bit errors.
8. A device comprising: a plurality of memory cells; and a controller communicatively coupled to the plurality of memory units via a plurality of channels, the controller being configured to: receiving a read command to access a first user data block (UDB) stored in one or more memory cells in a first memory cell group of the plurality of memory cells; In response to receiving the read command: reading, from the first memory cell group, subsets of a plurality of subsets corresponding to a plurality of UDBs, respectively, wherein the plurality of UDBs includes the first UDB and corresponds to a redundant array of independent disks (RAID); reading one or more subsets of the plurality of subsets corresponding to auxiliary data including error correction information from one or more second memory cells; reading one or more subsets of the plurality of subsets corresponding to parity data blocks (PDBs) including data recovery information from one or more third memory cells; and performing a data recovery operation on at least one of a plurality of subsets of the first UDB to recover the at least one subset using the data recovery information of the PDB and in response to the plurality of subsets being indicated as containing one or more bit errors; and An error correction operation is performed on the plurality of subsets using the error correction information to correct one or more bit errors remaining in at least two of the plurality of subsets after the data recovery operation.
9. The apparatus of claim 8, wherein the controller is configured to perform a double error correction (DEC) operation on the plurality of subsets as the error correction operation to correct: a single bit error in a first subset of the plurality of subsets for use in the data recovery operation; and a single bit error in a second subset of the plurality of subsets that is mirrored by the single bit error in the first subset, wherein the second subset is the subset recovered via the data recovery operation.
10. The apparatus of claim 8, wherein the controller is configured to perform the data recovery operation on each subset of the plurality of subsets regardless of whether each subset includes one or more bit errors.
11. The apparatus according to any one of claims 8 to 10, wherein: The auxiliary data further includes error detection information; and The controller is configured to: performing an error detection operation on a codeword comprising a plurality of UDBs corresponding to a RAID stripe, wherein the plurality of UDBs includes the first UDB; and The data recovery operation is performed in response to the error detection operation indicating one or more bit errors in a codeword.
12. The apparatus according to any one of claims 8 to 10, wherein: The auxiliary data further includes error detection information; and The controller is further configured to perform an error detection operation on the plurality of UDBs after the data recovery operation and determine whether the UDBs still contain one or more bit errors.
13. The apparatus of any one of claims 8 to 10, wherein the data recovery information corresponds to Redundant Array of Independent Disks (RAID) parity data.
14. A device comprising: a plurality of memory cells configured to store one or more user data blocks (UDBs); and a controller communicatively coupled to the plurality of memory units via a plurality of channels, the controller being configured to: receiving data corresponding to a first UDB as part of a write command to write the first UDB to one or more memory cells of the plurality of memory cells; generating error correction information based on the first UDB, wherein the error correction information is used to correct one or more bit errors in the first UDB; writing the first UDB to a first memory cell group of the plurality of memory cells; and The error correction information is written to a second memory cell of the plurality of memory cells.
15. The apparatus of claim 14, wherein the controller is configured to: writing the first UDB to the first memory cell group via one or more data input / output (DQ) pins; and The error correction information is written to the second memory cell via one or more data mask inversion (DMI) pins.
16. The apparatus according to any one of claims 14 to 15, wherein: The first UDB corresponds to a specific redundant array of independent disks (RAID) stripe, the specific RAID stripe including the first UDB, the second UDB, and the error correction information; and The controller is further configured to generate RAID parity data based on the data corresponding to the particular RAID stripe.
17. The apparatus of claim 16, wherein the controller is configured to: The error correction information is generated based on the first UDB and the second UDB. receiving a read command from the first memory cell group to access the first UDB; performing a data recovery operation on each of a plurality of subsets corresponding to the particular RAID stripe using the RAID parity data, wherein the plurality of subsets correspond to data read from the first group of memory cells and the second group of memory cells, respectively; and An error correction operation is performed on each subset of the plurality of subsets using the error correction information to correct one or more bit errors on at least two subsets of the plurality of subsets.