Error code correction consistency verification for ternary cell-based memory devices

By using a programming manager and an extended ECC algorithm in memory cells, the challenges of threshold voltage programming and inaccurate error correction in memory cells are solved, enabling more efficient data reading and correction and improving the reliability of the memory system.

CN120937083APending Publication Date: 2025-11-11MICRON TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480024795.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-18
Filing Date
2024-04-03
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently program threshold voltages to intermediate states within memory cells, leading to errors when reading data. Furthermore, existing error correction codes (ECCs) may incorrectly correct errors in ternary cells, impacting data integrity.

Method used

The program manager controls the bit line and word line drivers, and programs the threshold voltage of the memory cell through precise voltage pulses. Combined with the extended ECC algorithm, it uses redundant information to detect and correct random errors when reading memory cells, and adjusts the error location to improve decoding accuracy.

Benefits of technology

It improves the accuracy of data reading from memory cells, ensures the accuracy of the error correction process, reduces the data read error rate, and enhances the reliability of the memory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937083A_ABST
    Figure CN120937083A_ABST
Patent Text Reader

Abstract

In some embodiments, the technology described herein relates to a method that includes receiving a codeword having a first portion and a second portion, the first portion including user data and the second portion including synthetic data; detecting at least one error in the codeword at a first location using an ECC engine; and signaling an erroneous detection when the first location is within the second portion.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related applications

[0002] This application claims priority to U.S. Patent Application Serial No. 18 / 608,758, filed March 18, 2024, which claims priority to Provisional U.S. Patent Application Serial No. 63 / 496,780, filed April 18, 2023, the entire disclosure of which is hereby incorporated herein by reference. Technical Field

[0003] At least some of the embodiments disclosed herein generally relate to memory systems, and more specifically, but not limited to, techniques for configuring memory cells to store data. Background Technology

[0004] The memory subsystem may include one or more memory devices for storing data. Memory devices may be, for example, non-volatile memory devices and volatile memory devices. Generally, a host system may utilize the memory subsystem to store data at memory devices and retrieve data from memory devices.

[0005] A memory device may comprise a memory integrated circuit having an array of one or more memory cells formed on an integrated circuit die of a semiconducting material. A memory cell is the smallest unit of memory that can be individually used or operated to store data. Generally, a memory cell can store one or more data bits.

[0006] Different types of memory cells have been developed for memory integrated circuits, such as random access memory (RAM), read-only memory (ROM), dynamic random access memory (DRAM), static random access memory (SRAM), synchronous dynamic random access memory (SDRAM), phase change memory (PCM), magnetic random access memory (MRAM), NOR flash memory, electrically erasable programmable read-only memory (EEPROM), flash memory, etc.

[0007] Some integrated circuit memory cells are volatile and require power to maintain the data stored in the cells. Examples of volatile memory include Dynamic Random Access Memory (DRAM) and Static Random Access Memory (SRAM).

[0008] Some integrated circuit memory cells are non-volatile and retain stored data even when no power is supplied. Examples of non-volatile memories include flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electronically erasable programmable read-only memory (EEPROM). Flash memory includes NAND flash memory and NOR flash memory. NAND memory cells are based on NAND logic gates; and NOR memory cells are based on NOR logic gates.

[0009] Crosspoint memory (e.g., 3D Xpoint memory) uses an array of non-volatile memory cells. The memory cells in crosspoint memory are transistorless. Each of these memory cells may have selector devices and discretionary phase-change memory devices stacked together as a column in an integrated circuit. Such columns of memory cells are connected to the integrated circuit via two conductive layers extending in mutually perpendicular directions. One of these layers is above the memory cell; and the other is below the memory cell. Therefore, each memory cell can be individually selected at the intersection of the two conductive lines extending in different directions in the two layers. Crosspoint memory devices are fast and non-volatile and can be used as a unified memory cluster for processing and storage.

[0010] Non-volatile integrated circuit memory cells can be programmed to store data by applying voltage or voltage patterns to the memory cell during programming / writing operations. Programming / writing operations set the memory cell to a state corresponding to the data programmed / stored into the memory cell. Data stored in the memory cell can be retrieved during a read operation by checking the state of the memory cell. A read operation determines the state of the memory cell by applying a voltage and determining whether the memory cell becomes conductive under a voltage corresponding to a predefined state. Attached Figure Description

[0011] In the accompanying drawings, embodiments are illustrated by way of example and without limitation, wherein the same element symbols indicate similar elements.

[0012] Figure 1 A memory device configured to have a programming manager is shown according to one embodiment.

[0013] Figure 2 A memory cell having a bit line driver and a word line driver configured to apply voltage pulses is shown according to one embodiment.

[0014] Figure 3 This describes the distribution of threshold voltages of memory cells configured to represent one of three predetermined values ​​according to one embodiment.

[0015] Figures 4 to 6 This describes the application of voltage pulses, according to some embodiments, to configure memory cells for storing data.

[0016] Figure 7A This indicates that the message is split into two codewords.

[0017] Figure 7B This explains scenarios where physical errors in the ternary unit may affect the decoding process.

[0018] Figure 8A and 8B This explains the scenarios where the ECC engine incorrectly recommends error correction.

[0019] Figure 9 This is a flowchart illustrating a method for detecting incorrect error locations in an ECC algorithm performed on an extended ECC codeword.

[0020] Figure 10 This is a flowchart illustrating a method for increasing the correction capabilities of an ECC engine by using the correlation between two codewords.

[0021] Figure 11 This is a flowchart illustrating a method for detecting incorrect errors in an ECC algorithm executed on an extended ECC codeword.

[0022] Figure 12 This describes an example computing system having a memory subsystem according to some embodiments of the present disclosure.

[0023] Figure 13 This is a block diagram of an example computer system in which embodiments of the present disclosure may be operated. Detailed Implementation

[0024] At least some aspects of this disclosure relate to a memory subsystem configured to correct errors when reading one or more self-selected memory cells.

[0025] A memory subsystem can be used as a storage device and / or a memory module. Examples of storage devices, memory modules, and memory devices are described below with reference to the figures. A host system may utilize a memory subsystem that includes one or more components, such as a memory device for storing data. The host system can provide data to be stored in the memory subsystem and can request data to be retrieved from the memory subsystem.

[0026] Integrated circuit memory cells (e.g., memory cells in flash memory or cross-point memory) can be programmed to store data based on their state under voltages applied across the memory cell. For example, if a memory cell is configured or programmed to allow a large current to pass through its state at a voltage within a predefined voltage region, then the memory cell is considered configured or programmed to store a first bit value (e.g., 1 or 0); otherwise, the memory cell stores a second bit value (e.g., 0 or 1). Optionally, a memory cell can be configured or programmed to store more than one data bit by having a threshold voltage in one of more than two separate voltage regions.

[0027] A memory cell's threshold voltage is the voltage at which, when the applied voltage across the memory cell increases above the threshold voltage, the memory cell rapidly or abruptly changes, reverts, or jumps from a non-conductive state to a conductive state. The non-conductive state allows a small leakage current to pass through the memory cell; in contrast, the conductive state allows current exceeding the threshold amount to pass through. Therefore, memory devices can use sensors to detect changes or determine the conductive / non-conductive state of the memory device under one or more applied voltages to evaluate or classify the threshold voltage level of the memory cell and thus the data stored therein.

[0028] The threshold voltage of a memory cell is configured or programmed to be in different voltage zones to represent different data values ​​stored in the memory cell. For example, the threshold voltage of a memory cell can be programmed to be in any of three predefined voltage zones; and each of these zones can be used to represent the bit value of a different two-bit data item. Therefore, given a two-bit data item, one of the three voltage zones can be selected based on the mapping between the two-bit data item and the voltage zone; and the threshold voltage of the memory cell can be adjusted, programmed, or configured to be in the selected voltage zone to represent or store the given two-bit data item. To retrieve, identify, or read a data item from a memory cell, one or more read voltages can be applied across the memory cell to determine which of the three voltage zones contains the threshold voltage of the memory cell. Identification of the voltage zone containing the threshold voltage of the memory cell provides information about two-bit data items that have been stored, programmed, or written to the memory cell.

[0029] For example, a memory cell can be configured or programmed to store one data item in a single-level cell (SLC) mode, two data items in a multi-level cell (MLC) mode, three data items in a three-level cell (TLC) mode, four data items in a four-level cell (QLC) mode, or five data items in a five-level cell (PLC) mode.

[0030] The threshold voltage of a memory cell can change or drift over time, during use and / or read operations, and in response to certain environmental factors such as temperature variations. The rate of change or drift can increase as the memory cell ages. Change or drift can cause errors when determining, retrieving, or reading back data items from the memory cell.

[0031] Redundancy information can be used to detect and correct random errors when reading memory cells. Data to be stored in a memory cell can be encoded to include redundancy information to facilitate error detection and recovery. When data encoded with redundancy information is stored in the memory subsystem, the memory subsystem can detect errors in the data represented by the voltage region of the threshold voltage of the memory cell and / or recover the original data used to generate the threshold voltage for programming the memory cell. Recovery operations can be successful (or have a high probability of success) when the data represented by the threshold voltage of the memory cell and thus retrieved directly from the memory cell in the memory subsystem contains fewer errors, or when the bit error rate in the retrieved data is low, and / or when the amount of redundancy information is high. For example, techniques such as error correction codes (ECC), low-density parity-check (LDPC) codes, etc., can be used to perform error detection and data recovery, as will be discussed in more detail herein.

[0032] Efficiently programming memory cells into intermediate states, represented by the memory cell's threshold voltage falling within a voltage region assigned to represent a certain value and separated from high-voltage and low-voltage regions, is challenging. Programming the threshold voltage of a memory cell into the high-voltage and low-voltage regions is relatively easy. However, it is difficult to precisely program the threshold voltage of a memory cell into an intermediate region that lies between the high-voltage and low-voltage regions but does not overlap with them.

[0033] Figure 1 A memory device 130 configured to have a programming manager 113 is shown according to one embodiment.

[0034] exist Figure 1 In this memory device 130, there is an array 133 of memory cells (e.g., memory cells 101). The array 133 may be referred to as a chip; and the memory device (e.g., 130) may have one or more chips. Different chips may be operated in parallel in the memory device (e.g., 130).

[0035] For example, Figure 1 The memory device 130 described herein may have a cross-point memory with an array 133 having at least memory cells (e.g., 101).

[0036] In some implementations, the crosspoint memory uses a memory cell 101 having an element (e.g., a unique element) that functions as both a selector device and a memory device. For example, the memory cell 101 may use a monolithic alloy with variable threshold capability. Read / write operations of this memory cell 101 can suppress other cells at subthreshold bias based on thresholding the memory cell 101, similar to read / write operations of memory cells having a first element acting as a selector device and a second element acting as a phase-change memory device stacked together in a row. The selector device that can be used to store information may be referred to as a selector / memory device.

[0037] Figure 1 The memory device 130 includes a controller 131 that operates a bit line driver 137 and a word line driver 135 to access individual memory cells (e.g., 101) in the array 133.

[0038] For example, each memory cell (e.g., 101) in a voltage access array 133 driven by a pair of bit line drivers 147 and word line drivers 145, such as Figure 2 As explained in the text.

[0039] The controller 131 includes a programming manager 113 configured to implement programming pulses for counter control. The programming manager 113 may be implemented, for example, via logic circuitry and / or microcode / instructions. For example, to program the threshold voltage of memory cell 101 into a second voltage region adjacent to the first voltage region, the programming manager 113 may instruct the bit line driver 137 and word line driver 135 to initially apply voltage pulses configured to program the threshold voltage of memory cell 101 into the first voltage region. After the initial voltage pulses are completed, the programming manager 113 further instructs the bit line driver 137 and word line driver 135 to apply subsequent voltage pulses to move the threshold voltage of memory cell 101 from the first voltage region to an adjacent second voltage region separate from the first voltage region. The magnitude of the subsequent voltage pulses is dynamically controlled for a group of memory cells to be read together to obtain data items (e.g., codewords used for error detection and data recovery using error correction codes (ECC)). The programming manager 113 can instruct the bit line driver 137 and word line driver 135 to incrementally increase the applied magnitude until each memory cell to be programmed into the second voltage region is conductive at the applied magnitude. For example, a counter can be used to count the number of memory cells that are conductive at the current magnitude increment. When the magnitude is increased to the level that makes the value in the counter equal to the increment of the number of memory cells to be programmed into the codeword in the adjacent second voltage region, no further increment is applied to the magnitude of subsequent voltage pulses applied to the memory cells.

[0040] Figure 2 A memory cell 101 according to one embodiment is shown, having a bit line driver 147 and a word line driver 145 configured to apply voltage pulses. For example, the memory cell 101 may be... Figure 1 Typical memory cell 101 in memory cell array 133.

[0041] Controlled by the programming manager 113 of the controller 131 Figure 2 The bit line driver 147 and word line driver 145 selectively apply one or more voltage pulses to the memory cell 101.

[0042] Bit line driver 147 and word line driver 145 can apply voltages of different polarities to memory cell 101.

[0043] For example, when a voltage of a polarity (e.g., positive polarity) is applied, bit line driver 147 drives a positive voltage relative to ground on bit lines 141 of a row of memory cells in array 133; and word line driver 145 drives a negative voltage relative to ground on word lines 143 of a column of memory cells in array 133.

[0044] When a voltage of opposite polarity (e.g., negative polarity) is applied, bit line driver 147 drives a negative voltage on bit line 141; and word line driver 145 drives a positive voltage on word line 143.

[0045] Memory cell 101 is located both in the row connected to bit line 141 and in the column connected to word line 143. Therefore, memory cell 101 experiences the voltage difference between the voltage driven by bit line driver 147 on bit line 141 and the voltage driven by word line driver 145 on word line 143.

[0046] Generally, when the voltage driven by the bit line driver 147 is higher than the voltage driven by the word line driver 145, the memory cell 101 experiences a voltage of one polarity (e.g., positive polarity); and when the voltage driven by the bit line driver 147 is lower than the voltage driven by the word line driver 145, the memory cell 101 experiences a voltage of the opposite polarity (e.g., negative polarity).

[0047] In some embodiments, memory cell 101 is a self-selecting memory cell implemented using a selector / memory device. The selector / memory device has a chalcogenide (e.g., a chalcogenide material and / or a chalcogenide alloy). For example, the chalcogenide material may comprise a chalcogenide glass, such as, for instance, an alloy of selenium (Se), tellurium (Te), arsenic (As), antimony (Sb), carbon (C), germanium (Ge), and silicon (Si). The chalcogenide material may primarily comprise selenium (Se), arsenic (As), and germanium (Ge) and is referred to as a SAG alloy. The SAG alloy may comprise silicon (Si) and is referred to as a SiSAG alloy. In some embodiments, the chalcogenide glass may contain additional elements, each in atomic or molecular form, such as hydrogen (H), oxygen (O), nitrogen (N), chlorine (Cl), or fluorine (F). The selector / memory device has a top side and a bottom side. A top electrode is formed on the top side of the selector / memory device for connection to bit line 141; and a bottom electrode is formed on the bottom side of the selector / memory device for connection to word line 143. For example, the top and bottom electrodes may be formed of carbon material. For example, the chalcogenide material of memory cell 101 may be in a crystalline atomic configuration or an amorphous atomic configuration. The threshold voltage of memory cell 101 may depend on the ratio of crystalline to amorphous material in memory cell 101. This ratio may vary under various conditions (e.g., with different amounts and directions of current flowing through memory cell 101).

[0048] The self-selecting memory cell 101, having a selector / memory device, can be programmed to have a threshold voltage window. The threshold voltage window can be created by applying programming pulses of opposite polarity to the selector / memory device. For example, the memory cell 101 can be biased to have a positive voltage difference between the two sides of the selector / memory device, and alternatively, to have a negative voltage difference between the same two sides of the selector / memory device. When a positive voltage difference is considered positive polarity, a negative voltage difference is considered negative polarity, opposite to positive polarity. Reads can be performed with a given / fixed polarity. When programmed, the memory cell has a low threshold (e.g., lower than a cell that has been reset, or a cell that has been programmed to have a high threshold), such that during a read operation, the read voltage can cause the programmed cell to snap back and thus become conductive, while the reset cell remains non-conductive.

[0049] For example, to program the voltage threshold of memory cell 101, bit line driver 147 and word line driver 145 can drive voltage pulses to memory cell 101 with one polarity (e.g., positive polarity) to cause memory cell 101 to snap back, making memory cell 101 conductive. When memory cell 101 is conductive, bit line driver 147 and word line driver 145 continue to drive programming pulses to change the threshold voltage of memory cell 101 toward the voltage region representing the data or bit value to be stored in memory cell 101.

[0050] The controller 131 can be configured in an integrated circuit having multiple layers of memory cells. Each layer can be sandwiched between a layer of bit lines and a layer of word lines; and the memory cells in the layer can be arranged in an array 133. A layer can have one or more arrays or slabs. Adjacent layers of memory cells can share a layer of bit lines (e.g., 141) or a layer of word lines (e.g., 143). Bit lines are arranged to extend parallel in one direction within their layer; and word lines are arranged to extend parallel in another direction within their layer, orthogonal to the direction of the bit lines. Each of the bit lines is connected to a row of memory cells in the array; and each of the word lines is connected to a column of memory cells in the array. Bit line drivers 137 are connected to the bit lines in the layer; and word line drivers 135 are connected to the word lines in the layer. Thus, a typical memory cell 101 is connected to both bit line driver 147 and word line driver 145.

[0051] The threshold voltage of a typical memory cell 101 is configured high enough that when only one of its bit line driver 147 and word line driver 145 drives a voltage of either polarity while the other voltage driver holds the corresponding line to ground, the magnitude of the voltage applied across memory cell 101 is insufficient to cause memory cell 101 to become conductive. Therefore, memory cell 101 can be addressed for operation / selection by driving voltages of opposite polarity relative to ground via both the driver 147 and word line driver 145 of memory cell 101. Other memory cells connected to the same word line driver 145 can be deselected by holding their respective bit lines to ground via their respective bit line drivers; and other memory cells connected to the same bit line driver can be deselected by holding their respective word lines to ground via their respective word line drivers.

[0052] A group of memory cells (e.g., 101) connected to a common word line driver 145 can be selected for parallel operation by driving a voltage of one polarity to increase by its respective word line driver (e.g., 147) while the word line driver 145 also drives a voltage of the opposite polarity to increase. Similarly, a group of memory cells connected to a common bit line driver 147 can be selected for parallel operation by driving a voltage of one polarity to increase by its respective word line driver (e.g., 145) while the bit line driver 147 also drives a voltage of the opposite polarity.

[0053] At least some examples of crosspoint memory with self-selecting memory cells are disclosed herein. Other types of memory cells and / or memories with similar threshold voltage characteristics may also be used. For example, in at least some embodiments, memory cells and / or flash memory cells, each having a selector device and a phase-change memory device, may also be used.

[0054] Figure 3 This describes the distribution of threshold voltages for memory cells, each configured to represent one of three predetermined values, according to one embodiment. For example, it can be used... Figure 1 and 2 The programming manager 113 programs the threshold voltage of the memory cell 101 so that the probability distribution of its threshold voltage is as follows: Figure 3 The explanation is as follows.

[0055] The probability distribution of the threshold voltage of a memory cell can be illustrated using a normal quantile (NQ) plot, such as in... Figure 3 In the case of a threshold voltage programmed in a region (e.g., 151), the probability distribution (also known as a Gaussian distribution) is normal, its normal quantile (NQ) plot is considered to be aligned on a straight line (e.g., distribution 151).

[0056] The self-selectable memory cell (e.g., 101) may have a threshold voltage of negative polarity and a threshold voltage of positive polarity. When the voltage applied to the memory cell 101 with either polarity increases in magnitude to the threshold voltage of its corresponding polarity, the memory cell (e.g., 101) abruptly returns from a non-conductive state to a conductive state.

[0057] The threshold voltage of the negative-polarity memory cell 101 and the threshold voltage of the positive-polarity memory cell 101 may have different values. A memory cell programmed to have a large value in the positive-polarity threshold voltage may have a small value in the negative-polarity threshold voltage; and a memory cell programmed to have a small value in the positive-polarity threshold voltage may have a large value in the negative-polarity threshold voltage.

[0058] For example, memory cell 101 can be programmed to have a small threshold voltage value (e.g., zero) according to a positive polarity distribution 151; and therefore, its threshold voltage has a large value (e.g., zero) according to a negative polarity distribution 152. This is achieved by applying a positive polarity voltage pulse (e.g., as...). Figure 4 (As described in the text) to place memory cell 101 in a conductive state and cause a predetermined level of current (e.g., 120 mA) to pass through memory cell 101, the threshold voltages of the positive and negative polarities of memory cell 101 can be programmed to distributions 151 and 152.

[0059] Alternatively, memory cell 101 may be programmed to have a smaller threshold voltage value according to the negative polarity distribution 156 to represent another value (e.g., 2); and therefore, its threshold voltage may have a larger value according to the positive polarity distribution 155 to represent the same value (e.g., 2). By applying a negative polarity voltage pulse (e.g., as... Figure 5 (As described in the text) to place memory cell 101 in a conductive state and cause a predetermined level of current (e.g., 120 mA) to pass through memory cell 101, the threshold voltages of the positive and negative polarities of memory cell 101 can be programmed to distributions 155 and 156.

[0060] The states with threshold voltages in distributions 151 and 152, and the states with threshold voltages in distributions 155 and 156, are relatively easy to obtain. They can be used... Figure 4 and 5 The voltage pulses described herein are used to program memory cell 101 to these two states. The voltage regions 151, 152, 155, and 156 are mainly controlled by the polarity of the programming voltage pulse and the current level of memory cell 101 near the end of the programming voltage pulse.

[0061] To facilitate the storage of more than one data bit per memory cell, memory cell 101 can be programmed into an intermediate state between two states.

[0062] For example, memory cell 101 can be programmed to have a moderate value of threshold voltage according to positive polarity distribution 153 to represent another value (e.g., a); and therefore, its threshold voltage has a value according to negative polarity distribution 154 to represent the same value (e.g., a). The threshold voltages of positive and negative polarities of memory cell 101 can be programmed to distributions 153 and 154 by applying voltage pulses to move the threshold voltage of the memory from distributions 151 and 152, or from distributions 155 and 156.

[0063] In some implementations, more than one intermediate state can be used to program the threshold voltage of the positive polarity in one of the voltage regions of the four distributions, and the threshold voltage of the negative polarity in one of the voltage regions of the four distributions, in a similar manner. These four states can be used to represent two-bit data items stored in memory unit 101.

[0064] exist Figure 3 In this process, positive voltage distributions 151, 153, and 155 are separated by reading voltages V1 161 and V2 162. Therefore, it can be determined whether the threshold voltage of the positive memory cell 101 is in distribution 151 by testing whether the memory cell 101 is conductive under the positive reading voltage V1 161; and it can be determined whether the threshold voltage of the positive memory cell 101 is in distribution 155 by testing whether the memory cell 101 is non-conductive under the positive reading voltage V2 162. If the threshold voltage of the positive memory cell 101 is neither in distribution 151 nor in distribution 155, then it is in distribution 153, which represents the corresponding value (e.g., a).

[0065] Similarly, in Figure 3 In this process, negative polarity distributions 152, 154, and 156 are separated by reading voltages V3 163 and V4 164. Therefore, it can be determined whether the threshold voltage of the negative polarity memory cell 101 is in distribution 156 by testing whether the memory cell 101 is conductive under the negative polarity reading voltage V3 163; and it can be determined whether the threshold voltage of the negative polarity memory cell 101 is in distribution 152 by testing whether the memory cell 101 is non-conductive under the negative polarity reading voltage V4 164. If the threshold voltage of the negative polarity memory cell 101 is neither in distribution 152 nor in distribution 156, then it is in distribution 154, which represents the corresponding value (e.g., a).

[0066] Therefore, the determination of the state and thus the value represented by the state (e.g., the region of the threshold voltage) can be performed by reading the positive polarity memory cell 101 using read voltages V1 and V2, or by reading the negative polarity memory cell 101 using read voltages V3 and V4, or by a combination of reading the negative polarity memory cell 101 using read voltage V3 and reading the positive polarity memory cell 101 using read voltage V1.

[0067] In the following embodiments, a codeword can be read from a memory device. This single codeword can be physically stored in a ternary unit of the type described above. Generally, pairs of ternary units can be read together to produce a three-bit binary value. In various embodiments, it may be advantageous to segment the codeword based on the bit positions of these individual three-bit binary values.

[0068] Figure 7A This indicates that the message is split into two codewords.

[0069] In the illustrated embodiment, user data 702A comprises a set of k bits. The specific number of k is not limited, and the specifically described size of user data 702A is not limited. Generally, user data 702A may comprise any type of binary data.

[0070] In the first step, user data 702A is divided into three-bit blocks to form block user data 704A. In some implementations, the value of three is determined by underlying memory cell technology. For example, as used herein, a given memory device may utilize a tri-state ternary cell, and the selection of three for block division may be based on this underlying physical characteristic of the memory cell. Of course, other types of memory cells may change the block value; however, the value of three is used herein, and ternary cells are also used herein. Formally, user data 702A can be represented as:

[0071]

[0072] Next, in state 706A, the block-based user data 704A is divided into two separate codewords:

[0073] ;and

[0074]

[0075] In state 708A, the parity bit can be calculated independently for each codeword. Therefore, It can have its own associated parity bit ( ),and It can have its own parity bit ( Finally, the codeword and parity bit can be encoded into the ternary value 710A. As explained, each ternary value is generated by... One bit and from The two bits form the code. It is worth noting that, as illustrated, there is a correspondence between the values ​​of CWx and CWyz due to the construction of the codeword. In some embodiments, an encoding table can be used to map a three-bit string to a combination of ternary units. An example of such an encoding table is provided in a jointly owned application with Agent File No. 120426-063400, the entire contents of which are incorporated herein by reference.

[0076] During the reading and decoding process from the ternary unit, physical errors in the ternary unit may affect the decoding process. Figure 7B An example of this problem is described below. As illustrated, the paired ternary units 702B corresponding to the binary codeword are retrieved. Additionally, two ternary memory units ( and The system experienced read errors caused by the underlying physical memory structure (more fully described in the jointly owned application with Agent File No. 120426-063400, the entire contents of which are incorporated herein by reference). During decoding, these physical errors may propagate to errors in the binary decoded value 704B. Specifically, Errors in (exist An error, and The error in the middle caused three errors, one of which was in ( And two of them are in ( and )middle.

[0077] In some memory devices, a single ECC engine can be used. As illustrated, an ECC2 decoder 710B, such as the BCH-2 engine, is used; however, the specific algorithm used is not limited. As illustrated, 706B and Both 708B inputs are fed into the ECC2 decoder 710B. In this specific instance, 706B was decoded appropriately because it contained an error and did not cause the ECC2 decoder 710B to overflow. However, during detection and / or correction... When an error occurs in 708B, the ECC2 decoder overflows from 710B to 714B because... Error 708B contains three errors. Specifically, the parser generated by the ECC2 decoder 710B may lead to arbitrary corrections, resulting in incorrect data. If decoded... 712B and Decoding If the output combination of 708B is incorrect, the resulting data will be incorrect.

[0078] However, because a ternary unit is constructed from two codewords, therefore This aspect can be used to adjust before decoding. For example, the ECC engine can recognize those with corresponding... Any errors The incorrect positions are located and these bits are reversed. The result is shown in the partially reversed codeword 716B. Here, because... Contains errors and corresponding and Bit contains an error, therefore the ECC engine can be reversed. and The ECC2 decoder 710B attempts to decode the partially inverted codeword 716B. Since the partially inverted codeword 716B contains only one error, the ECC2 decoder 710B can decode the codeword (result 718B). Then, the result 718B can be compared with the decoded codeword. The 712B is combined to produce a correctly decoded value. Therefore, the aforementioned example can utilize data from... Information to correct The error is there. Of course. and The number of errors may vary and Figure 9 and 10 Provide a complete process to explain various error scenarios.

[0079] Figure 8A and 8B This explains the scenarios where the ECC engine incorrectly recommends error correction.

[0080] In the illustrated embodiment, the codeword may include a real part 802 and a parity portion 804. The real part 802 typically refers to user data stored in a memory device, while the associated parity portion 804 includes redundant parity data generated during writing by the ECC engine. The specific size of the real part 802 and the parity portion 804 is not limited.

[0081] A given ECC code (e.g., a BCH code) may have a fixed size. For example, the illustrated BCH code includes a user data portion 806 and a parity portion 808. In some implementations, the size of the ECC code may be chosen as the minimum size capable of storing the real part 802 of the codeword described above. For example, the size of the ECC code may be defined as... ,in This increases the number of user data bits by the number of parity bits. Less than or equal to ,and This indicates error correction capability. As explained, the total size of the ECC code is greater than the size of the real part 802 and the parity part 804. In fact, although the size of the parity part 804 is equal to that of the parity part 808, the user data part 806 is larger than the real part 802.

[0082] In some implementations, it may be desirable to reuse the ECC engine regardless of the user data size (e.g., to reduce the hardware complexity of the memory controller). However, the real part 802 and the parity section 804 will generally not be processed by an ECC engine not designed to operate on the size of the real part 802. To overcome this problem, composite data 810 can be added to the real part 802 before being input to the ECC engine. In some implementations, the composite data may include all zeros or all ones. In some implementations, the size of the composite data 810 is designed such that the total size of the composite data 810 and the real part 802 is equal to the expected user data size of the real part 806. In this way, the combination of the real part 802, the composite data 810, and the parity section 804 satisfies the requirements of the ECC engine used by the memory device. The combination of the real part 802, the composite data 810, and the parity section 804 is referred to as a “shortened” codeword.

[0083] However, the introduction of synthetic data 810 may introduce false positive errors detected by the ECC engine. Specifically, since the parity section 804 is generated based on the real part 802 during encoding and then the real part 802 and synthetic data 810 are used for decoding, the parity section 804 is no longer synchronized with the codeword. Figure 8B This decoding problem is explained. As illustrated, the shortened codeword is input into the ECC decoder 818. The ECC decoder may include, for example, an ECC2 decoder capable of detecting up to two errors (e.g., a BCH-2 decoder). In some embodiments, the shortened codeword may be generated before being input into the ECC decoder 818. In other embodiments, the ECC decoder 818 itself may add synthetic data 810 to form the shortened codeword.

[0084] As illustrated by the darkened location, the real part 802 contains three errors. The synthetic data 810 necessarily does not contain any real errors because it is synthetic data. However, when the shortened codeword is decoded using the ECC decoder 818, the ECC decoder 818 detects two errors: one in the real part 802 and one in the synthetic data 810 (also illustrated by the darkened location). Therefore, the ECC decoder 818 proposes to correct both errors; however, one error is incorrect. It is worth noting that when the ECC decoder 818 overflows, all proposed corrections may be incorrect. Additionally, in some embodiments, the use of synthetic data may overwhelm the ECC, and therefore the ECC may also incorrectly detect errors in the real part 802. As will be discussed, it can be ensured that errors in the synthetic data 810 are a sign of error overflow because errors are impossible to exist in the synthetic data 810. Generally, when the number of errors overwhelms the ECC, the resulting checksum can propose to correct two errors at potentially arbitrary locations, which may or may not coincide with the actual error location.

[0085] If the actual number of errors suppresses the ECC correction capability, then the probability of having two proposed corrections within the actual location can be defined probabilistically, and the probability is expressed as follows:

[0086]

[0087] here, This refers to the length of the real part 802 and This refers to the size of the shortened codeword input into the ECC engine (e.g., (Including the size of the synthesized data 810). For example, if the codeword size is 511 bits but the real part 802 is only 274 bits, then the probability of ECC correcting two true errors is as follows:

[0088]

[0089] Therefore, in this scenario, the ECC engine will incorrectly attempt to correct false errors in 72% of the considered codewords (containing more than two errors). This probability inevitably increases as the real part 802 occupies a smaller fraction of the total size of the shortened codeword. In all these scenarios, example embodiments provide techniques for resolving the consistency of detected errors given a shortened ECC codeword.

[0090] Figure 9 This is a flowchart illustrating a method for detecting incorrect error locations in an ECC algorithm performed on an extended ECC codeword.

[0091] In step 902, the method may include receiving codewords.

[0092] As discussed above, the codeword in step 902 may comprise a codeword including a first portion and at least one other portion. The following description utilizes the first portion and one other portion (“second” portion); however, this disclosure is not limited to a single other portion. In embodiments, the first portion may comprise actual data. As used herein, actual data refers to data written to a memory device or otherwise used by a computing system. In contrast, the second portion may comprise synthetic data. In embodiments, the codeword may also include a parity check portion generated using the first portion as input. In some embodiments, the synthetic data may comprise a random pattern of all zeros, all one-digit numbers, or zeros and one-digit numbers. In some embodiments, step 902 may be implemented within ECC circuitry or an algorithm. In other embodiments, step 902 may be implemented by a microcontroller or via software before the extended codeword is input into the ECC circuitry or algorithm.

[0093] In step 904, the method may include using an ECC engine (e.g., circuitry or algorithm) to detect the location of errors throughout the codeword.

[0094] In some implementations, the ECC engine can detect multiple bit errors. In some implementations, the ECC engine can implement existing ECC algorithms, such as Single Error Correction / Double Error Detection (SEC-DED) Schott code, Single Error Correction / Double Error Detection / Single Byte Error Detection (SEC-DED-SBD) Radidis code, Single Byte Error Correction / Double Byte Error Detection (SBC-DBD) finite field-based code, Double Error Correction / Triple Error Detection (DEC-TED) Dr.-Chowdhury-Hokungamme (BCH) code, or similar types of ECC. Generally, any ECC engine capable of detecting error locations can be used.

[0095] In some embodiments, the method may store the location of any errors detected in step 904. For example, the method may store the bit location of the error relative to the codeword in a volatile storage device (e.g., DRAM) for later use (e.g., in steps 908 and 910 discussed herein).

[0096] In step 906, the method may include determining whether any errors exist within the codeword. If none exist, the method may terminate because error correction is not required. As illustrated, in one embodiment, if even a single error is detected, and of course, if multiple errors are detected, the method may continue to step 908. As discussed in conjunction with Figure 8, since the codeword received in step 902 contains composite data that can be all zeros or all ones, errors can be detected either within the real part of the keyword (containing user data) or within this composite region. Since this composite region does not contain actual data, errors detected within this region are false positives.

[0097] In step 908, the method may include comparing the detected error location with the real part location of the codeword.

[0098] In some embodiments, the method may be configured to have a mapping of the actual position and the synthesized position of the codeword received in step 902. For example, in some embodiments, the method may store the length of the real part (starting from zero). Alternatively, in some embodiments, the method may store a bit mapping of the (e.g., non-contiguous) real bits of the codeword.

[0099] The method compares the detected error location with a list (or range) of real bit positions in the codeword. Then, in step 910, the method determines whether the ECC engine has appropriately detected the error. Specifically, in step 910, the method determines whether the detected error location corresponds to a bit of real data in the codeword. For example, if the codeword contains n bits and bits 0 to m comprise the real portion of the codeword (where...) If so, then step 910 may include determining the location of the error bit ( , … Whether it is located in a bit position between 0 and m.

[0100] In some scenarios, all detected bit errors may be within the real part. In this case, the method may proceed to step 912 and correct the error. In one embodiment, step 912 may include a correction engine running an ECC engine to correct the detected error. The specific operation of the ECC engine is not limited and is not discussed in detail herein. After correcting the error, the method may then return the codeword to the calling device in step 916.

[0101] In contrast, if the method determines (in step 910) that at least one error location is not in the real portion of the keyword (i.e., in the synthesized data), then the method may proceed to step 914, where it handles ECC false detections and false corrections before ending. In some embodiments, the method may signal that the error correction has failed due to the synthesized portion overwhelming the ECC engine. This method may be used in conjunction with any of the foregoing embodiments (e.g., as a flag indicating that this correction has been performed). Alternatively, a signal may be issued immediately, and the method may stop, indicating that remedial action is required. For example, a backup of the codeword may be read from a redundant memory device.

[0102] Use the above Figure 9 This method allows the codeword format to be reused by the ECC engine to process input codewords of varying lengths. However, the ECC engine will not properly detect errors in this "extended" codeword (such as...). Figure 8A and 8B (As explained in the document), therefore this method utilizes the structure of the codeword to ensure that only valid errors are detected. Using this structure allows standard ECC engines to be used with variable-length codewords and allows for the reuse of existing ECC engines, although it reduces user data.

[0103] Figure 10 and 11 This is a flowchart illustrating a method for increasing the correction capabilities of an ECC engine by using the correlation between two codewords.

[0104] The method is described in more detail below. At a higher level, the method may include receiving a codeword having a first part and a second part and using an ECC (e.g., ECC2) engine to detect at least one error or failure in the first part (step 1002). Based on the number of errors or failures in ECC2, the method may then perform error correction on the second part and invert zero or more bits of the first or second part, and then continue to perform error correction on both the first and second parts (step 1024). More specifically, if the ECC on the first part fails, the method may invert the bits of the first part based on the extended ECC performed on the second part (steps 1004, 1006, 1008). If there are no errors in the first part, the method may perform extended ECC on the second part and if three errors occur, the codeword is marked as uncorrectable; otherwise, the errors are corrected (steps 1010 and 1012). If an error is detected in the first part, the method may employ a combination of extended ECC detection and conformance checking on the second part to determine when to invert the bits of the second part (steps 1014, 1016, 1018, 1020, 1026, 1028). Finally, if two errors are detected in the first part, the method may perform an extended ECC operation on the second part and perform an alternative conformance check to determine when to invert the bits of the second part. Details of these operations are provided herein. While the foregoing method typically describes a maximum of three errors, the method can be generalized to include more than three errors.

[0105] In step 1002, the method may begin by receiving a codeword and using an ECC engine (e.g., circuitry or algorithm) to detect the location of an error within the first part of the codeword.

[0106] As discussed above, the codeword in step 1002 may comprise a codeword including a first portion (referred to as codeword X) and at least one other portion (referred to as codeword YZ). The following description utilizes the first portion and one other portion (“second” portion); however, the invention is not limited to the single other portion. In embodiments, the first portion may comprise actual data. As used herein, actual data refers to data written to a memory device or otherwise used by a computing system. In contrast, the second portion may comprise synthetic data. In some embodiments, the synthetic data may comprise a random pattern of all zeros, all one-digits, or zeros and one-digits. In some embodiments, step 1002 may be implemented within ECC circuitry or an algorithm. In other embodiments, step 1002 may be implemented by a microcontroller or via software before the extended codeword is input into the ECC circuitry or algorithm.

[0107] In some implementations, the ECC engine can detect multiple bit errors. In some implementations, the ECC engine can implement existing ECC algorithms, such as Single Error Correction / Double Error Detection (SEC-DED) Schott code, Single Error Correction / Double Error Detection / Single Byte Error Detection (SEC-DED-SBD) Radidis code, Single Byte Error Correction / Double Byte Error Detection (SBC-DBD) finite field-based code, Double Error Correction / Triple Error Detection (DEC-TED) Bosch-Chowdhury-Hokungamme (BCH) code, or similar types of ECC. Generally, any ECC engine capable of detecting error locations can be used. In some implementations, step 1002 may include utilizing an ECC2 engine.

[0108] As explained, the method in step 1002 may detect 0, 1, or 2 errors in codeword X, or it may fail, as described in the branch of step 1002. Generally, failure means that the ECC algorithm detects too many errors, to the point that it may be unable to correct them. In some implementations, failure may also refer to detecting errors as previously mentioned... Figure 9 The correction proposed in the synthesis region of the described codeword. This scenario is also known as suppressing or overpowering the error correction capability of ECC. An ECC algorithm or engine may consist of two phases: a detection phase or engine and a correction phase or engine. In various detection steps, only the detection engine can be utilized. In some implementations, the detection engine may not be able to identify a specific number of errors, but only identify error overflows that have occurred.

[0109] If the ECC2 engine fails in step 1002, the method proceeds to step 1004. In step 1004, the extended ECC2 engine is applied to the second part of the codeword (codeword YZ). As illustrated, in some embodiments, the extended ECC2 engine may detect zero to three errors or may fail, similar to the ECC2 engine discussed in step 1002. As illustrated, if the extended ECC2 engine detects one or two errors in codeword YZ, the method proceeds to step 1006 (described below). However, if the extended ECC2 engine does not detect any errors, detects three errors, or fails, the method marks the entire codeword (e.g., both codeword X and codeword YZ) as uncorrectable and fails in step 1008.

[0110] In step 1006, the method has determined that one or two errors exist in codeword YZ. In response, the method inverts one or more corresponding bits in codeword X. Specifically, in some embodiments, the method identifies which bits in codeword YZ are associated with the detected errors and inverts the corresponding bits in codeword X. As discussed above, a given bit in codeword YZ can (e.g., based on addressing) correspond to a corresponding bit in codeword X. Therefore, the extended ECC2 engine can indicate the address of the error in codeword YZ, and this address can be used to identify the corresponding bit in codeword X that should be inverted.

[0111] As explained, after the bits of codeword X are inverted in step 1006, the method continues to step 1024, where an appropriate ECC engine is used to correct the codeword. Specifically, an ECC2 engine is used to correct codeword X (utilizing the bit inversion applied in step 1006), while an extended ECC2 engine is used to correct codeword YZ. Notably, in step 1024, actual error correction is performed.

[0112] Returning to step 1002, in another scenario, the ECC2 engine in step 1002 may not detect any errors in codeword X. In this scenario, the method continues to step 1010. In step 1010, the method uses an extended ECC2 engine to detect errors in codeword YZ. As in step 1004, the method in step 1010 includes detecting zero to three errors (or, in some embodiments, zero to two errors) or failure. However, in step 1010, the method may include determining whether the number of detected errors is equal to three. If the extended ECC2 engine detects three errors (and in some embodiments, if the check fails), then the method continues to step 1012. In step 1012, the method (as in step 1008) marks the entire codeword as uncorrectable and fails. In contrast, if fewer than three errors are detected (or the extended ECC2 fails), then the method continues to step 1024, where an appropriate ECC engine is used to correct the codeword. Specifically, the ECC2 engine is used to correct codeword X, while the extended ECC2 engine is used to correct codeword YZ. It is worth noting that in step 1024, actual error correction is performed.

[0113] Returning to step 1002, in another scenario, the ECC2 engine in step 1002 can detect a single error in codeword X. In response to the detection of a single error in codeword X, the method may perform an extended ECC operation on the second part in step 1014. If the extended ECC operation in step 1014 indicates zero or one error, then the method may continue to 1026, where the appropriate ECC engine is used to correct the codeword. Specifically, the ECC2 engine is used to correct codeword X, while the extended ECC2 engine is used to correct codeword YZ. Notably, in step 1026, actual error correction is performed.

[0114] In contrast, if the extended ECC operation in step 1014 indicates three errors (and in some embodiments, if the extended ECC operation fails), the method may proceed to step 1020, where bits of codeword YZ are inverted based on the detected errors in codeword X. Specifically, in some embodiments, the method identifies which bits in codeword X are associated with the detected errors and inverts the corresponding bits in codeword YZ. As discussed above, given bits in codeword X can (e.g., based on addressing) correspond to corresponding bits in codeword YZ. Therefore, the ECC2 engine can indicate the address of the error in codeword X and can use this address to identify the corresponding bits in codeword YZ that should be inverted. After performing this inversion in step 1020, the method proceeds to step 1022, where a second extended ECC operation is performed on the second portion (codeword YZ). Here, if the second extended ECC operation indicates three errors, the method proceeds to step 1018 and marks the codeword as uncorrectable (similar to steps 1008 or 1012 discussed previously). In contrast, if the second extended ECC operation produces any other result (e.g., zero to two errors or failures), then the method continues to step 1028, where the codeword portion is corrected (similar to steps 1024 and 1026).

[0115] Finally, in some scenarios, the extended ECC operation in step 1014 may fail or indicate two errors. In this scenario, the method proceeds to step 1016, where a consistency check is performed. As used herein, a consistency check refers to a logical test that produces an OK or non-OK value (e.g., pass or fail). In some implementations, an OK value can be determined by determining whether at least one error in the second part (codeword YZ) matches at least one error in the first part (codeword X). In contrast, if the error location in the second part (codeword YZ) differs from the error location in the first part (codeword X), then a non-OK value can be returned. Therefore, this consistency check can be used to quickly confirm or reject errors by leveraging the relationship between codeword parts. If the consistency check passes, the method can correct the codeword in step 1026, as described above. In contrast, if the consistency check fails, the method can mark the codeword as uncorrectable in step 1018, as previously described.

[0116] Returning to step 1002 (final scenario), the ECC2 engine in step 1002 detects two errors in codeword X. In this scenario, step 1030 can be performed on codeword YZ, as will be discussed later in this article. Figure 11 The steps are described below.

[0117] Turn Figure 11 ,when Figure 10 This method can be invoked when it detects two errors in the first part of a codeword (codeword X). In step 1102, the method may include applying an extended ECC error correction operation to the second part of the codeword (codeword YZ).

[0118] In the first scenario, the method may not detect errors in the second part of the codeword. In this scenario, the method proceeds to step 1108, where errors in the codeword are corrected. In some implementations, an appropriate ECC engine may be used to correct the codeword. Specifically, an ECC2 engine is used to correct codeword X, while an extended ECC2 engine is used to correct codeword YZ. Notably, in step 1108, the actual error correction is performed.

[0119] In another scenario, the method can detect a single error in the second part of the codeword. In this scenario, the method performs a consistency check in step 1104. As in step 1016, this consistency check returns an OK or NOT OK value based on comparing the location of the error in each part of the codeword. The details of the consistency check are the same as in step 1016 and will not be repeated herein. If the consistency check passes (i.e., an OK value is generated), the method proceeds to step 1108 and corrects the error in the codeword. Step 1108 has been described above and will not be repeated herein. In contrast, if the consistency check fails (NOOK), the method proceeds to step 1106 and marks the codeword as uncorrectable.

[0120] In the third scenario, the method may detect two errors in the second part of the codeword or, alternatively, fail. In this scenario, the method may proceed to step 1110, where a second consistency check is performed on the codeword portion (the same as described in step 1104). If the consistency check passes (OK), the method proceeds to step 1108 and corrects the errors in the codeword. Step 1108 has been described above and will not be repeated herein. In contrast, if the consistency check fails (Not OK), the method proceeds to step 1112 as described herein.

[0121] Finally, if the extended ECC operation in step 1102 indicates three errors (and in some embodiments, if the extended ECC operation fails), or if the consistency check in step 1110 fails, the method may continue to step 1112, where bits of codeword YZ are inverted based on the errors detected in codeword X. Specifically, in some embodiments, the method identifies which bits in codeword X are associated with the detected errors and inverts the corresponding bits in codeword YZ. As discussed above, given bits in codeword X can (e.g., based on addressing) correspond to corresponding bits in codeword YZ. Therefore, the ECC2 engine can indicate the address of the error in codeword X and can use this address to identify the corresponding bits in codeword YZ that should be inverted. After performing this inversion in step 1112, the method continues to step 1114, where a second extended ECC operation is performed on the second portion (codeword YZ). Here, if the second extended ECC operation indicates three errors, the method continues to step 1106 and marks the codeword as uncorrectable. In contrast, if the second extended ECC operation produces any other result (e.g., zero to two errors or optionally, failure), then the method proceeds to step 1108, where the codeword portion is corrected (as previously discussed).

[0122] Figure 12This describes an example computing system 100 including a memory subsystem 110 according to some embodiments of the present disclosure. The memory subsystem 110 may include media, such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., ... Figure 1 (memory device 130) or a combination thereof.

[0123] The memory subsystem 110 may be a storage device, a memory module, or a combination of a storage device and a memory module. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage (UFS) drives, secure digital cards (SD cards), and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small form factor DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).

[0124] The computing system 100 may be, for example, a desktop computer, a laptop computer, a web server, a mobile device, a vehicle (e.g., an airplane, drone, train, car or other means of transport), a device with Internet of Things (IoT) capabilities, an embedded computer (e.g., an embedded computer contained in a vehicle, industrial equipment or networked commercial device), or a computing device that includes memory and processing devices.

[0125] The computing system 100 may include a host system 122 coupled to one or more memory subsystems 110. Figure 12 This describes an example of a host system 122 coupled to a memory subsystem 110. As used herein, “coupled to” or “coupled with” generally refers to a connection between components, which can be an indirect or direct communication connection (e.g., without an intermediary component), whether wired or wireless, including connections such as electrical, optical, magnetic, etc.

[0126] Host system 122 may include a processor chipset (e.g., processing device 118) and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., controller 116) (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). For example, host system 122 uses memory subsystem 110 to write data to and read data from memory subsystem 110.

[0127] Host system 122 can be coupled to memory subsystem 110 via a physical host interface. Examples of physical host interfaces include, but are not limited to, Serial Advanced Technology Attachment (SATA) interfaces, Peripheral Component Interconnect Fast (PCIe) interfaces, Universal Serial Bus (USB) interfaces, Fibre Channel, Serial Attached SCSI (SAS) interfaces, Double Data Rate (DDR) memory bus interfaces, Small Computer System Interface (SCSI), Dual In-line Memory Module (DIMM) interfaces (e.g., DIMM slot interfaces supporting Double Data Rate (DDR)), Open NAND Flash Interface (ONFI), Double Data Rate (DDR) interfaces, Low Power Double Data Rate (LPDDR) interfaces, or any other interfaces. The physical host interface can be used to transfer data between host system 122 and memory subsystem 110. When memory subsystem 110 is coupled to host system 122 via a PCIe interface, host system 122 can further utilize NVM Fast (NVMe) interface access components (e.g., Figure 1 (Memory device 130). The physical host interface provides an interface for transmitting control, address, data and other signals between the memory subsystem 110 and the host system 122. Figure 12 The memory subsystem 110 is described as an example. Generally, the host system 122 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or combinations of communication connections.

[0128] For example, the processing device 118 of the host system 122 may be a microprocessor, a central processing unit (CPU), a processor core, an execution unit, etc. In some examples, the controller 116 may be referred to as a memory controller, a memory management unit, and / or a starter. In one instance, the controller 116 controls communication via a bus coupled between the host system 122 and the memory subsystem 110. Generally, the controller 116 may send commands or requests to the memory subsystem 110 to perform desired access to memory devices 130, 140. The controller 116 may further include an interface circuitry for communicating with the memory subsystem 110. The interface circuitry may translate responses received from the memory subsystem 110 into information for the host system 122.

[0129] The controller 116 of the host system 122 can communicate with the controller 115 of the memory subsystem 110 to perform operations such as reading data, writing data, or erasing data at memory devices 130, 140, and other such operations. In some examples, the controller 116 is integrated within the same package as the processing device 118. In other examples, the controller 116 is packaged separately from the processing device 118. The controller 116 and / or the processing device 118 may include hardware such as one or more integrated circuits (ICs) and / or discrete components, buffer memory, cache memory, or combinations thereof. The controller 116 and / or the processing device 118 may be a microcontroller, a special-purpose logic circuit system (e.g., a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.), or another suitable processor.

[0130] Memory devices 130 and 140 may include different types of non-volatile memory components and / or any combination of volatile memory components. Volatile memory devices (e.g., memory device 140) may be, but are not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).

[0131] Examples of non-volatile memory components include NAND flash memory and in-situ write memory, such as three-dimensional crosspoint ("3D crosspoint") memory. Crosspoint arrays of non-volatile memory can perform bit storage based on variations in bulk resistance in conjunction with stackable cross-gate format data access arrays. Furthermore, compared to many flash-based memories, crosspoint non-volatile memory can perform in-situ write operations, where non-volatile memory cells can be programmed without prior erasing of the non-volatile memory cells. NAND flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).

[0132] Each of the memory devices 130 may include one or more arrays of memory cells. One type of memory cell (e.g., a single-level cell (SLC)) may store one bit per cell. Other types of memory cells (e.g., multi-level cell (MLC), three-level cell (TLC), four-level cell (QLC), and five-level cell (PLC)) may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more arrays of memory cells, such as SLC, MLC, TLC, QLC, PLC, or any combination thereof. In some embodiments, a particular memory device may include an SLC portion, an MLC portion, a TLC portion, a QLC portion, and / or a PLC portion of memory cells. The memory cells of the memory device 130 may be grouped into pages, which may refer to logical units of the memory device used for storing data. For some types of memory (e.g., NAND), pages may be grouped to form blocks.

[0133] Although non-volatile memory devices such as 3D crosspoint type and NAND type memory (e.g. 2D NAND, 3D NAND) are described, memory device 130 may be based on any other type of non-volatile memory, such as read-only memory (ROM), phase change memory (PCM), self-select memory, other chalcogenide-based memory, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin-transfer torque (STT)-MRAM, conductive bridged RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), NOR flash memory, and electrically erasable programmable read-only memory (EEPROM).

[0134] The memory subsystem controller 115 (or simply controller 115) can communicate with the memory device 130 to perform operations such as reading data, writing data, or erasing data at the memory device 130, and other such operations (e.g., in response to commands scheduled on the command bus by controller 116). Controller 115 may include hardware such as one or more integrated circuits (ICs) and / or discrete components, buffer memories, or combinations thereof. The hardware may include a digital circuit system having dedicated (e.g., hard-coded) logic for performing the operations described herein. Controller 115 may be a microcontroller, a dedicated logic circuit system (e.g., a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), etc.), or another suitable processor.

[0135] The controller 115 may include a processing means 117 (e.g., a processor) configured to execute instructions stored in local memory 119. In the illustrated example, the local memory 119 of the controller 115 includes embedded memory configured to store instructions for performing operations of the control memory subsystem 110 (including handling communication between the memory subsystem 110 and the host system 122).

[0136] In some embodiments, local memory 119 may include memory registers storing memory pointers, fetched data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Although already Figure 12 The instance memory subsystem 110 is described as including controller 115, but in another embodiment of this disclosure, memory subsystem 110 does not include controller 115 and may instead rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).

[0137] Generally, controller 115 can receive commands or operations from host system 122 and can translate these commands or operations into instructions or appropriate commands to achieve the desired access to memory device 130. Controller 115 may handle other operations such as wear leveling, discarded item collection, error detection and error correction code (ECC) operations, encryption, caching, and address translation between logical addresses (e.g., logical block addresses, namespaces) and physical addresses (e.g., physical block addresses) associated with memory device 130. Controller 115 may further include a host interface circuitry for communicating with host system 122 via a physical host interface. The host interface circuitry can translate commands received from the host system into command instructions to access memory device 130 and translate responses associated with memory device 130 into information for host system 122.

[0138] The memory subsystem 110 may also include additional circuitry or components not described. In some embodiments, the memory subsystem 110 may include caches or buffers (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that can receive addresses from the controller 115 and decode the addresses to access the memory device 130.

[0139] In some embodiments, memory device 130 includes a local media controller 131 that operates in conjunction with memory subsystem controller 115 to perform operations on one or more memory cells of memory device 130. An external controller (e.g., memory subsystem controller 115) may externally manage memory device 130 (e.g., perform media management operations on memory device 130). In some embodiments, memory device 130 is a managed memory device, which is a native memory device combined with a local controller (e.g., local media controller 131) for media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.

[0140] The controller 115 and / or memory device 130 may include a programming manager 113, such as those described above. Figures 1 to 6 The programming manager 113 is described herein. In some embodiments, a controller 115 in the memory subsystem 110 includes at least a portion of the programming manager 113. In other embodiments, or in combination, a controller 116 and / or processing device 118 in the host system 122 includes at least a portion of the programming manager 113. For example, controllers 115, 116, and / or processing device 118 may include a logic circuitry system implementing the programming manager 113. For example, controller 115 or processing device 118 of the host system 122 (e.g., a processor) may be configured to execute instructions stored in memory for performing the operations of the programming manager 113 described herein. In some embodiments, the programming manager 113 is implemented in an integrated circuit chip (e.g., memory device 130) mounted in the memory subsystem 110. In other embodiments, the programming manager 113 may be a portion of the firmware of the memory subsystem 110, an operating system of the host system 122, a device driver or application, or any combination thereof.

[0141] Figure 13 An example machine illustrating computer system 300 may execute within said computer system 300 a set of instructions for causing said machine to perform any or more of the methodologies discussed herein. In some embodiments, computer system 300 may correspond to a host system (e.g., Figure 12 The host system 122), which includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 12 The memory subsystem 110) or can be used to perform operations of the programming manager 113 (e.g., execute instructions to perform operations corresponding to the reference). Figure 12 and 13The operation of the described programming manager 113 is described. In alternative embodiments, the machine may connect (e.g., network) to other machines in a LAN, intranet, extranet, and / or the Internet. The machine may operate as a server or client machine in a client-server network environment, as a peer-to-peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.

[0142] A machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network device, a server, a network router, a switch, or a bridge, or any machine capable of (sequentially or otherwise) executing a set of instructions specifying actions to be taken by said machine. Furthermore, while describing a single machine, the term "machine" should also be considered as any collection of machines that individually or jointly execute a set (or more) of instructions to perform any or more of the methodologies discussed herein.

[0143] The example computer system 300 includes a processing device 302, a main memory 304 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), static random access memory (SRAM), etc.) and a data storage system 318, which communicate with each other via a bus 330 (which may include multiple buses).

[0144] Processing device 302 represents one or more general-purpose processing devices, such as a microprocessor, central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets, or multiple processors implementing combinations of instruction sets. Processing device 302 may also be one or more special-purpose processing devices, such as application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), network processors, or the like. Processing device 302 is configured to execute instructions 326 for performing the operations and steps discussed herein. Computer system 300 may further include a network interface device 308 for communication via network 320.

[0145] The data storage system 318 may include a machine-readable medium 324 (also referred to as computer-readable medium) on which one or more sets of instructions 326 or software embodying any or more of the methodologies or functions described herein are stored. The instructions 326 may also reside wholly or at least partially within main memory 304 and / or processing device 302 during execution by computer system 300, which also constitute machine-readable storage media. The machine-readable medium 324, data storage system 318, and / or main memory 304 may correspond to... Figure 12 The memory subsystem 110.

[0146] In one embodiment, instruction 326 includes implementations corresponding to programming manager 113 (e.g., reference...). Figures 1 to 6 The described programming manager 113) provides functional instructions. Although the machine-readable medium 324 is shown as a single medium in the exemplary embodiment, the term "machine-readable storage medium" should be considered as a single medium or multiple media containing one or more sets of instructions. The term "machine-readable storage medium" should also be considered as any medium capable of storing or encoding a set of instructions for machine execution and causing the machine to perform any or more of the methodologies of this disclosure. The term "machine-readable storage medium" should accordingly include, but is not limited to, solid-state memory, optical media, and magnetic media.

[0147] Some parts of the foregoing detailed description have been presented based on the algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are the means by which those skilled in the art of data processing most effectively communicate the essence of their work to others skilled in the art. The algorithms described herein are generally conceived as self-consistent sequences of operations that lead to desired results. An operation is an operation that requires the physical manipulation of physical quantities. Typically, although not always necessary, these quantities take the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. It has been shown that, primarily for common use, it is sometimes convenient to refer to these signals as bits, values, elements, symbols, characters, items, numbers, or similar terms.

[0148] However, it should be remembered that all these and similar terms should be associated with appropriate physical quantities and are merely convenient labels for application to those quantities. This disclosure may relate to the operation and processes of a computer system or similar electronic computing device that manipulate and transform data representing physical (electronic) quantities in the registers and memories of the computer system into other data similarly represented in the memory or registers of the computer system or other such information storage systems.

[0149] This disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specifically constructed for its intended purpose, or may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in a computer. This computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic cards or optical cards, or any type of media suitable for storing electronic instructions, each coupled to a computer system bus.

[0150] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general-purpose systems can be used in conjunction with programs based on the teachings herein, or it may prove convenient to construct more specialized devices to execute the methods. The architectures of many such systems will appear as described below. Furthermore, this disclosure is not described with reference to any particular programming language. It will be understood that various programming languages ​​can be used to implement the teachings of this disclosure as described herein.

[0151] This disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, the instructions being usable to program a computer system (or other electronic device) to perform processes according to this disclosure. Machine-readable media includes any means for storing information in a form readable by a machine (e.g., a computer). In some embodiments, machine-readable (e.g., computer-readable) media includes machine-readable storage media, such as read-only memory (“ROM”), random access memory (“RAM”), disk storage media, optical storage media, flash memory components, etc.

[0152] In this description, various functions and operations are described as being executed or caused by computer instructions for the sake of simplicity. However, those skilled in the art will recognize that such expressions mean that the functions originate from one or more controllers or processors, such as a microprocessor, executing computer instructions. Alternatively or in combination, functions and operations may be implemented using dedicated circuit systems, with or without software instructions, such as using application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). Embodiments may be implemented using hardwired circuit systems, either without or in combination with software instructions. Therefore, the technology is not limited to any particular combination of hardware circuit systems and software, nor to any particular source of instructions executed by a data processing system.

[0153] In the foregoing description, embodiments thereof have been described with reference to specific examples of the present disclosure. It will be understood that various modifications may be made to the present disclosure without departing from the broader spirit and scope of the embodiments set forth in the appended claims. Therefore, the description and drawings should be regarded in an illustrative rather than restrictive sense.

Claims

1. A method comprising: Receive codewords containing a first part and a second part; Use the ECC engine to detect errors at the first position in the codeword; and When the first position is within the second part, a signal is sent to notify of an error or false detection.

2. The method of claim 1, wherein the second portion comprises all ones or all zeros.

3. The method of claim 1, wherein the second portion comprises random data.

4. The method of claim 1, wherein determining whether the first position is within the first portion includes determining whether the first portion is less than the length of the first portion.

5. The method of claim 1, wherein detecting at least one error in the codeword at the first position comprises detecting two errors in the codeword and correcting the two errors if the two errors are in the first portion.

6. The method according to claim 1, wherein the codeword comprises a BCH codeword.

7. The method of claim 6, wherein the BCH codeword includes a parity check portion generated based on the first portion.

8. A method comprising: Receive a message containing a first codeword and a second codeword; Use the first ECC engine to detect errors in the position of the first codeword; Use a second ECC engine to analyze the second codeword to detect the location of errors within the second codeword; Based on the position of the error within the second codeword and the position thereon, invert at least one bit of the second codeword; and Correct the message.

9. The method of claim 8, wherein the first ECC engine and the second ECC engine comprise a single ECC engine.

10. The method of claim 8, wherein the method further comprises: Perform a consistency check on the first codeword and the second codeword; In the event that the consistency check fails, the message is marked as uncorrectable; and The message is corrected if the consistency check passes.

11. The method of claim 8, wherein analyzing the second codeword includes detecting three error locations by the second ECC engine.

12. The method of claim 11, wherein the method further comprises: The second codeword is input into the second ECC engine to detect the second error location; and When the number of the second error locations is less than three, the message is corrected.

13. A method comprising: Receive a message including the first codeword and the second codeword; Detect two errors in the first position of the first codeword; Analyze the second codeword to detect the location of errors within the second codeword; and The message is corrected based on the first position and the position of the error within the second codeword.

14. The method of claim 13, wherein analyzing the second codeword to detect the location of an error within the second codeword includes detecting an error, and the method further includes performing a consistency check on the second codeword, wherein the method corrects the message when the consistency check passes, and marks the message as uncorrectable when the consistency check fails.

15. The method of claim 13, wherein analyzing the second codeword to detect the location of an error within the second codeword includes not detecting an error.

16. The method of claim 13, wherein analyzing the second codeword to detect the location of an error within the second codeword includes detecting two errors or an ECC engine failure, the method further comprising: Perform a consistency check on the second codeword; The message is corrected when the consistency check passes; and When the consistency check fails, at least one bit of the second codeword is reversed based on the position of the error within the second codeword and the first position.

17. The method of claim 16, wherein the method further comprises inputting the second codeword into the ECC engine and correcting the message if the number of errors detected by the ECC engine is less than three.

18. The method of claim 13, wherein analyzing the second codeword to detect the location of an error within the second codeword includes detecting three errors, the method further comprising: Invert at least one bit of the second codeword based on the position of the error within the second codeword and the first position; and The second codeword is input into the second ECC engine and corrected if the number of errors detected by the second ECC engine is less than three.

19. A method comprising: Receive a message, the message having a first codeword and a second codeword; The first ECC engine was used to detect that error detection had failed when analyzing the first codeword; Use a second ECC engine to analyze the second codeword to detect the location of errors within the second codeword; Invert at least one bit of the first codeword based on the position of the error and the at least one position within the second codeword; and The message is corrected using the first ECC engine and the second ECC engine.

20. The method of claim 19, wherein the first ECC engine comprises an ECC2 engine.