Hard decoding method in data storage devices
By employing a multidimensional encoding method based on BCH component code and symmetric product code in NAND flash memory devices, the problem of insufficient error correction capability under high code rate conditions is solved, achieving more efficient error correction and reducing encoding complexity.
Patent Information
- Application Number
- CN202111611386.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-28
- Filing Date
- 2021-12-27
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-12-27
AI Technical Summary
Existing technologies in non-volatile memory devices, especially NAND flash memory devices, suffer from insufficient error correction capabilities. Particularly under high code rate conditions, conventional encoding methods are complex and costly, making it difficult to effectively correct errors caused by programming errors and read interference.
A component-code-based multidimensional encoding method is adopted, which uses BCH component codes and symmetric product codes to construct error correction codes. Through repeated decoding of component codes and multidimensional encoding, effective data correction is achieved.
It improves error correction capabilities, reduces encoding and decoding complexity, is suitable for NAND flash memory devices, and provides a more cost-effective error correction solution.
Smart Images

Figure CN114691413B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to systems and methods for hard-decoding to correct errors in data storage devices, and in particular, non-volatile memory devices. BACKGROUND
[0002] Flash memory devices (e.g., NAND flash memory devices) enable page reads based on voltage thresholds of the flash memory devices. Due to different noise (e.g., NAND noise) and interference sources during programming and reading, information bits stored in the flash memory devices can experience errors. Such errors can be due to one or more of programming errors, reads with non-optimal thresholds, retention / read interference stress, etc. Strong error correction codes (ECCs) can allow for fast programming (programming errors can be high) and reads under high stress conditions and / or by low complexity digital signal processing (DSP). Other causes of damage can result in complete erasure of a physical page, row, or block in a NAND flash memory device, which is known as a block becoming a "bad" block and no longer being readable. If damage is not detected during programming, encoding (e.g., RAID encoding) can be needed to allow for recovery of the unreadable area on the NAND flash memory device.
[0003] The code rate is defined by the ratio of the information content of a codeword (referred to as "payload") to the overall size of the codeword. For example, for a code containing k bits and r redundancy bits, the code rate R c is defined by Conventional encoding methods are not well suited to support codes with high code rates for both hard-decoding and soft-decoding. For example, conventional low-density parity-check (LDPC) codes with high code rates (e.g., 0.9) have a fairly long code length, resulting in a complex and costly implementation. SUMMARY
[0004] In some arrangements, systems, methods, and non-transitory processor-readable media involve decoding data read from a non-volatile storage device, including determining error candidates for the data based on component codes; determining whether at least one first error candidate is found from the error candidates based on two of the component codes agreeing on a same error candidate; determining whether at least one second error candidate is found based on two of the component codes agreeing on a same error candidate in response to implementing a suggested correction at one of the error candidates; and correcting an error in the data based on at least one of whether the at least one first error candidate is found or whether the at least one second error candidate is found. BRIEF DESCRIPTION OF DRAWINGS
[0005] Figure 1A block diagram showing an example of a system including a non-volatile storage device and a host, in accordance with some embodiments.
[0006] Figure 2 A process flow diagram illustrating an example encoding / decoding method, in accordance with some embodiments.
[0007] Figure 3 A diagram illustrating mapping in an encoding process using a folded product code (HFPC) structure, in accordance with various embodiments.
[0008] Figure 4 A diagram illustrating mapping in an encoding process using a group HFPC structure, in accordance with various embodiments.
[0009] Figure 5 A process flow diagram illustrating an example hard decoding method, in accordance with some embodiments.
[0010] Figure 6 A process flow diagram illustrating an example method for determining a candidate with a minimum probability of mis-correction, in accordance with some embodiments.
[0011] Figure 7 A process flow diagram illustrating an example method for determining a candidate, in accordance with some embodiments.
[0012] Figure 8 A diagram illustrating a decoding scenario in which two component codes indicate a consistent suggested correction, in accordance with various embodiments.
[0013] Figure 9 A diagram illustrating a decoding scenario in which two component codes indicate a consistent suggested correction following a test implementation of the suggested correction, in accordance with various embodiments.
[0014] Figure 10 A diagram illustrating a decoding scenario in which two component codes indicate a consistent suggested correction following a test implementation of the suggested correction, in accordance with various embodiments.
[0015] Figure 11 A diagram illustrating a decoding scenario in which two component codes indicate a consistent suggested correction following a test implementation of the suggested correction, in accordance with various embodiments.
[0016] Figure 12 A diagram illustrating a decoding scenario in which two component codes indicate a conflicting suggested correction, in accordance with various embodiments.
[0017] Figure 13 A process flow diagram illustrating an example method for performing look-ahead detection, in accordance with some embodiments.
[0018] Figure 14 is a process flow diagram illustrating an example method for performing hard decoding in accordance with some embodiments. DETAILED DESCRIPTION
[0019] In some arrangements, the code construction as described herein is based on simple component codes that can be efficiently implemented, such as but not limited to Bose-Chaudhuri-Hocquenghem (BCH) components. Component codes implement repetition decoding. Thus, the code construction has a more cost-effective implementation compared to conventional codes with complex and high-cost implementations, such as LDPC codes. This allows the code structure to be suitable for storage applications of flash memory devices, such as NAND flash memory devices and controllers thereof.
[0020] In some arrangements, the ECC structure uses multi-dimensional encoding. In multi-dimensional encoding, a data stream is passed through a set of multiple component encoders (implemented by or otherwise included by a controller) that together encode a full payload into a single codeword. BCH encoding can be performed by passing systematic data of a code through a shift register of a controller. Thus, as the shift register progresses, the systematic data can be modified only by passing through the component encoders of the controller. After the systematic data has completely passed through the shift register, the contents of the shift register are the redundancy of the code and are appended to the data stream. The same property applies to all component encoders across all dimensions. Multi-dimensional encoding can be obtained by product codes or symmetric product codes, and can provide improved capabilities. Such a structure creates a product of component codes to obtain a full codeword. As such, the decoding process can include repetition decoding of the component codes.
[0021] To assist in illustrating embodiments of the present invention, Figure 1 A block diagram illustrating a system including a non-volatile storage device 100 coupled to a host 101 in accordance with some embodiments is shown. In some examples, the host 101 can be a user device operated by a user. The host 101 can include an operating system (OS) configured to provision a file system and applications that use the file system. The file system communicates with the non-volatile storage device 100 (e.g., a controller 110 of the non-volatile storage device 100) via a suitable wired or wireless communication link or network to manage storage of data in the non-volatile storage device 100. In this regard, the file system of the host 101 sends data to and receives data from the non-volatile storage device 100 using a suitable interface to the communication link or network.
[0022] In some examples, the non-volatile storage device 100 is located in a data center (not shown for brevity). The data center can include one or more platforms, each of which supports one or more storage devices (such as, but not limited to, the non-volatile storage device 100). In some implementations, the storage devices within a platform are connected to a top-of-rack (TOR) switch and can communicate with each other via the TOR switch or another suitable in-platform communication mechanism. In some implementations, at least one router can facilitate communication among non-volatile storage devices in different platforms, racks, or cabinets via a suitable network connection architecture. Examples of the non-volatile storage device 100 include, but are not limited to, a solid-state drive (SSD), a non-volatile dual in-line memory module (NVDIMM), a universal flash storage (UFS), a secure digital (SD) device, and the like.
[0023] The non-volatile storage device 100 includes at least a controller 110 and a memory array 120. Other components of the non-volatile storage device 100 are not shown for brevity. The memory array 120 includes NAND flash memory devices 130a-130n. Each of the NAND flash memory devices 130a-130n includes one or more individual NAND flash dies, which are non-volatile memories (NVM) capable of holding data without power. Thus, the NAND flash memory devices 130a-130n refer to a plurality of NAND flash memory devices or dies within the flash memory device 100. Each of the NAND flash memory devices 130a-130n includes one or more dies, each of which has one or more planes. Each plane has a plurality of blocks, and each block has a plurality of pages.
[0024] While the NAND flash memory devices 130a-130n are shown as examples of the memory array 120, other examples of non-volatile memory technologies used to implement the memory array 120 include, but are not limited to, dynamic random access memory (DRAM), magnetic random access memory (MRAM), phase change memory (PCM), ferroelectric RAM (FeRAM), and the like. The ECC structures described herein can likewise be implemented on memory systems using such memory technologies and other suitable memory technologies.
[0025] Examples of the controller 110 include, but are not limited to, an SSD controller (e.g., a client SSD controller, a data center SSD controller, an enterprise SSD controller, and the like), a UFS controller, or an SD controller, and the like.
[0026] The controller 110 can combine raw data storage in multiple NAND flash memory devices 130a-130n such that those NAND flash memory devices 130a-130n are used as a single storage device. The controller 110 can include a microcontroller, buffers, an error correction system, a flash translation layer (FTL), and a flash interface module. Such functionality can be implemented in hardware, software, and firmware, or any combination thereof. In some arrangements, the software / firmware of the controller 110 can be stored in the non-volatile storage 120 or any other suitable computer-readable storage medium.
[0027] The controller 110 includes suitable processing and memory capabilities for performing the functions described herein, as well as other functions. As described, the controller 110 manages various features of the NAND flash memory devices 130a-130n, including but not limited to I / O processing, reads, writes / programs, erases, monitoring, logging, error handling, garbage collection, wear leveling, logical to physical address mapping, data protection (encryption / decryption), ECC capabilities, and so forth. Thus, the controller 110 provides visibility into the NAND flash memory devices 130a-130n.
[0028] The error correction system of the controller 110 can include or otherwise implement one or more ECC encoders and one or more ECC decoders, collectively referred to as ECC encoder / decoders 112. The ECC encoders of the ECC encoder / decoders 112 are configured to encode data (e.g., input payloads) to be programmed to the non-volatile storage devices 120 (e.g., the NAND flash memory devices 130a-130n) using the ECC structures described herein. The ECC decoders of the ECC encoder / decoders 112 are configured to decode encoded data in conjunction with read operations to correct programming errors, errors caused by reads through non-optimal thresholds, errors caused by retention / read disturb stress, and so forth. To achieve low complexity processing, the ECC encoder / decoders 112 are implemented on hardware and / or firmware of the controller 110.
[0029] In some implementations, host 101 includes an ECC encoder / decoder 102 that can use the ECC structures described herein. ECC encoder / decoder 102 is software running on host 101 and includes one or more ECC encoders and one or more ECC decoders. The ECC encoders of ECC encoder / decoder 102 are configured to encode data (e.g., input payloads) to be programmed to non-volatile storage 120 (e.g., NAND flash memory devices 130a-130n) using the ECC structures described herein. The ECC decoders of ECC encoder / decoder 102 are configured to decode encoded data in connection with read operations to correct errors. In some arrangements, one of ECC encoder / decoder 102 or ECC encoder / decoder 112 employs the ECC structures described herein. In some arrangements, one of ECC encoder / decoder 102 or ECC encoder / decoder 112 employs the hard-decoding methods described herein. In some implementations, the ECC encoders of ECC encoder / decoder 102 are configured to encode data (e.g., input payloads) to be written to multiple instances of non-volatile storage 100 with a redundancy code, including but not limited to erasure codes and RAID levels 0-6.
[0030] An encoding scheme, such as a HFPC encoding scheme, can be used to encode each of a plurality of short codewords. In some arrangements, the HFPC code structure is composed of a plurality of component codes. For example, each component code can be a BCH code. The number of component codes, n, can be determined by the correction capability and the decoding rate of each component code. For example, given a minimum distance D min of each component code, the correction capability t of each component code can be represented by:
[0031] t = (D min - 1) / 2 (1),
[0032] where D min of the linear block code is defined as the minimum Hamming distance between any pair of code vectors in the code. The number of redundancy bits, r, can be represented by:
[0033] r = Q · (D min - 1) / 2 (2),
[0034] where Q is a Galois field parameter of the BCH component code defined over GF(2 Q ). Given a code rate R and a payload length K bits, the number of component codes required can be determined by:
[0035]
[0036]
[0037] In some instances, the input payload bits (e.g., including information bits and signature bits) are arranged in a pseudo-triangular matrix form and folded encoding (e.g., folded BCH encoding) is performed for each component code. In some instances, each bit in the payload (e.g., each information bit) can be encoded by (at least) two component codes (also referred to as "code components"), and each component code intersects with all other component codes. That is, for a component code that encodes an information bit, the encoding process is performed such that the systematic bits of each component code are also encoded by all other component codes. The component codes together use the component codes to provide encoding for each information bit.
[0038] For example, Figure 2 is a process flow diagram illustrating an example of an encoding method 200 according to some embodiments. Referring to Figures 1-2 , the method 200 encodes an input payload to obtain a corresponding ECC. The input payload includes information bits.
[0039] At 210, one or more encoders of the ECC encoder / decoder 102 or 112 generate a signature for the input payload. The signature can be used during decoding to check whether the decoding is successful. In some instances, the signature can be generated by passing the information bits through a hash function. In some instances, the signature includes a cyclic redundancy check sum (CRC) generated from the information bits. In some instances, the signature can include other indications generated from the input payload in addition to the CRC. The CRC can be generated to have a specified length.
[0040] The length of the CRC can be determined based on factors such as, but not limited to, a target error detection probability of the codeword decoding, an error detection probability of the decoding process (alone, without CRC), and the like. The error detection probability of the codeword decoding refers to a probability of signaling "decoding successful" without considering that there is a decoding error. The error detection probability of the decoding process (alone, without CRC) refers to a probability of signaling "decoding failed" without considering that there is no decoding error. A certain level of decoding confidence can be provided using component code syndrome zeros, which in some cases can be sufficient to allow a zero-length CRC. Alternatively, the CRC can be used to combine the error detection decisions. For example, a longer length of CRC corresponds to a low error detection probability of the codeword decoding. On the other hand, a shorter length of CRC corresponds to a high target error detection probability of the codeword decoding.
[0041] At 220, one or more encoders of the ECC encoder / decoder 102 or 112 map each information bit of the input payload to two or more component codes. In some examples, the bits corresponding to the signature (e.g., CRC bits) can also be encoded (e.g., in arrangements where the ECC is regular HFPC, each CRC bit can be mapped to one or more component codes). That is, one or more encoders of the ECC encoder / decoder 102 or 112 implement a mapping function that maps each information bit of the input payload with the corresponding component code of the ECC. In arrangements where the ECC is regular HFPC (e.g., Figure 3 ) each information bit can be mapped to two component codes (e.g., i1and i2). In arrangements where the ECC is irregular HFPC, at least one information bit can be mapped to three or more component codes, thus resulting in an irregular encoding process.
[0042] In some examples, blocks 210 and 220 can be implemented simultaneously or in parallel. In other examples, blocks 210 and 220 can be implemented sequentially in any suitable order. The ECC code structure is composed of a plurality of component codes. For example, each component code can be a BCH code.
[0043] At 230, one or more encoders of the ECC encoder / decoder 102 or 112 update the encoded syndromes for each component code with the additional information bits. Thus, depending on the mapping function performed at 220, each component code encodes a portion of the input payload. The set of redundancy bits corresponding to the component codes is generated after encoding all payload bits (including information bits and signature bits) at each block 210-230.
[0044] At 240, in some arrangements, one or more encoders of the ECC encoder / decoder 102 or 112 encode the redundancy bits (in an additional encoding process). That is, the redundancy bits can be mapped to additional component codes. For example, the encoding can be obtained by a similar set of component codes. For example, for higher code rates, the set of component codes can be a set of sets that is smaller than the set of payload encoding sets. Each redundancy encoding component can receive a separate redundancy input bit for encoding. As such, parity check codes are generated.
[0045] In some examples where irregular codes are involved, 240 can not be performed (e.g., redundancy encoding can not be performed), such that the redundancy bits have a degree of protection, while the systematic information bits have a degree of two protection. Irregularity can also be obtained by performing the process of HFPC encoding with component codes having different correction capabilities and / or different lengths.
[0046] Figure 3 is a diagram illustrating mapping 300 in an encoding process using HFPC structure, in accordance with various embodiments. Reference is made to FIG. 2 for an overview of the encoding process.Figures 1-3 The mapping 300 corresponds to a HFPC encoding scheme and is an example implementation of the block 220. The controller 110 (e.g., one or more ECC encoders of the ECC encoder / decoder 112) or the host 101 (e.g., one or more ECC encoders of the ECC encoder / decoder 102) can include or otherwise implement a HFPC interleaver configured to organize (e.g., insert or map) the input bits 301 into the form of a pseudo-triangular matrix 310. In some examples, the input bits 301 include an input payload 302 and a signature bit D1303. The input payload 302 includes information bits. In some examples, the input payload 302 includes information bits and redundant bits introduced by the host 101 for RAID or erasure coding (e.g., by one or more ECC encoders of the ECC encoder / decoder 102). As described, an example of D1303 is an additional CRC bit. The bits of D1303 can also be referred to as "outer parity bits" on the condition that the CRC encoding can be considered an outer encoding process. The mapping from the input bits 301 to the pseudo-triangular matrix 310 is maintained by the controller 110.
[0047] As shown, the pseudo-triangular matrix 310 has an upper triangular form with rows 321-325 (where rows between rows 323 and 324 are omitted for clarity) and columns 331-335 (where columns between columns 333 and 334 are omitted for clarity). The pseudo-triangular matrix 310 is shown as having a plurality of blocks. Each block in the pseudo-triangular matrix 310 includes or otherwise represents two or more bits in the input bits 301. The number of input bits per block can be predetermined and equal for all blocks of the pseudo-triangular matrix 310. Thus, HFPC is obtained by allowing any pair of component codes to encode more than one bit (e.g., intersect at more than one bit). Conventionally, any pair of component HFPCs intersect at only one common (intersecting) bit. The disclosed implementations allow for two or more common bits at the intersection point of any pair of component codes. The pseudo-triangular matrix 310 is "pseudo" in the condition that each row has two or more bits (e.g., blocks) more than the row immediately below that row, and each column has two or more bits (e.g., blocks) more than the column immediately to the left of that column. Thus, each row or column of the pseudo-triangular matrix differs from an adjacent row or column by two or more bits.
[0048] In some implementations, the input bits 301 are mapped to the blocks in the pseudo-triangular matrix 310 consecutively (in any suitable order). For example, the rows 321-325 (in that order or in reverse order) can be filled consecutively with the input bits 301 from leftmost block of the row to the rightmost block of the row (or vice versa) block by block. In another example, the columns 331-335 (in that order or in reverse order) can be filled consecutively with the input bits 301 from the topmost block of the row to the bottommost block of the row (or vice versa) block by block. In some implementations, the input bits 301 are pseudo-randomly mapped to the pseudo-triangular matrix 310. In other implementations, the input bits 301 can be mapped to the pseudo-triangular matrix 310 using another suitable mapping mechanism. In one arrangement, the mapping is a one-to-one mapping, where each bit in the input bits 301 is mapped to one bit in the pseudo-triangular matrix 310, and the total number of bits in the pseudo-triangular matrix 310 is equal to the number of input bits 301. In another arrangement, the mapping can be one-to-many, where each bit in the input bits 301 is mapped to one or more bits in the pseudo-triangular matrix 310, and the total number of bits in the pseudo-triangular matrix 310 is greater than the number of input bits 301.
[0049] As shown, the upper triangular form has the same number of columns and the same number of rows. In the upper triangular form, the row 321 has the most bits of all the rows in the pseudo-triangular matrix 310. The row 322 has one less block than the row 321. The row 323 has one less block than the row 322, and so on. The row 324 has two blocks, and the row 325 is the lowermost row, having one block. In other words, any row (except the row 321) in the pseudo-triangular matrix 310 has one less block than the row immediately above it. Similarly, in the upper triangular form, the column 331 is the leftmost column, having one block. The column 332 has one more block than the column 331. The column 333 has one more block than the column 332, and so on. The column 335 is the rightmost column, having the most blocks of the columns in the pseudo-triangular matrix 310. In other words, any column (except the column 335) in the pseudo-triangular matrix 310 has one less block than the column immediately to the right of it.
[0050] Organizing or mapping the input bits 301, which include the bits of the input payload 302 and the signature bits D1 303, in an upper triangular form of the pseudo-triangular matrix 310 allows each component code to be associated in a described manner with bits in the row and column of the same size or nearly the same size. For example, R1 341 represents a redundancy bit corresponding to the first component code. The R1 341 redundancy bit is obtained by encoding (e.g., folded component encoding) the input bits 301 in the first row (e.g., the bits in row 321). The R2 342 redundancy bit is obtained by encoding (e.g., via folded component encoding) the input bits 301 in the first column (e.g., the bits in column 331) and the second row (e.g., the bits in row 322). The total number of bits encoded by R2 342 (e.g., the bits in column 331 plus the bits in row 322) is the same as the total number of bits encoded by R1 341 (e.g., the bits in row 321). The R3 343 redundancy bit is obtained by encoding (e.g., via folded component encoding) the input bits 301 in the second column (e.g., the bits in column 332) and the third row (e.g., the bits in row 323). The total number of bits encoded by R3 343 (e.g., the bits in column 332 plus the bits in row 323) is the same as the total number of bits encoded by R2 342 (and the total number of bits encoded by R1 341). This process continues to obtain the last redundancy bit Rn 345, which encodes (e.g., via folded component encoding) the input bits 301 in the last column (e.g., the bits in column 335). Thus, each component code encodes a row and a column in the pseudo-triangular matrix 310, providing folded component encoding. An example of folded component encoding is folded BCH encoding.
[0051] In other words, according to the mapping 300, the input bits 301 are mapped to component codes of the ECC and encoded as mapped component codes. For example, the encoding process organizes or maps the input bits 301 into a matrix (e.g., a pseudo-triangular matrix form) and performs folded BCH encoding for each component code. Each of the input bits 301 is encoded by two component codes. Each component code intersects with all other component codes. For a component code that encodes an input bit 301, the encoding process is performed such that the systematic bits of each component code are also encoded by all other component codes. The input bits encoded by any one of the component codes are also encoded by each other component code in the ECC in a non-overlapping manner.
[0052] For example, the bits encoded by the component code corresponding to R3 343 are also encoded by the other component codes corresponding to R1 341, R2 342, and R4-Rn 345. The bit at the intersection of row 321 and column 332 is also encoded by the component code corresponding to R1 341; the bit at the intersection of row 322 and column 332 is also encoded by the component code corresponding to R2 342; the bit at the intersection of row 323 and column 334 is also encoded by the component code corresponding to Rn-1 344; and the bit at the intersection of row 323 and column 335 is also encoded by the component code corresponding to Rn 345. Each block of bits encoded by any one of the component codes (e.g., the component code corresponding to R3 343) is encoded by that component code (e.g., the component code corresponding to R3 343) and by only one other of the component codes, thus in a non-overlapping manner. Thus, each component code is dependent on all the other component codes. The component codes together provide the encoding of each of the input bits 301 using two component codes. The component codes have the same code rate, conditional on each component code encoding the same number of bits.
[0053] In some implementations, the parity check bits can be generated via parity check encoding. For example, folded parity check encoding can be used to encode at least a portion of each of R1 341-Rn 345 into another component code (e.g., folded product code 350, which is a set of packets). The folded product code 350 is composed of parity check bits. This method of generating parity check bits can be efficient for a simple hardware encoding implementation of HFPC, as it can be repeatedly decoded using various hard or soft decoding methods.
[0054] In some examples, to provide an efficient structure, less than all of each of R1 341-Rn 345 is encoded to obtain the folded product code 350. This is because only the encoded version of the input bits 301 (e.g., the input payload 302) needs to be decoded, and decoding all of the redundancy bits R1 341-Rn 345 can lengthen the decoding time.
[0055] In some arrangements, the number of component codes used to encode the redundancy bits can vary depending on the code rate required for the redundancy bits and the intersection size. In some arrangements, the redundancy bits can not be encoded at all, resulting in an irregular degree of protection for the bits within the codeword. In some cases, an irregular degree of protection can be useful in its waterfall capability. In some arrangements, by encoding the folded product code with an irregularity, the degree of protection for some of the information bits can be greater than two. For example, in addition to encoding the reference Figure 3The described regularity can be applied beyond folded product codes, by applying an additional encoding process to some of the input bits 301, encoded with different component code sets. In some examples, encoding some of the input bits 301 with more than two component codes, while encoding other bits 301 with two component codes, causes irregularity in the encoding process, which results in unequal error protection of the bits within a codeword and improves correction capability (as applied to repetition decoding).
[0056] The redundant bits R1341-Rn-m345, produced by the described HFPC encoding process, can be encoded by another separate component code set to encode all of these redundant bits or a subset of these redundant bits with another component code set. This forms a folded product code encoding on the redundant bits R1341-Rn-m345, which together with the information bit encoding results in a low complexity encoding process. Figure 3 The redundant bits R1341-Rn-m345, produced by the described HFPC encoding process, can be encoded by another separate component code set to encode all of these redundant bits or a subset of these redundant bits with another component code set. This forms a folded product code encoding on the redundant bits R1341-Rn-m345, which together with the information bit encoding results in a low complexity encoding process.
[0057] As shown, during decoding in the ECC structure corresponding to the mapping 300, the bits of each component code depend on the bits of another component code.
[0058] For a regular semi-product code, each pair of component codes has only one common (intersecting) information bit. In some implementations, HFPC is obtained by using more than one information bit encoded by each pair of component codes. Thus, there can be two or more common (intersecting) bits for each pair of component codes.
[0059] In some implementations, the redundant bits produced from the HFPC encoding process described herein are encoded by a separate component code set. For example, a separate component code set encodes all of the redundant bits or a subset of the redundant bits to form a folded product code encoding on the redundant bits, which together with the information bit encoding results in a low complexity encoding process.
[0060] In some implementations, multiple component codes can be grouped together and act as a single element according to the HFPC structure, such that there is no dependency among the bits of the component codes within each group of component codes. Such an encoding scheme reduces the dependency of the HFPC structure and enables faster decoding implementation in hardware, under the condition that the encoding scheme is a low complexity encoding and decoding code structure obtained by defining groups, where each group contains non-dependent components.
[0061] In this regard, Figure 4 is a diagram illustrating a mapping 400 in an encoding process using a group HFPC structure, according to various implementations. Reference is made to Figures 1-4The mapping 400 corresponds to a group-HFPC encoding scheme and is an example implementation of the block 220. The HFPC interleaver of the controller 110 (e.g., one or more ECC encoders of the ECC encoder / decoder 112) or the host 101 (e.g., one or more ECC encoders of the ECC encoder / decoder 102) is configured to organize (e.g., insert) the input bits 401 into the form of the pseudo-triangular matrix 410. In some examples, the input bits 401 include an input payload 402 and a signature bit D1 403. The input payload 402 includes information bits. As described, an example of D1 403 is an extra CRC bit (an outer parity bit). The mapping from the input bits 401 to the pseudo-triangular matrix 410 is maintained by the controller 110.
[0062] As shown, the pseudo-triangular matrix 410 has an upper triangular form with rows 421-436 (where rows between rows 432 and 433 are omitted for clarity) and columns 441-456 (where columns between columns 452 and 453 are omitted for clarity). The pseudo-triangular matrix 410 is shown to have a plurality of blocks. Each block in the pseudo-triangular matrix 410 includes or otherwise represents two or more bits in the input bits 401. The number of input bits per block can be predetermined and equal for all blocks of the pseudo-triangular matrix 410. The disclosed implementations allow for two or more common bits at the intersection of any pair of component codes.
[0063] In some implementations, the input bits 401 are mapped consecutively (in any suitable order) to the blocks in the pseudo-triangular matrix 410. For example, the rows 421-436 (in that order or in reverse order) can be filled consecutively with the input bits 401 from leftmost block of a row to the rightmost block of a row (or vice versa) block-by-block. In another example, the columns 441-456 (in that order or in reverse order) can be filled consecutively with the input bits 401 from the topmost block of a column to the bottommost block of a column (or vice versa) block-by-block. In some implementations, the input bits 401 are pseudo-randomly mapped to the pseudo-triangular matrix 410. In other implementations, the input bits 401 can be mapped to the pseudo-triangular matrix 410 using another suitable mapping mechanism.
[0064] The blocks, rows, and columns in the pseudo-triangular matrix 410 can be grouped together. For example, the pseudo-triangular matrix 410 includes a first group of columns 441-444, a second group of columns 445-448, a third group of columns 449-452, and another group of columns 453-456. The pseudo-triangular matrix 410 includes a first group of rows 421-424, a second group of rows 425-428, a third group of rows 429-432, and another group of rows 433-436. Thus, the HFPC structure is divided into groups of 4 component codes. Each 4 component codes is encoded according to the HFPC guide. Although groups of 4 component codes (e.g., 4 rows / columns) are shown in Figure 4 any number (e.g., 2, 3, 6, 8, 10, 12, 16, etc.) of component codes can be grouped together.
[0065] As shown, the upper triangular form has the same number of columns and the same number of rows. The rows (e.g., rows 421-424) or columns (e.g., columns 441-444) in the same group of component codes have the same number of blocks and thus the same number of bits. In the upper triangular form, rows 421-424 contain the most bits of all the rows in the pseudo-triangular matrix 410. Each of rows 425-428 has one block group (4 blocks, corresponding to the group of columns 441-444) less than any of rows 421-424. Each of rows 429-432 has one block group (4 blocks, corresponding to the group of columns 445-448) less than any of rows 425-428, and so on. Each of rows 433-436 is the lowermost row, having a block group (e.g., 4 blocks). In other words, any row in the pseudo-triangular matrix 410 (except rows 421-424) has 4 blocks less than the row of the group immediately above. Similarly, in the upper triangular form, each of columns 441-444 is one of the leftmost columns, having a block group (e.g., 4 blocks). Each of columns 445-448 has one block group (4 blocks, corresponding to the group of rows 425-428) more than any of columns 441-444. Each of columns 449-452 has one block group (4 blocks, corresponding to the group of rows 429-432) more than any of columns 445-448, and so on. Each of columns 453-456 is the rightmost column, having the largest number of blocks. In other words, any column in the pseudo-triangular matrix 410 (except columns 453-456) has 4 blocks less than the column of the group immediately to the right.
[0066] Organizing or mapping the input bits 401 in the form of the pseudo-triangular matrix 410 allows each component code to be associated with bits in the rows and columns having the same size or nearly the same size in the described manner. The component codes within the same group encode separate sets of input bits 401 and are independent of each other.
[0067] R1 461-R4 464 are redundancy bits determined based on the same component code group. R1 461 represents a redundancy bit corresponding to the first component code and obtained by encoding (e.g., folded component encoding) the input bits 401 into the bits in the first row (e.g., the bits in row 421). R2 462, R3 463, and R4 464 represent redundancy bits corresponding to the additional component codes and obtained by encoding (e.g., folded component encoding) the input bits 401 into the bits in rows 422, 423, and 423, respectively. The bits used to determine each of R1 461-R4 464 do not overlap, and thus R1 461-R4 464 are determined independently.
[0068] R5 465, R6 466, R7 467, and R8 468 represent redundancy bits corresponding to the additional component codes and obtained by encoding (e.g., folded component encoding) the input bits 401 into the bits in column 444 and row 425, column 443 and row 426, column 442 and row 427, and column 441 and row 428, respectively. The bits used to determine each of R5 465-R8 468 do not overlap, and thus R5 465-R8 468 are determined independently.
[0069] R9 469, R10 470, R11 471, and R12 472 represent redundancy bits corresponding to the additional component codes and obtained by encoding (e.g., folded component encoding) the input bits 401 into the bits in column 448 and row 429, column 447 and row 430, column 446 and row 431, and column 445 and row 432, respectively. The bits used to determine each of R9 469-R12 472 do not overlap, and thus R9 469-R12 472 are determined independently.
[0070] This process continues until Rn-3 473, Rn-2 474, Rn-1 475, and Rn 476 are determined. Rn-3 473, Rn-2 474, Rn-1 475, and Rn 476 represent redundancy bits corresponding to the additional component codes and obtained by encoding (e.g., folded component encoding) the input bits 401 into the bits in column 456, column 455, column 454, and column 453, respectively. The bits used to determine each of Rn-3 473, Rn-2 474, Rn-1 475, and Rn 476 do not overlap, and thus Rn-3 473, Rn-2 474, Rn-1 475, and Rn 476 are determined independently. An example of folded component encoding is folded BCH encoding.
[0071] In the special case where the component codes are divided into two groups of independent component codes, the resulting coding scheme degenerates to a folded product code.
[0072] According to the map 400, the input bits 401 are mapped to component codes of the ECC and encoded as mapped component codes. For example, the encoding process organizes or maps the input bits 401 in a matrix (e.g., in a pseudo-triangular matrix form) and performs folded BCH encoding for each component code. Each of the input bits 401 is encoded by two component codes in different component code groups. Thus, any component code intersects with all other component codes in the same group as the component code. For the component codes that encode the input bits 401, the encoding process is performed such that the systematic bits of each component code are also encoded by all other component codes belonging to different groups, where dependencies within the component code groups are eliminated. The input bits encoded by a given component code among the component codes are also encoded in a non-overlapping manner by each other component code that is not in the same group as the component code. For example, the bits encoded by the component code corresponding to R9 469 are also encoded by the other component codes corresponding to R1 461-R8 468 and R11-Rn 476 that are not in the group of the component code corresponding to R9 469. Each block of bits encoded by any one of the component codes (e.g., the component code corresponding to R9 469) is encoded by the component code (e.g., the component code corresponding to R9 469) and only one other component code among the component codes, thus in a non-overlapping manner. As such, each component code is dependent on all other component codes that are not in the same group. The component codes together provide the encoding of each input bit 401 using two component codes.
[0073] In some embodiments, the parity bits can be generated via parity encoding. For example, folded parity encoding can be used to encode at least a portion of each of R1 461-Rn 476 into another component code (e.g., folded product code 480, which is a set of packets). The folded product code 480 (e.g., with Rp1-Rp3) is a parity bit. This method of generating parity bits can be efficient for a simple hardware encoding implementation of HFPC, as the method can be repeatedly decoded using various hard or soft decoding methods.
[0074] With respect to hard decoding HFPC as described herein, a HFPC repetition hard decoder can be employed. In a hard decoding process performed by the HFPC repetition hard decoder, a base sub-repetition includes attempting to decode all component codes. Hard decoding, also referred to as hard decision decoding, is a process that operates on code bits that can have a fixed set of values, such as 1 or 0 for binary codes. In contrast, soft decoding or soft decision decoding is a process that operates on code bits that can have a range of values between two values, as well as an indication of the reliability or probability that the value is correct. Component codes can include BCH codes that correct a number of errors, such as t < 4 per BCH component code, and thus decoding of each component code can be efficiently implemented in hardware, while obtaining high decoding reliability via repetition decoding. In NAND flash memory devices, read performance depends on decoder latency. Thus, high read performance requires fast decoding. When the number of errors is not too high, it is often possible to use repetition fast decoding only without advanced and / or intersection decoding to complete successful decoding at very low latency.
[0075] In response to determining that the inner decoding in a sub-repetition is not successful, other types of multi-dimensional decoding can be attempted. Some advanced types of decoding for BCH components include (t-1) limited correction per BCH component. This stage is also referred to as a "(t-1) decoding" stage and aims to minimize error correction by performing BCH correction with a lower probability of error correction. It is used in conjunction with hard decoding, where repetition hard decoding includes, for example, bounded distance decoding for each component code of the BCH component.
[0076] In response to determining that the BCH component has a decoding capability of t < 4, a direct solution from the correction sub can be applied, which enables efficient hardware implementation with high decoding throughput.
[0077] The simplified method includes performing a (t-1) decoding stage, where correction of up to one error is less than the BCH code correction capability t. For a BCH component with D min The correction capability is shown in expression (1) for a BCH component with D
[0078]
[0079] where N is the codeword length of the component code (including parity bits). Thus, limiting the number of corrections to m per component code can change each repetition in a way that the probability of error correction increases gradually.
[0080] The number of errors is limited to t-1, and multi-dimensional repetition decoding can be performed for each of M0>0 and M1>0, respectively. Although M0=0, no (t-1) decoding repetitions are performed. Such a configuration is efficient for fast decoding.
[0081] In response to determining that decoding so far has not been successful, other more advanced methods can be employed. In the case of using multi-dimensional codes, each input bit is encoded by multiple component codes. Thus, intersection point decoding can be useful at this point. In an example, in response to determining that there are still some unsolved decoder components, and there is no further progress of the bounded distance repetition hard repetition decoding, intersection point decoding can be employed.
[0082] An unsolved intersection point bit is defined as an information bit belonging to a distinct component code that is all unsolved (e.g., has misCorrection = 1). The more component codes that are used, the smaller the set of intersection points between the component codes. In the HFPC codes disclosed herein, the intersection point size is smallest by construction on regular codes in the case that each component bit is cross-encoded by all other component codes. Such HFPC properties result in the smallest intersection point size and enable low complexity enumeration for intersection point decoding. As described, the intersection point bit set length can vary based on the payload size of the component codes on the same dimension.
[0083] In intersection point decoding, a bit set is mapped (obtained by the intersection of component codes with non-zero (unsolved) correctors). If needed, the intersection point bit set list size is limited. The number of bits for enumeration can be determined. The enumeration complexity is bounded by where N b is the number of bits that are decoded simultaneously flipped per intersection point, and L b is the number of bits in a single bit set intersection point. For each selected intersection point set (another N b bits flipped) enumerated on the intersection point bits, decoding of the corresponding component code on the selected dimension is attempted. This enables correction of t+Nb errors for a single component code. If misCorrection = 0 after decoding for more than a certain threshold (zero threshold with respect to the number of component codes), the negation of the N b bits is accepted as a valid solution (of the intersection point set).
[0084] In general, after the intersection point decoding stream makes progress by providing valid solution candidates for decoding, the decoding stream of the sub-repetition can continue, and make greater decoding progress on the repetition decoding sub-repetition.
[0085] The arrangements disclosed herein relate to a hard-decoding method that improves endurance and equalizes read performance of NAND flash memory devices 130a-130n by implementing correction of high performance to high raw BER. Such a hard-decoding method is applicable to a general product code, where a plurality of sub-code decodings are employed with repetition. Furthermore, such a hard-decoding method has low complexity while being able to reduce the miss-correction probability of the sub-codes within the repetition decoding by applying a look-ahead detection method and thus obtaining higher decoding capability. In some examples, the hard-decoding method can be implemented on the controller 110 (e.g., executed by hardware and / or firmware of the controller 110, including but not limited to the ECC encoder / decoder 112). In some examples, the hard-decoding method can be implemented on the host 101 (e.g., executed by software of the host 101, including but not limited to the ECC encoder / decoder 102).
[0086] Figure 5 is a process flow diagram illustrating an example hard-decoding method 500 according to some embodiments. With reference to Figures 1-5 , the hard-decoding method 500 can be performed by a decoder of the ECC encoder / decoder 102 or a decoder of the ECC encoder / decoder 112 (referred to as the “decoder”). The hard-decoding method 500 employs a look-ahead algorithm and includes safe look-ahead (SLA) repetition. For example, the hard-decoding method 500 includes a plurality of “safe+sub-repetitions” that include SLA operations. The SLA operations allow the hard-decoding method 500 to achieve reliability improvement. In response to determining that decoding progress is not achieved through a safe+sub-repetition, the decoder can perform an intersection decoding. In response to determining that decoding progress is not achieved through the intersection decoding, the decoder can perform a one-degree decoding.
[0087] At 505, the decoder performs a safe+sub-repetition. The safe+sub-repetition corresponds to a method by which a candidate with a minimum error-correction probability is selected. The safe+sub-repetition includes various SLA operations. Examples of the safe+sub-repetition are disclosed in greater detail in Figure 6 and 7 In each safe+sub-repetition, the decoder attempts to independently decode all of the sub-codes.
[0088] At 510, the decoder determines whether decoding with the safe+sub-repetition is successful. The decoder can determine whether decoding is successful by checking the signature bits. In response to determining that decoding is successful (510: YES), the method 500 ends.
[0089] On the other hand, in response to determining that decoding was not successful (510:NO), the decoder determines at 515 whether decoding progress has been made. Decoding progress is considered to have been made in response to determining that at least one additional or different component code has been successfully decoded and solved in the current repetition 505, or in response to determining that at least one additional or different error candidate has been produced in the current repetition 505. In response to determining that decoding progress has been made (515:YES), the method 500 returns to 505 for the next repetition.
[0090] On the other hand, in response to determining that decoding progress has not been made (515:NO), the decoder performs intersection decoding at 520. Unresolved intersection bits are defined as information bits belonging to a distinct component code that are all unresolved (e.g., have misCorrection = 1). The more component codes that are used, the smaller the set of intersection bits between the component codes. In the HFPC codes disclosed herein, the intersection size is, by construction, minimal over regular codes, under the condition that each component bit is cross-coded by all other component codes. Such HFPC properties result in minimal intersection size and enable low complexity enumeration for intersection decoding. As described, the intersection bit set length can vary based on the payload size of the component codes in the same dimension.
[0091] In intersection decoding, first, the bit set is mapped (obtained by the intersection of component codes with non-zero (unresolved) corrections). Second, if needed, the intersection bit set list size is limited. Third, the number of bits for enumeration can be determined. The enumeration complexity is bounded by where N b is the number of decoded simultaneous flips per intersection, and L b is the number of bits in a single bit set intersection. Fourth, for each selected intersection set (another N b bits flipped) enumerated on the intersection bits, the decoding of the corresponding component code is attempted on the selected dimension. This enables correction of t+Nb errors for a single component code. If misCorrection = 0 after decoding for more than a certain threshold (zero threshold with respect to the number of component codes), the negation of the N b bits is accepted as the valid solution (of the intersection set).
[0092] At 525, the decoder determines whether decoding using intersection decoding was successful. In response to determining that decoding was successful (525:YES), the method 500 ends.
[0093] On the other hand, in response to determining that decoding was not successful (525:NO), the decoder determines at 530 whether decoding progress has been made using intersection decoding. In response to determining that decoding progress has been made (530:YES), the method 500 returns to 505 for the next repetition.
[0094] On the other hand, in response to determining that decoding progress has not been made (530:NO), at 535, one-time bit decoding is performed in some instances in which the irregular code is used. For example, the decoder or a different specialized decoder attempts to decode bits that have one-time encoding protection (e.g., redundancy bits).
[0095] At 540, the decoder determines whether decoding using one-time bit decoding was successful. In response to determining that decoding was successful (540:YES), the method 500 ends.
[0096] On the other hand, in response to determining that decoding was not successful (540:NO), the decoder determines whether decoding progress has been made using one-time bit decoding at 545. In response to determining that decoding progress has been made (545:YES), the method 500 returns to 505 for the next repetition.
[0097] On the other hand, in response to determining that decoding progress has not been made (545:NO), the decoder proceeds to next-stage decoding at 550, including but not limited to t-1 decoding, post-decoding, base decoding, and fast decoding. In the event that the next stage fails, the method 500 ends.
[0098] Figure 6 is a process flow diagram illustrating an example method 600 for determining a candidate with a minimum probability of error correction, in accordance with some embodiments. Reference is made to Figures 1-6 , the method 600 (also referred to as safe+ sub-repetition) is an example implementation of 505. The method 600 includes various test stages (referred to as "safe+ reject flow") applied on the decoding candidates to obtain a decision of which candidate to accept. In Figure 7 The safe+ reject flow is described in more detail in
[0099] For example, at 605, the decoder performs the safe+ reject flow. Examples of the safe+ reject flow include, but are not limited to, the method 700. In performing the safe+ reject flow, each component code is configured to produce a number of candidate errors that is less than or equal to t-1 (which is one less than the BCH code correction capability t). At 610, the decoder determines whether at least one candidate is found in 605. In response to determining that at least one candidate is found (610:YES) (e.g., 710:YES, 720:YES, 735:YES), the internal safe+ sub-repetition ends and the method 500 proceeds to the next safe+ sub-repetition (method 600) at 505. Those candidates are accepted (strong accept resolution).
[0100] On the other hand, in response to determining that at least one candidate was not found (610:NO) (e.g., 835:NO), the decoder performs a safe+ candidate reduction at 615. In instances where all candidate errors generated at 605 are marked as rejected (e.g., no strong accept solution was found in method 700), the decoder reevaluates each error candidate by assessing the number of aggressors for each error candidate. An aggressor is defined as component code that has a solution that changes the target component code. In other words, an aggressor is component code that produces at least one error candidate that is different from any of the error candidates produced by the target component code for the same bits. If the number of aggressors for the target component code exceeds a predetermined threshold, then as a reduction, the error candidates for the target component code are removed from the list of error candidates, provided that those error candidates can be mis-corrections.
[0101] Thus, in response to finding any reductions (620:YES), at 625, the decoder again performs a safe+ reject stream, where each component code is configured to produce a number of candidate errors that is less than or equal to t-1, and the reductions apply. That is, the decoder again performs method 700 without considering that one or more component codes have a number of aggressors that is higher than a predetermined threshold, thus removing from consideration error candidates that can be mis-corrections. On the other hand, in response to determining that no reductions were found (620:NO), method 600 proceeds to 635.
[0102] At 630, the decoder determines whether at least one candidate was found. In response to determining that at least one candidate was found (630:YES) (e.g., 710:YES, 720:YES, 735:YES), the inner safe+ sub-repetition ends and method 500 proceeds to the next safe+ sub-repetition (method 600) at 505. Those candidates are accepted (strong accept solutions).
[0103] On the other hand, in response to determining that at least one candidate was not found (630:NO) (e.g., 835:NO), the decoder performs a safe+ reject stream, where each component code is configured to produce a number of candidate errors that is less than or equal to t. That is, the decoder again performs method 700 with each component code producing up to the decoding capability t. While not as reliable as 605 and 625, performing method 700 with each component code producing up to the decoding capability t allows for more error candidates to be determined.
[0104] At 640, the decoder determines whether at least one candidate was found. In response to determining that at least one candidate was found (640:YES) (e.g., 710:YES, 720:YES, 735:YES), the inner safe+ sub-repetition ends and method 500 proceeds to the next safe+ sub-repetition (method 600) at 505. Those candidates are accepted (strong accept solutions).
[0105] On the other hand, in response to determining that at least one candidate is not found (640:No) (e.g., 835:No), the decoder performs a back-off evaluation and applies an aggressor at 645. In instances where no candidate is found (strong accept solution) (e.g., 610:No, 630:No, and 640:No), the decoder rejects all error candidates in the error candidate list. Additionally, the decoder can further evaluate or, in some cases, restore an error candidate that has been previously implemented (e.g., in a previous safety+ sub-repetition). For example, the decoder can calculate the number of aggressor component codes for each component code of a previously corrected error candidate. In response to determining that the number of aggressor component codes exceeds a predetermined threshold, the decoder flags the corresponding component code for back-off and restores the error candidate determined using the component code. The decoder can implement the suggested correction of the error candidate determined using the aggressor component code, thus achieving further progress in the safety+ stream. After 645, the inner safety+ sub-repetition ends and the method 500 proceeds to the next safety+ sub-repetition (method 600) at 505. Alternatively, after 645, the inner safety+ sub-repetition ends and the method 500 proceeds to the intersection decoding at 520.
[0106] Figure 7 is a process flow diagram illustrating an example method for determining candidates according to some embodiments. Referring to Figures 1-7 , the method 700 can be performed by a decoder of the ECC encoder / decoder 102 or a decoder of the ECC encoder / decoder 112 (referred to as a "decoder"). The method 700 (also referred to as a safety+ reject stream) is an example implementation of 605, 625, and 635. Generally, in the method 700, different thresholds or different input candidates are applied to obtain a useful candidate output list for implementing the embodiments.
[0107] At 705, the decoder performs a safety check. In the check stage, the decoder attempts to solve each of the component codes individually and saves any suggested candidate or solution to memory. At this point, no suggested solution is actually implemented, as implementing any suggested solution can affect the solution of other component codes. In the check stage, each component code is tested for a new solution. The suggested solution is denoted as:
[0108] {x i} i∈G (6),
[0109] where G is a group of indices each corresponding to a respective component code of the component codes that has a valid solution during the check. Each x i is an error vector candidate produced by solving during the check stage.
[0110] Once error candidates are ready, the decoder determines whether a strong accepted solution is found at 710. For example, the decoder determines whether a strong accepted solution is found by searching the same error candidates. In some instances, under the condition that all or almost all codeword bits are protected by two component codes, if the two component codes have a common / same error candidate, then it is expected that the solution corresponding to the common error candidate has a higher reliability of being the true solution (and not a mis-correction). Thus, such a solution is considered a "strong accepted" solution. The decoder labels the decoded component codes with a high reliability label (e.g., a "forced" status).
[0111] Provided Figure 8 to illustrate a strong accepted solution determined during the safety check at 705. Figure 8 is a diagram illustrating a decoding scenario in which two component codes indicate a consistent suggested correction, in accordance with various embodiments. Reference is made to Figures 1-8 , Figure 8 The ECC structure 800 shown in FIG. 8 can be implemented based on mappings such as, but not limited to, mappings 300 and 400, and is an HFPC. That is, the ECC structure 800 can be a result of mapping input bits (e.g., input bits 301 and 401) to a pseudo-triangular matrix (e.g., pseudo-triangular matrices 310 and 410). Input bits are encoded and decoded using interdependent component codes based on the ECC structure 800 similar to that described with respect to mappings 300 and 400. For example, input bits in row 811 are encoded / decoded using component code CI 810. Input bits in column 821 and row 822 are encoded / decoded using component code Ci 820. Input bits in column 831 and row 832 are encoded / decoded using component code Cj 830. Input bits in column 841 and row 842 are encoded / decoded using component code Ck 840. Input bits in column 851 and row 852 are encoded / decoded using component code Cm 850. Each of columns 821, 831, 841, and 851 is a column in the pseudo-triangular matrix 310 or 410. Each of rows 811, 822, 832, 842, and 852 is a row in the pseudo-triangular matrix 310 or 410. The ECC structure 800 is thus a pseudo-triangular matrix with m components. Each of CI 810, Ci 820, Cj 830, Ck 840, and Cm 850 can be a BCH component code. Other component codes (and rows and columns associated therewith) are omitted for clarity, except for those to which suggested corrections are directed.
[0112] As used herein, a proposed correction (e.g., a proposed bit flip) for one or more bits, referred to as an error candidate or error location, is illustratively indicated as a given block in the ECC structure 800. Each block contains a plurality of bits, and the proposed correction corresponds to one of the plurality of bits. Error detection using Ci 820 results in a proposed correction at blocks 823, 824, and 825. Error detection using Cm 850 results in a proposed correction at blocks 825, 833, and 843. Error detection using each component code is performed independently. Thus, the proposed corrections at block 825 by Ci 820 and Cm 850 agree, and a strong acceptance rule applies. All proposed corrections by Ci 820 and Cm 850 are mandatory and fixed (accepted), and are labeled as high reliability (strong accepted solutions). That is, in addition to the proposed correction at block 825, the proposed corrections at blocks 823, 824, 833, and 843 are also mandatory and fixed (accepted). The decoder implements the proposed corrections on all such error candidates, and proceeds to the next safety+ sub-repetition 505.
[0113] In response to determining that at least one strong accepted solution is found (710: YES), the decoder performs cross-component rescission at 740. Cross-component rescission refers to the decoder removing all other error candidates that are not strong accepted solutions.
[0114] On the other hand, in response to determining that at least one strong accepted solution is not found (710: NO), the decoder performs SLA detection at 715. By performing SLA detection, one or more proposed corrections determined during safety detection (at 705) with the lowest probability of incorrect correction are identified. Such identification is performed with the lowest degree of error candidate bias. In some arrangements, performing SLA detection includes evaluating each error candidate by the decoder determining whether the cross-component code can be solved by implementing the proposed correction at the error candidate. Implementing the proposed correction refers to flipping the bit at the proposed error candidate.
[0115] For example, in response to determining that the cross-component code has a suggested solution that results from a test implementation of an error candidate determined from the safety detection (e.g., at 705), the error candidate corresponding to the suggested solution is stored in memory for further evaluation. In the SLA detection, for each error candidate determined in the safety detection (referred to as an initiator error candidate), the decoder produces a derived error candidate by performing a single bit flip at the initiator error candidate and attempts to solve the associated cross-component code. An initiator component code is a component code that has an initiator error candidate whose test implementation (bit flip) produces a successful solution on the cross-component code (derived component code). The initiator error candidates for a single bit flip are those discovered using the initiator component code and saved during the safety detection at 705. For each initiator error candidate, the decoder saves a candidate solution to produce a successful evaluation of any cross-component code, referred to as a derived component code. After testing all initiator error candidates one at a time, a list of derived error candidates is produced. In some cases, this list can contain several error candidates for the same derived component code.
[0116] After producing the list, the decoder determines whether the derived error candidates (produced in the SLA at 715) are consistent with the initiator error candidates (produced in the safety detection at 705). In response to determining such consistency, the decoder accepts the suggested corrections determined based on the initiator component codes and derived component codes associated with the consistency. That is, the decoder accepts the suggested corrections on any initiator error candidate determined using the initiator component code and any derived error candidate determined using the derived component code. When such consistency exists, the initiator / derived error candidates for the relevant initiator / derived component codes are considered strong accepted solutions.
[0117] Provided Figure 9 to illustrate one type of strong accepted solution determined during the SLA detection at 715. Figure 9 is a diagram illustrating a decoding scenario in which two component codes indicate a suggested correction that is consistent after a test implementation or suggested correction, in accordance with various embodiments. Reference is made to Figures 1-9 , Figure 9 The ECC structure 800 shown in Figure 8 is similar to the ECC structure shown in
[0118] Error detection during safety detection (at 705) using Ci 820 (initiator component code) results in a proposed correction at block 923, 924, and 925 (initiator error candidates). Error detection during safety detection (at 705) using Cj 830 (initiator component code) results in a proposed correction at block 933, 934, and 935 (initiator error candidates). The initiator error candidates determined using Ci 820 and Cj 830 are inconsistent. Those initiator error candidates are saved in memory. During SLA detection at 715, the decoder tests each initiator error candidate saved in memory (e.g., by bit flipping) and attempts to resolve each cross component code (derivative component code) that intersects or crosses with the initiator component code to which the initiator error candidate belongs. For example, in response to bit flipping the initiator error candidate determined using initiator component code Ci 820 (e.g., block 925), the decoder attempts to resolve at least one cross component code, including cross component code Cm 850 (derivative component code) that intersects with Ci 820 at the flipped initiator error candidate. When error candidate 925 is bit flipped (implementing the proposed correction at error candidate 925), resolutions of three derivative error candidates at 933, 943, and 953 are found. As shown, initiator component code Cj 830 and derivative component code Cm 850 agree on the proposed correction at block 933, and the strong acceptance rule applies. All proposed corrections by initiator component code Cj 830 and derivative component code Cm 850, as well as initiator component code Ci 820, are mandatory and fixed (accepted), and are labeled as high reliability (strong accepted resolution). That is, in addition to the proposed correction at block 933, the proposed corrections at blocks 923, 924, 925, 934, 935, 925, 943, and 953 are also mandatory and fixed (accepted). The decoder implements the proposed corrections on all such error candidates, and proceeds to the next safety+ subiteration 505. All other resolutions of any of the blocks (e.g., bits therein) are rejected.
[0119] In some instances in which the decoder determines that any derivative error candidate produced by testing implementation of the proposed correction at the initiator error candidate does not agree with any initiator error candidate, the decoder can seek to determine whether there is agreement between two derivative error candidates produced during SLA detection at 715, e.g., by flipping bits of the initiator error candidate. In response to determining that two derivative component codes produce two derivative error candidates that are the same after testing implementation (bit flipping) of the initiator error candidate determined using the initiator component code, the initiator error candidate determined using the initiator component code is accepted as a high reliability fixed point (e.g., strong accepted resolution).
[0120] To this end, there is provided Figure 10 to illustrate another type of strong solution determined during SLA detection at 715. Figure 10 is a diagram illustrating a decoding scenario in which two component codes indicate a suggested correction that is consistent after the suggested correction is implemented in testing, in accordance with various embodiments. Reference is made to Figures 1-10 , Figure 10 The ECC structure 800 shown in Figure 8 and 9 but with different suggested error candidates.
[0121] Error detection during safety detection (at 705) using initiator component code Ci 820 results in suggested corrections at blocks 1023, 1024, and 1025 (initiator error candidates). Error detection during safety detection (at 705) using initiator component code Cj 830 results in suggested corrections at blocks 1033, 1034, and 1035 (initiator error candidates). The initiator error candidates determined using Ci 820 and Cj 830 are inconsistent. Those initiator error candidates are saved in memory. During SLA detection at 715, the decoder tests each initiator error candidate saved in memory (e.g., by bit flipping) and attempts to resolve other component codes. For example, in response to bit flipping the initiator error candidate determined using Ci 820 (e.g., block 1025), the decoder attempts to resolve all cross component codes that intersect with Ci 820, including deriving component code Cm 850. When initiator error candidate 1025 is bit flipped (implementing the suggested correction at initiator error candidate 1025), resolutions of three derived error candidates at 1053, 1044, and 1054 are found. Additionally, in response to bit flipping the initiator error candidate determined using Cj 830 (e.g., block 1033), the decoder attempts to resolve all cross component codes that intersect with Cj 830, including deriving component code Ck 840. When initiator error candidate 1033 is bit flipped (implementing the suggested correction at initiator error candidate 1033), resolutions of three derived error candidates at 1043, 1044, and 1045 are found. As shown, derived component codes Ck 840 and Cm 850 agree on the suggested correction at block 1044, and the strong acceptance rule applies. All suggested corrections by derived component codes Ck 840 and Cm 850 and initiator component codes Ci 820 and Cj 830 are mandatory and fixed (accepted), and are labeled as high reliability (strong accepted resolutions). That is, in addition to the suggested correction at block 1044, the suggested corrections at blocks 1023, 1024, 1025, 1033, 1034, 1035, 1043, 1045, 1053, and 1054 are also mandatory and fixed (accepted). The decoder implements the suggested corrections on all such error candidates, and proceeds to the next safety+ subiteration 505. All other resolutions of any of the blocks (e.g., bits therein) are rejected.
[0122] In some instances in which the decoder determines that (1) any derived error candidate resulting from applying the suggested correction at the test-implement initiator error candidate does not agree with any initiator error candidate; and (2) two of the derived error candidates resulting from applying the suggested correction at each of the test-implement initiator error candidates do not agree with each other, the decoder can seek to determine whether there is agreement among the subsequent derived error candidates resulting from, for example, flipping the bits of the derived error candidates determined during SLA detection at 715.
[0123] In response to determining that the two subsequent derived error candidates resulting from the two component codes are the same after the test-implement (bit flipping) determines additional error candidates using the component code (not the initiator component code) determined during SLA detection, the associated subsequent derived error candidates, derived error candidates, and initiator error candidates are accepted as high-reliability fixed points (e.g., strong-accept solutions).
[0124] To this end, there is provided Figure 11 to illustrate another type of strong-accept solution determined during SLA detection at 715. Figure 11 is a diagram illustrating a decoding scenario in which two component codes indicate that the suggested correction is consistent after the test-implement applies the suggested correction, in accordance with various embodiments. Reference is made to Figures 1-11 , Figure 11 The ECC structure 800 shown in Figures 8-10 is similar to the ECC structure shown in
[0125] Error detection during safe detection (at 705) using the initiator component code Ci 820 results in suggested corrections at blocks 1123, 1124, and 1125 (initiator error candidates). Error detection during safe detection (at 705) using the initiator component code Cj 830 results in suggested corrections at blocks 1133, 1134, and 1135 (initiator error candidates). The initiator error candidates determined using Ci 820 and Cj 830 do not agree. Those initiator error candidates are saved in memory.
[0126] During the first round of SLA detection at 715, the decoder tests implementing (e.g., by bit flipping) each initiator error candidate saved in memory and attempts to resolve other component codes. For example, in response to bit flipping the initiator error candidate determined using Ci 820 (e.g., block 1125), the decoder attempts to resolve all cross-component codes intersecting Ci 820, including deriving component code Cm 850. When initiator error candidate 1125 is bit flipped (implementing the suggested correction at initiator error candidate 1125), the resolutions of three derived error candidates at 1153, 1154, and 1155 are found. Additionally, in response to bit flipping the initiator error candidate determined using Cj 830 (e.g., block 1133), the decoder attempts to resolve all cross-component codes intersecting Cj 830, including deriving component code Ck 840. The derived error candidates are saved in memory. When error candidate 1133 is bit flipped (implementing the suggested correction at initiator error candidate 1133), the resolutions of three derived error candidates at 1143, 1144, and 1145 are found. As shown, no error candidates 1123, 1124, 1125, 1133, 1134, 1135, 1143, 1144, 1145, 1153, 1154, and 1155 agree.
[0127] During the second round of SLA detection at 715, the decoder tests implementing (e.g., by bit flipping) each derived error candidate saved in memory and attempts to resolve other component codes. For example, in response to bit flipping the derived error candidate determined using Cm 850 (e.g., block 1153), the decoder attempts to resolve all cross-component codes intersecting Cm 850, including subsequent derived component code Ci 810. When derived error candidate 1153 is bit flipped (implementing the suggested correction at derived error candidate 11553), the resolutions of three derived error candidates at 1113, 1114, and 1143 are found. All suggested corrections by subsequent derived component code Ci 810, derived component codes Ck 840 and Cm 850, and initiator component codes Ci 820 and Cj 830 are mandatory and fixed (accepted), and are marked as high reliability (strong accepted resolutions). That is, in addition to the suggested correction at block 1143, the suggested corrections at blocks 1113, 1114, 1123, 1124, 1125, 1133, 1134, 1135, 1144, 1145, 1153, 1154, and 1155 are also mandatory and fixed (accepted). The decoder implements the suggested corrections on all such error candidates, and proceeds to the next safety+ sub-repetition 505. All other resolutions of any block (e.g., bits therein) are rejected.
[0128] Thus, one or more rounds of SLA detection can be performed by the decoder at 715. In each round, the decoder tests one or more error candidates to resolve other component codes to generate additional error candidates. The decoder then determines whether any of the additional error candidates is consistent with another of the additional error candidates or with any previously saved error candidate. In response to determining a consistency, any error candidate corresponding to the relevant component code is accepted as a strong acceptance solution. On the other hand, in response to determining that there is no consistency, the additional error candidates generated in the current round are saved in memory and become part of the previously saved error candidates for the next SLA round.
[0129] In some instances, a given derived component code can have more than one derived error candidate during SLA detection at 715. In such instances, the number of derived candidates for a given derived component code are saved and a search for potential consistency is performed over all derived error candidates for the same derived component code. In instances where multiple consistencies are found (different consistent error candidates), the decoder selects the error candidate with the largest number of consistencies or selects the first detected error candidate in case multiple error candidates have the largest number of consistencies. Consistency refers to a single bit that is present on both an error candidate for a component code and an error candidate for a cross component code (intersecting the component code). First detected error candidate refers to the error candidate that is detected first in time according to any suitable search order.
[0130] In some arrangements, after each round in SLA detection, only a single error candidate is saved per component code even when there are multiple error candidates for the same component code. The single error candidate is used as the initial condition for the next round in SLA detection.
[0131] In some instances, blocks 705, 710, 715, and 720 are only performed once per safety+ sub-repetition 505.
[0132] The decoder determines whether a strong acceptance solution was found in the SLA detection at 715 at 720. In response to determining that a strong acceptance solution was found (720: YES), the decoder performs cross-component rescission at 740.
[0133] On the other hand, in response to determining that a strong acceptance solution was not found (720: NO), the decoder performs a safe rejection at 725. For example, the decoder can determine that a strong acceptance solution was not found in response to determining that a strong acceptance solution was not found after a maximum number of SLA rounds. In another example, the decoder can determine that a strong acceptance solution was not found in response to determining that no new additional error candidates were found in the current SLA round.
[0134] In the safety rejection flow at 725, the decoder determines whether to accept any error candidates determined and saved at 715 before proceeding to the next sub-repetition (505).
[0135] In some arrangements, in response to determining that implementing the suggested correction on an error candidate of a given component code would cause modification of a resolved component code or a component code with a candidate solution, the decoder rejects the error candidate in the current sub-repetition. This does not apply to error candidates identified in 715 (e.g., error candidates other than the initiator error candidate identified at 705) in the condition that the error candidates identified in 715 are error candidates of the initiator error candidate identified at 705.
[0136] In some arrangements, in response to determining that implementing the suggested correction on a later determined error candidate of a given component code would cause modification of an earlier detected resolved component code or an earlier detected component code with a candidate solution, the resolved component code or the component code with a candidate solution is rejected. This applies to error candidates identified in 715 (e.g., error candidates other than the initiator error candidate identified at 705) and the initiator error candidate identified at 705, thus allowing for a higher rejection rate and a reduced probability of mis-correction. In this regard, Figure 12 A scenario in which this rule can be implemented is illustrated.
[0137] Figure 12 is a diagram illustrating a decoding scenario in which two component codes indicate conflicting suggested corrections according to various embodiments. Reference is made to Figures 1-12 , Figure 12 The ECC structure 800 shown in Figures 8-11 is similar to the ECC structure shown in
[0138] Error detection during safety detection (at 705) using initiator component code Ci 820 results in suggested corrections at blocks 1223, 1224, and 1225 (initiator error candidates). Error detection during safety detection (at 705) using initiator component code Cj 830 results in suggested corrections at blocks 1233, 1234, and 1235 (initiator error candidates). The initiator error candidates determined using Ci 820 and Cj 830 are inconsistent. Those initiator error candidates are saved in memory.
[0139] During SLA detection at 715, the decoder tests each initiator error candidate (e.g., by bit flipping) saved in memory and attempts to solve the other component code. For example, when the initiator error candidate is bit flipped, solutions 1243, 1244, and 1245 for three derived error candidates for Ck 840 are found. When the initiator error candidate is bit flipped, solutions 1253, 1254, and 1255 for three derived error candidates for Cm 850 are found. Assuming no strong accepted solution is found (720: No), during safe rejection at 725, in some instances, error candidate 1255 for Cm 850 determined later suggests rejection of error candidates 1233, 1234, 1235 for Cj 830 under the condition that the correction at the error candidate different from the error candidates 1233, 1234, 1235 for Cj 830 determined during safe detection at 705.
[0140] At 730, the decoder performs a first-solved mandate. For example, the decoder can determine that an error candidate is accepted at safe rejection (725), that the total number of error candidates accepted at safe rejection, including error candidates, is t - 1 or less, and that the error candidate is first-solved during decoding (the error candidate was not previously found to be a prior error candidate in method 500). A total number of error candidates found for each component code that is t - 1 or less corresponds to a lower probability of mis-correction. In response to identifying such an error candidate, the decoder marks this error candidate as a high-reliability fixed point (mandate). In other words, in response to determining that such an error candidate exists, a strong accepted solution is found (735: Yes), and method 700 proceeds to 740.
[0141] After block 740 or in response to determining that no strong accepted solution is found (735: No), method 700 ends (repetition of safe+ rejection flow at 605, 625, or 635). In response to determining that at least one strong accepted solution (including at least one error candidate) is found in method 700 (e.g., 710: Yes, 720: Yes, or 735: Yes), the suggested correction on any strong accepted solution is accepted, and the candidate is found (610: Yes, 630: Yes, or 640: Yes). In other words, in response to determining that at least one strong accepted solution (including at least one error candidate) is found in method 700 (e.g., 710: Yes, 720: Yes, or 735: Yes), method 500 proceeds to the next safe+ sub-repetition at 505.
[0142] Figure 13 FIG. 13 is a process flow diagram illustrating an example method 1300 for performing proactive detection, in accordance with some embodiments. Reference is made to FIG. 1. Figures 1-13 Method 1300 is an example implementation of 710. Method 1300 can be performed by a decoder of ECC encoder / decoder 102 or a decoder of ECC encoder / decoder 112 (referred to as a "decoder").
[0143] For each SLA round, at 1310, the decoder tests each error candidate implementing at least one component code. In instances where the current SLA round is the first round, the at least one component code refers to any initiator component code having at least one initiator error candidate. In instances where the current SLA round is not the first round, the at least one component code refers to any new component code (e.g., derived component code, subsequent derived component, etc.) having at least one error candidate determined in the previous round.
[0144] At 1320, the decoder determines at least one additional error candidate by solving at least one intersection component code intersecting each of the at least one component code. Any additional error candidate determined in the current round corresponds to a component code determined in the current round. The additional error candidate can be a derived error candidate or a subsequent derived error candidate.
[0145] At 1330, the decoder determines whether the at least one additional error candidate is identical to one of the previously determined error candidates. As described, any error candidate determined in 705 and 715 is stored in memory. Thus, a list of previously determined error candidates is available.
[0146] In response to determining that one of the additional error candidates is identical (e.g., identifies with) one of the previously determined error candidates (1330: YES), a strong acceptance solution is found at 1340 (720: YES). On the other hand, in response to determining that none of the at least one additional error candidate is identical to any of the previously determined error candidates, the at least one additional error candidate is added to the list of previously determined error candidates at 1350, and method 1300 proceeds to the next SLA round.
[0147] Figure 14 is a process flow diagram illustrating an example method 1400 for performing proactive detection in accordance with some embodiments. Reference is made to Figures 1-14 Methods 500, 600, 700, and 1300 are particular implementations of one or more aspects of method 1400. Method 1400 can be performed by a decoder of ECC encoder / decoder 102 or a decoder of ECC encoder / decoder 112 (referred to as a "decoder").
[0148] At 1410, the decoder determines error candidates for the data based on the component codes. In some instances, determining error candidates for the data based on the component codes includes determining error candidates by decoding the codeword based on the component codes. The codeword corresponds to an input payload having input bits. The input bits are organized into a pseudo-triangular matrix. Each row or column of the pseudo-triangular matrix differs from an adjacent row or column by two or more bits. The input bits are encoded using the component codes to map each of the input bits to two or more of the component codes based on the pseudo-triangular matrix. The pseudo-triangular matrix includes a plurality of blocks. Each of the plurality of blocks includes two or more of the input bits. Two component codes encode a same block of the plurality of blocks.
[0149] At 1420, the decoder determines whether at least one first error candidate is found from the error candidates based on the two of the component codes agreeing on a same error candidate.
[0150] At 1430, the decoder determines whether at least one second error candidate is found based on the two of the component codes agreeing on the same error candidate in response to implementing the proposed correction at one of the error candidates. In some instances, determining whether at least one second error candidate is found includes performing one or more rounds of detection. A current round of the one or more rounds of detection includes: testing each error candidate implementing at least one component code; determining at least one additional error candidate by solving at least one cross-component code intersecting each of the at least one component code (the at least one additional error candidate corresponding to the at least one component code determined in the current round); and determining whether the at least one additional error candidate is the same as one of the previously determined error candidates.
[0151] The current round of the one or more rounds of detection additionally includes: determining that at least one second error candidate is found in response to determining that one of the at least one additional error candidate is the same as one of the previously determined error candidates. The current round of the one or more rounds of detection additionally includes: proceeding to a next round of the one or more rounds in response to determining that none of the at least one additional error candidate is the same as any of the previously determined error candidates. The at least one additional error candidate is added to the previously determined error candidates for the next round.
[0152] At 1440, the decoder corrects the error in the data based on whether at least one of the at least one first error candidate is found or whether at least one second error candidate is found. Correcting the error in the data based on whether at least one of the at least one first error candidate is found or whether at least one second error candidate is found includes one or more of: accepting the at least one first error candidate in response to determining that the at least one first error candidate is found; or accepting the at least one second error candidate in response to determining that the at least one second error candidate is found.
[0153] In some examples, the method 1400 additionally includes determining that the at least one first error candidate is found in response to determining that two of the component codes agree on a same error candidate. The at least one first error candidate includes one of the same error candidate and one or more of: at least one error candidate determined using a first of the two of the component codes; or at least one error candidate determined using a second of the two of the component codes.
[0154] In some examples, the method 1400 additionally includes determining whether the at least one second error candidate is found in response to determining that no two of the component codes agree on a same error candidate.
[0155] In some examples, determining error candidates for the data based on the component codes includes determining initiator error candidates based on the initiator component codes. Determining whether the at least one second error candidate is found includes testing the implemented proposed correction at one of the initiator error candidates (the one of the initiator error candidates determined using a first of the initiator component codes); determining derived error candidates using at least one derived component code intersecting the first initiator component code by implementing the proposed correction; and determining whether one of the derived error candidates is the same as one of the initiator error candidates. Determining whether the at least one second error candidate is found additionally includes determining that the at least one second error candidate is found in response to determining that one of the derived error candidates is the same as one of the initiator error candidates. The one of the derived error candidates is determined using a first of the at least one derived component code. The one of the initiator error candidates is determined using a second of the initiator component codes. The at least one second error candidate includes the one of the derived error candidates that is the same as one of the initiator error candidates and one or more of: at least one error candidate determined using the first initiator component code; at least one error candidate determined using the first derived component code by implementing the proposed correction; or at least one error candidate determined using the second initiator component code.
[0156] In some examples, determining the error candidates for the data based on the component codes includes determining initiator error candidates based on initiator component codes. Determining whether the at least one second error candidate is found includes testing implementation of a first suggested correction at a first initiator error candidate of the initiator error candidates (the first initiator error candidate is determined using a first initiator component code of the initiator component codes), determining first derived error candidates using at least one first derived component code by implementing the first suggested correction, testing implementation of a second suggested correction at a second error candidate of the initiator error candidates (the second error candidate is determined using a second initiator component code of the initiator component codes), determining second derived error candidates using at least one second derived component code by implementing the second suggested correction, and determining whether one of the first derived error candidates is the same as one of the second derived error candidates. Determining whether the at least one second error candidate is found additionally includes determining that the at least one second error candidate is found in response to determining that the one of the first derived error candidates is the same as the one of the second derived error candidates. The at least one second error candidate includes the one of the first derived error candidates that is the same as the one of the second derived error candidates and one or more of: at least one error candidate determined using the first initiator component code, at least one error candidate determined using the second initiator component code, at least one error candidate determined using the first derived component code by implementing the first suggested correction, or at least one error candidate determined using the second derived component code by implementing the second suggested correction.
[0157] In some examples, determining the error candidates for the data based on the component codes includes determining initiator error candidates based on initiator component codes. Determining whether the at least one second error candidate is found includes testing implementation of a first suggested correction at a first initiator error candidate of the initiator error candidates, determining the first initiator error candidate using a first initiator component code of the initiator component codes; determining a first derived error candidate using a first derived component code by implementing the first suggested correction; testing implementation of a second suggested correction at the first derived error candidate; determining at least one first subsequent derived error candidate using a first subsequent derived component code by implementing the second suggested correction; testing implementation of a third suggested correction at a second error candidate of the initiator error candidates (the second error candidate determined using a second initiator component code of the initiator component codes); determining a second derived error candidate using at least one second derived component code by implementing the third suggested correction; and determining whether one of the first subsequent derived error candidates is the same as one of the second derived error candidates. Determining whether the at least one second error candidate is found additionally includes determining that the at least one second error candidate is found in response to determining that the one of the first subsequent derived error candidates is the same as the one of the second derived error candidates. The at least one second error candidate includes the one of the first subsequent derived error candidates that is the same as the one of the second derived error candidates and one or more of: at least one error candidate determined using the first initiator component code; at least one error candidate determined using the second initiator component code; at least one error candidate determined using the first derived component code by implementing the first suggested correction; at least one error candidate determined using the first subsequent derived component code by implementing the second suggested correction; or at least one error candidate determined using the second derived component code by implementing the third suggested correction.
[0158] Accordingly, the methods described herein improve error correction capability for hard decoding. NAND flash memory devices implementing such methods can achieve correction of higher BERs at low implementation complexity with hard decoding and high throughput encoding / decoding, improving read performance when applying a single page read for each host request.
[0159] Low complexity hard decoding results in higher reliability under the condition that the probability of mis-correction of component code is reduced within the repetitive decoding of the general product code. This is achieved by performing multiple rounds of repetitive SLA component code decoding. In some instances, a non-dependent consistent solution is caused by a de-escalation to a high reliability (strong accept) solution. Solutions that cause SLA consistency are accepted with the addition of a high reliability flag. The method additionally allows for the rejection of suspect and contradictory solutions of individual component codes found in the SLA round. Furthermore, the method includes the use of forced bits for the solution of components to obtain another safeguard against mis-correction.
[0160] Additionally, such embodiments allow for efficient hardware implementations. To decode irregular product code structures, a one-time bit decoding algorithm is used, such that a dedicated decoder handles bits that are applied with one-time encoding protection.
[0161] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other aspects. Thus, the claims are not intended to be limited to the aspects presented herein, but are to be accorded the full scope consistent with the language of the claims, wherein reference to an item means at least one, unless otherwise indicated. The terms "some" and "another" are defined as one or more unless explicitly stated otherwise. All structural and functional equivalents to the aspects described throughout this specification that are known or later come to be known to those of ordinary skill in the art are expressly incorporated herein by reference and intended to be encompassed by the claims. Moreover, nothing disclosed herein is intended to be dedicated to the public regardless of whether these disclosure is explicitly recited in the claims. No claim element is to be construed as a means plus function unless the element is expressly recited using the phrase "means for."
[0162] It is understood that the specific order or hierarchy of steps in the processes disclosed is an example of illustrative approaches. Based upon design preferences, it is understood that the specific order or hierarchy of steps in the processes can be rearranged, while remaining within the scope of the previous description. The accompanying method claim is intended to cover all steps of the process, regardless of the order or hierarchy, in which the steps are presented herein. The accompanying method claim is intended to cover at least all possible combinations of the steps outlined herein.
[0163] The foregoing description of the disclosed implementations is provided as an enabling teaching of the disclosed subject matter in order to enable any person skilled in the art to
[0164] The various examples illustrated and described are provided merely as examples to illustrate various features of the claims appended hereto. However, features shown and described in relation to any given example are not necessarily limited to that associated example, and can be used or combined with other examples shown and described. Moreover, the claims are not intended to be limited by any one example.
[0165] The above method descriptions and process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the steps of the various examples must be performed in the order presented. As will be appreciated by one of skill in the art, the order of steps in the foregoing examples can be performed in any order. Words such as "thereafter," "then," "next," etc. are not intended to limit the order of the steps; these words are simply used to guide the reader through the description of the methods. Furthermore, any reference to claim elements in the singular, for example, using the articles "one," "the," "said," and "the," is not
[0166] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the examples disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0167] The hardware used to implement the various illustrative logics, logical modules, and circuits described in connection with the examples disclosed herein can be implemented or performed with a general purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Alternatively, some steps or methods can be performed by circuitry that is specific to a given function.
[0168] In one or more exemplary embodiments, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a non-transitory computer-readable medium. The steps of a method or algorithm disclosed herein can be embodied in a processor-executable software module which can reside on a non-transitory computer- or processor-readable storage medium. Non-transitory computer- or processor-readable storage media can be any storage media that can be accessed by a computer or a processor. By way of example but not limitation, such non-transitory computer- or processor-readable storage media can include RAM, ROM, EEPROM, FLASH memory, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Disk and disc, as used herein, includes compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks and blu-ray discs where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of non-transitory computer- or processor-readable media. Additionally, the operations of a method or algorithm can reside in one or any combination of the above memory hardware, which can be incorporated into a computer program product.
[0169] The previous description of the disclosed examples is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these examples will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to some other examples without departing from the spirit or scope of the disclosure. Thus, the present disclosure is not intended to be limited to the examples shown herein but is to be accorded the widest scope consistent with the claims and the principles and novel features disclosed herein.
Claims
1. A method for decoding data read from a non-volatile storage device, comprising: determining, by a decoder, error candidates for the data based on component codes; determining, by the decoder, whether at least one first error candidate is found from the error candidates based on two of the component codes agreeing on a same error candidate; when two of the component codes disagree on the same error candidate, implementing a proposed correction at one of the error candidates; determining, by the decoder, whether at least one second error candidate is found based on two of the component codes agreeing on the same error candidate in response to implementing the proposed correction at the one of the error candidates; correcting an error in the data based on at least one of whether the at least one first error candidate is found or whether the at least one second error candidate is found.
2. The method of claim 1, wherein determining the error candidates for the data based on the component codes comprises determining the error candidates by decoding codewords based on the component codes; the codewords correspond to an input payload having input bits; the input bits are organized into a pseudo-triangular matrix; each row or column of the pseudo-triangular matrix differs from an adjacent row or column by two or more bits; and the input bits are encoded using the component codes based on mapping each of the input bits to two or more of the component codes based on the pseudo-triangular matrix.
3. The method of claim 2, wherein the pseudo-triangular matrix comprises a plurality of blocks; each of the plurality of blocks comprises two or more of the input bits; and the two component codes encode a same block of the plurality of blocks.
4. The method of claim 1, wherein correcting the error in the data based on at least one of whether the at least one first error candidate is found or whether the at least one second error candidate is found comprises one or more of: in response to determining that the at least one first error candidate is found, accepting the at least one first error candidate; or in response to determining that the at least one second error candidate is found, accepting the at least one second error candidate.
5. The method of claim 1, further comprising determining that the at least one first error candidate is found in response to determining that the two of the component codes agree on the same error candidate, wherein the at least one first error candidate comprises one of the same error candidates and one or more of: at least one error candidate determined using a first of the two of the component codes; or at least one error candidate determined using a second of the two of the component codes.
6. The method of claim 1, further comprising, in response to determining that no two component codes of the component codes agree on the same error candidate, determining whether the at least one second error candidate is found.
7. The method of claim 1, wherein determining whether the at least one second error candidate is found comprises performing one or more rounds of detection; a current round of the one or more rounds of detection comprises: testing each error candidate of at least one component code; determining at least one additional error candidate by solving at least one cross-component code that intersects each of the at least one component code, the at least one additional error candidate corresponding to the at least one component code determined in the current round; and and determining whether the at least one additional error candidate is the same as one of the previously determined error candidates.
8. The method of claim 7, the current round of the one or more rounds of detection further comprising: in response to determining that one of the at least one additional error candidate is the same as one of the previously determined error candidates, determining that the at least one second error candidate is found; and in response to determining that none of the at least one additional error candidate is the same as any of the previously determined error candidates, proceeding to a next round of the one or more rounds, wherein the at least one additional error candidate is added to the previously determined error candidates for the next round.
9. The method of claim 1, wherein determining the error candidates for the data based on the component codes comprises determining initiator error candidates based on initiator component codes; determining whether the at least one second error candidate is found comprises: testing an advised correction at one of the initiator error candidates, determining one of the initiator error candidates using a first initiator component code of the initiator component codes; determining a derived error candidate by implementing the advised correction, using at least one derived component code that intersects the first initiator component code; and determining whether one of the derived error candidates is the same as one of the initiator error candidates.
10. The method of claim 9, wherein determining whether the at least one second error candidate is found further comprises: in response to determining that one of the derived error candidates is the same as the one of the initiator error candidates, determining that the at least one second error candidate is found, wherein the one of the derived error candidates is determined using a first derived component code of the at least one derived component code, and the one of the initiator error candidates is determined using a second initiator component code of the initiator component codes; and in response to determining that none of the derived error candidates is the same as any of the initiator error candidates, proceeding to a next round of the one or more rounds, wherein the at least one derived error candidate is added to the previously determined error candidates for the next round. the at least one second error candidate includes one of the one of the derived error candidates that is the same as the one of the initiator error candidates and one or more of: at least one error candidate determined using the first initiator component code; at least one error candidate determined using the first derived component code by implementing the first suggested correction; or at least one error candidate determined using the second initiator component code.
11. The method of claim 1, wherein determining the error candidates for the data based on the component codes includes determining initiator error candidates based on initiator component codes; determining whether the at least one second error candidate is found includes: testing implementation of a first suggested correction at a first initiator error candidate of the initiator error candidates, the first initiator error candidate determined using a first initiator component code of the initiator component codes; determining a first derived error candidate using at least one first derived component code by implementing the first suggested correction; testing implementation of a second suggested correction at a second error candidate of the initiator error candidates, the second error candidate determined using a second initiator component code of the initiator component codes; determining a second derived error candidate using at least one second derived component code by implementing the second suggested correction; and determining whether one of the first derived error candidates is the same as one of the second derived error candidates.
12. The method of claim 11, wherein determining whether the at least one second error candidate is found additionally includes: in response to determining that the one of the first derived error candidates is the same as the one of the second derived error candidates, determining that the at least one second error candidate is found, wherein the at least one second error candidate includes the one of the first derived error candidates that is the same as the one of the second derived error candidates and one or more of: at least one error candidate determined using the first initiator component code; at least one error candidate determined using the second initiator component code; at least one error candidate determined using the first derived component code by implementing the first suggested correction; or at least one error candidate determined using the second derived component code by implementing the second suggested correction.
13. The method of claim 1, wherein determining the error candidates for the data based on the component codes includes determining initiator error candidates based on initiator component codes; determining whether the at least one second error candidate is found includes: determining the first initiator error candidate using a first initiator component code of the initiator component codes; determining a first export error candidate using a first export component code based on implementing the first suggested correction; testing implementation of a second suggested correction at the first export error candidate; determining at least one first subsequent export error candidate using a first subsequent export component code based on implementing the second suggested correction; determining a second error candidate of the initiator error candidates at which to test implementation of a third suggested correction using a second initiator component code of the initiator component codes; determining a second export error candidate using at least one second export component code based on implementing the third suggested correction; and determining whether one of the first subsequent export error candidates is identical to one of the second export error candidates.
14. The method of claim 13, wherein determining whether the at least one second error candidate is found further comprises: in response to determining that the one of the first subsequent export error candidates is identical to the one of the second export error candidates, determining that the at least one second error candidate is found, wherein the at least one second error candidate comprises the one of the first subsequent export error candidates that is identical to the one of the second export error candidates and one or more of: at least one error candidate determined using the first initiator component code; at least one error candidate determined using the second initiator component code; at least one error candidate determined using the first export component code based on implementing the first suggested correction; at least one error candidate determined using the first subsequent export component code based on implementing the second suggested correction; or at least one error candidate determined using the second export component code based on implementing the third suggested correction.
15. An error correction system comprising processing circuitry configured to: determine error candidates for data based on component codes; determine whether at least one first error candidate is found from the error candidates based on two of the component codes agreeing on a same error candidate; implement a suggested correction at one of the error candidates when the two of the component codes do not agree on the same error candidate; determine whether at least one second error candidate is found based on two of the component codes agreeing on a same error candidate in response to implementing the suggested correction at the one of the error candidates; correct an error in the data based on at least one of whether the at least one first error candidate is found or whether the at least one second error candidate is found.
16. A non-transitory computer-readable medium storing computer-readable instructions such that, when executed, cause processing circuitry to decode data stored in a non-volatile storage device by: determining error candidates for the data based on component codes; determining whether at least one first error candidate is found from the error candidates based on two of the component codes agreeing on a same error candidate; when two of the component codes disagree on the same error candidate, implementing a proposed correction at one of the error candidates; in response to implementing the proposed correction at the one of the error candidates, determining whether at least one second error candidate is found based on two of the component codes agreeing on a same error candidate; correcting an error in the data based on at least one of whether the at least one first error candidate is found or whether the at least one second error candidate is found.
17. The non-transitory computer-readable medium of claim 16, wherein correcting the error in the data based on at least one of whether the at least one first error candidate is found or whether the at least one second error candidate is found comprises one or more of: in response to determining that the at least one first error candidate is found, accepting the at least one first error candidate; or in response to determining that the at least one second error candidate is found, accepting the at least one second error candidate.
18. The non-transitory computer-readable medium of claim 16, wherein the processing circuitry is further configured to, in response to determining that the two of the component codes agree on the same error candidate, determine that the at least one first error candidate is found, wherein the at least one first error candidate comprises one of the same error candidate and one or more of: at least one error candidate determined using a first one of the two of the component codes; or at least one error candidate determined using a second one of the two of the component codes.
19. The non-transitory computer-readable medium of claim 16, wherein the processing circuitry is further configured to, in response to determining that no two of the component codes agree on the same error candidate, determine whether the at least one second error candidate is found.
20. The non-transitory computer-readable medium of claim 16, wherein determining whether the at least one second error candidate is found comprises performing one or more rounds of detection; a current round of the one or more rounds of detection comprises: testing each error candidate of at least one component code; determining at least one additional error candidate by solving at least one cross-component code that intersects each of the at least one component code, the at least one additional error candidate corresponding to at least one component code determined in the current round; and and determining whether the at least one additional error candidate is the same as one of the previously determined error candidates.
Citation Information
Patent Citations
Decoding scheme for error correction code structure
US20200293399A1