Machine learning based llr generation for early soft decoding of unsupervised reading
By using hardware decoding to generate LLRs in solid-state storage, combined with machine learning and pattern matching operations, the problem of insufficient information in early software decoding is solved, improving the error correction capability and decoding efficiency of the storage system.
Patent Information
- Application Number
- CN202111178117.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-01-15
- Filing Date
- 2021-10-09
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2041-10-09
AI Technical Summary
Existing technologies struggle to effectively generate the log-likelihood ratio (LLR) required for early software decoding in solid-state storage, and conventional software decoding methods are complex and difficult to obtain sufficient information, resulting in low data decoding efficiency.
A hard-decoding-based method is used to generate LLRs. By using machine learning to utilize hard-read data, checksums, and 1 counting, combined with pattern matching operations and deep neural networks, the LLRs required for early soft decoding are generated, avoiding the use of auxiliary reading.
It improves the error correction capability of solid-state storage, simplifies the data decoding process, and enhances the performance and efficiency of storage systems.
Smart Images

Figure CN114765462B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates generally to a system and method for storage devices, and in particular, to improving the performance of non-volatile memory devices such as solid state drives (SSDs). BACKGROUND
[0002] Solid state memory is widely used in a variety of electronic systems, including, for example, consumer electronics devices (e.g., mobile phones, cameras, computers) and enterprise computing systems (e.g., hard drives, random access memory (RAM)). Solid state memory is becoming more prevalent than mechanical or other memory storage technologies due to latency, throughput, shock resistance, packaging, and other factors.
[0003] To increase storage density, the use of multi-bit, multi-level memory cells is increasing. As the density increases, the margin for error decreases. Thus, in solid state memory, error correction codes have become indispensable. Therefore, there is a great need for efficient and effective techniques for performing error correction. SUMMARY
[0004] In embodiments of the present invention, a method is provided that generates log-likelihood ratios (LLRs) using information from hard decoding without the use of additional reads (ARs) used in soft decoding. In embodiments, the LLRs are generated using hard read data, a checksum from the hard read data, and a count of ones to support early soft decoding. This method provides several advantages over conventional soft decoding processes. First, conventional soft reads typically involve identifying a center or optimal read threshold voltage, and deriving additional read threshold voltages for ARs that are designed to obtain sufficient soft information to determine LLRs. Thus, conventional soft decoding using ARs is a more complex process. Additionally, in conventional decoding flows, it is difficult to obtain sufficient information to generate an LLR table to support early soft decoding.
[0005] According to some embodiments of the present application, a method of decoding a low density parity check (LDPC) codeword, the method comprising: performing, by a system comprising an LDPC decoder, a hard decoding of memory cells of a given page associated with a word line (WL), the hard decoding comprising a first hard read using a predetermined hard read threshold voltage and one or more re-reads. The method further comprises determining, by the system, that the hard decoding based on the hard read has failed. The method further comprises determining, by the system, whether the hard read is a first hard read or a re-read of the given page. Upon determining that the hard read is a first hard read, the system continues to perform a hard decoding of another page. Upon determining that the hard read is a re-read of the given page, the method comprises grouping the memory cells of the given page into bins based on read threshold voltages associated with the hard read and previous hard reads of the given page. The method further comprises determining a parity sum and a count of ones of the memory cells in each bin and using machine learning to compute a LLR for each bin based on the read data, the parity sum and the count of ones of each bin. The generated LLRs can then be used to perform a soft read and soft decoding of the given page.
[0006] In some embodiments, the method further comprises detecting whether the hard read is a first hard read or a re-read by using a pattern matching operation between the read data of the current hard read and the previous hard read from the given page. In some embodiments, the pattern matching operation comprises performing a sum operation as follows:
[0007] SUM(XOR(incoming_data,saved_data)),
[0008] wherein:
[0009] incoming_data is the read data of the current hard read from the given page;
[0010] saved_data is the read data of the previous hard read from the given page;
[0011] XOR is an exclusive OR operation; and
[0012] SUM is an operation that determines a sum of bits that are 1.
[0013] In some embodiments, the parity sum is based on weights of non-zero syndromes of the codeword.
[0014] In some embodiments, the machine learning comprises using a neural network (NN). In some embodiments, the NN is a deep neural network (DNN) that receives the parity sum and the count of ones as input and determines weighting factors for computing the optimal LLR.
[0015] According to some embodiments of the application, there is provided a method of determining LLRs for soft decoding based on information obtained from hard decoding in a memory system configured to perform hard decoding and soft decoding of an LDPC codeword. The method comprises performing hard decoding of the codeword in a page, the hard decoding including a first hard read using a predetermined hard read threshold voltage and one or more re-reads, and grouping memory cells in the page into a plurality of bins based on read threshold voltages used for the first hard read and the one or more re-reads. The method further comprises calculating a parity check sum and a count of ones of the memory cells in each bin, and determining the LLRs of the memory cells in each bin based on the read data, the parity check sum and the count of ones of each bin.
[0016] In some embodiments, the method further comprises determining the LLRs of each bin using machine learning.
[0017] In some embodiments, the machine learning comprises a NN.
[0018] In some embodiments, the parity check sum comprises weights of non-zero syndromes of the codeword for LDPC decoding.
[0019] In some embodiments, the count of ones of a given bin comprises a number of memory cells in the bin whose cell values are ones.
[0020] In some embodiments, the method further comprises determining the LLRs without using AR, wherein AR comprises determining additional read threshold voltages for determining the LLRs from read data from the hard read.
[0021] In some embodiments, the method further comprises detecting whether a hard read is a first hard read or a re-read by using a pattern matching operation between read data from a current hard read and a previous hard read of a given page.
[0022] In some embodiments, the method further comprises determining the LLRs for soft decoding based on information obtained from the hard decoding after determining that a hard read of a given page is a re-read of the given page.
[0023] According to some embodiments of the application, a memory system includes a memory cell and a memory controller coupled to the memory cell for controlling operation of the memory cell, the operation including hard decoding and soft decoding of an LDPC codeword. The memory controller is configured to perform hard decoding of the codeword in a page, the hard decoding including a first hard read using a predetermined hard read threshold voltage and one or more re-reads. The memory controller is further configured to group the memory cells in the page into a plurality of bins based on the read threshold voltages used for the first hard read and the one or more re-reads. The memory controller is further configured to compute a parity sum and a count of ones of the memory cells in each bin and determine LLRs of the memory cells in each bin based on the read data, the parity sum and the count of ones of each bin.
[0024] In some embodiments of the memory system, the memory controller is further configured to determine the LLRs of each bin using a DNN.
[0025] In some embodiments of the memory system, the parity sum includes weights of non-zero syndromes of the codeword for the LDPC decoding.
[0026] In some embodiments of the memory system, the count of ones of a given bin includes a number of memory cells in the bin whose values are ones.
[0027] In some embodiments, the memory system further includes:
[0028] a re-read detection unit configured to detect whether a hard read is a first hard read or a re-read by using a pattern matching operation between read data from a current hard read and a previous hard read of a given page; and
[0029] an LLR generation unit configured to determine LLRs for soft decoding based on information obtained from the hard decoding after determining that the hard read of the given page is a re-read of the given page.
[0030] In some embodiments, the re-read detection unit is configured to use a summing operation as follows for the pattern matching operation:
[0031] SUM(XOR(incoming_data,saved_data)),
[0032] where:
[0033] incoming_data is read data from a current hard read of a given page;
[0034] saved_data is read data from a previous hard read of the given page;
[0035] XOR is an exclusive OR operation; and
[0036] SUM is an operation that determines the sum of bits that are 1. BRIEF DESCRIPTION OF DRAWINGS
[0037] An understanding of the nature and advantages of various embodiments can be realized by reference to the following drawings. In the drawings, like reference numerals can designate similar structures or features. Further, various components of the same type can be distinguished by adding a dash and a second numeral to the reference designation of the component with the second numeral referring to a different drawing figure. If only the first numeral is used to designate the component throughout the specification, the description is true of any or all of the like components and like suffixes can refer to like components.
[0038] Figure 1 An exemplary high-level block diagram of an error correction system is shown in accordance with certain embodiments of the present disclosure;
[0039] Figure 2A An example parity check matrix is shown, and Figure 2B An exemplary bipartite graph corresponding to a parity check matrix is shown in accordance with certain embodiments of the present disclosure;
[0040] Figure 3 An example diagram for terminating LDPC iterative decoding based on syndrome and maximum number of iterations is shown in accordance with certain embodiments of the present disclosure;
[0041] Figure 4 An exemplary architecture of a computing system 400 is shown in accordance with certain embodiments of the present disclosure;
[0042] Figure 5 is a simplified diagram showing a cell voltage profile of a memory device with 3-bit triple-level cell (TLC) in a flash memory device in accordance with certain embodiments of the present disclosure;
[0043] Figure 6 is a simplified diagram showing determination of LLR based on a cell voltage profile of a memory device with adjacent program voltage (PV) levels in a flash memory device in accordance with certain embodiments of the present disclosure;
[0044] Figure 7 is a simplified flow diagram showing a method for operating a storage system in accordance with certain embodiments of the present disclosure;
[0045] Figure 8 is a simplified block diagram showing an LLR generator in accordance with certain embodiments of the present disclosure;
[0046] Figure 9is a simplified flowchart showing a method for generating LLRs implemented in decoding a LDPC codeword according to certain embodiments of the present disclosure;
[0047] Figure 10 Three simplified diagrams showing examples of generating LLRs using a checksum and a count of ones according to certain embodiments of the present disclosure are shown;
[0048] Figure 11 Block diagrams of a serial DNN-based LLR generator and a parallel DNN-based LLR generator according to certain embodiments of the present disclosure are shown;
[0049] Figure 12 is a simplified block diagram showing a DNN unit that can also be used to Figure 8 generate LLRs according to certain embodiments of the present disclosure;
[0050] Figure 13 is a simplified block diagram of a solid state storage system according to certain embodiments of the present disclosure; and
[0051] Figure 14 is a simplified block diagram showing a device that can be used to implement various embodiments according to certain embodiments of the present disclosure. DETAILED DESCRIPTION
[0052] Error correction codes are often used for communication and for reliable storage in media such as CDs, DVDs, hard disks and RAM, flash memory, etc. Error correction codes can include LDPC codes, turbo product codes (TPC), Bose-Chaudhuri-Hocquenghem (BCH) codes, Reed-Solomon codes, etc.
[0053] Figure 1 is a high level block diagram showing an exemplary LDPC error correction system according to certain embodiments of the present disclosure. As Figure 1 shown, an LDPC encoder 110 of an error correction system 100 can receive information bits, which include data intended to be stored in a storage system 120. LDPC encoded data can be generated by the LDPC encoder 110 and can be written to the storage system 120. The encoding can use an encoder-optimized parity check matrix H' (112).
[0054] In various embodiments, the storage system 120 can include multiple storage types or media. Errors can occur in the data storage or communication channel. For example, errors can be caused by, for example, inter-cell interference and / or coupling. When the stored data is requested or otherwise needed (e.g., by an application or user that stored the data), the detector 130 can receive the data from the storage system 120. The received data can include some noise or errors. The detector 130 can include a soft output detector and a hard output detector, and can perform detection on the received data and output decisions and / or reliability information.
[0055] For example, a soft output detector outputs reliability information as well as a decision for each detected bit. On the other hand, a hard output detector outputs a decision for each bit without providing corresponding reliability information. As an example, a hard output detector can output a decision that a particular bit is a “1” or a “0” without indicating the confidence or certainty of the detector in that decision. In contrast, a soft output detector outputs a decision as well as reliability information associated with the decision. In general, a reliability value indicates the confidence of the detector in a given decision. In one example, a soft output detector outputs an LLR, where the sign indicates the decision (e.g., a positive value corresponds to a decision of “1” and a negative value corresponds to a decision of “0”), and the magnitude indicates the confidence of the detector in that decision (e.g., a larger magnitude indicates a higher reliability or confidence).
[0056] The decisions and / or reliability information can be passed to an LDPC decoder 140 that can use the decisions and / or reliability information to perform LDPC decoding. A soft LDPC decoder can utilize both the decisions and the reliability information to decode the codeword. A hard LDPC decoder can utilize only the decision values from the detector to decode the codeword. The decoded bits generated by the LDPC decoder 140 can be passed to the appropriate entity (e.g., the user or application that requested it). The decoding can utilize a parity check matrix H 142, which can be optimized for the LDPC decoder 140 by design. With proper encoding and decoding, the decoded bits will match the information bits. In some embodiments, the parity check matrix H 142 can be the same as the encoder-optimized parity check matrix H’ 112. In some embodiments, the encoder-optimized parity check matrix H’ 112 can be modified from the parity check matrix H 142. In some embodiments, the parity check matrix H 142 can be modified from the encoder-optimized parity check matrix H’ 112.
[0057] LDPC codes are typically represented by a bipartite graph comprising two sets of nodes. One set of nodes, the variable or bit nodes, correspond to elements of a codeword, and the other set of nodes, the check nodes, correspond to a set of parity check constraints that the codeword satisfies. Connections between variable nodes and check nodes are defined by a parity check matrix H (e.g., Figure 1
[0058] Further details of LDPC decoding can be found in U.S. Patent Application No. 15 / 903,604, entitled “Min-Sum Decoding of LDPC Codes,” filed February 23, 2018, now U.S. Patent No. 10,680,647, which is assigned to the assignee of the present application and is expressly incorporated by reference herein in its entirety.
[0059] In various embodiments, the illustrated systems can be implemented using a variety of technologies, including application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), and / or general purpose processors (e.g., Advanced RISC Machines (ARM) cores).
[0060] LDPC codes are typically represented by a bipartite graph. One set of nodes, the variable or bit nodes, correspond to elements of a codeword, and the other set of nodes, the check nodes, correspond to a set of parity check constraints that the codeword satisfies. Typically, the edge connections are chosen at random. The error correction capability of an LDPC code can be improved if short length cycles are avoided in the graph. In (r, c) regular codes, each of n variable nodes (V1, V2,..., Vn) is connected to r check nodes, and each of m check nodes (C1, C2,..., Cm) is connected to c bit nodes. In irregular LDPC codes, the check node degrees are not uniform. Similarly, the variable node degrees are not uniform. In QC-LDPC codes, the parity check matrix H is constructed as blocks of p x p matrices, such that one bit in a block participates in only one check equation in the block, and each check equation in a block involves only one bit in the block. In QC-LDPC codes, a cyclic shift of a codeword by p brings another codeword. Here p is the size of the rectangular matrix, which can be a zero matrix or a circulant matrix. This is a generalization of cyclic codes, where a cyclic shift of a codeword by 1 brings another codeword. The blocks of p x p matrices can be zero matrices or cyclic shift unit matrices of size p x p.
[0061] Figure 2A An exemplary parity check matrix H 200 is shown, and Figure 2B An exemplary bipartite graph corresponding to parity check matrix H 200 is shown in accordance with certain embodiments of the present disclosure. In this example, parity check matrix H 200 has six column vectors and four row vectors. Network 202 shows a network corresponding to parity check matrix H 200 and represents a bipartite graph. Various types of bipartite graphs are possible, including, for example, Tanner graphs.
[0062] Generally, variable nodes in network 202 correspond to column vectors in parity check matrix H 200. Check nodes in network 202 correspond to row vectors of parity check matrix H 200. Interconnections between nodes are determined by the values of parity check matrix H 200. Specifically, a "1" indicates that the corresponding check node and variable node have a connection. A "0" indicates that there is no connection. For example, the "1" in the leftmost column vector and the second row vector from the top in parity check matrix H 200 corresponds to a connection between variable node 204 and check node 210.
[0063] Message passing algorithms are generally used to decode LDPC codes. There are several variants of message passing algorithms in the prior art, such as the MS algorithm, scaled MS algorithm, etc. Generally, any variant of a message passing algorithm can be used in an LDPC decoder without departing from the teachings of the present disclosure. Message passing uses a network of variable nodes and check nodes as shown in Figure 2B As shown in Figure 2A The connections between variable nodes and check nodes are described by and correspond to the values of parity check matrix H 200.
[0064] A hard decision message passing algorithm can be performed. In a first step, each of the variable nodes sends a message to the one or more check nodes to which it is connected. In this case, the message is the value that each of the variable nodes believes to be the correct value.
[0065] In a second step, each of the check nodes uses the information previously received from the variable nodes to compute a response to be sent to the variable nodes to which it is connected. The response message corresponds to the value that the check node believes the variable node should have based on the information received from the other variable nodes connected to that check node. This response is computed using a parity check equation that forces the sum of the values of all variable nodes connected to a particular check node to be zero (modulo 2).
[0066] At this point, if all the equations for all the check nodes are satisfied, the decoding algorithm declares that the correct codeword has been found and terminates. If the correct codeword has not been found, the iteration continues with another update of the variable nodes using the messages they received from the check nodes to decide whether the bits in their positions are zero or one according to the majority rule. The variable nodes then send this hard decision message to the check nodes connected to them. As further illustrated in the following figure, the iteration continues until the correct codeword is found, a certain number of iterations is performed according to the syndrome of the codeword (e.g., the decoded codeword), or a maximum number of iterations is performed without finding the correct codeword. It should be noted that soft decision decoders work similarly; however, each of the messages passed between the check nodes and the variable nodes also includes the reliability of each bit.
[0067] An exemplary message passing algorithm can be performed. In this example, L(q ij ) represents the message sent by a variable node v i to a check node c j ; L(r ji ) represents the message sent by a check node c j to a variable node v i ; and L(c i ) represents the initial LLR value for each variable node v i . The variable node processing for each L(q ij ) can be accomplished by the following steps:
[0068] (1) Read L(c i ) and L(r ji ) from memory.
[0069] (2) Compute
[0070] (3) Compute each L(Qi-sum) - L(r ji ).
[0071] (4) Output L(Qi-sum) and write back into memory.
[0072] (5) If this is not the last column of memory, go to step 1, increment i by 1.
[0073] (6) Compute parity sums (e.g., syndrome). If they are all equal to zero, the number of iterations reaches a threshold, and, the parity sums are greater than another threshold, or the number of iterations equals a maximum limit, stop; otherwise, perform check node processing.
[0074] The check node processing for each L(r ji ) can be performed as follows:
[0075] (1) Read a column of qij from memory.
[0076] (2) Compute L(Rj-sum) as follows:
[0077]
[0078] a ij = sign(L(q ij )), b ij = |L(q ij )|,
[0079]
[0080] (3) Compute the
[0081]
[0082] (4) Write L(rji) back to memory.
[0083] (5) If this is not the last row of memory, go to the first step, increment j by 1.
[0084] Figure 3 An example diagram 300 for terminating LDPC iterative decoding based on syndrome and maximum number of iterations according to certain embodiments of the present disclosure is shown. Termination occurs according to a syndrome of a codeword being zero or a number of iterations reaching a maximum number.
[0085] As shown in diagram 300, let x = [x0, x1,..., x N-1 ] be a bit vector, and H = [h i,j ] be an M x N LDPC matrix with binary value h i,j at the intersection of the i-th row and j-th column. Each row of H then provides a parity check on x. If x is a codeword of H, then xH T = 0 due to the LDPC code structure. Assume x is transmitted over a noisy channel, and the corrupted channel output is y = [y0, y1,..., y N-1 ] and its hard decision is z = [z0, z1,..., z N-1 ]. The syndrome of z is a binary vector of weight ||s|| computed by s = [s0, s1,..., s N-1 ] = zH T . The weight ||s|| represents the number of unsatisfied check nodes, and is also called the check sum because . Assume z (j) = [z0, z1,..., z N-1 ] is the hard decision of the j-th iteration, and the syndrome vector of the j-th iteration is Then ||s (j) is the checksum for the jth iteration.
[0086] As further shown in FIG. 300, the iterative decoding terminates when the checksum is zero (indicated by s ( j) = 0) or when the checksum is non-zero and the number of iterations reaches a predetermined maximum number of iterations (indicated by j = It max , where It max is the maximum number of iterations). Otherwise, the iterative decoding is repeated.
[0087] Figure 4 An exemplary architecture of a computing system 400 is shown in accordance with certain embodiments of the present disclosure. In an example, the computer system 400 includes a host 410 and one or more SSDs 420. Each solid state drive (SSD) can be a storage system that can include a memory unit and a memory controller coupled to the memory unit for controlling operations of the memory unit. Examples of the storage system are described below in connection with Figure 9 FIGS. 300-305. Each memory unit can be an m-layer cell (MLC), where m is an integer. In some embodiments, as shown below in Figure 3 、 Figure 4 and Figure 6 , the memory unit is arranged in m pages, each of the m bits of a given memory unit providing data for a respective page of the m pages.
[0088] The host 410 represents a customer storing data in the SSDs 420. The data is stored in the SSDs as ECC-protected codewords. For example, the SSDs can include an ECC encoder (e.g., the LDPC encoder 110 of Figure 1 ).
[0089] The host 410 can receive client requests for the client data stored in the SSDs 420. In response, the host sends data read commands 412 to the SSDs 420 as applicable. Each of the SSDs 420 processes the received data read commands and sends a response 422 to the host 410 upon completion of the processing. The response 422 can include read data and / or a decoding failure. In an example, each of the SSDs includes an ECC decoder (e.g., the LDPC decoder 140 of Figure 1 ). Processing the data read commands and sending the response 422 includes decoding, by the ECC decoder, the codewords stored in the SSD to output the read data and / or the decoding failure. The ECC decoder can include multiple decoders implementing a message passing algorithm, such as the MS algorithm.
[0090] Generally, SSDs can be storage devices that store data persistently or cache data temporarily in non-volatile semiconductor memory, and are intended for use in storage systems, servers (e.g., servers within data centers), and direct-attached storage (DAS) devices. Increasingly, applications require higher data throughput and lower transaction latency, and SSDs are used as a viable storage solution to improve performance, efficiency, reliability, and reduce overall operating expenses. SSDs generally use NAND flash memory, which can provide higher performance and lower power consumption than rotating hard disk drives (HDDs). NAND flash memory has a number of inherent issues associated with it; two of the most important aspects include: a limited expected lifetime due to NAND flash cell wear-out during repeated writes; and a naturally occurring error rate. SSDs can be designed and manufactured according to a set of industry standards that define specific performance specifications, including latency specifications, to support more strenuous write workloads, more extreme environmental conditions, and recovery from higher bit error rates (BERs) than client SSDs (e.g., personal computers, laptops, and tablets).
[0091] In the following description, techniques for improving determination of LLRs in multi-layer memory devices are described. These techniques are applicable to any soft decoder that uses LLRs in decoding.
[0092] Figure 5 is a simplified diagram 500 illustrating a cell voltage distribution of a memory device with 3-bit TLC in a flash memory device according to some embodiments of the present application. Flash memory modulates cells into different states or PV levels by using a program operation, such that each cell stores multiple bits. Data can be read from NAND flash memory by applying a read reference voltage to the control gate of each cell to sense the threshold voltage of the cell. For TLC flash memory, there are 8 PV levels, each level corresponding to a unique 3-bit tuple. As shown, the first, second, and third bits of the cell are grouped into a least significant bit (LSB) page, a center significant bit (CSB) page, and a most significant bit (MSB) page, respectively. Figure 5
[0093] In Figure 5 , the target cell PV for the erase state is shown as "PV0", and the PVs for the seven program states are shown as "PV1" through "PV7". The distribution of cell voltages or cell threshold voltages for each of the eight data states is represented as a bell curve associated with each PV. Cell threshold voltage spread can be caused by differences in cell characteristics and operation history. In Figure 5 , each cell is configured to store eight data states represented by the following three bits: MSB, CSB, and LSB. Figure 5 The seven read thresholds, labeled "Vrl," "Vr2,"..., "Vr7," are also shown in the middle as the reference voltages used to determine the data stored in the memory cells. For example, two thresholds Vrl and Vr5 are used to read the MSB. If the voltage stored by the cell (PV) is less than Vrl or greater than Vr5, the MSB is read as 1. If the voltage is between Vrl and Vr5, the MSB is read as 0. Two thresholds Vr3 and Vr7 are used to read the LSB. If the voltage stored by the cell is less than Vr3 or greater than Vr7, the LSB is read as 1. If the voltage is between Vr3 and Vr7, the LSB is read as 0. Similarly, three thresholds Vr2, Vr4, and Vr6 are used to read the CSB.
[0094] Figure 6 is a simplified diagram showing determination of LLRs based on the cell voltage distribution of memory devices with adjacent PV levels in a flash memory device according to some embodiments of the application. For example, in Figure 6 , the cells of level 0 and level 1 PVs are shown as distributions 601 and 602, respectively. Multiple read operations are performed using different AR threshold voltages (Ar1 to Ar7) to divide the cells into different bins, numbered 0 to 7. The AR threshold voltages can be selected to facilitate determination of the LLRs. It can be assumed that flash memory cells falling into the same bin have the same threshold voltage and thus map to the same LLR value corresponding to the respective voltage sub-region. In Figure 6 , an example where there are eight bins corresponding to the respective voltage sub-regions, the LLRs can be represented in three bits. In some embodiments, the 3-bit LLR values can be represented by 000, 001, 010,..., and 111.
[0095] Figure 7is a simplified flowchart illustrating a method for decoding in a storage system. In a NAND flash system, after a read command is received, a series of data recovery steps are typically run for the purpose of retrieving noiseless data from the NAND flash system. In 710, these data recovery steps can include hard decoding processing and soft decoding processing. The hard decoding processing can include a series of hard reads, which can include a first hard read and one or more hard re-reads. The first hard read attempts what is sometimes referred to as a “history read.” In an example, the history read uses the threshold Vt used in a previous successful read in which the decoder successfully recovered noiseless data from the NAND page flash system. The history read information is maintained separately per physical block or physical die, and will be updated if the decoding fails and a different Vt is used in a subsequent step that will result in successful decoding. Thus, the history read represents the first read attempt in response to a new read command. If the history read fails, a re-read or read retry is performed. This re-read or read retry is sometimes referred to as a high priority read retry (HRR).
[0096] In 720, the HRR can include a re-read that uses a series of predetermined fixed threshold Vt that remains the same throughout the life cycle of the NAND flash system. For example, five to ten HRR read attempts can be performed before proceeding to the next step. For each HRR read, a decoding operation is performed.
[0097] The system can perform multiple reads to find the best center Vt for soft reads. For example, the system can find the center Vt at the minimum of the valley in the read data distribution. The hard decoding can be performed, for example, using MS hard decoding or bit flipping (BF) hard decoding.
[0098] If all HRR reads fail, it can be determined that the hard decoding has failed and soft decoding 730 is started. In the first part of soft read and soft decoding (SR / SD) 730, the system finds the center Vt, which is the best Vt that separates the two states, and then takes additional Vt around the center Vt to perform additional ARs to generate LRRs for each bin. The AR threshold voltage can be identified. All read attempts before the soft read can generate noisy hard read information, which can be used with the AR information to generate the LLRs for soft decoding. The successful Vt is updated as the history read, as indicated by marker 732.
[0099] In embodiments of the present invention, a method of using information from hard decoding to generate LLRs without using AR in soft decoding is described. At each read attempt before soft read, a particular Vt is used to read data from NAND. A method for generating LLR tables using information generated during hard read to support early soft decoding is described. Previous hard read data is combined with current hard read data to generate LLRs and fed to a soft input decoder. This processing can be employed for hard reads other than the first hard read (history read) and improves the error correction capability of the decoder in hard read.
[0100] There are two problems in existing systems that support the above early soft decoding. The first challenge is that there is no simple way to derive LLR tables to support early soft decoding in the existing decoding flow. The second problem is that in order to distinguish the first read or those re-reads (second read and after), the data path must provide an interface signal to inform the LLR generation module. This complicates the data path design and reduces the modularity of the LLR generation and ECC modules, and makes it more difficult to support different applications.
[0101] Embodiments of the present invention include a machine learning based LLR generation scheme with re-read detection. A machine learning based approach is used for early LLR table generation with re-read detection. The best LLR table is selected based on the Vt used in previous read attempts, syndrome, and the count of ones information. Further, a command detection module is used to detect whether the read command is the first read or a re-read without the need to send an additional signal from the data path to the LLR generation module.
[0102] The inventors have observed that the best Vt between all different word lines (WLs) varies greatly for different physical locations on the wafer, retention conditions, and read disturb counts. For example, in TLC storage devices, as described above, each physical page is divided into three logical pages: MSB, CSB, and LSB. To read the voltage in LSB, 11000011 requires two threshold voltages: V2 and V6. The inventors have observed that V2 and V6 vary greatly depending on the page location, starting age, remaining age, erase- write count, etc. Therefore, the read voltage needs to be optimized for each page throughout the lifetime of the storage device.
[0103] Because of this difference, when performing hard read according to historical read, different bit errors can be obtained depending on which WL is being read. Large variations have also been observed for re-reads following a predetermined read threshold voltage. Therefore, a static LLR table is not sufficient. In embodiments of the present invention, the LLR table can be updated by what is observed in each individual specific WL. In some embodiments, a WL is associated with a cell in a page.
[0104] In soft decoding, soft information such as LLR is generated using AR and the read threshold voltage selected for valid LLR generation. For early soft decoding, LLR generation is difficult to perform since AR is not available at early read. Therefore, PV distributions at different valleys are mixed together without AR. The shape and location of each PV play a role in deciding the best LLR table. Also, it is not guaranteed that one Vt will always be on the left / right side of another Vt. Further, the Vt used in early hard read has randomness. In early hard read, the read threshold Vt is decided by historical read and HRR entries, where HRR entries are pre-selected, can be arbitrary, and can not reflect the current state of the cell related to the actual PV. Without knowing the actual PV distribution, it is difficult to determine a better LLR table given the specific Vt used in previous reads.
[0105] Some embodiments of the present invention provide a method for error correction decoding including generating LLR for soft decoding using only information from hard decoding without AR used in conventional soft decoding, and a storage system. In some embodiments, the storage system includes a memory cell and a memory controller coupled to the memory cell, the memory controller configured to control operations of the memory cell including hard decoding and soft decoding of a LDPC codeword. Examples of such storage system are described below in connection with Figure 13 and 14 A memory controller is configured to perform hard decoding of a codeword in a page including a first hard read using a predetermined hard read threshold voltage and one or more re-reads. The memory controller is also configured to group memory cells in the page into a plurality of bins based on read threshold voltages used for the first hard read and the one or more re-reads, and to compute a parity sum and a count of ones for memory cells in each bin. The memory controller is further configured to determine LLR for memory cells in each bin based on read data, parity sum, and count of ones for each bin. In some embodiments, as described below in connection with Figure 8 A memory controller can include an LLR generation block. The method of generating LLR is further illustrated below in connection with Figure 9
[0106] Figure 8 This is a simplified block diagram illustrating an LLR generator according to some embodiments of the present invention. Figure 8 A system 800 for generating LLRs is shown; system 800 can be part of a decoder for a storage system. For example... Figure 8 As shown, the LLR generator 800 includes a reread detection unit 810, a DNN unit 820, an LLR table generation unit 830, an interval label generation unit 840, and an LLR generation unit 850. (Refer to...) Figure 9 The flowchart in the document further illustrates the operation of the LLR generator 800 used to generate LLRs.
[0107] Figure 9 This is a simplified flowchart illustrating a method for generating an LLR during the decoding of LDPC codewords, according to some embodiments of the present invention. (Refer to below) Figure 8 LLR generator 800 and Figure 10 The example shown illustrates method 900. For example... Figure 8 As shown, the method for decoding LDPC codewords includes: at 910, a system including an LDPC decoder performs hardware decoding on a given page of a memory cell associated with WL. Combined with... Figures 1 to 4 and Figures 13 to 14 An example system including an LDPC decoder is described. Hard decoding may include a first hard read using a predetermined hard read threshold voltage, followed by one or more rereads.
[0108] In 920, the method includes the system determining that hard decoding based on a hard read has failed. If hard decoding is successful, the system can continue reading and decoding other pages in the storage system. On the other hand, if hard decoding fails, the conventional method is usually to perform a reread or hard read retry, followed by software decoding. Software decoding typically involves determining the LLR using the AR after a soft read and soft decoding. However, as mentioned above, it is often difficult to obtain information for software decoding in the early stages of decoding. In embodiments of the present invention, the LLR can be generated based on hard reread information.
[0109] At 930, the system determines whether a hard read is the first hard read or a reread of a given page. This is because a first hard read has not yet generated enough information for effective LLR generation. Therefore, this LLR generation method is only applicable to hard rereads. In this regard, Figure 8The re-read detection unit 810 in the read retry detection unit 810 determines whether the current hard read is a first hard read or a re-read of a given page. In some embodiments, the system has a "current read" buffer that stores hard read data for a read that failed in a previous read and has not yet been re-read. To detect whether a read command is a first read or a re-read, a pattern matching process is used to compute a similarity between the incoming data and the data stored in the current read buffer. In embodiments, the pattern matching process can be implemented using the following summation operation:
[0110] SUM(XOR(incoming_data,saved_data)), where:
[0111] incoming_data is the read data from the current hard read of a given page;
[0112] saved_data is the read data from a previous hard read of the given page;
[0113] XOR is the exclusive OR operation; and
[0114] SUM is an operation that determines the sum of bits that are 1.
[0115] As an example, a page can have 4K bytes of memory cells and a codeword can have 256 bits. Then the pattern matching expression is as follows:
[0116] SUM(XOR(incoming_data(0:255),saved_data(0:255))).
[0117] The pattern matching operation effectively computes the sum of the number of 1s in the comparison of the incoming data to the saved data. In other words, the sum is the number of bits that match between the incoming data and the saved data. Typically, the original BER of a page is less than 1%. Thus, the data pattern in the re-read data should be similar across multiple reads. On the other hand, a random codeword will likely match the data in the current read buffer with about 50% probability. Thus, if the incoming data is for a different page than the saved data, the sum can be 128 bits or 50% in the case of a codeword length of 256 bits. If the incoming data is for a re-read of the same page as the saved data, then the sum should be lower, e.g., 1%. In embodiments, a pattern match can be declared if the percentage of matching bits is higher than, e.g., 75% or 192 out of 256 bits. Once a pattern match is declared, the read count is updated to indicate how many reads have been performed for that codeword.
[0118] At 940, upon determining that the hard read was a first read and not a re-read, the system proceeds with a re-read. As noted above, the first hard read did not produce enough information for valid generation of LLRs. Because the hard read has failed, the system can proceed with a re-read. Optionally, the system can take other actions. In some embodiments, for the first read, the system can perform a hard decoding of the page, where the sign of the LLR is determined from the read data and the magnitude is set to some fixed value.
[0119] At 950, upon determining that the hard read was a re-read of the given page, the memory cells of the given page are grouped into bins based on the read threshold voltages associated with the previous hard read and this hard read of the given page. Figure 8 The bin marker generation unit 840 in 1000 is used to generate a marker for each bin.
[0120] Figure 10 Three simplified diagrams showing examples of generating LLRs using checksums and counts of ones, according to some embodiments of the application, are shown. Figure 10 1010 of 1000 shows a cell voltage distribution for a memory device with 3-bit TLC in a flash memory device, according to some embodiments of the application. As with 1000, Figure 5 Similar to the TLC cell voltage distribution in 1000, Figure 10 1010 of 1000 shows eight PV levels. As explained above in connection with 1000, Figure 5 Two or three read threshold voltages are used to determine the bit values of the MSB, CSB and LSB. In 1010 of 1000, Figure 10 In 1010 of 1000, the two read threshold voltages marked VT0 are used for the first hard read, the two read threshold voltages marked VT1 are used for the first re-read, and the two read threshold voltages marked VT2 are used for the second re-read. VT0, VT1 and VT2 are pre-selected for the hard read. After the second re-read, the cells can be divided into eight groups according to their PV levels, denoted by cells A, B, C, D, E, F, G and H, respectively.
[0121] At 960, the system determines the checksums of parity and counts of ones for the memory cells in each bin. In embodiments of the application, the checksums of parity are based on the weights of the non-zero syndromes of the codeword. For a noisy parity check matrix, all the parity check equations that do not satisfy produce "1" and those parity check equations that satisfy the parity check produce "0". In linear codes such as LDPC, even if the decoding is not successful, useful information can still be derived from the checksums that can provide information about how many errors exist. For example, given two reads, both of which do not produce a correct codeword, the information that one codeword has more errors than the other can be used for decoding, e.g., to determine the next Vt, to compute LLRs, etc. The above is explained in connection withFigures 1 to 3 More details of the parity sum are described.
[0122] The count of ones is the number of cells with a value of one in each interval. In a storage system that randomizes data before writing the data, the number of ones and the number of zeros is expected to be approximately 50% of the data bits. Both the count of ones and the parity sum can be used to determine the LLR value for a given interval. For example, a smaller parity sum indicates fewer errors and can represent a higher likelihood. Further, a ratio of the count of ones close to 50% can represent a higher likelihood.
[0123] In Figure 10 In the example of 1010 of TLC Page, three reads have been performed on one of the TLC pages. The hard read information is used to generate a 3-bit LLR, which can have a value from -3 to +3. Figure 10 1020 of shows an example of applying three conceptual Vts to a single layer cell (SLC) model. Each Vt is associated with its parity sum and a percentage of ones count. As Figure 10 As shown in 1020 of, a particular cell will fall into the same conceptual interval and thus is assigned the same interval marker. For example, cell C and cell H are grouped into interval #0, cell A and cell F are grouped into interval #1, cell B and cell E are grouped into interval #2, and cell D and cell G are grouped into interval #3. As shown, the read value of 111 can be in either cell C or cell H, and in a conventional LLR generation method used for soft decoding, without AR, there is not enough information to distinguish between cell C and cell H.
[0124] In Figure 10 In the example of 1020 of, the first hard read associated with read threshold voltage VT0 is characterized by a parity sum (CS) of 470 and a percentage of ones count of 48%. Similarly, the first re-read associated with read threshold voltage VT1 is characterized by a parity CS of 500 and a percentage of ones count of 55%. Further, the second re-read associated with read threshold voltage VT2 is characterized by a parity CS of 500 and a percentage of ones count of 45%. As described above, a smaller parity CS represents a higher likelihood, and a ratio of the count of ones close to 50% represents a higher likelihood. Because the order of Vts can be different at different valleys, the values of Vts are not used to generate the LLR values.
[0125] Figure 10 1030 of is an example of LLR values generated using the above described method. In Figure 10 In 1030 of, eight cells are listed as A, B, C, D, E, F, G, and H, and three hard operations are designated as R0, R1, and R2. Figure 10Section 1030 also lists the interval number BIN and LLR value. The hard read data for each cell is determined by using the read threshold voltage VT0 for hard read operation R0, the read threshold voltage VT1 for hard read operation R1, and the read threshold voltage VT2 for hard read operation R2. Figure 10 The hard read data for each cell in 1030 is listed as follows: A(011), B(010), C(111), D(000), E(010), F(110), G(000), and H(111). Using the information provided by the checksum and the count of 1s, the LLR value for each cell can be determined as follows: A(-1), B(1), C(-2), D(+3), E(1), F(-1), G(+2), and H(-3).
[0126] Therefore, in embodiments of the invention, information generated during hard reads can be used to generate LLR values without using AR. LLR can be estimated using information such as read data, parity sum, and the count of 1s (or the percentage of 1s). For example, the statistical distribution of cell values can be associated with read data, checksum, and the count of 1s. This information can be used to estimate LLR. The results can be included in a lookup LLR table for use in soft decoding.
[0127] like Figure 9 As shown in flowchart 970, embodiments of the present invention include a method for calculating the LLR for each interval based on read data, checksum, and count of 1s for each interval, using machine learning. An example of machine learning is based on neural network (NN) learning. Figure 8 In the example, DNN unit 820 receives the read count of the current read, the count of 1s, and the checksum to generate an LLR value. Therefore, it stores the count of 1s and the checksum of all previous read attempts, and increments the read count by 1 for each read.
[0128] In this example, DNN unit 820 is used to determine the LLR value that needs to be assigned to a given interval marker at a given read count. The count of 1s and the dimension of the checksum may differ depending on the current read count. At the start of the LLR generation process, an inference is performed once for each interval marker value, and the association between the interval marker and the LLR value is stored in an LLR table. During LLR generation, the LLR table is repeatedly applied during runtime to generate new LLR values and feed them to the decoder. In some embodiments, the NN can be applied to perform offline machine learning. See below for reference. Figure 11 An example describing a neural network.
[0129] Re-reference Figure 8 The flowchart above shows that at 980, soft reading and soft decoding of a given page are performed. (Refer to the above...) Figures 1 to 4An example of software decoding is described.
[0130] exist Figure 8 In the proposed approach, multiple DNN inferences are performed, with each inference generating an LLR value for a specific interval label from the DNN. The advantage of this approach is that it reduces the size of the DNN and can be applied to product lines with lenient Quality of Service (QoS) requirements. For enterprise applications with stringent QoS requirements, a parallel DNN-based LLR generator may be desirable. Figure 11 Some examples are shown.
[0131] Figure 11 Block diagrams of an LLR generator based on a serial DNN and an LLR generator based on a parallel DNN, according to embodiments of the present invention, are shown. Figure 11 As shown, 1110 is an LLR generator based on a serial DNN, which uses interval markers as input, along with a count of 1s and a checksum, to compute the LLR for each interval marker. Figure 11 The 1120 is an LLR generator based on a parallel DNN that simultaneously computes the LLR for all intervals. The DNN 1120 receives a count of 1s and a checksum as input. Parallel DNN-based LLR generators are faster but more complex.
[0132] Figure 12 The embodiments shown in this paper can also be used to illustrate the invention. Figure 8 A block diagram of an exemplary two-layer feedforward NN for generating LLR in the DNN unit 820. Figure 12 In the example shown, the feedforward NN 1200 includes an input port 1210, a hidden layer 1220, an output layer 1230, and an output port 1240. In this network, information moves in only one direction, forward, from the input node, through the hidden nodes, and then to the output node. Figure 12 In this context, W represents the weighted vector, and b represents the bias factor.
[0133] In some embodiments, the hidden layer 1220 may have sigmoid neurons, and the output layer 1230 may have softmax neurons. A sigmoid neuron has an output relationship defined by a sigmoid function, which is a mathematical function with a characteristic S-shaped curve or sigmoid curve. The sigmoid function has a domain containing all real numbers, and depending on the application, the return value typically increases monotonically from 0 to 1 or optionally from -1 to 1. Various sigmoid functions can be used as activation functions for artificial neurons, including logistic functions and hyperbolic tangent functions.
[0134] In the output layer 1230, the softmax neurons have an output relationship defined by a softmax function. The softmax function, or normalized exponential function, is a generalization of the logistic function that "squashes" an arbitrary real-valued K-dimensional vector z into a real-valued K-dimensional vector σ(z) where each entry is in the range (0, 1) and all entries sum to 1. The output of the softmax function can be used to represent a classification distribution, i.e., a probability distribution over K different possible outcomes. The softmax function is commonly used in the final layer of a NN-based classifier. In Figure 12 where W represents a weight vector and b represents a bias factor.
[0135] A NN with many hidden layers is sometimes referred to as a DNN. In some embodiments, the NN is a DNN that receives a checksum and a count of ones as input and determines weighting factors for computing a best LLR.
[0136] To achieve reasonable classification, ten or more neurons can be allocated in the first hidden layer. If more hidden layers are used, any number of neurons can be used in the additional hidden layers. Given more computational resources, more neurons or layers can be allocated. By providing enough neurons in its hidden layers, performance can be improved. More complex networks, such as convolutional NNs or recurrent NNs, can also be applied to obtain better performance. Given enough neurons in its hidden layers, this can do an arbitrary classification of vectors well.
[0137] Figure 13 is a simplified block diagram illustrating a solid state storage system in accordance with certain embodiments of the present disclosure. As shown, the solid state storage system 1300 can include a solid state storage device 1350 and a storage controller 1360. The storage controller 1360 (also referred to as a memory controller) is one example of a system that performs the techniques described herein. In some embodiments, the storage controller 1360 can be implemented on a semiconductor device such as an ASIC or an FPGA. Some functionality can also be implemented in firmware or software.
[0138] The controller 1304 can include one or more processors 1306 and memory 1308 for performing the control functions described above. The storage controller 1360 can also include a lookup table 1310, which can include a table of degraded blocks and a table of bad blocks, among others. Registers 1314 can be used to store data for control functions, such as a threshold for degraded block count, among others.
[0139] The controller 1304 can be coupled to the solid state storage device 1350 through the storage device interface 1302. An error correction decoder 1312 (e.g., an LDPC decoder or a BCH decoder) can perform error correction decoding on the read data and send the corrected data to the controller 1304. The controller 1304 can identify to a garbage collector 1316 pages that failed to read, which performs correction processing (e.g., by copying the data to a new location with or without error correction decoding) on those pages.
[0140] Figure 14 is a simplified block diagram illustrating a device that can be used to implement various embodiments in accordance with the present disclosure. Figure 14 The foregoing description of embodiments of the application has been presented for the purposes of illustration and description. It is not intended to be exhaustive or to limit the application to the precise form disclosed. Many modifications and variations are possible in light of the above teaching. In one embodiment, computer system 1400 generally includes monitor 1410, computer 1420, user output device 1430, user input device 1440, communication interface 1450, etc.
[0141] As shown in Figure 14 , computer 1420 can include a processor 1460 that communicates with a number of peripheral devices via a bus subsystem 1490. These peripheral devices can include user output device 1430, user input device 1440, communication interface 1450, and storage subsystems such as RAM 1470 and disk drives 1480. By way of example, a disk drive can include an SSD implemented with non-volatile memory devices such as the SSD 420 depicted in FIG. 4 above with the features described above. Figure 4
[0142] User input device 1440 includes all possible types of devices and mechanisms used to input information to computer system 1420. These can include a keyboard, a keypad, a touch screen incorporated into display 1410, audio input devices such as voice command systems, and other types of input devices. In various embodiments, user input device 1440 is typically implemented as a computer mouse, a track ball, a touch pad, a joystick, a wireless remote, a graphics tablet, a voice command system, an eye tracking system, etc. User input device 1440 usually allows the user to select objects, icons, text, etc. that appear on the monitor 1410 via commands that are either physical, e.g., by clicking a button, or via voice commands.
[0143] User output device 1430 includes all possible types of devices and mechanisms used to output information from computer system 1420. These can include a display, e.g., monitor 1410, non-visual displays such as audio output devices, etc.
[0144] Communication interface 1450 provides an interface to other communication networks and devices. Communication interface 1450 can be used to receive data from and transmit data to other systems. Embodiments of communication interface 1450 typically include an Ethernet card, a modem (telephone, satellite, cable TV, ISDN), (asynchronous) digital subscriber line (DSL) unit, FireWire interface, USB interface, etc. For example, communication interface 1450 can connect to a computer network, FireWire bus, etc. In other embodiments, communication interface 1450 can be physically integrated with the motherboard of computer 1420, and can be a software program such as SoftDSL, etc.
[0145] In various embodiments, computer system 1400 can also include software that enables communications among the components over various networks, such as Hypertext Transfer Protocol (HTTP), Transmission Control Protocol and Internet Protocol (TCP / IP), Real-Time Streaming Protocol and Real-Time Transport Protocol (RTSP / RTP), etc. In alternative embodiments, other communication software and transfer protocols can be used, such as Internet Packet Exchange (IPX), User Datagram Protocol (UDP), etc. In some embodiments, computer 1420 includes one or more Xeon microprocessors from Intel as processor 1460. Further, in one embodiment, computer 1420 includes a UNIX-based operating system.
[0146] RAM 1470 and disk drive 1480 are examples of tangible media configured to store data such as computer-executable programs, human-readable code, etc., including embodiments of the present application. Other types of tangible media include floppy disks, removable hard disks, optical storage media such as CD-ROMs, DVDs and bar codes, semiconductor memories such as flash memories, non-transitory read-only memories (ROM), battery-backed volatile memories, network storage devices, etc. RAM 1470 and disk drive 1480 can be configured to store basic programming and data constructs that provide the functionality of the present application.
[0147] Software code modules and instructions that provide the functionality of the present application can be stored in RAM 1470 and disk drive 1480. These software modules are executed by processor 1460. RAM 1470 and disk drive 1480 can also provide a repository for storing data used in accordance with the present application.
[0148] RAM 1470 and disk drive 1480 may include multiple memories, including main RAM for storing instructions and data during program execution and ROM for storing fixed, non-transitory instructions. RAM 1470 and disk drive 1480 may include a file storage subsystem that provides persistent (non-volatile) storage for program and data files. RAM 1470 and disk drive 1480 may also include removable storage systems, such as removable flash memory.
[0149] Bus subsystem 1490 provides mechanisms for enabling the various components and subsystems of computer 1420 to communicate with each other as intended. Although bus subsystem 1490 is schematically shown as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 1490 may be a high-speed PCI bus that can be implemented using the PCIe PHY embodiments of this disclosure.
[0150] Figure 14 This is a representative computer system capable of implementing the present invention. It will be apparent to those skilled in the art that many other hardware and software configurations are suitable for the present invention. For example, the computer may be a desktop, portable, rack-mount, or tablet computer configuration. Additionally, the computer may be a network of networked computers. Furthermore, the use of other microprocessors such as Pentium is contemplated. TM ) or Itanium TM The microprocessor, Opteron from Advanced Micro Devices, Inc. TM ) or Athlon (XP) TM Microprocessors, etc. Furthermore, it is anticipated that other types of operating systems, such as those from Microsoft Corporation, will be used. Examples include Solaris from Sun Microsystems, Linux, and UNIX. In other embodiments, the above techniques can be implemented on a chip or auxiliary processing board.
[0151] Various embodiments of the application can be implemented in the form of logic that is either software or hardware or a combination of both. The logic can be stored in a computer-readable or machine-readable non-transitory storage medium as a set of instructions adapted to direct a processor of a computer system to perform a set of steps disclosed in embodiments of the application. The logic can form part of a computer program product adapted to direct an information processing apparatus to perform a set of steps disclosed in embodiments of the application. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will appreciate other ways and / or methods to implement the application.
[0152] The data structures and code described herein can be stored in part or in whole on a computer-readable storage medium and / or hardware module and / or hardware device. Computer-readable storage media includes, but is not limited to, volatile memory, non-volatile memory, magnetic and optical storage devices such as disk drives, magnetic tape, CDs (compact discs), DVDs (digital versatile discs or digital video discs), or other media during analog communication or digital communication of data. Hardware modules or devices include, but are not limited to, application- specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), dedicated or shared processors, and / or other hardware modules or devices now known or later developed.
[0153] The methods and processes described herein can be implemented in part or in whole as code and / or data stored in a computer-readable storage medium or a machine-readable storage medium. When the code and / or data stored is accessed by one or more computer systems or one or more hardware devices, the computer systems or hardware devices execute the code and / or data to perform the associated methods and processes. The code and / or data can be executed by one or more computer systems or one or more hardware devices to cause the computer systems or hardware devices to perform the associated methods and processes. The code and / or data can be executed by one or more computer systems or one or more hardware devices to cause the computer systems or hardware devices to perform the associated methods and processes.
[0154] The scope of embodiments disclosed herein is not limited to the specific embodiments described herein. Various modifications can be made to embodiments of the application as described above with reference to the preceding description and accompanying drawings without departing from the spirit of the application. Further, while some embodiments of the application have been described above in the context of particular implementations in a particular environment for a particular purpose, those of ordinary skill in the art will appreciate that the application is not limited to this, and embodiments of the application can be advantageously implemented in any number of environments for any purpose.
Claims
1. A method for decoding low-density parity-check codewords, i.e., LDPC codewords, the method comprising: A system including an LDPC decoder performs hard decoding on a memory cell of a given page associated with a word line, i.e., WL, the hard decoding including a first hard read using a predetermined hard read threshold voltage and one or more rereads; The system determines that the hard decoding based on the hard read has failed; The system determines whether the hard read is the first hard read or a reread of the given page; If it is determined that the hard read is the first hard read, continue to perform hard decoding of another page; When it is determined that the hard read is a reread of the given page, The memory cells of the given page are grouped into intervals based on the read threshold voltage associated with the hard read and previous hard read of the given page; Determine the parity sum and the count of 1s for the memory cells in each interval; Based on the data read, checksum, and 1 count for each interval, machine learning is used to calculate the log-likelihood ratio (LLR) for each interval. and Use the calculated LLR to perform software decoding on the given page.
2. The method according to claim 1, further comprising: The hard read is determined to be either the first hard read or the reread by using a pattern matching operation between the current hard read from the given page and the read data from the previous hard read.
3. The method of claim 2, wherein the pattern matching operation includes performing a summation operation as follows: SUM(XOR(incoming_data,saved_data)), in: incoming_data is the read data from the current hard read of the given page; saved_data is the read data from the previous hard read of the given page; XOR is the exclusive OR operation; and SUM is an operation that sums up all bits that are 1.
4. The method of claim 1, wherein the parity check and the weights are based on the non-zero correctors of the codeword.
5. The method of claim 1, wherein using the machine learning includes using a neural network, i.e., an NN.
6. The method of claim 5, wherein the NN is a deep neural network, i.e., a DNN, the DNN receiving the checksum and the count of 1 as input and determining a weighting factor for calculating the optimal LLR.
7. A method for determining an LLR for software decoding in a storage system, the LLR for software decoding being determined based on information obtained from hardware decoding, the storage system performing hardware decoding and software decoding of LDPC codewords, the method comprising: Perform hardware decoding on the codewords in the page, the hardware decoding including a first hard read using a predetermined hard read threshold voltage and one or more rereads; The memory cells in the page are grouped into multiple intervals based on the read threshold voltage used for the first hard read and the one or more rereads; Calculate the parity sum and the count of 1s for the memory cells in each interval; The LLR of the memory cell in each interval is determined based on the read data, checksum, and count of 1s for each interval.
8. The method of claim 7, further comprising: Machine learning is used to determine the LLR for each interval.
9. The method of claim 8, wherein the machine learning includes a neural network (NN).
10. The method of claim 8, wherein the parity check sum includes weights for the non-zero correctors of the codewords used for LDPC decoding.
11. The method of claim 8, wherein the count of 1s in a given interval includes the number of memory cells in the interval whose cell value is 1.
12. The method of claim 7, further comprising: The LLR is determined without the use of an auxiliary read, i.e., AR, wherein the AR includes determining an additional read threshold voltage for determining the LLR based on read data from a hard read.
13. The method of claim 7, further comprising: The hard read is determined to be either the first hard read or the reread by using a pattern matching operation between the current hard read from a given page and the read data from a previous hard read.
14. The method of claim 13, further comprising: After determining that a hard read of a given page is a reread of the given page, an LLR for the soft decoding is determined based on information obtained from the hard decoding.
15. A storage system comprising a memory cell and a memory controller, the memory controller being coupled to the memory cell for controlling the operation of the memory cell, the operation including hardware decoding and software decoding of LDPC codewords, wherein the memory controller: Perform hardware decoding on the codewords in the page, the hardware decoding including a first hard read using a predetermined hard read threshold voltage and one or more rereads; The memory cells in the page are grouped into multiple intervals based on the read threshold voltage used for the first hard read and the one or more rereads; Calculate the parity sum and the count of 1s for the memory cells in each interval; The LLR of the memory cell in each interval is determined based on the read data, checksum, and count of 1s for each interval.
16. The storage system of claim 15, wherein the memory controller further uses a deep neural network, i.e., a DNN, to determine the LLR of each interval.
17. The storage system of claim 15, wherein the parity check includes weights for non-zero correctors of codewords used for LDPC decoding.
18. The storage system of claim 15, wherein the count of 1s in a given interval includes the number of memory cells in the interval whose cell value is 1.
19. The storage system according to claim 15, further comprising: The reread detection unit detects whether a hard read is the first hard read or the reread by using a pattern matching operation between the current hard read from a given page and the read data from a previous hard read. as well as The LLR generation unit, after determining that a hard read of a given page is a reread of the given page, determines the LLR for the soft decoding based on information obtained from the hard decoding.
20. The storage system of claim 19, wherein the pattern matching operation includes performing a summation operation as follows: SUM(XOR(incoming_data,saved_data)), in: incoming_data is the read data from the current hard read of the given page; saved_data is the read data from the previous hard read of the given page; XOR is the exclusive OR operation; and SUM is an operation that sums up all bits that are 1.
Citation Information
Patent Citations
Min-sum decoding for LDPC codes
US10680647B2
Min-sum decoding for LDPC codes
US20190097656A1
Incremental llr generation for flash memories
CN105989890A
VSS LDPC decoder with improved throughput for hard decoding
CN106997777A