Advanced ultra-low power error correction code encoder and decoder
Through reinforcement learning and grid interpolation technology, the threshold table is generated and updated, and the problem of suboptimal threshold in existing ultra-low power LDPC decoders is solved, which improves error correction efficiency and reduces power consumption, and achieves faster decoding speed and higher accuracy.
Patent Information
- Application Number
- CN202380043905.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-18
- Filing Date
- 2023-09-26
- Publication Date
- 2025-07-08
AI Technical Summary
The threshold table of existing ultra-low power LDPC decoders is suboptimal due to the existence of cycles and ignoring the sum of correction subweights in the generated graph, which leads to suboptimality, affecting error correction efficiency and power consumption.
The threshold table is generated using reinforcement learning and mesh interpolation methods. Through soft quantization and mesh interpolation technology, the threshold value is dynamically adjusted to optimize the error correction process, and combined with soft quantization and mesh interpolation technology, the threshold table is generated and updated.
Improve error correction efficiency, reduce power consumption, and achieve faster decoding speed and higher error correction accuracy.
Smart Images

Figure CN120283225A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of U.S. Non - Provisional Application No. 17 / 968,249, filed on October 18, 2022, entitled "ADVANCED ULTRA LOW POWER ERROR CORRECTING CODE ENCODERS AND DECODERS", and the contents thereof are hereby incorporated by reference in their entirety for all purposes. Background of the Invention
[0003] Ultra - low - power (ULP) ECC (error - correcting code) decoders are fast and efficient ECC engines that improve on LDPC (low - density parity - check) schemes compared to various other ECC schemes. Thus, ULP ECC is widely used in various data storage devices.
[0004] The ULP ECC scheme is iterative. Its convergence time is not deterministic and depends on the bit - error rate (BER) and the underlying LDPC code (generating graph). In current ULP decoders, a table containing thresholds for each clock is pre - computed offline and provided to the storage controller. This table is used in the decoding scheme. For each variable bit, the number of unsatisfied equations (sum of the check - node values connected to the bit) is calculated, and based on the threshold, it is decided whether to invert the bit value.
[0005] However, when pre - computing the table thresholds, due to the cycles present in the generating graph and due to ignoring the syndrome weight (SW, sum of all check nodes), the thresholds in the table may be sub - optimal. Thus, there is a need to generate optimal or near - optimal thresholds for ULP LDPC decoders. Summary of the Invention
[0006] This application generally relates to error - correcting codes and more specifically to advanced ultra - low - power error - correcting codes. More specifically, the systems and methods described herein disclose generating thresholds for inverting bits for a ULP LDPC decoder using soft quantization, lattice interpolation, and reinforcement learning at least partially based on the clock and syndrome weight.
[0007] Accordingly, the present application describes a method for generating multiple thresholds for use by an ultra-low power (ULP) low density parity check (LDPC) decoder. In an example, this includes performing multiple simulations for decoding stored data. For each of the multiple simulations, determining a syndrome weight (SW) for the current iteration of the simulation; performing a trellis interpolation on multiple SW values over a subset of trellis points that are adjacent to the SW for the current iteration of the simulation; and generating multiple thresholds for multiple trellis points based at least in part on multiple distances between the subset of trellis points and the SW for the current iteration of the simulation. A table of the multiple thresholds may then be stored.
[0008] The present application also describes a storage device that includes: a non-volatile memory device that includes: multiple memory cells; and a controller communicatively coupled to the non-volatile memory device. The controller includes an ultra-low power (ULP) low density parity check (LDPC) decoder and is configured to decode data stored in the multiple memory cells. Based on determining that the data stored in the multiple memory cells includes an error, the controller performs error correction code (ECC) decoding on the data using a threshold table. In an example, the threshold table identifies which data bits and corresponding error code bits of the data stored in the multiple memory cells are to be inverted. The threshold table is generated based at least in part on: multiple distances between a subset of trellis points and a determined syndrome weight (SW) that is associated with an iteration of a simulation among multiple simulations in which sample data is decoded; and a trellis interpolation determined for multiple SW values over the subset of trellis points that are associated with the iteration of the simulation in which the sample data is decoded.
[0009] Also described is a method for generating multiple thresholds for use by an ultra-low power (ULP) low density parity check (LDPC) decoder. In an example, the method includes: using reinforcement learning to perform multiple simulations to train a model for decoding stored data. For each of the multiple simulations, determining a syndrome weight (SW) for the current iteration of the simulation; performing a trellis interpolation on multiple SW values over a subset of trellis points that are adjacent to the SW for the current iteration of the simulation; and generating multiple thresholds for multiple trellis points based at least in part on multiple distances between the subset of trellis points and the SW for the current iteration of the simulation. Storing a table of the multiple thresholds and using the multiple thresholds to update the model.
[0010] The present invention content is provided to introduce some concepts in a simplified form, which will be further described in the detailed implementation below. The present invention content is not intended to identify the key features or essential features of the subject matter protected by the claims, nor is it intended to limit the scope of the subject matter protected by the claims. Brief Description of the Drawings
[0011] The following non - restrictive and non - exhaustive examples are described with reference to the accompanying drawings.
[0012] Figure 1 is a perspective view of a storage system including a three - dimensional (3D) stacked non - volatile memory.
[0013] Figure 2 is a functional block diagram of an example storage device (such as Figure 1 a 3D stacked non - volatile storage device).
[0014] Figure 3 is a block diagram of an example storage system that depicts more details of the controller of the storage system.
[0015] Figures 4A to 4C Illustrates a simplified version of decoding using a ULP ECC scheme according to at least one example.
[0016] Figure 5 Illustrates a process for training a threshold table for use with an error - correcting code in at least one example. Detailed Description
[0017] This application generally relates to error - correcting codes and more particularly to advanced ultra - low - power error - correcting codes. More specifically, the systems and methods described herein recite the use of soft quantization, trellis interpolation, and reinforcement learning to generate a threshold that is used by an LDPC decoder to determine whether to invert the bits of stored data and which bits of the stored data to invert when performing error correction.
[0018] For example, when data is written to memory, errors can occur. This is especially true when writing to flash memory. However, double - checking each bit and comparing it to its original position is very time - consuming and power - consuming. Therefore, an error - correcting code (ECC) can be used to detect and correct these errors. In some examples, this includes generating a redundant bit code that encodes the exclusive - OR result of multiple different bits into an error - code bit. In many cases, the multiple different bits of each error - code bit are not contiguous. The multiple different bits of the error - code bit are selected such that when the multiple different bits are exclusive - ORed together, the result is zero. Exclusive - OR is performed on each bit in more than one error - code bit to help identify the error. If a bit is associated with two error - code bits that are both 1, this indicates that the single bit may be incorrect.
[0019] Therefore, the system considers the number of unsatisfied XOR equations for that bit. The more unsatisfied XOR equations for that bit, the higher the likelihood of an error for that bit. The question then becomes how many unsatisfied XOR equations are required for the system to invert that bit? When that bit is inverted, all the XOR equations associated with the inverted bit also have their error code bits inverted. The number of unsatisfied XOR equations depends on the number of XOR equations connected to that bit, the number of times the system accesses that bit during the process, and the syndrome weight, which is the total number of unsatisfied XOR equations in the stored data being checked.
[0020] In many cases, density evolution is used to generate a threshold table to find the threshold number of unsatisfied XOR equations to indicate that the system should invert a specific bit. However, density evolution requires that the graph being analyzed does not contain cycles. That is, density evolution assumes that the graph is tree-like. However, many storage solutions using graphs are not tree-like, so there are problems with using density evolution.
[0021] To solve the above problems, a reinforcement learning process can be used to find an appropriate and better threshold. In addition, the use of reinforcement learning also allows for faster decoding compared to traditional ULP LDPC decoders.
[0022] Reinforcement learning (RL) is a field of machine learning that involves how intelligent agents act in an environment to maximize the concept of cumulative reward. RL is one of the three basic machine learning paradigms in addition to supervised learning and unsupervised learning.
[0023] RL differs from supervised learning in that it does not require providing labeled input / output pairs and does not require explicit correction of suboptimal actions. Instead, the focus is on finding a balance between exploration (of unknown areas) and exploitation (of existing knowledge). Some partially supervised RL algorithms can combine the advantages of supervised learning and RL algorithms.
[0024] The environment is usually described in the form of a Markov decision process (MDP) because many reinforcement learning algorithms used in this context use dynamic programming techniques. The main difference between classical dynamic programming methods and reinforcement learning algorithms is that the latter do not assume knowledge of the exact mathematical model of the MDP and aim at large MDPs (where exact methods become infeasible).
[0025] RL trains an agent in an environment. The agent takes actions, and the environment returns rewards and observations about the current state of the environment. Basic RL is modeled as a Markov decision process (MDP), which includes: a set S of the environment and agent states; a set A of the agent's actions; P a (s, s′) = Pr(s t+1 = s′|s t = s, a t= a) is the probability of transitioning from state s to state s′ under action a (at time t); and R a (s, s′) is the immediate reward after transitioning from s to s′ in the case of action a.
[0026] The goal of RL is for the agent to learn an optimal or near-optimal policy that maximizes the "reward function" or other reinforcement signals provided by users accumulated from immediate rewards. The basic RL agent AI interacts with its environment at discrete time steps. At each time t, the agent receives the current state st and a reward rt. The agent then selects an action from the set of available actions, which is then transmitted to the environment. When the transition (st, at, st+1) is determined, the environment switches to the new state st+1 and the associated reward rt+1. The goal of the RL agent is to learn a policy: π: A × S → [0, 1], π(a, s) = Pr(a t | s t = s), which maximizes the expected cumulative reward.
[0027] Formulating the problem as an MDP assumes that the agent directly observes the current environmental state; in this case, the problem is said to have full observability. When comparing the performance of the agent with that of an agent that takes optimal actions, the difference in performance gives rise to the concept of regret. To act near-optimally, the agent must reason about the long-term consequences of its actions (i.e., maximize future rewards), even though the associated immediate rewards may be negative.
[0028] Two elements make RL powerful: using samples to optimize performance and using function approximation to handle large environments. Thanks to these two key components, RL can be used in large environments in the following cases: the environmental model is known, but an analytical solution is not available; only a simulation model of the environment is given (the subject of simulation-based optimization); and the only way to collect information about the environment is to interact with it.
[0029] The first two of these problems can be considered planning problems (since some form of model is available), while the last problem can be considered a true learning problem. However, RL transforms these two planning problems into machine learning problems.
[0030] Figures 1 to 3 An example of a storage system that can be used to implement the techniques disclosed herein is described. Figure 1is a perspective view of a storage system including a three-dimensional (3D) stacked non-volatile memory. The storage device 100 includes a substrate 101. Above and on the substrate are example memory cell blocks in which memory cells are formed, including BLK0 and BLK1 (non-volatile memory elements). Also on the substrate 101 is a peripheral region 104 having support circuits for the blocks. The substrate 101 may also carry circuits under the blocks, along with one or more lower metal layers that are patterned in conductive paths to carry signals of the circuits. The blocks are formed in an intermediate region 102 of the storage device 100. In an upper region 103 of the storage device 100, one or more upper metal layers are patterned in conductive paths to carry signals of the circuits. Each memory cell block includes a stacked region of memory cells, where the alternating layers of the stack represent word lines. Although only two blocks are depicted as examples, additional blocks extending in the x direction and / or y direction may be used.
[0031] In one example embodiment, the length of the plane in the x direction represents the direction in which the word line signal path extends (e.g., the word line direction or the source / drain extreme select gate (SGD) line direction), and the width of the plane in the y direction represents the direction in which the bit line signal path extends (e.g., the bit line direction). The z direction represents the height of the storage device 100.
[0032] Figure 2 is a functional block diagram of an example storage device (such as Figure 1 the 3D stacked non-volatile storage device 100). Figure 2The components depicted in [Figure 0] are a circuit. The storage device 100 includes one or more memory dies 108. Each memory die 108 includes a three-dimensional memory structure 126 of memory cells (e.g., a 3D array of memory cells), control circuitry 110, and read / write circuitry 128. In other examples, a two-dimensional array of memory cells may be used. The memory structure 126 is addressable by word lines via a decoder 124 (e.g., a row decoder) and is addressable by bit lines via a column decoder 132. The read / write / erase circuitry 128 includes a plurality of sense blocks 150, including SB1, SB2, … SBp (e.g., sense circuitry) and allows parallel reading or programming of a page of memory cells. In some systems, the controller 122 is included in the same storage device 100 (e.g., a removable memory card) as one or more memory dies 108. However, in other systems, the controller may be separate from the memory dies 108. In some examples, the controller may be located on a die different from the memory die. In some examples, one controller 122 may communicate with a plurality of memory dies 108. In other examples, each memory die 108 has its own controller. Commands and data are transferred between the host 140 and the controller 122 via a data bus 120 and between the controller 122 and one or more memory dies 108 via lines 118. In one example, the memory die 108 includes an input and / or output (I / O) pin set connected to the lines 118.
[0033] The memory structure 126 may include one or more arrays of memory cells, including 3D arrays. The memory structure may include a monolithic 3D memory structure where multiple memory levels are formed above (e.g., rather than within) a single substrate such as a wafer, without an intervening substrate. The memory structure may include any type of non-volatile memory that is monolithically formed in one or more physical levels of an array of memory cells having an active region disposed above a silicon substrate. The memory structure may be in a non-volatile memory device having circuitry associated with the operation of the memory cells, whether the associated circuitry is above or within the substrate.
[0034] The control circuitry 110 cooperates with the read / write circuitry 128 to perform memory operations (e.g., erase, program, read, etc.) on the memory structure 126, and includes a state machine 112, an on-chip address decoder 114, and a power control module 116. The state machine 112 provides chip-level control of the memory operations. The temperature detection circuit 113 is configured to detect temperature and can be any suitable temperature detection circuit known in the art. In one example, the state machine 112 can be programmed by software. In other examples, the state machine 112 does not use software and is implemented entirely in hardware (e.g., electrical circuitry). In one example, the control circuitry 110 includes registers, ROM fuses, and other devices for storing default values such as reference voltages and other parameters.
[0035] The on-chip address decoder 114 provides an address interface between the addresses used by the host 140 or the controller 122 to the hardware addresses used by the decoders 124 and 132. The power control module 116 controls the power and voltage supplied to the word lines and bit lines during memory operations. The power control module may include drivers for word line layers in a 3D configuration, select transistors (e.g., SGS transistors and SGD transistors), and source lines. The power control module 116 may include a charge pump for generating voltage. The sense block includes bit line drivers. The SGS transistor is a select gate transistor at the source end of the NAND string, and the SGD transistor is a select gate transistor at the drain end of the NAND string.
[0036] Any one or any combination of the control circuitry 110, the state machine 112, the decoders 114 / 124 / 132, the temperature detection circuit 113, the power control module 116, the sense block 150, the read / write circuitry 128, and the controller 122 can be considered as one or more control circuits or management circuits that perform some or all of the functions described herein.
[0037] The controller 122 (in one example, a circuit that can be on-chip or off-chip) can include one or more processors 122c, a ROM 122a, a RAM 122b, a memory interface 122d, and a host interface 122e, all of which are interconnected. The one or more processors 122c are an example of control circuitry. Other examples can use state machines or other custom circuits designed to perform one or more functions. Devices such as the ROM 122a and the RAM 122b can include code such as an instruction set, and the processor 122c can be operative to execute the instruction set to provide some or all of the functions described herein. Alternatively or additionally, the processor 122c can access code from a memory device in a memory structure, such as a reserved area of memory cells connected to one or more word lines. The memory interface 122d, which communicates with the ROM 122a, the RAM 122b, and the processor 122c, is a circuit that provides an electrical interface between the controller 122 and the memory die 108. For example, the memory interface 122d can change the format or timing of signals, provide buffers, isolate electrical surges, latch I / O, etc. The processor 122c can issue commands via the memory interface 122d to the control circuitry 110 or any other component of the memory die 108. The host interface 122e, which communicates with the ROM 122a, the RAM 122b, and the processor 122c, is a circuit that provides an electrical interface between the controller 122 and the host 140. For example, the host interface 122e can change the format or timing of signals, provide buffers, isolate electrical surges, latch I / O, etc. Commands and data from the host 140 are received by the controller 122 via the host interface 122e. Data transmitted to the host 140 is sent via the host interface 122e.
[0038] A plurality of memory elements in the memory structure 126 can be configured such that they are connected in series or such that each element is individually accessible. As a non-limiting example, flash memory devices in a NAND configuration (e.g., NAND flash memory) typically contain memory elements connected in series. A NAND string is an example of a series connection of memory cells and select gate transistors.
[0039] A NAND flash memory array can be configured such that the array includes a plurality of NAND strings, where a NAND string includes a plurality of memory cells that share a single bit line and are accessed as a group. Alternatively, the memory elements can be configured such that each element is individually accessible (e.g., a NOR memory array). The NAND and NOR memory configurations are exemplary, and the memory cells can be configured in other ways.
[0040] Memory cells can be arranged in an ordered array at a single memory device level, such as arranged in multiple rows and / or columns. However, the memory elements can be arranged in a non-regular configuration or a non-orthogonal configuration, or in a structure that is not regarded as an array.
[0041] Some three-dimensional memory arrays are arranged such that the memory cells occupy multiple planes or multiple memory device levels, thereby forming a three-dimensional structure (e.g., x, y, and z directions, where the z direction is substantially vertical, and the x direction and the y direction are substantially parallel to the main surface of the substrate).
[0042] As a non-limiting example, a 3D memory structure can be vertically arranged as a stack of multiple 2D memory device levels. As another non-limiting example, a 3D memory array can be arranged as multiple vertical columns (e.g., columns extending substantially perpendicular to the main surface of the substrate such as in the y direction), where each column has multiple memory cells. The vertical columns can be arranged in a two-dimensional arrangement of memory cells, where the memory cells are located on multiple vertically stacked memory planes. Other configurations of three-dimensional memory elements can also constitute a 3D memory array.
[0043] As a non-limiting example, in a 3D NAND memory array, the memory elements can be coupled together to form vertical NAND strings that traverse multiple horizontal memory device levels. Other 3D configurations can be envisioned, where some NAND strings contain memory elements in a single memory level, while other strings contain memory elements spanning multiple memory levels. A 3D memory array can also be designed to be in a NOR configuration and in a ReRAM configuration.
[0044] One of ordinary skill in the art will recognize that the techniques described herein are not limited to a single specific memory structure, but cover many related memory structures within the spirit and scope of the techniques described herein and as understood by one of ordinary skill in the art.
[0045] Figure 3 is a block diagram of an example storage system 100, depicting more details of the controller 122. In one example, Figure 3The system is a solid state drive (SSD). As used herein, a flash memory controller is a device that manages data stored on flash memory and communicates with a host (such as a computer or other electronic device). In addition to the specific functions described herein, a flash memory controller may have various functions. For example, a flash memory controller may format the flash memory to ensure proper operation of the memory, map out bad flash memory cells, and allocate spare memory cells to replace failed memory cells in the future. Some of the spare memory cells among the spare memory cells may be used to accommodate firmware to operate the flash memory controller and implement other features. During operation, when the host reads data from or writes data to the flash memory, the host will communicate with the flash memory controller. If the host provides a logical address for the data to be read / written, the flash memory controller may convert the logical address received from the host into a physical address in the flash memory. Alternatively, in some examples, the host may provide the physical address. The flash memory controller may also perform various memory management functions, such as but not limited to wear leveling (e.g., allocating writes to avoid wearing out specific memory blocks that may otherwise be repeatedly written to) and garbage collection (e.g., after a block is full, only moving valid data pages to a new block so that the full block can be erased and reused). Non-volatile memory other than flash may have a non-volatile memory controller similar to a flash memory controller.
[0046] The communication interface between the controller 122 and the non-volatile memory die 108 can be any suitable flash interface, such as a toggle mode. In one example, the storage subsystem 100 can be a card-based system, such as a Secure Digital card (SD) or a Micro Secure Digital (Micro SD) card. In another example, the storage system 100 can be part of an embedded storage system. For example, the flash memory can be embedded within a host, such as in the form of a solid state drive installed in a personal computer.
[0047] In some examples, the storage system 100 includes a single channel between the controller 122 and the non-volatile memory die 108. However, the subject matter described herein is not limited to having a single memory channel. For example, in some storage system architectures, there may be two, four, eight, or more channels between the controller and the memory die 108 (e.g., depending on the capabilities of the controller). In any of the examples described herein, even if a single channel is shown in the drawings, there may be more than one single channel between the controller and the memory die 108.
[0048] As Figure 3 depicted, the controller 122 includes a front-end module 208 that interfaces with the host, a back-end module 210 that interfaces with one or more non-volatile memory dies 108, and various other modules that perform the functions described herein.
[0049] Figure 3 The components of the controller 122 depicted in
[0049] may take the form of a packaged functional hardware unit (e.g., circuitry) designed to be used with other components, a portion of program code (e.g., software or firmware) executable by a processor or processing circuitry (e.g., one or more processors) typically performing a particular function or related functions, or a stand-alone hardware or software component interfacing with a larger system. For example, each module may include an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), circuitry, digital logic circuitry, analog circuitry, discrete circuits, gates, or any other type of hardware in combination, or a combination thereof. Alternatively or in addition, each module may include software stored in a processor-readable device (e.g., memory) to program one or more processors of the controller 122 to perform the functions described herein. Figure 3 The architecture depicted in Figure 3 is one example implementation that may or may not use Figure 2 the components (e.g., RAM, ROM, processor, interface) of the controller 122 depicted in
[0049] .
[0050] Referring again to the modules of the controller 122, the buffer management / bus controller 214 manages buffers in the random access memory (RAM) 216 and controls internal bus arbitration of the controller 122. The read only memory (ROM) 218 stores system boot code. Although Figure 3 illustrated as being located separately from the controller 122, in other examples, one or both of the RAM 216 and the ROM 218 may be located within and outside of the controller 122. Additionally, in some implementations, the controller 122, the RAM 216, and the ROM 218 may be located on separate semiconductor dies.
[0051] The front end module 208 includes a host interface 220 and a physical layer interface 222 (PHY) that provide an electrical interface to a host or a next-level storage controller. The type of the host interface 220 may be selected depending on the type of memory used. Examples of the host interface 220 include, for example, SATA, SATA Express, SAS, Fibre Channel, USB, PCIe, and NVMe. The host interface 220 may be a communication interface facilitating the transfer of data, control signals, and timing signals.
[0052] The backend module 210 includes an Error Correction Controller (ECC) engine 224 that encodes data bytes received from a host and decodes and corrects error data bytes read from the non-volatile memory. A command sequencer 226 generates command sequences, such as programming command sequences and erase command sequences, to send to the non-volatile memory die 108. A RAID (Redundant Array of Independent Dies) module 228 manages the generation of RAID parity and the recovery of failed data. RAID parity can be used as an additional level of integrity protection for data written to the storage system 100. In some cases, the RAID module 228 can be part of the ECC engine 224. Note that RAID parity can be added as an additional one or more dies, or can be added within an existing die (e.g., as an additional plane, additional block, or additional WL within a block). The ECC engine 224 and the RAID module 228 can compute redundant data that can be used for recovery in case of an error and can be considered examples of redundant encoders. The ECC engine 224 and the RAID module 228 can be considered together to form a combined redundant encoder 234. A memory interface 230 provides command sequences to the non-volatile memory die 108 and receives status information from the non-volatile memory die 108. In some examples, the memory interface 230 can be a Double Data Rate (DDR) interface. A flash control layer 232 controls the overall operation of the backend module 210.
[0053] The backend module 210 further includes an XOR engine 250. The XOR engine 250 performs many of the parity protection methods for managing the non-volatile memory 108 in various parity protection methods, as shown and described below with respect to Figures 4A to 5 as shown and described.
[0054] Figure 3 Exemplary additional components of the storage system 100 include a media management layer 238 that performs wear leveling of the memory cells of the non-volatile memory die 108. The storage system 100 also includes other discrete components 240, such as an external electrical interface, external RAM, resistors, capacitors, or other components that can interface with the controller 122. In other examples, one or more of the physical layer interface 222, the media management layer 238, and the buffer management / bus controller 214 are optional components that are not necessary in the controller 122.
[0055] The flash translation layer (FTL) or media management layer (MML) 238 can be integrated as part of flash memory management that can handle flash memory errors and interface with the host. Specifically, the MML can be a module in flash memory management and can be responsible for the interior of NAND management. Specifically, the MML 238 can include an algorithm in the storage device firmware that converts writes from the host into writes to the flash memory structure 126 of the memory die 108. The MML 238 can be used because, for example, flash memory may have limited durability, the flash memory structure 126 can only be written in multiples of pages, or the flash memory structure 126 may not be writable unless it is erased as a block (e.g., a block can be considered the minimum erase unit and such non-volatile memory can be considered block erasable non-volatile memory). The MML 238 is configured to operate under these potential limitations of the flash memory structure 126, which may be invisible to the host. Thus, the MML 238 attempts to convert writes from the host into writes to the flash memory structure 126.
[0056] The controller 122 can interface with one or more memory dies 108. In one example, the controller 122 and multiple memory dies 108 (e.g., together constituting the storage system 100) implement an SSD that can be used as a NAS device, etc., to simulate, replace, or substitute for a hard disk drive inside a host device. Additionally, the SSD does not need to be used as a hard disk drive.
[0057] Figure 4A The first state 400 of an error correction code according to at least one example is illustrated. The first state 400 represents the result of encoding or a state where there are no errors.
[0058] In Figure 4A , a plurality of data bits 405 are shown on the left side and a plurality of error code bits 410 are shown on the right side. The error code bits 410 are calculated by performing an exclusive OR on groups of the data bits 405. The exclusive OR equation for groups of the data bits 405 is used to determine whether there may be an error in the data bits 405.
[0059] The first state 400 illustrates the absence of errors in the transfer or write of data in the memory. Before transfer, the data to be transferred is encoded to generate an error correction code (ECC) as shown on the right side. In some examples, the ECC is generated by the ECC engine 224, the RAID module 228, and / or the flash translation layer (FTL) or media management layer (MML) 238 (all in Figure 3encoded by one or more of those shown in [the figure]. ECC is based on selecting a subset (or group) of data bits 405 such that those bits 405 will XOR to zero, and the corresponding error code bits 410 are thus zero. Accordingly, these XOR equations are passed to the system that performs decoding. In some examples, the ECC is decoded by one or more of an ECC engine 224, a RAID module 228, and / or a flash translation layer (FTL) or media management layer (MML) 238. In some examples, the stored bits associated with the XOR equations are set or predetermined by such as the system or by a passing method. In other examples, the bits 405 associated with the XOR equations are determined before and after the passing time.
[0060] Figure 4B illustrates a second state 420 of an error correction code according to at least one example. Contrary to the first state 400 ( Figure 4A shown), the second state 420 includes two errors in the first data bit and the second data bit 405. These two errors cause the first error code bit, the third error code bit, and the fourth error code bit 410 to be set to one, while the second error code bit and the fifth error code bit 410 are set to zero. In this example, the first error code bit to the fourth error code bit 410 have XOR equations that include at least one of the first two data bits 405. Since the equation for the second error code bit includes both data bits 405 with errors, these errors cancel out and the second error code bit 410 remains zero.
[0061] To correct this problem, the system needs to determine which data bits 405 are in error and invert those data bits. For each data bit among the data bits 405, it is determined how many XOR equations are not satisfied. For the first bit 405, there are two XOR equations and only one XOR equation is not satisfied. For the second bit 405, there are three XOR equations and two XOR equations are not satisfied. For the third bit 405, there are two XOR equations and only one XOR equation is not satisfied. For the fourth bit 405, there are three XOR equations and one XOR equation is not satisfied. For the fifth and sixth bits 405, there are two XOR equations each and one XOR equation each is not satisfied.
[0062] Based on the unsatisfied equations, the percentage of unsatisfied XOR equations for the second bit 405 is the highest. Accordingly, the system inverts the second bit 405. The system also inverts the three error code bits 410 connected to the second bit 405 via the XOR equations.
[0063] Figure 4C illustrates a third state 440 of an error correction code according to at least one example. After inverting the second state 420 ( Figure 4BAfter the second bit 405 and its connected error code bits 410 (as shown), the system reaches the third decoding state 440. At this time, the first bit 405 is still in the first state 400( Figure 4A shown) in the error, and the two error code bits 410 connected to the first bit 405 are set to one. Since the two error code bits 410 of the first data bit 405 are set to one, the system selects to invert the first data bit 405 and its two connected error code bits 410. After inverting these bits, the system has removed all errors and returns to the first state 400.
[0064] Figures 4A to 4C Illustrates a simplified version of decoding a very small sample of data bit 405 and error code bits 410 using the ULP ECC scheme. As the number of bits increases, the problems also increase, including which bit to invert next. In a large system with a large number of bits and multiple errors, there may be problems in determining which bit to invert next and when to stop inverting bits. In some cases, the system may enter a loop of repeatedly inverting the same bit. Therefore, the systems and methods described herein provide reinforcement learning-based systems and methods for efficient and effective decoding.
[0065] In the RL problem for decoding, each segment can be considered as decoding a word (bit sequence). In each segment step, the RL agent decides the threshold of the current variable (which is the action a). The reward r for each clock is -1. The segment ends when the word is successfully decoded (SW equals zero) or the system reaches max_clocks (a predetermined number of clocks).
[0066] If the algorithm fails to converge, the segment ends in failure and a new segment will start. The LDPC engine has more complex and time-consuming schemes that can be used when the faster ULP engine fails and the algorithm fails to converge for decoding. In these cases, the reward is deducted using the number of clocks these algorithms run until successful decoding.
[0067] The systems and methods described herein quantify the state space and use grid interpolation in the algorithm, enabling the new RL scheme to perform generalization. This improves the performance of the sparse reward problem. The system quantifies the clocks and SW. The clocks are hard (or soft) quantized into groups of variables with the same degree. The SW is soft quantized. The system learns piecewise linear functions (for each clock group) instead of learning step functions.
[0068] This approach allows the learned function to be represented and constrained as a monotonic function. That is, when SW is higher, the quality of a specific threshold is lower because on average it will take more time to successfully decode it.
[0069] The method includes two main steps. The first step is to formulate the ULP threshold as an RL problem. The second step is to perform state soft quantization and grid interpolation. This can be applied to many other sparse state RL problems.
[0070] Quantization is very important because both the SW and clock counts are very high, so the RL scheme converges extremely slowly. However, normal quantization will have problems in terms of reducing the algorithm's learning and generalizing the experience collected from one quantization value to another. Therefore, a soft quantization and interpolation scheme is proposed, which significantly improves the convergence during the learning phase (training) and the ability to generalize learning.
[0071] During the fragment executed in parallel, the soft quantization and grid interpolation algorithm is performed for multiple SW values at grid points adjacent to the currently computed SW. Soft quantization is performed by assigning weights inversely proportional to the normalized distances to two quantization points for two quantization points between which the point lies. The algorithm used for soft quantization is the PL1 algorithm in the paper "The Pairwise Piecewise-Linear Embedding for Efficient Non-Linear Classification", Ofir Pele, Ben Taskar, Amir Globerson, Michael Werman, ICML 2013, which is incorporated herein by reference. An extended d-dimensional soft quantization is given in the paper "Interpolated Discretized Embedding of Single Vectors and Vector Pairs for Classification, Metric Learning and Distance Approximation", Ofir Pele, Yakir Ben-Aliz arXiv 2016, which is also incorporated herein by reference. The embedded high-dimensional sparse vectors are then used as feature vectors with a linear Q-learning algorithm that learns the high-dimensional weight vectors for each quantization / grid point. This allows interpolation calculations to be performed later during the inference phase on-site, for example, after the learning phase is completed and the desired Q(st,a) table has converged.
[0072] The proposed soft quantization and grid interpolation algorithm can be implemented either offline, where training is performed in the laboratory and the converged Q(st,a) table is downloaded to a storage device during product image download, or online, where Q(st,a) is computed on-the-fly during operation and convergence is achieved. After convergence, the complete threshold table can be computed via the activation of the inference function (dot product of the weight vector and the embedding vector) for each clock and each possible syndrome weight.
[0073] The proposed online version of soft quantization and grid interpolation has the major advantage that the quantization resolution and the state span st(Clock,SW) will be adjusted during the lifetime of the device according to the observed average BER, which is typically small at the beginning of life (BOL) and grows until the end of life (EOL). Thus, the Q(st,a) table can be computed at a higher resolution and with a more limited span in the BOL with smaller SW values and continuously modified during the lifetime to increasing SW values. Therefore, the grid state space can be adapted via the online version of the proposed soft quantization and grid interpolation scheme (via online learning of the weights).
[0074] After convergence, the inference of the learned model will also use grid interpolation and significantly improve the generalization ability of the scheme via the proposed interpolation scheme.
[0075] Figure 5 Illustrated in at least one example is a process 500 for training a threshold table for use with an error correction code. In the example, process 500 is performed by one or more of the ECC engine 224, the RAID module 228, and / or the flash translation layer (FTL) or media management layer (MML) 238 (all shown in Figure 3 ). In some examples, process 500 is performed offline, where training is performed in the laboratory and the converged Q(st,a) table is downloaded to a storage device during product image download. In other examples, process 500 is performed online, where Q(st,a) is computed on-the-fly during operation and convergence. In some examples, process 500 can be performed each time a memory write is performed. In other examples, process 500 is performed on a periodic basis, such as but not limited to daily, weekly, monthly, every 100 memory writes, or other time periods or number of memory writes, which can be based on the expected usage and / or expected lifetime of the memory being written to or written from the memory.
[0076] In some examples, in addition to the XOR equations, the sender memory system may send a signal that instructs the receiver memory system to update the threshold table online. In some examples, the system may determine to update the threshold table based on a threshold number of error code bits 410 exceeded in a single memory write or based on multiple memory writes cumulatively. For example, the system may initiate process 500 after detecting 1000 error code bits 410 across multiple memory writes.
[0077] In one example, the system generates 505 multiple grid points as the starting points for the Q(st,a) table. Multiple grid points are generated for different SW (syndrome weight) values, where the different SW (syndrome weight) values are the sum of all error code bits 410 (shown in FIG. 4). For example, Figure 4A the SW value of Figure 4B is zero, Figure 4C the SW value of
[0078] The system then begins to execute (510) the first simulation segment. In an example, the training system includes multiple simulation segments for decoding. In the online version, the simulation segments may be based on the past state of the storage device. The system determines (515) the current clock, where each cycle of the simulation segment increments the clock by one. The system also determines (520) whether the current clock is greater than or equal to the maximum clock. If the current clock is greater than or equal to the maximum clock, the simulation segment ends and fails, and the system determines (550) the reward (positive or negative) for that simulation segment. The system then determines (555) whether convergence has been reached. If convergence has been reached, method 500 ends. However, if convergence has not been reached, method 500 begins (560) the next simulation segment.
[0079] However, if the current clock is less than the maximum clock, the system determines (525) the current SW for that simulation segment by adding together all the error code bits (e.g., error code bits 410 ( Figure 4A )) or by counting the number of non-zero error code bits. The system then determines (530) whether SW = 0 holds. If SW = 0 holds, there are no errors indicated by the error code bits, and the simulation segment is complete. The system determines (550) the reward for that simulation segment. An example of the reward function can be the number of clocks (including more complex algorithmic clocks) until successful decoding multiplied by negative one. For an unsuccessful or incorrect decoding, the reward can also be a very low (negative) number. The system may also determine (555) whether convergence has been reached. If convergence has been reached, method 500 is complete. If convergence has not been reached, method 500 begins (560) the next simulation segment.
[0080] If SW is not equal to zero, the system calculates (535) weights for multiple grid points based on the distances to the current SW, based on the distances between the corresponding grid points and the current SW. The weights are calculated via a linear Q-learning algorithm, where the feature vectors are large but sparse embedding vectors. The system updates (540) the threshold table based on the adjusted weights. The system can then determine which storage bit 405 to invert next (as Figures 4A to 4B shown) and perform (545) the bit inversion, where the determined data bit 405 and the connected error code bit 410 are all inverted. The clock is advanced, and the system continues to perform operation 515.
[0081] According to the above, an example of the present application describes a method for generating multiple thresholds for use by an ultra-low power (ULP) low density parity check (LDPC) decoder, the method comprising: performing multiple simulations for decoding stored data; for each of the multiple simulations: determining a syndrome weight (SW) for the current iteration of the simulation; performing trellis interpolation on multiple SW values over a subset of grid points of multiple grid points, the subset of grid points being adjacent to the SW for the current iteration of the simulation; and generating multiple thresholds for the multiple grid points based at least in part on multiple distances between the subset of grid points and the SW for the current iteration of the simulation; and storing a table of the multiple thresholds.
[0082] In some examples, performing grid interpolation for the plurality of SW values includes: calculating a plurality of weights for the plurality of grid points based at least in part on the plurality of distances between the subset of grid points and the SW of the current pass used for the simulation. In some examples, the method further includes: using the plurality of weights to adjust one or more of the plurality of thresholds. In some examples, the method further includes: providing the table of the plurality of thresholds to a controller of a storage device, where the controller includes the ULP LDPC decoder. In some examples, the plurality of thresholds provide information for determining which data bits in the stored data will be inverted during the decoding process. In some examples, the method further includes: soft quantizing a clock input. In some examples, the method further includes: using reinforcement learning to perform the plurality of simulations to train a model for decoding the stored data. In some examples, the method further includes: determining a reward for the model and the current pass of the simulation when the SW is equal to zero, where the reward is a positive reward. In some examples, the method further includes: ending execution of the current pass of the simulation when the clock is equal to or exceeds a maximum clock. In some examples, the method further includes: determining a reward for the model and the current pass of the simulation when the clock is equal to or exceeds the maximum clock, where the reward is a negative reward. In some examples, each of the plurality of simulations includes a plurality of error code bits connected to a plurality of data bits in a memory, where the plurality of data bits are XORed together to generate the plurality of error code bits. In some examples, an error code bit among the plurality of error code bits is selected based at least in part on the number of unsatisfied error code bits of the data bit compared to the total number of error code bits connected to the data bit.
[0083] This application also describes a storage device including: a non-volatile memory device including a plurality of memory cells; and a controller communicatively coupled to the non-volatile memory device, the controller including an ultra-low power (ULP) low density parity check (LDPC) decoder and configured to: decode data stored in the plurality of memory cells; and perform error correction code (ECC) decoding on the data using a threshold table based on determining that the data stored in the plurality of memory cells includes an error, the threshold table identifying which data bits and corresponding error code bits of the data stored in the plurality of memory cells will be inverted, generating the threshold table based at least in part on: a plurality of distances between a subset of grid points among a plurality of grid points and a determined syndrome weight (SW), the determined syndrome weight (SW) being associated with a pass of a simulation among a plurality of simulations in which sample data is decoded; and a grid interpolation determined for a plurality of SW values on the subset of grid points among the plurality of grid points, the plurality of SW values being associated with the pass of the simulation among the plurality of simulations in which the sample data is decoded.
[0084] In an example, the threshold table is further generated by the following steps: calculating weights for the plurality of grid points at least in part based on the plurality of distances between the subset of grid points and the determined SW; generating additional thresholds for the threshold table at least in part based on the calculated weights; performing the plurality of simulations until the SW is equal to zero; and updating the threshold table based on the additional thresholds. In some examples, the threshold table is further generated by the following steps: using reinforcement learning to perform the plurality of simulations to train a model for decoding the stored data; and determining a reward for the model when the SW is equal to zero, where the reward is a positive reward. In some examples, for the storage device according to claim 15, wherein the threshold table is further generated by the following steps: ending the execution of the simulation when the clock is equal to or exceeds the maximum clock; and determining a reward for the model when the clock is equal to or exceeds the maximum clock, where the reward is a negative reward.
[0085] This application also describes a method for generating a plurality of thresholds used by an ultra-low power (ULP) low density parity check (LDPC) decoder, the method comprising: using reinforcement learning to perform the plurality of simulations to train a model for decoding stored data; for each of the plurality of simulations: determining a syndrome weight (SW) for the current iteration of the simulation; performing grid interpolation on a plurality of SW values on a subset of grid points adjacent to the SW for the current iteration of the simulation; and generating a plurality of thresholds for a plurality of grid points at least in part based on the plurality of distances between the subset of grid points and the SW for the current iteration of the simulation; storing a table of the plurality of thresholds; and updating the model using the plurality of thresholds.
[0086] In an example, the method further comprises: providing the table of the plurality of thresholds to the ULP LDPC decoder. In an example, the method is performed in an online setting, and the method further comprises: adjusting one or more of the plurality of thresholds at least in part based on the grid interpolation; and updating the table of the plurality of thresholds based on the one or more adjusted thresholds. In an example, the method further comprises: determining a reward for the model when the SW is equal to zero, where the reward is a positive reward.
[0087] As used herein, the term computer-readable medium may include computer storage media. Computer storage media may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, or program modules. Computer storage media may include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage devices, magnetic cassettes, magnetic tape, magnetic disk storage devices or other magnetic storage devices, or any other article that can be used to store information and that can be accessed by a computing device. Any such computer storage media may be part of a computing device. Computer storage media does not include carrier waves or other propagated or modulated data signals.
[0088] Additionally, the examples described herein may be discussed in the general context of computer-executable instructions, such as program modules, that reside on some form of computer-readable storage media and that are executed by one or more computers or other devices. By way of example, and not limitation, computer-readable storage media may include non-transitory computer storage media and communication media. In general, program modules include routines, programs, objects, components, data structures, etc. that perform particular tasks or implement particular abstract data types. The functionality of the program modules may be combined or distributed as desired in various examples.
[0089] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and includes any information delivery media. The term "modulated data signal" may describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wire connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0090] The description and illustration of one or more aspects provided in this disclosure are not intended to limit or in any way restrict the scope of the disclosure in any way. The various aspects, examples, and details provided in this disclosure are considered sufficient to convey ownership and the best mode for others to make and use the disclosure protected by the claims.
[0091] The disclosure protected by the claims should not be construed as limited to any aspect, example, or detail provided in this disclosure. Various features (both structural and method features), whether shown and described in combination or separately, are intended to be selectively rearranged, included, or omitted to yield examples having a particular set of features. Having provided the description and illustration of this application, those skilled in the art can envision variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application and do not depart from the broader scope of the disclosure protected by the claims.
[0092] Aspects of the present disclosure have been described below with reference to schematic flowcharts and / or schematic block diagrams of methods, apparatuses, systems, and computer program products according to examples of the present disclosure. It should be understood that each block of the schematic flowcharts and / or schematic block diagrams, and combinations of blocks in the schematic flowcharts and / or schematic block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, such that the instructions executed via the processor or other programmable data processing device create a means for implementing the functions and / or actions specified in one or more blocks of the schematic flowcharts and / or schematic block diagrams.
[0093] The use of names such as "first" and "second" herein to refer to elements generally does not limit the number or order of those elements. Instead, these names can be used as a way to distinguish between two or more elements or instances of an element. Thus, referring to a first element and a second element does not mean that only two elements can be used, nor does it mean that the first element must precede the second element. Additionally, unless otherwise specified, a set of elements can include one or more elements.
[0094] The term in the form of "at least one of A, B, or C" or "any combination of A, B, C, or them" used in the specification or claims means "A or B or C or any combination of these elements". For example, this term can include A, or B, or C, or A and B, or A and C, or A and B and C, or 2A, or 2B, or 2C, or 2A and B, etc. As an additional example, "at least one of A, B, or C" is intended to cover A, B, C, AB, AC, BC, and ABC, as well as multiples of the same members. Similarly, "at least one of A, B, and C" is intended to cover A, B, C, AB, AC, BC, and ABC, as well as multiples of the same members.
[0095] Similarly, as used herein, a phrase that refers to a list of items linked by "and / or" means any combination of the items. As an example, "A and / or B" is intended to cover A alone, B alone, or A and B together. As another example, "A, B, and / or C" is intended to cover A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0096] This written description uses examples to disclose the invention (including the best mode), and also enables any person skilled in the art to practice the invention (including making and using any device or system and performing any method incorporated herein). The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. If such other examples have structural elements that are not different from the literal language of the claims, or if such other examples include equivalent structural elements that are not materially different from the literal language of the claims, then such other examples are expected to be encompassed within the scope of the claims.
Claims
1. A method for generating multiple thresholds for use by an ultra-low power (ULP) low density parity check (LDPC) decoder, the method comprising: Performing multiple simulations for decoding stored data; For each of the multiple simulations: Determining a syndrome weight (SW) for a current iteration of the simulation; Performing grid interpolation on multiple SW values over a subset of grid points adjacent to the SW for the current iteration of the simulation; And Generating multiple thresholds for the multiple grid points based at least in part on multiple distances between the subset of grid points and the SW for the current iteration of the simulation; And Storing a table of the multiple thresholds.
2. The method according to claim 1, wherein performing grid interpolation on the multiple SW values comprises calculating multiple weights for the multiple grid points based at least in part on the multiple distances between the subset of grid points and the SW for the current iteration of the simulation.
3. The method according to claim 2, the method further comprising adjusting one or more of the multiple thresholds using the multiple weights.
4. The method according to claim 1, the method further comprising providing the table of the multiple thresholds to a controller of a storage device, wherein the controller includes the ULP LDPC decoder.
5. The method according to claim 1, wherein the multiple thresholds provide information for determining which data bits in the stored data will be inverted during a decoding process.
6. The method according to claim 1, the method further comprising soft quantizing a clock input.
7. The method according to claim 1, the method further comprising using reinforcement learning to perform the multiple simulations to train a model for decoding the stored data.
8. The method according to claim 7, the method further comprising determining a reward for the model and the current iteration of the simulation when the SW is equal to zero, wherein the reward is a positive reward.
9. The method according to claim 7, the method further comprising ending execution of the current iteration of the simulation when a clock is equal to or exceeds a maximum clock.
10. The method according to claim 9, the method further comprising determining a reward for the model and the current iteration of the simulation when the clock is equal to or exceeds the maximum clock, wherein the reward is a negative reward.
11. The method according to claim 1, wherein each of the multiple simulations includes multiple error code bits connected to multiple data bits in a memory, wherein the multiple data bits are XORed together to generate the multiple error code bits.
12. The method according to claim 11, wherein an error code bit among the multiple error code bits is selected based at least in part on a number of unsatisfied error code bits of the data bits compared to a total number of error code bits connected to the data bits.
13. A storage device, the storage device comprising: A non-volatile memory device including multiple memory cells; And A controller communicatively coupled to the non-volatile memory device, the controller including an ultra-low power (ULP) low density parity check (LDPC) decoder and configured to: Decode data stored in the plurality of memory cells; And Based on determining that the data stored in the plurality of memory cells includes an error, perform error correction code (ECC) decoding on the data using a threshold table that identifies which data bits and corresponding error code bits of the data stored in the plurality of memory cells are to be inverted, the threshold table being generated at least in part based on: A plurality of distances between a subset of grid points among a plurality of grid points and a determined syndrome weight (SW) associated with a pass of an analog in a plurality of analogs in which sample data is decoded; And Grid interpolation determined for a plurality of SW values on the subset of grid points among the plurality of grid points, the plurality of SW values being associated with the pass of the analog in the plurality of analogs in which the sample data is decoded.
14. The storage device according to claim 13, wherein the threshold table is further generated by the following steps: Calculate weights for the plurality of grid points at least in part based on the plurality of distances between the subset of grid points and the determined SW; Generate additional thresholds for the threshold table at least in part based on the calculated weights; Perform the plurality of analogs until the SW equals zero; And Update the threshold table based on the additional thresholds.
15. The storage device according to claim 13, wherein the threshold table is further generated by the following steps: Use reinforcement learning to perform the plurality of analogs to train a model for decoding the sample data; and Determine a reward for the model when the SW equals zero, wherein the reward is a positive reward.
16. The storage device according to claim 15, wherein the threshold table is further generated by the following steps: End the execution of the analog when the clock is equal to or exceeds a maximum clock; and Determine a reward for the model when the clock is equal to or exceeds the maximum clock, wherein the reward is a negative reward.
17. A method for generating a plurality of thresholds used by an ultra-low power (ULP) low density parity check (LDPC) decoder, the method comprising: Use reinforcement learning to perform a plurality of analogs to train a model for decoding stored data; For each of the plurality of analogs: Determine a syndrome weight (SW) for the current pass of the analog; Perform grid interpolation on a plurality of SW values on a subset of grid points adjacent to the SW for the current pass of the analog; and Generate a plurality of thresholds for a plurality of grid points at least in part based on a plurality of distances between the subset of grid points and the SW for the current pass of the analog; Store a table of the plurality of thresholds; And Update the model using the plurality of thresholds.
18. The method according to claim 17, the method further comprising providing the table of the plurality of thresholds to the ULP LDPC decoder.
19. The method according to claim 17, wherein the method is performed in an online setting, and the method further comprises: adjusting one or more of the plurality of thresholds at least partially based on the trellis interpolation; and updating the table of the plurality of thresholds based on the one or more adjusted thresholds.
20. The method according to claim 17, the method further comprising determining a reward of the model when the SW is equal to zero, wherein the reward is a positive reward.