Data storage device with efficient decoder pool and method for in-run decoder initialization

Through parallel verification sub-computation and bit error rate estimation scanning operations, the problem of line head blocking in the shared decoder pool is solved, the decoder throughput and efficiency of the data storage device is improved, and the QoS is improved.

CN120256191APending Publication Date: 2025-07-04SANDISK TECHNOLOGIES LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410495633.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-04-24
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The shared decoder pool is susceptible to head blockage in the data storage device, resulting in damage to the decoder pipelined operation, affecting the decoder throughput and QoS.

Method used

Through parallelized verification sub-computation and bit error rate estimation scanning operations, the verification sub-computation is performed in parallel using a buffer to monitor the input status, avoiding line head blockage, while maintaining pipelined operations of the decoder.

Benefits of technology

Improve the throughput and efficiency of the decoder pool, reduce the clock frequency and silicon area of ​​the decoder, improve decoder performance, especially in the case of high bit error rates, significantly reduce the delay of the bit error rate estimation scan operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256191A_ABST
    Figure CN120256191A_ABST
Patent Text Reader

Abstract

A shared decoder pool is susceptible to thread end blockage, at which time decoding of a given data block delays decoding of other pipelined data blocks in the decoder. While this problem can be avoided by not using pipeline operations, the pipelining benefits will be compromised. In one embodiment provided herein, a syndrome of an error pattern is computed in parallel with data being written in an input buffer of the decoder. Parallelizing the syndrome calculation and padding of the input buffer of the decoder can avoid the thread end blocking problem mentioned above while still achieving pipelining benefits. In another embodiment, similar techniques are used in bit error rate estimation scan (BES) operations. Other embodiments are provided.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A host can write data to and / or read data from a memory in a data storage device. The data storage device may include one or more decoders for use in correcting errors in data read from the memory. Brief Description of the Drawings

[0002] Figure 1A is a block diagram of a data storage device according to an embodiment.

[0003] Figure 1B is a block diagram showing a storage module according to an embodiment.

[0004] Figure 1C is a block diagram showing a hierarchical storage system according to an embodiment.

[0005] Figure 2A is a block diagram showing components of a controller of a data storage device shown in Figure 1A according to an embodiment.

[0006] Figure 2B is a block diagram showing components of a data storage device shown in Figure 1A according to an embodiment.

[0007] Figure 3 is a block diagram of a host and a data storage device according to an embodiment.

[0008] Figure 4 is an illustration of a shared decoder pool according to an embodiment.

[0009] Figure 5 is an illustration of a shared decoder pool with head-of-line blocking according to an embodiment.

[0010] Figure 6 is an illustration of a shared decoder pool without head-of-line blocking according to an embodiment.

[0011] Figure 7A is a timing diagram of non-parallelized syndrome calculation according to an embodiment.

[0012] Figure 7B is a timing diagram of parallelized syndrome calculation according to an embodiment.

[0013] Figure 8 is a diagram showing parallelized syndrome calculation according to an embodiment.

[0014] Figure 9 is a diagram showing parallelized bit error rate estimation scan (BES) syndrome calculation in an embodiment where parallel syndrome weight calculation is used.

[0015] Figure 10A and 10BA diagram showing the BES operation with parallel syndrome weight calculation of an embodiment.

[0016] Figure 11 A graph showing the decoding throughput comparison of an embodiment. Detailed implementation

[0017] The following embodiments generally relate to a data storage device having an efficient decoder pool and a method for in-operation decoder initialization. In one embodiment, a data storage device is provided that includes a memory; a decoder; an input buffer; and one or more processors. The one or more processors are configured, individually or in combination, to: monitor the input buffer; and in response to detecting that data read from the memory is being written into the input buffer, cause the decoder to perform syndrome calculation in parallel with the data being written into the input buffer to calculate the syndrome of the data, wherein at least a portion of the syndrome calculation is completed by the time the data is completely written into the input buffer (e.g., the syndrome calculation is almost completed or some calculations are completed).

[0018] In another embodiment, a method is provided that is performed in a data storage device including a memory, a decoder, and an input buffer. The method includes: identifying a plurality of hypothesized read thresholds of the memory; and after the plurality of hypothesized read thresholds have been identified: calculating the syndrome weight for each of the plurality of hypothesized read thresholds; and using the calculated syndrome weights for each of the plurality of hypothesized read thresholds in a bit error rate estimation scan (BES) operation.

[0019] In yet another embodiment, a data storage device is provided that includes: a memory; a decoder; an input buffer; and means for causing the decoder to: in response to detecting that data read from the memory is being written into the input buffer, perform syndrome calculation in parallel with the data being written into the input buffer to calculate the syndrome of the data, wherein at least a portion of the syndrome calculation is completed by the time the data is completely written into the input buffer.

[0020] Other embodiments are possible, and each of the embodiments can be used alone or in combination. Accordingly, the various embodiments will now be described with reference to the accompanying drawings.

[0021] Embodiment

[0022] The following embodiments relate to a data storage device (DSD). As used herein, "data storage device" refers to a non-volatile device that stores data. Examples of DSDs include, but are not limited to, hard disk drives (HDDs), solid state drives (SSDs), tape drives, hybrid drives, etc. Details of an example DSD are provided below.

[0023] Examples of data storage devices suitable for aspects implementing these embodiments are shown in Figures 1A - 1C the following. It should be noted that these are only examples, and other implementations may be used. Figure 1A FIG. is a block diagram showing a data storage device 100 according to an embodiment. Referring to Figure 1A , the data storage device 100 in this example includes a controller 102 coupled to a non-volatile memory, and the non-volatile memory may be composed of one or more non-volatile memory dies 104. As used herein, the term die refers to a collection of non-volatile memory cells formed on a single semiconductor substrate and associated circuitry for managing the physical operations of those non-volatile memory cells. The controller 102 interfaces with a host system and transmits command sequences for read, program, and erase operations to the non-volatile memory die 104. Also, as used herein, the phrase "communicate with" or "coupled to" may represent direct communication / coupling, or indirect communication / coupling via one or more components that may or may not be shown or described herein. The communication / coupling may be wired or wireless.

[0024] The controller 102 (which may be a non-volatile memory controller (e.g., flash, resistive random access memory (ReRAM), phase change memory (PCM), or magnetoresistive random access memory (MRAM) controller)) may include one or more components that are individually or collectively configured to perform certain functions, including but not limited to the functions described herein and shown in the flowcharts. For example, as Figure 2A shown in, the controller 102 may include one or more processors 138 that are individually or collectively configured to perform functions (e.g., but not limited to the functions described herein and shown in the flowcharts) by performing computer-readable program code stored in one or more non-transitory memories 139 internal and / or external to the controller 102 (e.g., random access memory (RAM) 116 or read-only memory (ROM) 118). As another example, the one or more components may include circuitry such as, but not limited to, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers.

[0025] In one example embodiment, the non-volatile memory controller 102 is a device that manages data stored on non-volatile memory and communicates with a host such as a computer or an electronic device having any suitable operating system. The non-volatile memory controller 102 may have various functions in addition to the specific functionality described herein. For example, the non-volatile memory controller may format the non-volatile memory to ensure proper operation of the memory, list defective non-volatile memory cells, and allocate spare cells to replace cells that may fail in the future. Some portions of the spare cells may be used to store firmware (and / or other metadata for housekeeping and tracking) to operate the non-volatile memory controller and implement other features. In operation, when the host needs to read data from or write data to the non-volatile memory, it may communicate with the non-volatile memory controller. If the host provides a logical address to which the data is to be read / written, the non-volatile memory controller may translate the logical address received from the host into a physical address in the non-volatile memory. The non-volatile memory controller may also perform various memory management functions, such as but not limited to wear leveling (distributing writes to avoid wearing out specific memory blocks that would otherwise be repeatedly written to) and garbage collection (after a block is full, only moving valid data pages to a new block so that the full block can be erased and reused).

[0026] The non-volatile memory die 104 may include any suitable non-volatile storage medium, including resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), phase change memory (PCM), NAND flash memory cells, and / or NOR flash memory cells. The memory cells may take the form of solid-state (e.g., flash) memory cells and may be once programmable, few times programmable, or multiple times programmable. The memory cells may also be single-level cells (SLCs), multi-level cells (MLCs) (e.g., dual-level cells, triple-level cells (TLCs), quad-level cells (QLCs), etc.), or use other memory cell level technologies known now or developed in the future. Also, the memory cells may be fabricated in two-dimensional or three-dimensional fashions.

[0027] The interface between the controller 102 and the non-volatile memory die 104 may be any suitable flash interface, such as Toggle Mode 200, 400, or 800. In one embodiment, the data storage device 100 may be a card-based system, such as a Secure Digital (SD) or Micro Secure Digital (Micro-SD) card. In an alternative embodiment, the data storage device 100 may be part of an embedded data storage device.

[0028] Although in Figure 1AIn the example shown, data storage device 100 (sometimes referred to herein as a storage module) includes a single channel between controller 102 and non-volatile memory die 104, but the subject matter described herein is not limited to having a single memory channel. For example, in some architectures (such as the architectures shown in Figure 1B and 1C ), depending on the controller capabilities, there may be two, four, eight, or more memory channels between the controller and the memory device. In any of the embodiments described herein, there may be more than a single channel between the controller and the memory die, but a single channel is shown in the figures.

[0029] Figure 1B FIG. shows a storage module 200 that includes a plurality of non-volatile data storage devices 100. Thus, storage module 200 may include a storage controller 202 that interfaces with a host and interfaces with data storage devices 204, which include a plurality of data storage devices 100. The interface between storage controller 202 and data storage device 100 may be a bus interface, such as Serial Advanced Technology Attachment (SATA), Peripheral Component Interconnect Express (PCIe) interface, Double Data Rate (DDR) interface, or Serial Attached SCSI (SAS / SCSI) interface. In one embodiment, storage module 200 may be a solid state drive (SSD) or a non-volatile dual in-line memory module (NVDIMM), such as those seen in a server PC or a portable computing device (e.g., a laptop computer and a tablet computer).

[0030] Figure 1C FIG. is a block diagram showing a hierarchical storage system. Hierarchical storage system 250 includes a plurality of storage controllers 202, each of which controls a corresponding data storage device 204. Host system 252 may access the memory within storage system 250 via a bus interface. In one embodiment, the bus interface may be a Non-Volatile Memory Express (NVMe) or Fibre Channel over Ethernet (FCoE) interface. In one embodiment, Figure 1C the system shown in FIG. may be a rack-mountable mass storage system accessible by multiple host computers, such as those seen in a data center or other locations that require large-capacity storage.

[0031] Referring again to Figure 2A, the controller 102 in this example also includes a front-end module 108 interfacing with the host, a back-end module 110 interfacing with the one or more non-volatile memory dies 104, and various other components or modules, such as but not limited to a buffer manager / bus controller module that manages buffers in the management RAM 116 and controls the internal bus arbitration of the controller 102. The module may include one or more processors or components, as discussed above. The ROM 118 may store system boot code. Although Figure 2A shown as being located separately from the controller 102 in FIG. Figure 2A , in other embodiments, one or both of the RAM 116 and the ROM 118 may be located within the controller 102. In still other embodiments, portions of the RAM 116 and the ROM 118 may be located both within and outside the controller 102.

[0032] The front-end module 108 includes a host interface 120 that provides an electrical interface with the host or the next-level storage controller and a physical layer interface (PHY) 122. The choice of the type of the host interface 120 may depend on the type of memory being used. Examples of the host interface 120 include but are not limited to SATA, SATA Express, Serial Attached SCSI (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 generally facilitates the transfer of data, control signals, and timing signals.

[0033] The back-end module 110 includes an Error Correction Code (ECC) engine 124 that encodes data bytes received from the host and decodes and corrects errors in data bytes read from the non-volatile memory. A command sequencer 126 generates command sequences to be transmitted to the non-volatile memory die 104, such as programming and erase command sequences. A Redundant Array of Independent Drives (RAID) module 128 manages the generation of RAID parity and the recovery of failed data. The RAID parity can be used as an additional level of integrity protection for writing data to the memory device 104. In some cases, the RAID module 128 may be part of the ECC engine 124. A memory interface 130 provides the command sequences to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104. In one embodiment, the memory interface 130 may be a Double Data Rate (DDR) interface, such as a toggle mode 200, 400, or 800 interface. The controller 102 in this example also includes a media management layer 137 and a flash control layer 132 that controls the overall operation of the back-end module 110.

[0034] The data storage device 100 also includes other discrete components 140, such as an external electrical interface, external RAM, resistors, capacitors, or other components that can interface with the controller 102. In an alternative embodiment, one or more of the physical layer interface 122, RAID module 128, media management layer 138, and buffer management / bus controller 114 are optional components that are not necessary in the controller 102.

[0035] Figure 2B is a block diagram that more particularly illustrates the components of the non-volatile memory die 104. The non-volatile memory die 104 includes a peripheral circuit 141 and a non-volatile memory array 142. The non-volatile memory array 142 includes non-volatile memory cells for storing data. The non-volatile memory cells can be any suitable non-volatile memory cells, including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in two-dimensional and / or three-dimensional configurations. The non-volatile memory die 104 further includes a data cache memory 156 for caching data. The peripheral circuit 141 in this example includes a state machine 152 that provides status information to the controller 102. The peripheral circuit 141 may also include one or more components that are configured, individually or in combination, to perform certain functions, including but not limited to the functions described herein and shown in the flowcharts. For example, as Figure 2B shown, the memory die 104 may include one or more processors 168 that are configured, individually or in combination, to execute computer-readable program code stored in one or more non-transitory memories 169, stored in the memory array 142, or stored external to the memory die 104. As another example, the one or more components may include circuitry, such as but not limited to logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers.

[0036] As a supplement to or alternative to the one or more processors 138 (or more generally, components) in the controller 102 and the one or more processors 168 (or more generally, components) in the memory die 104, the data storage device 100 may include another collection of one or more processors (or more generally, components). Generally speaking, regardless of where they are located and regardless of how many there are, the one or more processors (or more generally, components) in the data storage device 100 can be configured, individually or in combination, to perform various functions, including but not limited to the functions described herein and shown in the flowcharts. For example, the one or more processors (or components) may be in the controller 102, the memory device 104, and / or other locations in the data storage device 100. Also, different processors (or components) or combinations of processors (or components) may be used to perform different functions. In addition, the means for performing the functions may be implemented using a controller that includes one or more components (such as the processors or other components described above).

[0037] Returning again to Figure 2A , the flash control layer 132 (which will be referred to herein as the flash translation layer (FTL)) processes flash errors and interfaces with the host. Specifically, the FTL, which can be an algorithm in the firmware, is responsible for internal memory management and translates writes from the host into writes to the memory 104. Because the memory 104 may have limited durability, may only be written in multiple pages, and / or may not be writable unless erased as a block, the FTL may be required. The FTL is aware of these potential limitations of the memory 104, which may not be visible to the host. Therefore, the FTL attempts to translate writes from the host into writes to the memory 104.

[0038] The FTL may include a logical to physical address (L2P) mapping (sometimes referred to herein as a table or data structure) and an allocated cache memory. In this way, the FTL translates the logical block address ("LBA") from the host into a physical address in the memory 104. The FTL may include other features, such as but not limited to power loss recovery (such that the data structure of the FTL can be recovered in the event of a sudden power loss) and wear leveling (such that wear on the memory blocks is more evenly distributed to prevent some blocks from wearing out excessively, which would lead to a greater probability of failure).

[0039] Turning again to the figures, Figure 3FIG. 0 is a block diagram of a host 300 and a data storage device 100 of an embodiment. The host 300 can take any suitable form, including but not limited to a computer, a mobile phone, a tablet computer, a wearable device, a digital video recorder, a surveillance system, etc. The host 300 (here, a computing device) in this embodiment includes one or more processors 330 and one or more memories 340. In one embodiment, the computer-readable program code stored in the one or more memories 340 configures the one or more processors 330 to perform the actions described herein as being performed by the host 300. Thus, the actions performed by the host 300 are sometimes referred to herein as being performed by an application (computer-readable program code) running on the host 300. For example, the host 300 can be configured to send data (e.g., initially stored in the memory 340 of the host) to the data storage device 100 for storage in the memory 104 of the data storage device.

[0040] With the development of flash storage solutions (especially in enterprises and data centers), the quality of service (QoS) requirements from flash controllers have become more stringent, and creative solutions may be needed to ensure that the latency is constant under different conditions caused by process variations and changing operating conditions (e.g., programming / erase (P / E) cycles, temperature, data retention, interference effects, etc.) exhibited by flash memory. The following embodiments provide an effective shared decoder pool (PoD) optimized to ensure high performance and quality of service.

[0041] In recent years, as product requirements have increased, the limit of the maximum decoding throughput achievable from a single decoder has been reached. In addition, using a single decoding engine can create "head-of-line blocking" situations, in which high BER events result in high decoding latency and block the servicing of other decoding requests, which can degrade QoS and may be unacceptable in QoS-driven systems. Therefore, multiple decoder cores, called "decoder pools", have been used. Figure 4 FIG. 8 is a diagram of a shared decoder pool 400 of an embodiment. As Figure 4 shown, in this example, the shared decoder pool 400 includes multiple decoders (each having a corresponding multiple input buffers 410), which are coupled to a DMA input module 420 and a DMA output module 430. This system performs the abstraction of the decoder into a simple interface while performing input-output (I-O) operations and internal arbitration.

[0042] However, problems can occur with the QoS-optimized pool. For example, state-of-the-art error correction codes (ECC) use iterative decoders. Iterative decoding algorithms (and specifically, low-density parity-check codes (LDPC)) involve iterating over the data until the decoder converges to the correct codeword - correcting all errors. These codes offer excellent performance and error-correction capabilities approaching the Shannon limit. On the other hand, a disadvantage of using an iterative scheme is that the decoding time is unpredictable and depends on the number of errors and the specific error pattern, which are unknown until the decoding operation is complete.

[0043] To ensure fast and efficient operation, the decoding operation can be pipelined. As such, the decoder pool architecture can experience head-of-line (HoL) blocking situations. Subsequent ECC blocks ("eblocks") are loaded into the decoding engine before the previous decoding operation is complete, and can thus get stuck in the case where the previous decoding operation takes a long time to decode (e.g., due to a high BER event). This can severely impact QoS, as the eblocks that are stuck waiting for the previous decoding to complete will thus exhibit high read latency, whereas they could have been decoded quickly and easily in another decoder engine (since they are likely to exhibit low BER). As such, in this PoD design, a single high BER event can cause multiple eblocks to be delayed and exhibit high latency.

[0044] Figure 5 A shared decoder pool with three decoder cores vulnerable to HoL blocking is shown. More specifically, Figure 5 An example timing diagram is presented that depicts the input and decoding times of eblocks through the system. Due to the pipelined operation, subsequent eblocks are loaded before the previous eblock is complete. As such, when it will complete is unknown. In this example, eblock #4 is significantly delayed because it is loaded into engine #1 and inappropriately follows eblock #1, which has an extremely long decoding time. In some embodiments, the eblock pipeline can be deeper, and several eblocks can be loaded into the engine before the first eblock is complete, making the HoL blocking even worse.

[0045] To address HoL blocking, one solution is to avoid loading the next eblock until the previous eblock decoding is complete; thus avoiding HoL blocking. However, this essentially means unpipelining the operation, and unpipelining the operation can significantly degrade the decoding throughput measured in terms of the number of eblock decoders per second or the number of decoded bytes per second, thus requiring more decoder cores / higher clock frequencies to achieve the same throughput requirements. This is in Figure 6Shown herein. In this example, an eblock is loaded into the engine only after the previous eblock has been decoded and it is ensured that the decoder is free for a new eblock. It can be seen that this cancels the pipelined operation and reduces the throughput, as is obvious from the idle time of the input and decoder engine operations.

[0046] The following embodiments observe that when using the HoL decoder pool as described above, the decoder engine is idle during the input operation. This is necessary to avoid the HoL blocking situation as described above. The following embodiments observe that a common method in LDPC decoding is to first compute the syndrome of the error pattern. This can be used to estimate the BER and is a necessary part of the decoding operation in the case of using the bit flip decoding algorithm prevalent in memory ECC solutions. The computation of the syndrome can be read only from the memory and can also be done in the order of the data. Using this method, the following embodiments propose to parallelize the syndrome computation and the data input (filling of the input buffer of the decoder). This is only feasible in the HoL blocking setting because in other cases, there is a pipeline between eblocks and the decoder is occupied by the decoding of the previous eblock when the next eblock is input.

[0047] Due to increased parallelism and / or a higher clock, the decoder is typically much faster than the input. Also, the decoder rate can be lower than the input rate, in which case, when the input is complete, some of the syndrome computation is complete and some latency is still reduced. This means that the decoder can catch up with the input rate. When the data input is complete, the syndrome computation is also complete. This syndrome computation typically takes one decoding iteration. Figure 7A is a timing diagram of the non - parallelized syndrome computation of the embodiment, and Figure 7B is a timing diagram of the parallelized syndrome computation of the embodiment. At the performance checkpoint, the decoder typically performs only a few iterations (e.g., two to three iterations), so the reduction in the overall iterations is quite substantial and significantly increases the decoding throughput.

[0048] One embodiment provides a parallelized syndrome computation implementation. The parallelized syndrome computation is implemented by adding a buffer access management unit that monitors the state of the input buffer. When new data is written, the unit signals the decoder to read the new data and add it to the syndrome computation. Since the decoder is typically faster than the input process, the syndrome computation can be performed substantially in parallel with the input and should be complete once the last of the input data is written. Figure 8 An example implementation of the parallelized syndrome computation of the embodiment is presented in Figure 8As shown, the decoder 800 has a syndrome calculation module 810. The decoder 800 provides the data read from the memory to the input buffer 820. The pointer management module 830 provides a write pointer 840 to the input buffer 820 and a read pointer 850 to the syndrome calculation module 810. In this example, a data row is processed once the read pointer is incremented, rather than starting processing after all data rows are valid.

[0049] In another embodiment, the techniques discussed above can be applied to read threshold calibration. In a NAND memory, each logical page contains a combination of several read thresholds. As part of the read threshold calibration, different read threshold combinations are identified, and a syndrome weight (SW) is calculated for each read threshold combination. The read threshold can also be simulated by calculating the bit values for each voltage interval from a limited set of reads. The simulation generates hypotheses for each combination in order to converge on the correct read threshold position. This operation can be represented as a bit error rate scan ("BES") operation. By using the above circuit, the SW calculation can start immediately after the start of the hypothesis calculation. In Figure 9 , 10A and 10B illustrate this embodiment.

[0050] Figure 9 is a diagram showing parallelized hypothesis calculation and syndrome calculation for an embodiment in which parallel syndrome weight calculation is used. Figure 9 Similar to Figure 8 , but in Figure 9 , the read threshold hypothesis calculation 900 is provided to the input buffer 820. Figure 10A and 10B are diagrams showing a BES operation with parallel syndrome weight calculation for an embodiment. As can be seen in these diagrams, the use of this embodiment greatly increases the efficiency of the BES operation and can reduce the BES latency by up to two times.

[0051] Figure 11 is a graph comparing the decoding throughput in megabytes per second for an embodiment, which shows the decoding performance of a single decoder in MB / second as a function of the mean BER. This graph compares the reference system with the systems of these embodiments. In the relevant BER performance range of one example implementation, the throughput gain can be approximately 10% to 50%.

[0052] There are several advantages associated with these embodiments. For example, these embodiments can significantly increase throughput and the efficiency of the decoder pool, with a negligible impact on silicon cost (since they reuse existing or other idle decoder resources to compute syndromes in parallel with decoder I / O operations). This increase can translate into decoder performance improvements (for a persistent target performance at higher BER), a reduction in decoder clock frequency (which reduces silicon area and power), and a reduction in the number of decoder engines in the pool (decoder parallelism reduction, decoder clock reduction, and decoder hierarchical buffer reduction, all of which reduce silicon area). These embodiments can also improve the BES, as they can significantly reduce the latency of BES operations. This is particularly important for X4 (four bits per cell) where the BES latency can be extremely high and / or for enterprise products that are QoS-sensitive.

[0053] Finally, as mentioned above, any suitable type of memory can be used. Semiconductor memory devices include: volatile memory devices such as dynamic random access memory (“DRAM”) or static random access memory (“SRAM”) devices; non-volatile memory devices such as resistive random access memory (“ReRAM”); electrically erasable programmable read-only memory (“EEPROM”); flash memory (which can also be considered a subset of EEPROM); ferroelectric random access memory (“FRAM”) and magnetoresistive random access memory (“MRAM”); and other semiconductor elements capable of storing information. Each type of memory device can have a different configuration. For example, flash memory devices can be configured in NAND or NOR configurations.

[0054] Memory devices can be formed from passive and / or active elements in any combination. By way of non-limiting example, passive semiconductor memory elements include ReRAM device elements which, in some embodiments, include resistive-switching memory elements such as antifuses, phase change materials, etc., and optionally steering elements such as diodes, etc. Further by way of non-limiting example, active semiconductor memory elements include EEPROM and flash memory device elements which, in some embodiments, include elements containing charge storage regions such as floating gates, conductive nanoparticles, or charge storage dielectric materials.

[0055] Multiple memory elements can be configured such that they are connected in series or such that each element can be accessed individually. By way of non-limiting example, flash memory devices in a NAND configuration (NAND memory) typically contain memory elements connected in series. A NAND memory array can be configured such that the array consists of multiple memory strings, where a string consists of multiple memory elements that share a single bit line and are accessed as a group. Alternatively, the memory elements can be configured such that each element can be accessed individually, such as in a NOR memory array. NAND and NOR memory configurations are examples, and the memory elements can be configured in other ways.

[0056] Semiconductor memory elements located within and / or on a substrate can be arranged in two or three dimensions, such as a two-dimensional memory structure or a three-dimensional memory structure.

[0057] In a two-dimensional memory structure, the semiconductor memory elements are arranged in a single plane or a single memory device level. Typically, in a two-dimensional memory structure, the memory elements are arranged in a plane that extends generally parallel to the main surface of the substrate that supports the memory elements (e.g., in the x-z plane). The substrate can be a wafer that is a layer on or in which the memory elements are formed, or it can be a carrier substrate that is attached to the memory elements after the memory elements are formed. By way of non-limiting example, the substrate can comprise a semiconductor such as silicon.

[0058] The memory elements can be arranged in a single memory device level in an ordered array such as multiple rows and / or columns, for example. However, the memory elements can be arranged in an irregular or non-orthogonal configuration. Each memory element can have two or more electrodes or contact lines, such as bit lines and word lines.

[0059] A three-dimensional memory array is arranged such that the memory elements occupy multiple planes or multiple memory device levels, thereby forming a structure in three dimensions (i.e., in the x, y, and z directions, where the y direction is generally perpendicular to the main surface of the substrate, and the x and z directions are generally parallel to the main surface of the substrate).

[0060] By way of non-limiting example, a three-dimensional memory structure can be arranged vertically as a stack of multiple two-dimensional memory device levels. As another non-limiting example, a three-dimensional memory array can be arranged as multiple vertical columns (e.g., generally perpendicular to the main surface of the substrate, i.e., columns extending in the y direction), where each column has multiple memory elements in each column. The columns can be arranged in a two-dimensional configuration (e.g., in the x-z plane), thereby resulting in a three-dimensional arrangement of memory elements having elements on multiple vertically stacked memory planes. Other configurations of memory elements in three dimensions can also constitute a three-dimensional memory array.

[0061] By way of non-limiting example, in a three-dimensional NAND memory array, memory elements can be coupled together to form NAND strings within a single horizontal (e.g., x-z) memory device tier. Alternatively, memory elements can be coupled together to form vertical NAND strings that traverse multiple horizontal memory device tiers. Other three-dimensional configurations can be envisioned, where some NAND strings contain memory elements in a single memory tier, while other strings contain memory elements spanning multiple memory tiers. Three-dimensional memory arrays can also be designed in NOR configurations and ReRAM configurations.

[0062] Generally, in an integrated three-dimensional memory array, one or more memory device tiers are formed above a single substrate. Optionally, the integrated three-dimensional memory array can also have one or more memory layers at least partially within the single substrate. By way of non-limiting example, the substrate can comprise a semiconductor such as silicon. In an integrated three-dimensional array, the layers that make up each memory device tier of the array are typically formed on the layers of the underlying memory device tier of the array. However, the layers of adjacent memory device tiers of the integrated three-dimensional memory array can be shared, or there can be intervening layers between the memory device tiers.

[0063] Thus, again, two-dimensional arrays can be formed separately and then packaged together to form a non-integrated memory device having multiple memory layers. For example, a non-integrated stacked memory can be constructed by forming memory tiers on separate substrates and then stacking the memory tiers on top of each other. The substrates can be thinned or removed from the memory device tiers prior to stacking, but since the memory device tiers are initially formed above separate substrates, the resulting memory array is not an integrated three-dimensional memory array. Additionally, multiple two-dimensional memory arrays or three-dimensional memory arrays (integrated or non-integrated) can be formed on separate chips and then packaged together to form a stacked chip memory device.

[0064] Associated circuitry is generally required for the operation of the memory elements and for communication with the memory elements. By way of non-limiting example, a memory device can have circuitry for controlling and driving the memory elements to perform functions such as programming and reading. This associated circuitry can be on the same substrate as the memory elements and / or on a separate substrate. For example, a controller for memory read and write operations can be located on a separate controller chip and / or on the same substrate as the memory elements.

[0065] Those skilled in the art will recognize that the present invention is not limited to the two-dimensional and three-dimensional structures described, but encompasses all relevant memory structures as described herein and as understood by those skilled in the art within the spirit and scope of the present invention.

[0066] It is to be understood that the foregoing detailed description is to be taken as illustrative of selected forms of the invention, and not as limiting of the invention. The scope of the invention is defined only by the appended claims (including all equivalents). Finally, it should be noted that any aspect of any of the embodiments described herein may be used alone or in combination with each other.

Claims

1. A data storage device, comprising: a memory; a decoder; an input buffer; and one or more processors, individually or in combination, configured to: monitor the input buffer; and in response to detecting that data read from the memory is being written in the input buffer, cause the decoder to perform syndrome calculation in parallel with the data being written in the input buffer to calculate the syndrome of the data, wherein at least a portion of the syndrome calculation is completed by the time the data is completely written in the input buffer.

2. The data storage device according to claim 1, wherein the one or more processors are individually or in combination further configured to: generate a write pointer pointing to a location in the input buffer where the data is to be written.

3. The data storage device according to claim 1, wherein the one or more processors are individually or in combination further configured to: generate a read pointer pointing to a location in the input buffer from which the data is to be read.

4. The data storage device according to claim 3, wherein the one or more processors are individually or in combination further configured to: in response to incrementing the read pointer rather than after all data rows in the input buffer have been verified, cause the data row to be included in the syndrome calculation.

5. The data storage device according to claim 1, wherein the decoder is configured to perform iterative decoding.

6. The data storage device according to claim 5, wherein the iterative decoding involves low density parity check codes.

7. The data storage device according to claim 1, further comprising at least one additional decoder, wherein the decoder and the at least one additional decoder form a decoder pool.

8. The data storage device according to claim 7, wherein the one or more processors are individually or in combination further configured to pipeline data to the decoder pool.

9. The data storage device according to claim 1, wherein the one or more processors are individually or in combination further configured to use the syndrome to estimate the bit error rate.

10. The data storage device according to claim 1, wherein the syndrome calculation consumes one decoding iteration.

11. The data storage device according to claim 1, wherein the memory comprises a three-dimensional memory.

12. A method, in a data storage device comprising a memory, a decoder, and an input buffer, the method comprising: identifying a plurality of hypothesized read thresholds of the memory; and after the plurality of hypothesized read thresholds have been identified: calculating the syndrome weight for each of the plurality of hypothesized read thresholds; and using the calculated syndrome weight for each of the plurality of hypothesized read thresholds in a bit error rate estimation scan (BES) operation.

13. The method according to claim 12, further comprising monitoring the input buffer; and In response to detecting that data read from the memory is being written into the input buffer, causing the decoder to perform syndrome calculation in parallel with the data being written into the input buffer to calculate the syndrome of the data, wherein by the time the data is completely written into the input buffer, the syndrome calculation is completed.

14. The method according to claim 13, further comprising, in response to incrementing a read pointer rather than after all data rows in the input buffer have been verified, causing the data rows to be included in the syndrome calculation.

15. The method apparatus according to claim 12, wherein the decoder is configured to perform iterative decoding.

16. The method according to claim 15, wherein the iterative decoding involves low-density parity-check codes.

17. The method according to claim 12, further comprising at least one additional decoder, wherein the decoder and the at least one additional decoder form a decoder pool.

18. The method according to claim 12, wherein the syndrome weight calculation consumes one decoding iteration.

19. The method according to claim 12, wherein the memory includes a three-dimensional memory.

20. A data storage device, comprising: a memory; a decoder; an input buffer; and means for causing the decoder to perform syndrome calculation in parallel with data being written into the input buffer to calculate the syndrome of the data in response to detecting that data read from the memory is being written into the input buffer, wherein at least a portion of the syndrome calculation is completed by the time the data is completely written into the input buffer.