Systems and methods for configurable hardware architecture for read operations of flash memories
By dynamically adjusting the read threshold in NAND flash memory and using machine learning models to optimize read operations, the problem of NAND flash memory read output errors is solved, improving read performance and accuracy while reducing system power consumption and complexity.
Patent Information
- Application Number
- CN202511323538.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-21
- Filing Date
- 2025-09-16
- Publication Date
- 2026-03-17
AI Technical Summary
During the programming and reading process of NAND flash memory, errors may occur in the read output due to different stress conditions. Existing technologies are unable to effectively improve decoding capabilities, thus affecting read performance.
A configurable hardware block is provided that dynamically adjusts the read threshold for each row, uses machine learning models such as DNN to optimize the read threshold estimation, and combines lookup tables and K-means search to reduce the read retry rate and improve read performance.
This achieves a reduction in read failure probability, an improvement in read accuracy and throughput, and a reduction in system gate count and power consumption without affecting read stream performance.
Smart Images

Figure CN121687149A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 695,132, filed September 16, 2024; U.S. Provisional Patent Application No. 63 / 695,114, filed September 16, 2024; and U.S. Non-Provisional Utility Patent Application No. 18 / 921,321, filed October 21, 2024, the entire contents of which are incorporated herein by reference for all purposes. Technical Field
[0003] This embodiment generally relates to systems and methods for performing operations on flash memory, and more specifically to systems and methods for providing configurable hardware blocks to perform read operations on flash memory. Background Technology
[0004] As the number and types of computing devices continue to expand, so does the demand for memory used in such devices. Memory includes volatile memory (e.g., RAM) and non-volatile memory. A popular type of non-volatile memory is flash memory, or NAND flash memory. A NAND flash memory array consists of rows and columns (strings) of cells. Cells may include transistors.
[0005] Errors may occur in the programming and reading outputs due to varying stress conditions (e.g., NAND noise and interference sources) during NAND flash memory programming and / or reading. Improvements in decoding capabilities under such a wide range of stress conditions in NAND flash memory devices remain desirable. Summary of the Invention
[0006] This embodiment relates to a system and method for providing a configurable hardware block to perform read operations on a flash memory.
[0007] According to certain aspects, embodiments provide a method for performing operations on a non-volatile memory comprising one or more blocks, each block comprising a plurality of rows of cells. The method can comprise generating, by a first processor, a first configuration comprising a pointer to a first set of pre-defined configurations of a plurality of pre-defined configurations for performing read operations on the non-volatile memory. The method can comprise generating, in a memory, the first set of pre-defined configurations in response to generating the first configuration. The method can comprise performing, by a controller, a first operation according to the first set of pre-defined configurations generated in the memory. The method can comprise generating, by a second processor, a second configuration comprising a pointer to a second set of pre-defined configurations of the plurality of pre-defined configurations. The method can comprise generating, in the memory, the second set of pre-defined configurations in response to generating the second configuration. The method can comprise performing, by the controller, a second operation according to the second set of pre-defined configurations generated in the memory.
[0008] According to other aspects, embodiments provide a flash memory system comprising a non-volatile memory, a controller for performing operations on the non-volatile memory, and a plurality of processors comprising a first processor and a second processor. The non-volatile memory can comprise one or more blocks, each block comprising a plurality of rows of cells. The first processor generates a first configuration comprising a pointer to a first set of pre-defined configurations of a plurality of pre-defined configurations for performing read operations on the non-volatile memory. In response to generating the first configuration, the controller can generate, in a memory, the first set of pre-defined configurations. The controller can perform a first operation according to the first set of pre-defined configurations generated in the memory. The second processor can generate a second configuration comprising a pointer to a second set of pre-defined configurations of the plurality of pre-defined configurations. In response to generating the second configuration, the controller generates, in the memory, the second set of pre-defined configurations. The controller can perform a second operation according to the second set of pre-defined configurations generated in the memory. BRIEF DESCRIPTION OF DRAWINGS
[0009] These and other aspects and features of the present embodiments will become apparent to those ordinarily skilled in the art from a review of the following description of specific embodiments, taken in conjunction with the accompanying drawings, in which:
[0010] Figure 1 An example of a voltage threshold distribution is shown in accordance with some embodiments;
[0011] Figure 2 An example process illustrating a read flow in a conventional flash device is described;
[0012] Figure 3An example of a fully connected (FC) deep neural network (DNN) for a row-to-row (R2R) estimator is shown, in accordance with some embodiments;
[0013] Figure 4 An example of a read flow employing an R2R estimator for all stages in the read flow is shown, in accordance with some embodiments;
[0014] Figure 5 is a diagram showing an example static random access memory (SRAM) structure, in accordance with some embodiments;
[0015] Figure 6 is a diagram showing an example codebook database, in accordance with some embodiments;
[0016] Figure 7 is a diagram showing an example row-to-row (R2R)-lookup table (LUT) database, in accordance with some embodiments;
[0017] Figure 8 is a diagram showing an example weight coefficient matrix, in accordance with some embodiments;
[0018] Figure 9 is a diagram showing an example of DNN parameter placement in memory, in accordance with some embodiments;
[0019] Figure 10 is a diagram showing an example scheme of a set of register file (regfile) configurations, in accordance with some embodiments;
[0020] Figure 11 is a diagram showing an example fully connected (FC) DNN, in accordance with some embodiments;
[0021] Figure 12 is a diagram showing an example entity embedding scheme, in accordance with some embodiments;
[0022] Figure 13 is a diagram showing an example of fixed point computation of neurons in a layer, in accordance with some embodiments;
[0023] Figure 14 is a diagram showing an example architecture of a fixed point computation scheme, in accordance with some embodiments;
[0024] Figure 15 is a diagram showing an example DNN computation scheme, in accordance with some embodiments;
[0025] Figure 16 is a diagram showing an example DNN engine master interface, in accordance with some embodiments;
[0026] Figure 17 is a diagram showing an example R2R engine master interface, in accordance with some embodiments;
[0027] Figure 18 FIG. 1 is a diagram illustrating an example instance of a per-threshold computation building block, according to some embodiments;
[0028] Figure 19 FIG. 2 is a diagram illustrating an example of an iterative K-means search based on x3 instances of the per-threshold computation building block, according to some embodiments;
[0029] Figure 20 FIG. 3 is a diagram illustrating an example of a K-means search engine home interface, according to some embodiments;
[0030] Figure 21 FIG. 4 is a diagram illustrating an example of a top-level architecture of a read operation hardware, according to some embodiments;
[0031] Figure 22 FIG. 5 is a diagram illustrating an example scheme for loading a mem-regfile, according to some embodiments;
[0032] Figure 23 FIG. 6 is a diagram illustrating an example output mapping logic, according to some embodiments;
[0033] Figure 24 FIG. 7 is a diagram illustrating an example of chipping and DC offset adjustment, according to some embodiments;
[0034] Figure 25 FIG. 8 is a diagram illustrating an example system with a pipelined hardware utilizing multiple CPU operations, according to some embodiments;
[0035] Figure 26 FIG. 9 is a diagram illustrating an example hardware implementation of HT-GET-DNN for a first stage read operation, according to some embodiments;
[0036] Figure 27 FIG. 10 is a diagram illustrating an example hardware implementation of HT-GET-DNN using HT-codebook (CB) indexing with R2R DNN for target row threshold estimation, according to some embodiments;
[0037] Figure 28 FIG. 11 is a diagram illustrating an example hardware implementation of HT-GET-LUT for using HT-CB indexing with R2R lookup table (LUT) for target row threshold estimation, according to some embodiments;
[0038] Figure 29 FIG. 12 is a diagram illustrating an example hardware implementation for general DNN operations, according to some embodiments;
[0039] Figure 30 FIG. 13 is a diagram illustrating an example hardware implementation using LUT for R2R target row to reference row threshold estimation, according to some embodiments;
[0040] Figure 31 This is a diagram illustrating an example hardware implementation of R2R reference row to target row threshold estimation using a LUT according to some embodiments;
[0041] Figure 32 This is a diagram illustrating an example hardware implementation of an HT-Set that uses K-means search to compute the CB index given an input threshold, according to some embodiments;
[0042] Figure 33 This is a block diagram illustrating an example flash memory system based on some arrangement.
[0043] Figure 34 This is a flowchart illustrating an example method for providing a configurable hardware block to perform a flash memory read operation, according to some embodiments. Detailed Implementation
[0044] According to some aspects, embodiments of this disclosure relate to techniques for providing configurable hardware blocks to perform read operations on flash memory.
[0045] In conventional flash memory systems (e.g., the controller in a NAND flash device), a simplified read stream can be implemented, where fixed thresholds are used at the start of lifetime (SOL). These thresholds are referred to as default thresholds, first-phase read thresholds, or normal read thresholds. In the event of a failure, predetermined thresholds from a lookup table (LUT) can be used to perform read retries. If the retry succeeds, these thresholds can be used for all other reads from the same block. This simple and straightforward approach has limitations and is typically implemented in firmware (FW) without degrading system read performance. However, when more complex read streams are introduced, there is a risk of performance degradation. This is due to the increased latency associated with executing more complex algorithms, which may necessitate optimizing read thresholds on a per-command basis, potentially impacting overall system read performance.
[0046] To address these issues, embodiments of this disclosure relate, according to certain aspects, to systems and methods for improving the performance of read operations having a configurable hardware architecture for read operations in NAND flash memory (e.g., read digital signal processor hardware (RDSP-HW) operations). In some embodiments, the flash memory system (e.g., a NAND flash device) can provide a generic block that enables RDSP operations in a controller of the NAND flash device. In some embodiments, the flash memory system (“System”) can dynamically adapt read thresholds during a read flow on a per-row, per-stress basis, replacing conventional fixed default thresholds.
[0047] In some embodiments, the system can provide an RDSP-HW architecture that allows for the real-time calculation and application of per-row optimized thresholds without any degradation in read performance. In the event of a read failure, the system can calculate an estimated optimized read threshold for the failed row, which is then used to reread the failed row. In some embodiments, the system can search an offline-prepared quantized database to identify a compressed version of the index used as a reference row for the read threshold. In some embodiments, the system can save or store this index for future read commands from different rows within the same block.
[0048] In some embodiments, the system may provide one or more RDSP-HW blocks that perform read operations with minimal latency and high throughput (e.g., optimizing the estimation of read thresholds, identifying and saving compressed versions of read thresholds, etc.) to ensure that performance requirements are met.
[0049] In some embodiments, the system may provide a centralized focus block within the system that wraps and / or implements one or more RDSP algorithms for managing read stream operations and read retry stream operations. The one or more RDSP algorithms may include, for example, row-to-row (R2R) estimation of the read threshold, estimation of the read threshold using a machine learning model (e.g., a DNN), fast threshold tracking (QT), a compressed version of K-means search for the read threshold, etc.
[0050] In some embodiments, the system can provide hardware acceleration that enables complex RDSP operations to run with short latency. Compared to conventional algorithms implemented in firmware, the system can provide a more accurate read threshold and thus reduce the read retry rate (RRR) without impacting read stream performance (e.g., performance measured in operations per second (IOPS) or throughput) during the start of lifetime (SOL).
[0051] In some embodiments, the system may provide a reusable architecture (e.g., RDSP-HW) that can reuse the same or shared engines (e.g., circuitry, firmware, software, or combinations thereof) for different RDSP algorithms (e.g., R2R, DNN, QT, K-means search). The same or shared engines can reduce gate count and power consumption in the flash memory system.
[0052] In some embodiments, the system can provide a highly configurable architecture (e.g., RDSP-HW) that uses different register files (“regfiles”) to enable different streams and / or parameters to efficiently support current and future flash devices with tuned algorithms and / or parameters. In some embodiments, the system can enable multiple processors (e.g., CPUs) to access read operation blocks (e.g., RDSP-HW blocks) simultaneously. Each CPU can be aware of or use the read operation block as a different virtual machine, with a dedicated register file configuration space (e.g., a per-CPU regfile in memory). The system can provide a highly configurable architecture designed to synchronize and manage tasks originating from different CPUs. This architecture allows multiple CPUs to interact with the same read operation block orthogonally, thereby eliminating the need to copy this read operation block to each CPU.
[0053] In some embodiments, the system can provide a hardware architecture for rapid configuration and read state, which can perform read operations (e.g., RDSP operations) on the read stream during SOL without performance degradation. Even before the first retrace without performance degradation, the system can achieve high read performance by reducing the probability of read failures by adjusting the read threshold applicable to SOL conditions. This can be achieved by replacing the traditional default read with a row-to-row (R2R) estimator used during the first phase of the read.
[0054] In some embodiments, the system can provide a single general-purpose DNN hardware engine (e.g., a DNN engine) that can be used with different algorithms using different parameters (e.g., R2R, QT). For example, for R2R estimation, the DNN engine can receive stress conditions and a target row as input and compute a target threshold to be used for the target row under the current stress conditions as output. In this way, the system can achieve high read performance and allow real-time estimation of the target page read threshold for each read operation. For QT operations, the DNN engine can receive stress conditions and histograms of several simulated reads with fixed threshold rows as input and compute the estimated optimal read threshold for the current row as output. The estimated threshold can be configured as a NAND for read retries.
[0055] According to some aspects, embodiments of this disclosure relate to a method for performing operations on a non-volatile memory comprising one or more blocks, each block comprising multiple rows of cells. The method may include generating a first configuration by a first processor, the first configuration including pointers to a first set of predefined configurations among a plurality of predefined configurations for performing read operations on the non-volatile memory. The method may include generating the first set of predefined configurations in memory in response to generating the first configuration. The method may include performing a first operation by a controller based on the first set of predefined configurations generated in memory. The method may include generating a second configuration by a second processor including pointers to a second set of predefined configurations among a plurality of predefined configurations. The method may include generating a second set of predefined configurations in memory in response to generating the second configuration. The method may include performing a second operation by a controller based on the second set of predefined configurations generated in memory.
[0056] According to some aspects, embodiments of this disclosure relate to a flash memory system including non-volatile memory, a controller for performing operations on the non-volatile memory, and a plurality of processors including a first processor and a second processor. The non-volatile memory may include one or more blocks, each block including multiple rows of cells. The first processor may generate a first configuration including pointers to a first set of predefined configurations among a plurality of predefined configurations for performing read operations on the non-volatile memory. In response to generating the first configuration, the controller may generate the first set of predefined configurations in the memory. The controller may perform a first operation based on the first set of predefined configurations generated in the memory. The second processor may generate a second configuration including pointers to a second set of predefined configurations among a plurality of predefined configurations. In response to generating the second configuration, the controller may generate the second set of predefined configurations in the memory. The controller may perform a second operation based on the second set of predefined configurations generated in the memory.
[0057] The embodiments of this disclosure have at least the following advantages and benefits. First, the embodiments of this disclosure can provide a highly configurable architecture that enables different flows / parameters to efficiently support current and future flash devices with adapted algorithms / parameters. For example, the system can achieve rapid configuration by copying the configuration set from static random access memory (SRAM) to a register file in memory (e.g., mem-regfile), which is faster than conventional Advanced Peripheral Bus (APB) configuration and reduces CPU configuration time. The system can achieve area savings because the configurable architecture is instantiated using only a single mem-regfile, rather than being copied in a per-CPU regfile for each CPU. The system can also minimize APB traffic because only pointers are configured on the APB, and the hardware copies the configuration from SRAM to the mem-regfile without loading the APB bus.
[0058] Second, the embodiments in this disclosure can provide hardware acceleration that enables complex RDSP operations to run with low latency, providing higher read threshold accuracy compared to simple algorithms implemented in firmware, and thus the system can reduce RRR without impacting read stream performance. The system can provide a method for rapid configuration and read state, thereby facilitating read operations that do not degrade the read stream performance during Start of Life (SOL). Even before the first retrace without performance degradation, the system can achieve high read performance by reducing the probability of read failures by adjusting the read threshold to suit SOL conditions. This can be achieved by replacing the conventional default read with a row-to-row (R2R) estimator used during the first phase of reading.
[0059] Third, the embodiments in this disclosure can provide a reusable architecture for reusing the same or shared engines (e.g., hardware, firmware, software, or a combination thereof) for different RDSP algorithms, enabling shared hardware engines to reduce gate count and power consumption in flash memory systems. For example, a single general-purpose DNN hardware can be used for different algorithms using different parameters (e.g., R2R-DNN, QT, K-means search).
[0060] Fourth, the embodiments in this disclosure can provide an architecture that enables multiple CPUs to access read operation blocks (e.g., RDSP-HW blocks) simultaneously. Each CPU can use and / or be aware of the read operation block as a different virtual machine with a dedicated register file configuration space (e.g., a per-CPU regfile in memory). The system can synchronize and manage tasks originating from different CPUs and allows multiple CPUs to interact with the same read operation block orthogonally, thereby eliminating the need to copy such read operation blocks for each CPU. The system can enable multiple CPUs to access read operation blocks or units (e.g., RDSP-HW units) simultaneously for configuration, reading status, or polling. This architecture allows multiple CPUs to access a single RDSP-HW unit. Therefore, this architecture can also reduce area because there is a single RDSP-HW unit working with multiple CPUs, rather than a per-CPU dedicated RDSP-HW unit. This architecture can configure only pointers to SRAM (instead of configuring the full configuration), thereby minimizing APB traffic. The RDSP-HW logic can retrieve a predefined configuration in SRAM based on the pointer and copy the predefined configuration in SRAM to the mem-regfile.
[0061] refer to Figures 1 to 34 This describes and illustrates embodiments of a system and method for dynamically adjusting the read threshold based on the optimal threshold representation for each row in this solution.
[0062] Figure 1 An example of a voltage threshold distribution 100 according to some embodiments is shown. Figure 1 This describes the voltage threshold distribution for a 4-bit per cell (bpc) flash memory device (i.e., a four-level cell (QLC) with 16 programmable states). The voltage threshold (VT) distribution comprises 16 lobes. Lower page reads require thresholds T1, T3, T6, and T12. Middle pages are read using read thresholds T2, T8, T11, and T13. Upper pages are read using read thresholds T4, T10, and T14. Top pages are read using thresholds T5, T7, T9, and T15. The lowest lobe (0) is referred to as the erase level. Hold, programming / erasing cycles, and read interference can alter the voltage threshold distribution in various ways (e.g., Figure 1 The voltage threshold distribution shown in the diagram is used to generate various bit error ratio (BER) conditions. For each condition, a different read threshold can be selected to achieve the lowest BER after a read operation. Therefore, the read threshold of the target page in the NAND device is repeatedly estimated during the device's lifetime to maintain high read performance and benefit from an efficient read stream with low latency, which avoids SB decoding (soft bit decoding) as much as possible.
[0063] Figure 2 This section describes a simplified example of the read stream process in a conventional flash device. Figure 2 Typical stages of read retry in fault conditions are described. By default, the flash memory system (e.g., the controller of a NAND flash device) can perform a first-stage read, which refers to a read with a pre-configured (or predefined) initial default threshold (step 202). In some embodiments, the read operation (e.g., a read digital system processor) and error correction code (ECC) operation can be implemented in the controller of the NAND flash device.
[0064] The system (e.g., the controller of a NAND flash memory device) can decode reads by a hard bit (HB) decoder (e.g., a decoder that operates on binary inputs) (step 204). In the event of a decoding failure, the controller can refer to a shift table that stores several threshold candidates. These candidate thresholds are also referred to as a “retry fixed threshold table.” Upon a first (read) failure on a page, the controller can select or choose a first table entry, configure the NAND threshold based on the first entry, read the same page again, and perform HB decoding (step 206). In the event of a second failure, this process can be repeated with other shift table candidates until HB decoding succeeds. Upon successful HB decoding, shift table entries (e.g., threshold candidates for reads corresponding to successful HB decoding) can be stored in a table available for each block called the history table (HT). Pointers to the HT can be used for future recurring reads from the same block to allow the controller to use the same threshold compatible with the current stress of that block. If decoding with all shift table candidates fails, the controller can perform fast threshold tracking (QT) to estimate the optimal threshold for the current row (step 208). QT can perform several simulated reads using a fixed threshold, from which a histogram is computed. An estimator (e.g., a controller or software, firmware, hardware, or a combination thereof) can use the histogram to estimate the current threshold. The estimator can be a linear estimator or a DNN-based estimator. The controller can configure the estimated threshold into the NAND and perform read retries, followed by HB decoding (step 210). If HB decoding fails, the controller can perform higher-complexity threshold tracking (step 212), such as pre-soft tracking (PST), followed by sampling and / or soft decoding (step 214).
[0065] In some embodiments of the invention, the system (e.g., a NAND flash device or its controller) can perform row-to-row (R2R) estimation. Depending on the physical characteristics of the NAND, each NAND row of each block exhibits a typical voltage threshold (VT) probability distribution. In 3D NAND, there may be a typical distribution for each word line (WL), where rows within a given WL can have similar VT distributions (referred to as row-VT distributions). Therefore, if a threshold for a target row is known as a result of activating an estimation process on that row, it may be useful to use that result and estimate the threshold of any other row from that given row (e.g., the target row) and the threshold of that given row by using a typical row-VT distribution, thereby saving the cost and / or overhead of per-row threshold estimation.
[0066] According to some embodiments of this disclosure, when the controller performs a first-stage read, a row-to-row (R2R) estimator can be trained to provide a minimized retry probability. The R2R estimator can receive the target row as input and provide an optimal shift for the first-stage read shift application (e.g., optimal in terms of reducing the retry probability). In some embodiments, the first-stage read shift can be a zero shift of a default threshold. The R2R estimator can be implemented in various ways, including (1) a lookup table (LUT) that provides a shift per threshold and per row; (2) a linear-based estimator; and / or (3) a DNN-based estimator. In some embodiments, the LUT-based R2R estimator for the first-stage read can be fully optimized to support all required stresses to provide the lowest RRR for the first-stage read using the LUT (e.g., a LUT that provides a shift per threshold and per row). As NAND density increases, blocks can become larger due to more layers and strings per block. An advantage of using a DNN-based R2R estimator is the relatively small memory requirements for such large blocks. Therefore, a DNN-based R2R estimator can efficiently perform LUT compression. This DNN-based compression can also be extended to future NAND devices.
[0067] As another embodiment of this disclosure, an R2R estimator can be trained for a fixed set of thresholds used within a read retry process (or read retry procedure / operation). That is, the R2R estimator can have a specific training configuration for each entry in a retry fixed threshold table, where each entry represents another subset of stress conditions supported by the controller. For example, in the case of data retention (DR) stress, thresholds can be optimized on a specific row referred to as a "reference row." The table (e.g., a LUT for R2R) can also be optimized on this stress to transform the reference row thresholds to every other row under this DR stress.
[0068] In some embodiments, the R2R estimator can be more formally described as:
[0069] TH r2r (row, ShiftIdx) = TH ref (ShiftIdx)+LUT shiftIdx (row) (Equation 1)
[0070] For each shift index, a LUT can be defined per row to provide a target threshold. The shift index can be an index of a retry fixed threshold table that serves as an index to the retry fixed threshold table. An "index" or "shift index" refers to a retry pointer stored for each block. The retry pointer can be associated with stress conditions. Maintaining a LUT for each shift index means that a different R2R estimator exists for each read retry. The row index can be a pointer to an entry in the LUT. This can adjust the R2R estimate based on stress conditions. In some embodiments, a first-stage read may correspond to ShiftIdx = 0. This LUT-based implementation may be memory-inefficient. In a LUT implementation, a suboptimal solution for memory saving can use a common LUT for all shift indices, as follows:
[0071] TH r2r (row, ShiftIdx) = TH ref (ShiftIdx) + LUT(row) (Equation 2)
[0072] The same LUT can be used for all shifted indexes. For reads after Fast Threshold Tracking (QT), the LUT can also be the same table. The reference threshold for reads after QT can be mapped from the failed row to the (common) reference row using the LUT, and the threshold can then be compressed by clustering to the nearest cluster (e.g., using K-means clustering), with only the index cluster centers saved as ShiftIdx. This compression significantly reduces the memory requirements per threshold tracking operation, allows the use of a compact history table (HT) to save the state of blocks after a failure, and / or allows for near-optimal thresholds for all rows using an R2R estimator with mapped ShiftIdx after QT.
[0073] As another embodiment of this disclosure, the R2R estimator can be implemented by a DNN that receives ShiftIdx as input features, concatenates row indices (e.g., the row index of the target row), and provides a threshold to be used to read the target row. ShiftIdx can be obtained from the history table of each block.
[0074] TH r2r (row)=DNN(ShiftIdx, row) (Equation 3)
[0075] Figure 3An example of a fully connected (FC) deep neural network (DNN) 300 for a row-to-row (R2R) estimator is shown according to some embodiments. The example DNN may include an input layer 302, one or more hidden layers 303, and / or an output layer 304. Figure 3 In the example DNN shown, input layer 302 may include a target row index (e.g., an index pointing to the target row) and a shift index. Output layer 304 may include an estimated threshold for the target row.
[0076] In some embodiments of this disclosure, row indices can be represented by entity embeddings (EEs), which are the result of training a DNN estimator (e.g., a DNN-based R2R estimator) with one-hot inputs. In some embodiments, the entity embeddings of row indices can be implemented or obtained by training the row indices of several neurons (e.g., neuron 305) fully connected to the DNN with one-hot inputs. The entity embedding values for each row can be stored in a LUT that is used as input instead of a one-hot input. For example, a LUT can map row indices to values of neurons connected to the original one-hot inputs. LUTs can be used to provide neuron values for each row index instead of one-hot inputs and fully connected neuron weights. This saves significant memory and reduces implementation complexity. LUT-based implementations of entity embeddings (EEs) are very robust for large NAND blocks with multiple rows. Because entity embedding (EE) implementations save memory and reduce implementation complexity, EEs can be used for large NAND blocks. EEs can be an alternative form for implementing row index encoding of neuron values.
[0077] As another embodiment of this disclosure, a DNN (or a DNN-based R2R estimator) can be trained using an input threshold corresponding to either the optimal threshold of the reference row selected in (1) or the QT threshold of the reference row selected in (2). In some embodiments, the R2R threshold obtained by the DNN-based R2R estimator can be given by the following formula:
[0078] TH r2r (row)=DNN(ShiftIdx,row,TH HT-ref (Equation 4)
[0079] ShiftIdx (shift index) can be a pointer to a stage / retry period in the history table. The shift index can correspond to the number of retries or the current stress condition (e.g., a retry index). This retry index can be a subset of the history table (HT). The HT can be a generalized form where each block stores thresholds corresponding to different stress conditions. ShiftIdx can be a pointer to a generalized HT. The initial few entries of the HT (e.g., low index values) can correspond to stresses in several ordered life start (SOL) groups, so the shift index can be used as input to the DNN. The input can be a reference threshold from the HT that is closest to the estimated threshold through QT operations while being read in a flow-dependent manner.
[0080] Figure 4 An example of a read process 400 employing an R2R estimator for all stages of the read process, according to some embodiments, is shown. Figure 4 A reading process is illustrated that applies an R2R transformation to the input threshold based on the reading stage. According to some embodiments of this disclosure, the R2R threshold can be obtained from or acquired by an R2R estimator. Figure 4 An exemplary read flow is described, which includes row-to-row (R2R) estimation within normal reads and shift retries. Figure 4 The process of retrying the read is also described. Figure 4The read flow shown includes receiving and / or executing a read command for the target page (step 402). A History Table (HT)-Get operation can extract the state of the holding block and point to an HT index (e.g., an index to the history table) of the type of read on the first stage (e.g., a first-stage read) (step 404). By default, the flash memory system (e.g., one or more controllers of a NAND flash device) can perform a first-stage read (step 406), which refers to a read with a pre-configured initial default threshold. The system can also apply an R2R estimator (step 408) to adapt to the target row. In some embodiments, all read operations (e.g., RDSP operations) and error correction code (ECC) operations can be implemented in one or more controllers. Reads can be decoded via a hard bit (HB) decoder (e.g., a decoder that operates on binary inputs) (step 420). In the event of decoding failure, the controller can refer to a "shift table" that stores several threshold candidates. In the event of a first failure, the controller can select the first entry of the shift table and jointly configure the NAND threshold with the R2R estimator (step 410) to the target row. The controller can then read the same page again and perform HB decoding (step 412). In the event of a second failure, the controller can repeat this process with other shift table candidates and R2R estimators until HB decoding is successful. Upon successful HB decoding, the corresponding shift table entry can be stored in a table called the History Table (HT) available for each block (step 414). A pointer to the HT can be used for future reads from the same block (step 416) to allow the controller to use the same threshold compatible with the current stress of that block.
[0081] In some embodiments, if decoding with all shift table candidates fails, the controller may perform a Fast Threshold Tracking (QT) (step 422) to estimate the optimal threshold for the current row. In some embodiments, the QT may perform several simulated reads with a fixed threshold, from which a histogram is computed. The histogram can be used to estimate the current optimal read threshold. In some embodiments, the threshold estimator may be a linear estimator or a DNN-based estimator. The current estimated threshold may be configured into the NAND for retrying reads and HB decoding. An R2R operation (e.g., a LUT-based R2R operation or a DNN-based R2R operation) may transfer the currently estimated threshold to a reference row threshold, and the reference row threshold may be used to update the HT table. In some embodiments, the flash system (e.g., the controller) may perform an HT-Set operation (step 414) that compresses the threshold into an index pointer (e.g., HTindex) of the HT table. The HTindex may point to the HT threshold closest to the estimated read threshold and may be used for subsequent reads from the same block (step 416). If HB decoding fails after QT (step 424), then perform higher complexity threshold tracking (step 426), for example, pre-soft tracking (PST), followed by soft decoding (step 4128).
[0082] In the following sections, the hardware architecture according to some embodiments will be described in a bottom-up manner, including, for example, (1) the main hardware database in SRAM, (2) the main hardware engine, (3) the RDSP-HW top layer, (4) the main RDSP algorithm, (5) the read stream, and (6) the read retry stream in this order. The main hardware database in SRAM (see [link to documentation]) Figure 5 ) may include a codebook (see Figure 6 R2R LUT (see) Figure 7 K-means weights (see...) Figure 8 ), DNN parameters (see Figure 9 ) and regfile configuration set (see Figure 10 The main hardware engine may include a DNN engine (see...). Figures 11 to 16 LUT engine (see) Figure 17 ) and K-means search engine (see Figures 18 to 20 RDSP-HW top-level configuration may include connections to the engine / database / configuration (see...). Figure 21 ), regfile configuration (see Figure 21 ), per CPU regfile (see Figure 21 mem-regfile (for quick configuration of predefined configuration sets; see also...) Figure 22 ), mapping features (for efficient and short read states; see Figure 23), clipping, and DC offset (see Figure 24 It can also operate with multiple CPUs simultaneously (e.g., task scheduling and arbitration; see...). Figure 25 The main RDSP algorithm may include a SOL read stream, a MOL (Middle of Lifetime) / EOL (End of Lifetime) read stream, and an HT-SET (e.g., an HT-SET including QT, R2R, and / or K-means search). The SOL read flow may include codebook (CB) reading (e.g., extracting read thresholds from the codebook without additional RDSP operations), an R2R LUT (e.g., using a dedicated LUT for the first / second / third HT index to perform R2R LUT-based operations to provide a target row threshold for each HT index), and an R2R-DNN (e.g., using a dedicated DNN for the first / second / third HT index to perform DNN-based operations to provide a target row threshold for each HT index). The MOL / EOL read flow may include R2R-HT (e.g., extracting read thresholds from the codebook and performing R2RLUT-based operations and / or R2R DNN-based operations to provide a target row threshold for each set of input thresholds on a reference row). The read stream may include DNN-based HT-GET (e.g., R2R normal, shift, HT; see [link]). Figure 26 and Figure 27 ) and LUT-based HT-GET (e.g., R2R normal, shift, HT; see Figure 28 The read retry process can include QT (see...). Figure 29 ) and HT-SET (see Figure 32 ).
[0083] In some embodiments, the system may include a master hardware database in SRAM, allowing SRAM to be used in a hardware architecture (e.g., RDSP-HW). SRAM is generally more efficient than registers (e.g., flip-flops) used for large-scale data storage due to its higher density and smaller physical coverage. In hardware architectures according to some embodiments (e.g., RDSP-HW), SRAM can be used to store various databases accessed during RDSP operation.
[0084] Figure 5This is a diagram illustrating an example static random access memory (SRAM) structure 500 according to some embodiments. In some embodiments, to improve access efficiency, the SRAM can be implemented as multiple physical SRAM cuts 501, 502, 503, 504 instead of a single large SRAM block. This approach increases the bandwidth available for accessing the SRAM because each physical SRAM cut can be accessed simultaneously. This design can be particularly effective when sequentially accessing a database. For example, the SRAM can be implemented in four SRAM cuts 501, 502, 503, 504, each SRAM cut having a width of 16 bytes, and can be used to store a database (e.g., database #1, database #2). The database can be distributed across the four memories (e.g., the four SRAM cuts) in a ping-pong-like configuration. As a result, sequential reads from the database can involve accessing different physical memory cuts in each clock cycle. This arrangement can allow efficient sequential read transactions from two different clients, effectively utilizing the full bandwidth (16 bytes per port per cycle). To ensure that two clients do not access the same SRAM block simultaneously, an arbitration mechanism (e.g., a memory arbitrator 510) can be implemented. This arbitration can verify access requests and manage the order in which clients access SRAM blocks, thus preventing conflicts. At the start of a transaction, for a single client, a minimum initial latency (one pushback) can be expected for only one client at the start of the transaction.
[0085] In some embodiments, the flash system (e.g., the controller) may include a master hardware database in SRAM to support dynamic allocation of databases in SRAM, allowing flexibility in algorithm optimization and trade-offs. Databases may be initialized in SRAM before the database is first used (typically after power-on). In some embodiments, these databases are not static and may be scaled from one setup to another. For example, a database may be scaled up at the expense of another to better support a particular flash device. Different sets of databases may be allocated in different setups to optimize RDSP algorithms for another flash device. One constraint in this dynamic allocation may be that the total size of all databases must fit within the available SRAM memory budget. In summary, databases according to some embodiments may be stored in SRAM and used during RDSP operation. This configuration can provide efficiency and density. SRAM can be used for registers (e.g., flip-flops) for large-scale data storage due to its higher density and smaller footprint. In some embodiments, the system may provide multiple physical SRAM shards for increased bandwidth, rather than a single large SRAM block. This design increases the bandwidth for accessing SRAM because each shard can be accessed simultaneously (e.g., particularly efficient for sequential access, which can be used in RDSP operation). In some embodiments, the system can perform dynamic database allocation in SRAM, thereby enabling flexibility to allow for optimization of the RDSP algorithm and to allow the database to be adjusted from one setting to another to better support different flash devices.
[0086] Figure 6 This is a diagram illustrating an example codebook database 600 according to some embodiments. A database for a hardware architecture (e.g., RDSP-HW) according to some embodiments may include one or more codebooks (CBs). In some embodiments, the database may include a codebook table comprising an m1 set of read thresholds (e.g., HT CodeBook[0], ..., HTCodeBook[m1-1]). Each row in the CB contains n read thresholds, where n is the number of read thresholds. For a three-level cell (TLC), n = 7. For a four-level cell (QLC), n = 15, as... Figure 6As shown. CB content in the database can be characterized offline and initialized in SRAM before the first activation of the hardware architecture (e.g., RDSP-HW). CB characterization can be performed, for example, based on weighted K-means clustering and / or vector quantization methods. In some embodiments, HT-indexes (referred to as specific codebook entries) can be stored in the system memory for each block to point to (or specify) a set of read thresholds that should be used for that block, depending on the stress conditions of the block. In this way, only a small amount of information (HT indexes) for each block can be stored in memory (e.g., firmware memory). In some embodiments, the regfile configuration can point to (or specify) the CB start address and CB size in SRAM.
[0087] Figure 6 An example of a logical view of a codebook database is shown. In this example, a first CB index (CB index 0) may hold or store a normal read threshold (e.g., the default read threshold used in SOL), and CB indices 1 through 3 may hold or store hold-shift table read thresholds that represent read thresholds associated with common stress conditions (e.g., optical DR (data retention)). The read thresholds associated with common stress conditions can be used in the event of a read failure with a normal read threshold. Other CB indices in the codebook (e.g., CB index 4 or higher) may represent read thresholds (e.g., n read thresholds) that have been characterized offline to meet different sets of different stresses.
[0088] Figure 7 This is a diagram illustrating an example row-to-row (R2R) lookup table (LUT) database 700 according to some embodiments. In some embodiments, the database for a hardware architecture according to some embodiments may include a table (referred to as an R2R table or R2RLUT) that includes row-to-row offsets (R2R offsets). In some embodiments, the R2R table may include m2 groups (or rows), and each row may include n offsets from a reference row read threshold (e.g., the read threshold of a reference row). For example, for TLC, n = 7, and for QLC, n = 15. Each entry in the R2R LUT may represent an offset from the reference row threshold. The contents of the R2R-LUT may be characterized offline and initialized in SRAM prior to the first activation of the hardware architecture (e.g., RDSP-HW). In some embodiments, the regfile configuration may point to (or specify) the R2R table start address and / or R2R table size in SRAM. Figure 7This describes an example of a logical view of an R2R LUT database. In this example, the first row of the R2R LUT may hold or store an offset from a reference row read threshold to a first row read threshold. In some embodiments, when a read command for the first row of block X is received, the system (e.g., a controller) may first extract the reference row read threshold from the codebook based on an HT index present in the firmware memory for block X. The system may then perform a linear operation to estimate the first row read threshold based on the reference row read threshold, obtain the corresponding offset from the first row of the R2R LUT, and apply the obtained offset (e.g., the offset from the reference row read threshold) to the (estimated) first row read threshold. For example, the offset from the R2R LUT to the first row may be applied to the reference row threshold to provide an estimated threshold for the first row used to read the first row of block X.
[0089] Figure 8 This is a diagram illustrating an example weight coefficient matrix 800 according to some embodiments. In some embodiments, a database for a hardware architecture according to some embodiments may include K-means search weights (e.g., W(0,0), ..., W(0,14), ..., W(14,14)). In some embodiments, the K-means search operation may find an index from the nearest centroid (e.g., row) in the codebook to a reference row read threshold based on a weighted MSE (mean squared error) metric. In some embodiments, weights (per row, per read threshold) may be used during the K-means search. The coefficient matrix may be used to compute an estimated weight for each reference row. The contents of the coefficient matrix may be initialized in SRAM prior to the first activation of the hardware architecture (e.g., RDSP-HW). For example, in the TLC case, the size of the weight coefficient matrix may be [7Bx7B] = 49 bytes, and the size of the weight coefficient matrix may be [15Bx15B] = 225 bytes in the QLC case. For each K-means search operation, the weight of a particular row may be computed based on the input row threshold and the weight coefficient matrix. In some embodiments, the regfile configuration can point to (or specify) the starting address of the weight coefficient matrix in SRAM and the size of the matrix.
[0090] Figure 9 This is a diagram illustrating an example of DNN parameter placement 900 in a memory according to some embodiments. Figure 9 An example placement of DNN parameters in memory is shown, scattered across four physical SRAM dices 901, 902, 903, and 904. In some embodiments, the database used for the hardware architecture may include DNN parameters (or parameters of any other machine learning model). In some embodiments, DNN parameters may include weights (e.g., W...). 0,0 W 1,0Weights include weights, biases (e.g., B0, B1, etc.), EE (entity embedding) LUT elements (e.g., EE0, EE1, etc.), and / or scaling parameters (e.g., S0, S1, etc.). Weights can be used during MAC (multiplication and accumulation) computation. Bias can be used after the MAC stage is completed. EE can be used to efficiently represent categorical input features in the DNN input layer. Scaling parameters can be used to better utilize the dynamic range of weights / biases / activations during neuron computation. Each network (e.g., a neural network) can have its own set of parameters generated during the offline training process (depending on the network usage). In some embodiments, the regfile configuration can point to (or specify) the starting address of each parameter in SRAM (e.g., EE start pointer 910, weight start pointer 911, bias start pointer 912, scaling start pointer 913) and the size of each parameter. The regfile configuration can also describe the network architecture (e.g., the number of hidden layers, the width per layer, etc.).
[0091] In some embodiments, the database for a hardware architecture, according to some embodiments, may include a regfile configuration set. In conventional architectures, the register file (regfile) is configured by firmware (FW) to define specific uses of the hardware engine. Each CPU typically accesses its own regfile, referred to as the per-CPU regfile. Hardware architectures according to some embodiments (e.g., RDSP-HW architecture) introduce a more efficient approach by using a "regfile configuration set" pre-stored in SRAM.
[0092] In some embodiments, such as during power-up, the regfile configuration set can be prepared and stored (or initialized) offline. When a specific read operation (e.g., an RDSP operation) is invoked or indicated, the system (e.g., CPU or firmware) can configure a pointer to the appropriate regfile configuration set in SRAM in the per-CPU regfile, and the system (e.g., hardware or circuitry) can extract the regfile configuration set from SRAM as a mem-regfile (e.g., a regfile loaded from memory, not the APB interface). The mem-regfile can be loaded before performing the read operation. In this way, a specific engine (e.g., RDSP-HW) within the hardware architecture according to some embodiments can be activated with the corresponding regfile configuration set.
[0093] In some embodiments, the flash system may support simultaneous access by multiple CPUs to their per-CPU regfile, wherein each CPU treats the hardware architecture (e.g., RDSP-HW) as a different virtual machine that includes a dedicated regfile configuration space (per-CPU regfile in memory).
[0094] According to some embodiments, regfile configuration can have the following advantages. First, the system can achieve shorter CPU configuration time. According to some embodiments, firmware-based regfile configuration can introduce latency, especially when compared to hardware processes that quickly fetch data from SRAM. For example, a CPU typically configures hardware units via an Advanced Peripheral Bus (APB), which is designed to interface slower peripherals with the main processor or core in a System-on-Chip (SoC). The overall latency of a read operation (e.g., an RDSP operation) can include three steps: (1) firmware configuration, (2) hardware processing, and (3) firmware read state. To meet system performance requirements (e.g., requirements in metrics such as IOPS or throughput), especially during read stream operations at SOL, minimizing read operation latency can be critical. Using a regfile configuration set can help reduce this overall latency by shortening CPU configuration time. When multiple CPUs are connected to the same regfile via a single APB structure, the configuration of one CPU can prevent other CPUs from accessing their per-CPU regfile. In some embodiments, the system can allow firmware to configure a single pointer as a regfile configuration set, thereby improving efficiency and reducing APB traffic.
[0095] Secondly, area savings can be achieved in the regfile configuration according to some embodiments for at least the following reasons. In conventional setups, where multiple CPUs access hardware for read operations, each CPU can maintain a copied set of regfile configurations in its virtual machine. For example, if a specific set of regfile configurations is required for DNN operations, then that specific set of regfile configurations must be copied for each CPU. On the other hand, in some embodiments of this disclosure, only a single set of regfile configurations can be loaded in the mem-regfile, while all available configurations for that set are stored in SRAM, eliminating the need for copying and thus saving area. This architecture allows multiple CPUs to access a single RDSP-HW unit. Therefore, the architecture also reduces the area because there is a single RDSP-HW unit working with multiple CPUs, rather than a dedicated RDSP-HW unit for each CPU. This architecture can (via APB) configure only pointers to SRAM (instead of configuring the complete configuration), thus minimizing APB workload. The RDSP-HW logic can retrieve a predefined configuration from SRAM based on the pointer and copy the predefined configuration from SRAM to the mem-regfile.
[0096] In summary, in some embodiments of this disclosure, the system can achieve rapid configuration by copying the configuration set from SRAM to a register file in memory (e.g., mem-regfile), which is faster than conventional Advanced Peripheral Bus (APB) configuration, reduces CPU configuration time, and minimizes APB traffic. The system achieves area savings because the configurable architecture is instantiated using only a single mem-regfile, rather than being copied in a per-CPU regfile for each CPU, and because multiple CPUs use a single unit (e.g., an RDSP-HW unit) instead of a dedicated unit for each CPU (e.g., an RDSP-HW unit).
[0097] Figure 10 This is a diagram illustrating an example system environment 1000 of a hardware engine 1020 (e.g., reading a DSP HW architecture or engine) that implements a register file (regfile) configuration set according to some embodiments. Figure 10 As shown, each CPU (e.g., CPU-0, CPU-1, CPU-2, CPU-3) can simultaneously configure its own per-CPU regfile (e.g., RF-0, RF-1, RF-2, RF-3). This configuration may involve setting dedicated registers within each CPU regfile and assigning pointers from the per-CPU regfile (e.g., per-CPU RF 1001) to a specific set of regfile configurations stored in SRAM (e.g., RegFile CFG set 1002). When the hardware selects a specific RDSP operation to perform from a specific per-CPU regfile (e.g., per-CPU RF 1001), the hardware can retrieve the corresponding configuration data from SRAM into a mem-regfile (e.g., mem-regfile 1003). The selected per-CPU regfile, combined with the updated mem-regfile, can form the complete configuration required to perform the RDSP operation.
[0098] Figure 11This is a diagram illustrating an example fully connected (FC) deep neural network (DNN) 1100 according to some embodiments. The example DNN may include an input layer 1101, one or more hidden layers 1102, and / or an output layer 1103. Hidden layers 1102 may include multiple neurons in each layer, such as neurons 1111 and 1112 in layer 0, neurons 1113 in layer 1, neurons 1114 in layer 1, etc. In a hardware architecture according to some embodiments, the flash memory system may include a DNN engine as one of the main hardware engines. The DNN engine can be used to perform inference tasks using the DNN (e.g., DNN 1100). The engine may include a series of processing elements that perform DNN computations in parallel. The main data path may include parallel multiply-accumulate units (MAC) and nonlinear activation functions (ReLU), enabling the DNN engine to perform nonlinear computations quickly and efficiently with low latency and low power consumption compared to conventional firmware / software implementations. In some embodiments, the DNN engine may be highly configurable and may be used for several tasks such as QT, R2R, and other tasks. A DNN engine can compute inference results for a fully connected (FC) network (in the output layer) based on the input (in the input layer) and the network architecture (e.g., network length, width). Network parameters (e.g., weights, biases, scaling parameters) can be configurable and can be stored in SRAM.
[0099] Figure 12 This is a diagram illustrating an example entity embedding scheme 1200 according to some embodiments. In some embodiments, the weights and biases stored in SRAM can be used for different DNNs. For example, different coefficients can be used for QT or R2R, and different DNN architectures can be used. For example, the number of layers and / or the number of neurons per layer can be different for various estimation tasks. The DNN HW engine includes multiple configurable multiplication-accumulation (MAC) modules, and they are used in parallel depending on the network configuration.
[0100] In some embodiments, the DNN engine can be configured for read operations performed in streaming mode, which means that maximum read throughput can be achieved, and the DNN engine can perform operations such as HT-Get and R2R per-read commands within the data path to provide an optimized threshold for per-page reads.
[0101] In some embodiments, the DNN engine may support entity embedding (EE) techniques as described below. EE is a technique used to represent categorical features in a DNN. Categorical features are those features that take a finite set of values, such as row numbers or WL (word line) numbers. Typically, categorical features can be represented using one-hot encoding, which requires a large number of input neurons and can be computationally expensive. Entity embedding (EE) can address this problem by using intermediate layers (called "EE layers," e.g., EE layer 1210) connected to neurons that represent one-hot representations. For efficient hardware implementations, the intermediate layers (e.g., EE layers) of neurons can be computed offline for each value of the one-hot input features in the form of a LUT, and the EE layers (e.g., LUTs) can be stored in SRAM.
[0102] exist Figure 13 and Figure 14 The diagram illustrates the basic computational unit for ReLU neuron computation, which involves multiplying the input set by the corresponding value. Figure 13 This is a diagram illustrating an example of fixed-point calculation of 1300 for neurons in a layer according to some embodiments. The following describes the lth layer (e.g., Figure 11 Fixed-point calculation of the k-th neuron in neuron 1114):
[0103] Hidden layer:
[0104] M(l) and P(l) are scaling factors calculated offline to better utilize the confidence dynamic range.
[0105] Figure 14 This is a diagram illustrating an example architecture of a fixed-point computation scheme 1400 according to some embodiments. To perform this fixed-point computation in high bandwidth (see Equation 5), the fixed-point computation scheme according to some embodiments can use the following bit widths defined according to a quantization strategy to perform arithmetic computations for each neuron in the neural network:
[0106] Q A = 10 [bits] (used for rounding and the result of the limiting activation function (e.g., ReLU); Figure 14 1403 in the middle);
[0107] Q W = 8 [bits] (used for quantization and scaling) 1401 Figure 14 middle);
[0108] Q B = 16 [bits] (used for quantization and scaling) Figure 14 (1402 in the middle).
[0109] QA, QW, and QB indicate the bit widths of activation, weight, and bias, respectively.
[0110] Figure 15 This is a diagram illustrating an example DNN computation scheme 1500 according to some embodiments. In some embodiments, the DNN engine can perform arithmetic computations on each neuron in the (neural) network. In some embodiments, all network parameters (e.g., weights / biases) can be stored in SRAM 1501. Data can be retrieved from the SRAM sequentially (using an aligner 1502) and provided to the DNN engine according to the ordered computation progress. The DNN engine can have registers (e.g., triggers or FFs) in two arrays (e.g., previous layer 1510, current layer 1520) that respectively store the activation values of the previous layer and the activation values of the current layer, such as... Figure 15 As shown.
[0111] In some embodiments, the DNN engine can be based on two layers (e.g., ... Figure 15 The computational processing is performed using the current layer 1520 and the previous layer 1510 shown. Data used for computational processing (e.g., activations, biases, weights, metadata) can be stored in registers (e.g., ...). Figure 15 In some embodiments, the number of registers can be determined based on the maximum layer width. In some embodiments, the DNN engine can perform all computational steps in a pipeline. In some embodiments, the input to the computational processing can include neurons from previous layers and / or related data (e.g., activations, biases, weights, metadata). In some embodiments, the order of computational processing can be determined based on the order of processing layers (e.g., layer-by-layer order) and / or the order of processing neurons within the same layer (e.g., neuron-by-neuron order within each layer). In some embodiments, the DNN engine can perform layer switching such that, in order to move to the next layer of computation, the current layer is switched to the previous layer, and the next layer is switched to the current layer. Figure 15 The registers in the current layer (e.g., FF) and the registers in the previous layer are shown.
[0112] In some embodiments, the DNN engine can perform MAC (multiplication and accumulation) operations or parts thereof in parallel. In some embodiments, the DNN engine can determine the engine bandwidth or bandwidth requirements that the system (e.g., a NAND flash memory device) can achieve, and determine the number of multipliers (as a parallelism factor) based on the engine bandwidth (see [link to documentation]). Figure 14(Multipliers and adders in a network). For example, if the bandwidth requirement is to perform 16 MAC operations in one cycle, the DNN engine can initiate 16 multipliers. In some embodiments, the DNN engine can extract the weights required for each multiplication operation from SRAM. Storing weights in SRAM offers two advantages. First, SRAM provides efficient area and / or power consumption for storing this information compared to storing weights in registers. Second, weights can be read in high bandwidth according to a parallelism factor. For example, a single line in SRAM can store multiple weights required to perform MAC operations according to a parallelism factor. In some embodiments, multiple SRAMs can be implemented to provide as many weights as needed per cycle.
[0113] Figure 16 This is a diagram illustrating an example DNN engine main interface 1600 according to some embodiments. In some embodiments, the DNN engine may receive or obtain network configuration 1601 (e.g., from a register file configured by the CPU). The DNN engine may receive or obtain input features 1602 at the input layer (e.g., from a register file configured by the CPU). The DNN engine may receive or obtain network parameters 1603 from SRAM (e.g., the SRAM may be configured offline by the CPU). The DNN engine may compute output features 1604 with high bandwidth and provide output (e.g., output to a register file that can be read by the CPU).
[0114] Figure 17 This is a diagram illustrating an example R2R engine main interface 1700 (e.g., an R2R-LUT engine) according to some embodiments. In some embodiments, the R2R engine (or R2R-LUT engine) may perform linear R2R transformations based on LUTs in SRAM. In some embodiments, an offset LUT storing a voltage threshold offset 1701 may be stored in SRAM and used for transformations from R2R_IN (e.g., input 1702 of the R2R transformation) to R2R_OUT (e.g., output 1703 of the R2R transformation).
[0115] In some embodiments, the content of the k-th row of the LUT can be an offset from the reference row (for each threshold). In this case, the R2R engine can calculate R2R_Out as follows:
[0116] R2R_OUT=R2R IN ±R2R LUT[Target_Row] (Equation 6)
[0117] In some embodiments, R2R input 1702 may include one of the following: (1) a reference row read threshold from the codebook (according to the HT-index) or (2) a target row read threshold from the regfile. For R2R configuration 1704, the flash system (e.g., CPU or firmware) may configure the regfile to include pointers to the start address of the R2R-LUT in SRAM and the size of the R2R-LUT. The regfile may include an R2R transformation direction bit indicating (1) a direction from the target row to the reference row (addition), or (2) a direction to the target row from the reference row (subtraction). The addition or subtraction may be related to the estimation process using R2R. For direction (1), the system may use the target-to-threshold to estimate the reference row threshold. For direction (2), the system may use the reference row threshold to estimate the target row threshold. These are two opposite R2R directions.
[0118] Figure 18 This is a diagram illustrating an example instance of calculating building block 1800 per threshold according to some embodiments. Figure 19 This is a diagram illustrating an example of an iterative K-means search (or K-means search engine) 1900 that calculates three instances of building blocks based on each threshold according to some embodiments.
[0119] In some embodiments, the flash memory system may include a K-means search engine 1900. The K-means search engine or operation can find the index of the nearest centroid in the codebook 1901 to a reference threshold based on a weighted MSE (e.g., MSE 1801 multiplied by 1802 multiplied by the weight). The K-means search engine can calculate the codebook index as follows: In the first step, the K-means search engine can calculate weights based on the target row. In the second step, for each K-means search operation, the K-means search engine can calculate the weight of a specific row based on the input row threshold and the weight coefficient matrix. For example, the K-means weight can be calculated as follows:
[0120]
[0121] Where N = 7 (TLC); N = 15 (QLC).
[0122] In the third step, the K-means search engine can find the HT index (e.g., the line number or CB index in the codebook) that represents the read threshold closest to (in terms of added error) the reference line's read threshold by searching or scanning across all entries in the codebook and comparing the read threshold to the reference line. The calculation can be performed as follows:
[0123]
[0124] Where N = 7 (TLC); N = 15 (QLC).
[0125] In some embodiments, to accelerate the K-means search operation, the K-means search engine can use per-threshold building block 1800 (for each index i) to calculate the weighted distance to the reference row threshold (see [link to K-means search engine]). Figure 19 ).like Figure 19 As shown, multiple instantiations (or instances) of this per-threshold computation building block (e.g., 3 blocks 1800-1, 1800-2, 1800-3) can each compute weighted distances from different thresholds in parallel. The increased instantiation of the per-threshold computation building block of the K-means search engine has the lower latency of the K-means search algorithm (a trade-off between gate counting and K-means search latency), which the K-means search engine can implement. In some embodiments, the results of all weighted distances can be accumulated until all thresholds are computed (e.g., 7 thresholds for a single row or CB entry, 15 thresholds for the TLC, and 1902 thresholds for the QLC). Figure 19 As shown, the best candidate 1903 can be a variable initialized to MAX_VAL and representing the value of the minimum weighted distance sum of all thresholds. Once the weighted distance sum of all thresholds 1904 has been calculated, the weighted distance sum of all thresholds 1904 can be compared with the best candidate 1903 1905, and the weighted distance sum of all thresholds and CB entries (which is the HT index) can only be saved (or updated) if the current value of the weighted distance sum of all thresholds is less than the best candidate value.
[0126] In some embodiments, K-means search can traverse all clusters in the codebook. In some embodiments, the K-means search engine can perform an ArgMin search, which calculates the weighted Euclidean distance for each cluster in the codebook. The number of operations per cluster can be 7 (for TLC) or 15 (for QLC). The throughput of K-means search can be 3 operations / cycle (see [link to K-means search]). Figure 19 In this case, the total delay can be calculated as follows:
[0127] Total latency~O(NumOfRowsInCB*n / NumOfBuldingBlocks) (Equation 9),
[0128] Where n = 7 (TLC) or 15 (QLC).
[0129] In addition, weights can be calculated based on the input reference rows.
[0130] Figure 20This is a diagram illustrating an example of a K-means search engine main interface 2000 according to some embodiments. In some embodiments, input 2001 to the K-means search engine may include: (1) a reference line (e.g., per CPU regfile) provided by the regfile interface, and / or (2) a target line (e.g., per CPU regfile) provided by the regfile interface. In some embodiments, the system (e.g., CPU or firmware) may configure pointers 2002 in the regfile, (1) pointers to the start address of the codebook in the SRA and / or the size of the codebook in the SRAM, and / or (2) pointers to the start address of the weight coefficient matrix in the SRAM and / or the size of the weight coefficient matrix in the SRAM. In some embodiments, in response to obtaining codebook content 2003 from the SRAM, the K-means search engine may output HT indexes 2004 (e.g., codebook entries).
[0131] Figure 21 This is a diagram illustrating an example system environment 2100 of a top-level architecture for implementing read operation hardware or read hardware engine 2120 (e.g., read DSP hardware engine / architecture) according to some embodiments. Figure 22 This is a diagram illustrating an example scheme 2200 for loading a mem-regfile according to some embodiments.
[0132] Figure 21 The top-level RDSP-HW with algorithms for TLC and QLC is described. The read hardware engine 2120 may include SRAM 2150 storing RegFile CFG sets 2151 and RegFile CFG sets 2159, a codebook for TLC 2152, a codebook for QLC 2153, an R2R LUT for TLC 2154, an R2R LUT for QLC 2155, a first set of DNN parameters 2956, and / or a first set of DNN parameters 2157. The read hardware engine 2120 may include a K-means search engine 2125 (e.g., K-means search engine 2000), an R2R engine 2126 (e.g., R2R engine 1700), and / or a DNN engine 2127 (e.g., DNN engine 1600). When performing tasks associated with TLC operations, the CPU may configure appropriate configurations and appropriate pointers to register profile configuration sets. Figure 21As shown, each CPU 2110-0, 2110-1, 2110-2, and 2110-3 can simultaneously configure its own per-CPU regfile 2122-0, 2122-1, 2122-2, and 2122-3 via APB 2112. In other words, each CPU can independently and simultaneously configure its own per-CPU regfile. This configuration may involve setting dedicated registers within each CPU regfile and assigning pointers from the per-CPU regfile (e.g., per-CPU regfile 2122-0) to a specific regfile configuration set (e.g., RegFile CFG set 2152) stored in SRAM 2150. When the arbitration and management control (system) 2230 selects a specific RDSP operation to execute from a specific per-CPU regfile (e.g., per-CPU RF 2122-0), the hardware can extract the corresponding configuration data (e.g., regfile CFG set 2151) from SRAM 2150 into a mem-regfile (e.g., mem-regfile 2123) in memory (e.g., DRAM 3310). The mem-regfile may include one or more registers. The selected per-CPU regfile 2122-0, combined with the updated mem-regfile 2123, can form the complete configuration required to execute the RDSP operation.
[0133] The top layer of the hardware architecture according to some embodiments (e.g., RDSP HW) may be composed of engines (e.g., K-means search engine 2125, R2R-LUT engine 2126, DNN engine 2127), SRAM 2150, per-CPU regfiles 2122-0, 2122-1, 2122-2, 2122-3, mem-regfile 2123, and arbitration and management control 2230. Each engine may execute a read operation algorithm (e.g., an RDSP algorithm). Each engine may be connected to memory (e.g., SRAM 2150) to obtain relevant parameters during RDSP operation. Each engine may obtain a reg-file configuration (e.g., regfile CFG set 2151) to define the precise algorithm.
[0134] In some embodiments, SRAM 2150 may contain database parameters. The SRAM may include multiple instantiations representing database parameters for different algorithms. The SRAM may contain a set of regfile configurations (e.g., regfile CFG set 2151). The SRAM may include multiple instantiations representing regfile configuration sets for different algorithms. The SRAM may be constructed from multiple physical blocks (e.g., slits 501, 502, 503, 504) and / or arbitration logic (e.g., memory arbitrator 510) to provide high bandwidth based on engine processing bandwidth.
[0135] In some embodiments, each CPU (e.g., CPUs 2110-0, 2110-1, 2110-2, 2110-3) can independently and simultaneously configure its own per-CPU regfile. In some embodiments, once a task is selected from a particular CPU (e.g., via arbitration and management control 2230), before performing RDSP operations, a mem-regfile (e.g., mem-regfile 2123) can be loaded from SRAM according to a pointer in the per-CPU regfile (e.g., per-CPU regfile 2122-0).
[0136] In some embodiments, arbitration and management control (e.g., arbitration and management control 2230) can be performed. A CPU (e.g., CPU 2110-0) can configure and / or activate a specific task via its per-CPU regfile (e.g., per-CPU regfile 2122-0). In some embodiments, an activation bit in the regfile (not shown) can trigger hardware and notify the task configuration that it is ready to execute. The RDSP-HW (e.g., controller 3320, arbitration and management control 2230) can select ready tasks for execution based on one of the following options. As a first option, a First-Come-First-Served (FCFS) strategy can be used to schedule tasks in the order they arrive. For example, the system can configure short tasks with higher priority to improve system performance. As a second option, tasks can be scheduled based on arrival order and / or task priority (as configured in a configuration file). For example, the system can configure a specific task with higher priority to be implemented before other tasks with lower priority.
[0137] The top layer of the hardware architecture according to some embodiments (e.g., RDSP HW) may consist of several key components, including engines (e.g., K-means search engine 2125, RR LUT engine 2126, DNN engine 2127), SRAM 2150, per-CPU regfile, mem-regfile, and / or arbitration and management control (e.g., arbitration and management control 2230). Each component may play an important role in the operation of the RDSP-HW.
[0138] In some embodiments, each engine may be responsible for executing a specific RDSP algorithm. Engines may be connected to memory (e.g., SRAM 2150) to retrieve necessary parameters during RDSP operation. Engines may receive a regfile configuration defining the exact algorithm to be executed.
[0139] In some embodiments, the SRAM (e.g., SRAM 2150) may store database parameters (and may include multiple instantiations representing different algorithms used for the same engine). The SRAM may store a regfile configuration set (e.g., regfileCFG set 2151). The SRAM may store a regfile configuration set (and may include multiple instantiations representing different algorithms used for the same engine).
[0140] In some embodiments, the SRAM may be constructed from multiple physical blocks (e.g., blocks 501, 502, 503, 504) and accompanying arbitration logic (e.g., memory arbitrator 510) to provide high bandwidth aligned with the processing power of the engine.
[0141] In some embodiments, each CPU (e.g., CPU 2110-0) may have the ability to independently and simultaneously configure its own per-CPU regfile (e.g., per-CPU regfile 2122-0). Each CPU can initiate and activate a specific task through its per-CPU regfile. An activation bit (not shown) within the per-CPU regfile can trigger hardware to indicate that the task configuration is ready for execution.
[0142] In some embodiments, the CPU can monitor the completion of its tasks by polling its per-CPU regfile. Alternatively, an interrupt signal can be used to notify the CPU when a task completes. Polling may be preferred for tasks with short latency because it avoids potential performance degradation due to context switching overhead. Once a task is complete, the CPU can retrieve the task result from the status register within the per-CPU regfile.
[0143] In some embodiments, the mem-regfile (e.g., mem-regfile 2123) may be a single configuration structure loaded from SRAM (e.g., SRAM 2150) when a task from a particular CPU is selected. This loading may occur before RDSP operations are performed and is guided by pointers in the per-CPU regfile. The flash system may perform arbitration and management control (e.g., arbitration and management control 2230). Each CPU can configure and activate a particular task through its per-CPU regfile. Activation bits in the regfile can trigger hardware to signal that the task configuration is ready for execution.
[0144] In some embodiments, the RDSP-HW (e.g., read hardware engine 2120, arbitration and management control 2230) may select ready tasks for execution based on the following options: (1) First-Come-First-Served (FCFS) – tasks may be scheduled in the order they arrive; for example, the system may assign higher priority to shorter tasks to improve overall system performance; and (2) Priority-based scheduling – tasks may be scheduled based on both their arrival order and priority configured in the regfile of each CPU; for example, the system may assign higher priority to specific tasks to ensure that they are executed before other lower priority tasks.
[0145] like Figure 22As shown, SRAM 2250 may include a first registration file configuration set 2251, a second registration file configuration set 2252, a first DNN(1) parameter 2254, and / or a second DNN(2) parameter 2256. A first CPU (CPU(1)) (e.g., CPU 2110-0) may configure a first DNN task (using DNN(1) parameter 2254) in its per-CPU regfile, and a second CPU (CPU(2)) (e.g., CPU 2110-0) may configure a second DNN task (using DNN(2) parameter 2256) in its per-CPU regfile. Assuming that a task from CPU(1) is selected for execution, a pointer in its per-CPU regfile 2220 may direct RDSP-HW to RegFileSet(1) 2251. RDSP-HW may then copy the contents from RegFileSet(1) 2251 to MemRegFile 2220. As a result, the MemRegFile register can be configured with DNN pointers 2221, 2222, and 2223 for weights, biases, and scaling parameters, specifically pointing to DNN(1) parameter 2254. Once the DNN engine (e.g., DNN engine 2127) is activated, the DNN engine can access DNN(1) parameter 2254 according to the configuration stored in MemRegFile. Similarly, when a task from the CPU(2) is selected for execution, RegFileSet(2) 2252 can be copied to mem-regfile 2220, etc.
[0146] Figure 23 This is a diagram illustrating an example scheme 2300 of output mapping logic 2310 according to some embodiments. In some embodiments, after a task is completed, the CPU (e.g., CPU 2110-0) can retrieve the task result from the status register within the per-CPU regfile. This mapping logic 2310 can be particularly useful for minimizing APB traffic (e.g., APB 2112) and reducing overall RDSP operation latency (especially during R2R operations in SOL scenarios). Figure 23The R2R operation and threshold mapping are illustrated. In some embodiments, the R2R operation, whether performed via a DNN engine or an R-LUT engine, can generate read thresholds (e.g., TLC: T0-T6, QLC: T0-T14). Typically, each read threshold can be mapped to a separate status register. For example, each threshold from CB 2301, LUT 2302, and / or DNN 2303 can be mapped to a corresponding status register in perCPU regfile 2320 using mapping logic 2310. However, in some systems, depending on the bit width, multiple read thresholds can be packed into a single status register. For example, if each read threshold is 8 bits wide and the status register is 32 bits wide, and the mapping logic 2310 defines the mapping layers (e.g., multiplexers) 2312-0, 2312-1, 2312-2, 2312-3, then the R2R output can be mapped as follows: Status register 0 (2322-0) uses mapping 2312-0 to hold RdThrest[3] to RdThrest[0]; Status register 1 (2322-1) uses mapping 2312-1 to hold RdThrest[7] to RdThrest[4]; Status register 2 (2322-2) uses mapping 2312-2 to hold RdThrest
[11] to RdThrest[8]; Status register 3 (2322-3) uses mapping 2312-3 to hold RdThrest
[15] to RdThrest
[12] .
[0147] In some embodiments, the mapping can be performed as follows. Thresholds can be generated in ascending order. For example, DNN and R2R engines can generate read thresholds in a predefined ascending order (e.g., RdThresholdT1 to T15). The system can include a mapping layer that is configurable to transform these thresholds before writing them to the output registers. This provides flexibility in the output format without affecting the operation of the internal engine. The system can perform output register multiplexing, allowing the mapping layer to allow specific read thresholds to be multiplexed to specific output registers. The system can perform regfile configuration, allowing the 16 first outputs from the [CB / LUT / DNN] engine to be mapped to corresponding 16 status register bits.
[0148] This mapping process can offer the following benefits. First, mapping enables selective reading, allowing the CPU to map relevant outputs to a single or a few status registers, thereby reducing the number of APB reads required. Second, mapping improves efficiency, eliminating the need for firmware to process all read thresholds from all status registers, selecting only the required thresholds and rearranging them in the necessary order.
[0149] For example, each page of a QLC device, DN-R2R (15 thresholds, a DNN with 15 outputs), can be configured as follows: (1) lower page: read thresholds T0, T2, T5, T11; (2) middle page: read thresholds T1, T7, T10, T12; (3) upper page: read thresholds T3, T9, T13; and (4) top page: read thresholds T4, T6, T8, T14. In some embodiments, the effective mapping can be configured as follows: (1) the lower page can map read thresholds T0, T2, T5, T11 to status register 0 (2322-0); (2) the middle page can map read thresholds T1, T7, T10, T12 to status register 0; (3) the upper page can map read thresholds T3, T9, T13 to status register 0; and (4) the top page can map read thresholds T4, T6, T8, T14 to status register 0.
[0150] Figure 24 This is a diagram illustrating an example of limiting and DC offset adjustment 2400 according to some embodiments. The flash memory system may perform additional operations on the read threshold (e.g., limiting 2410 and / or DC offset adjustment 2420). In some embodiments, R2R operations (whether LUT-based or DNN-based) can generate an estimated read threshold for the target row. In some cases, additional operations may be performed on these thresholds. Performing such additional operations in the RDSP-HW can reduce the overall RDSP operation latency. On the other hand, if this additional operation is implemented in the firmware, then performing this additional operation can have an impact on the controller read performance.
[0151] In some embodiments, the system may perform clipping 2410 as follows. For example, in certain edge cases, the R2R algorithm may generate results outside the expected range, particularly when dealing with stresses not considered during offline training. To ensure valid results, a clipping operation can be applied to the estimated threshold. This ensures that the calculated read threshold remains within predetermined limits.
[0152] In firmware-based implementations, limiting each threshold can introduce latency. For example, in a QLC device, there are up to four read thresholds per page, each of which may need to be checked and clipped within defined upper and lower limits, requiring multiple CPU operations. However, using the RDSP-HW architecture, all limiting checks can be performed simultaneously, completing the process in a single clock cycle. The following formula describes read threshold limiting:
[0153]
[0154] When Thr[k]_clipped is the read threshold result from the R2R engine, max_clip_thr[k] and min_clip_thr[k] are parameters configured offline.
[0155] In some embodiments, the system can adjust the DC offset 2420, which can be added to the read threshold, in real time. RDSP algorithm characterization can be performed offline based on a VT scan database already generated from a representative flash device. For various reasons, the VT scan database may not precisely match the actual NAND device in real time. This gap can be addressed by applying a fixed offset to the threshold used on the actual flash device to improve accuracy. However, in firmware-based implementations, such an operation might require adding operations per threshold, whereas for hardware architectures (e.g., RDSP-HW architecture), all DC offset operations can be performed simultaneously within a single clock cycle.
[0156] Another example of an operation that can be applied in real time is adding a DC offset to a read threshold. RDSP algorithms can be characterized offline based on a Vt scan derived from a representative flash device. However, due to various reasons (e.g., variations in manufacturing), the actual Vt distribution in a NAND device may deviate from the original Vt scan database used for algorithm development. This gap can be mitigated by applying a fixed DC offset to the read threshold when working with a real-time flash device. In firmware-based implementations, this operation would require adding a fixed offset individually to each threshold, introducing latency, whereas for the RDSP-HW architecture, all DC offset adjustments can be applied to the thresholds simultaneously within a single clock cycle.
[0157] In some embodiments, the system may provide a hardware architecture that can operate with multiple CPUs (e.g., RDSP-HW). In typical memory controllers, random read operations can pose significant performance challenges. Unlike sequential read commands, which typically involve large blocks of data (e.g., 16KB), each random read command may be associated with small blocks of data (e.g., 4KB) scattered across different dies / blocks / pages. For each 4KB of data sent to the host, the controller needs to perform all associated management tasks (e.g., command parsing, logical-to-physical address translation). This can increase latency and limit throughput, especially during random reads.
[0158] To improve performance, a common solution is to use multiple CPUs to distribute the workload across different read commands. However, the R2R operation for estimating the optimal read command for a particular row is often too time-sensitive to be handled efficiently in firmware (FW), especially during start-of-life (SOL) when system performance is critical. This is because hardware architectures, such as RDSP-HW blocks, excel at providing low-latency R2R operations to optimize read performance, according to some embodiments. A simple approach would be to attach the RDSP-HW block to each CPU, but this significantly increases gate count, memory footprint, and power consumption, as attaching the RDSP-HW block may require copying it for each CPU. This copying can lead to undesirable inefficiencies in complex system architectures.
[0159] The hardware architecture according to some embodiments can provide an optimized solution by allowing multiple CPUs to access a single RDSP-HW block simultaneously. Each CPU can interact with the RDSP-HW block as if each CPU were a different virtual machine, due to the dedicated register file configuration space (e.g., a per-CPU regfile). This design allows RDSP-HW to manage and synchronize tasks originating from various CPUs, eliminating the need to copy the RDSP-HW block to each CPU. In this way, RDSP-HW throughput can be improved.
[0160] In some embodiments, the flow configuration allows RDSP-HW to maximize throughput by operating in a pipeline: while one task is configured (mem-regfile configuration), another task can be executed on the engine (task execution). Simultaneously, the CPU is idle to interact with its corresponding per-CPU regfile for configuration or status checks, thus allowing continuous task management in the background.
[0161] In some embodiments, a pipelined workflow can be defined such that RDSP-HW operates in two main pipelined steps: pipelined step (1) and pipelined step (2). In pipelined step (1), the system can configure the mem-regfile such that RDSP-HW selects a task activated by the CPU, loads the mem-regfile from SRAM (based on a per-CPU regfile pointer), and samples the configuration for the next stage when it is available. In pipelined step (2), the system can perform task execution such that RDSP-HW activates the corresponding engine based on the sampled configuration, and updates the status register with the engine once the task is complete.
[0162] The example system workflow is as follows. In step 1, each CPU can independently configure and activate its per-CPU regfile for a specific task. This can be done simultaneously by all CPUs. In step 2, RDSP-HW can select a task according to the scheduling policy, load the required configuration into the mem-regfile, and begin executing the task on the engine. In step 3, while the engine is executing a task, other CPUs can continue to configure their per-CPU regfiles in parallel or check for task completion.
[0163] Figure 25 This is a diagram illustrating an example system 2500 of pipeline hardware utilizing multiple CPUs according to some embodiments. Figure 25 A system with three CPUs 2510-0, 2510-1, and 2510-2 is illustrated. These three CPUs work with a single RDSP-HW block (e.g., HW engine 2530) and a mem-regfile (e.g., mem-regfile 2520). In some embodiments, system 2500 can perform back-to-back task execution. For example, task execution on the RDSP engine can occur sequentially with minimal latency. In some embodiments, system 2500 can perform background configuration and status checks. For example, CPU configuration (e.g., configuration 2501) and status read operations (e.g., Done Rd status 2502) can be performed in parallel, independent of engine task execution. In some embodiments, the system can achieve optimal CPU utilization. For example, CPU latency (e.g., as...) can be minimized. Figure 25 The time from task configuration T1 to completion T2 (as shown) can be effectively used for other firmware tasks, thereby enhancing overall system performance. In this way, the hardware architecture according to some embodiments can significantly improve system efficiency, minimize latency, and optimize area and power usage by allowing multiple CPUs 2510-0, 2510-1, 2510-2 to share the same RDSP-HW block 2530 instead of copying the RDSP-HW block for each CPU. Figure 25 As shown, in some embodiments, each CPU can independently and simultaneously execute its configuration (e.g., configuration 2551 by CPU-1 and configuration 2552 by CPU-0), while sequentially accessing and / or loading mem-regfile (e.g., mem-regfile loading 2561 and mem-regfile loading 2562) and sequentially executing tasks using the same hardware engine (e.g., RDSP-HW block) (e.g., task execution 2571 and task execution 2572).
[0164] In some embodiments, the hardware engine according to some embodiments may implement algorithms used in the read process and read retry process, as described in the following sections.
[0165] In some embodiments, the engine, according to some implementations, can perform HT-GET operations based on R2R operations. In some embodiments, on each read command, the HT-Get operation can use an HT index to determine whether the reference row threshold is a default threshold, a retry fixed threshold read, or even a post-QT threshold. In some embodiments, depending on the read command, the reference row threshold in the target block can be extracted during the HT-Get operation, and then the system (e.g., firmware) can perform an R2R operation to calculate the read threshold of the target row using RDSP-HW, and the target row threshold can be provided to the NAND read command in real time. During SOL, system read performance may require performing R2R with a very short latency to avoid impacting system read performance.
[0166] In some embodiments, the engine, according to some embodiments, can perform HT-SET operations based on QT, R2R (target-to-reference), and / or K-means search. In some embodiments, in the event of HB decoding failure (normal read and all shift table reads retried), the system can apply Fast Threshold Tracking (QT) to perform threshold tracking and estimate the optimal read threshold. In some embodiments, QT can perform several simulated reads with a fixed threshold, from which a histogram is computed. The histogram can be used by an estimator to estimate the current threshold. The estimator can be a linear estimator (using a DNN with zero hidden layers) or a DNN-based estimator. The current estimated threshold can be configured as a NAND for retry reads and HB decoding.
[0167] In some embodiments, for future reads, the currently estimated threshold can be transferred to a reference row threshold via an R2R operation (e.g., a LUT-based R2R operation or a DNN-based R2R operation), and the reference row threshold can be used to update the HT table. Performing an HT-Set compresses the threshold into an index pointer HTindex of the HT table. The HTindex can point to the HT threshold closest to the reference row threshold associated with the estimated read threshold and can be used for subsequent reads from the same block. The process of finding the HTindex can be performed via an exhaustive search using a K-means search engine. Note that without the RDSP-HW engine according to some embodiments, the HT index search can be performed in a non-exhaustive method (such as using a binary search tree) by taking into account the search latency.
[0168] In some embodiments, the engine according to some embodiments may execute a read process as follows. An HT-GET stream may be initiated by a controller (e.g., firmware) for each read command. In this HT-GET stream, the controller may provide an HT-index associated with the target block and a target row number. In return, the HT-GET stream may generate or return an estimated read threshold for the specified target row.
[0169] Figures 26 to 32 The diagram illustrates several different stream uses with hardware blocks (e.g., RDSP-HW (Read Digital Signal Processor Hardware) blocks). In each diagram, active inputs, engines, and / or outputs are highlighted in bold and thick lines.
[0170] Figure 26 This is a diagram illustrating an example hardware implementation of an HT-GET-DNN for a first-stage read operation according to some embodiments. Figure 26 The general architecture of hardware block (or hardware engine) 2600 is described. The hardware block may include an engine (e.g., one or more circuits or processors) for DNN 2610, R2R (estimator) 2620, or K-means search 2630. In some embodiments, such hardware block may be replaced or combined with software, firmware, or a combination thereof. The hardware block may also include a database (e.g., one or more memories or storage devices 2640, 2650) for codebook and / or R2R estimation, which may be computed offline and can be initialized once upon power-up. In some embodiments, the inputs to the hardware block may include (1) input features 2601, (2) target rows 2602, and / or (3) CB (codebook) index 2603 for use by DNN 2610 and / or R2R estimator 2620. Input features may include a threshold -In 2605, which may be used as input to the DNN 2610, the R2R estimator 2620, or the K-means search 2630. Input features may include additional inputs 2606, such as a set of rows, a loop range, temperature at programming and / or reading, etc. In some embodiments, the output of the hardware block may include (1) an estimated reading threshold 2607 and / or (2) a CB index 2608 (e.g., a CB index as the output of the K-means search).
[0171] Figure 26A flow or hardware block for implementing or activating R2R-DNN operations (or R2R-DNN engine 2600) for first-stage reading is illustrated. The inputs to the hardware block may include (in a per-CPU regfile) a CB index 2602 and / or a target row 2604. The outputs of the hardware block may include (in a per-CPU regfile) a reading threshold 2651 for the target row. In some embodiments, the DNN's input layer does not include the reading threshold as an input feature, and instead, a constant input reading threshold 2601 may be embodied in or included in other network parameters. In some embodiments, the DNN's input layer may include additional parameters 2603 arriving from the per-CPU regfile (e.g., loop count, row set, programming and / or temperature at read time, etc.). The DNN 2611 may compute the reading threshold 2611 for the target row 2604.
[0172] Figure 27 This is a diagram illustrating an example hardware implementation of an HT-GET-DNN using an HT-codebook (CB) index with an R2RDNN for target row threshold estimation, according to some embodiments. Figure 27 A flow or hardware block for implementing or activating HT-GET-DNN operations (or HT-GET-DNN engine 2700) is illustrated. The inputs to the hardware block may include (in a per-CPU regfile) a CB index 2701 and / or a target row 2702. The outputs of the hardware block may include (in a per-CPU regfile) a read threshold 2751 for the target row 2702. In some embodiments, the read threshold associated with a reference row can be read from codebook 2640 based on the CB index 2701. In some embodiments, the input layer of the DNN may include a reference row read threshold, and optionally additional parameters 2712 (e.g., loop count, row set (row collection), programming and / or temperature at read time, etc.). The DNN 2610 may calculate the read threshold 2713 for the target row 2702.
[0173] Figure 28 This is a diagram illustrating an example hardware implementation of an HT-GET-LUT with an HT-CB index having an R2R lookup table (LUT) for target row threshold estimation according to some embodiments. Figure 28A flow or hardware block for implementing or activating an R2R LUT-based operation (or R2R-LUT engine 2800) is illustrated. The inputs to the hardware block may include (in a per-CPU regfile) a CB index 2801 and / or a target line 2802. In some embodiments, a reference line may be extracted from codebook 2640. The outputs of the hardware block may include (in a per-CPU regfile) a read threshold 2851 for the target line 2602. In some embodiments, a read threshold associated with the reference line (e.g., a reference line read threshold) may be read from codebook 2640 based on the CB index 2801. In some embodiments, an offset from the reference line to the target line may be read or obtained from R2R estimator 2620 based on the target line. In some embodiments, an R2R transformation may be performed based on the reference line read threshold and the offset.
[0174] According to some embodiments, the engine can perform a read retry process as follows. In some embodiments, the read and retry process can be performed after the HB fails to decode all shift indices. In this case, a read threshold for the failed page can be estimated (e.g., by performing QT) and used to read the failed page. Alternatively, an HT_set operation can be performed, where the system can transform the estimated threshold of the target row into a common reference row. The system can then compress the common reference row by assigning the closest HT table threshold and storing the corresponding HTIndex in the HT table.
[0175] Figure 29 This is a diagram illustrating an example hardware implementation for general DNN operations according to some embodiments. Figure 29 A flow or hardware block for implementing or activating a general DNN operation / engine 2900 (e.g., a DNN operation / engine that can be used for QT-DNN operations) is shown. Various DNN operations can be implemented based on different DNN parameters. The inputs to the hardware block may include (in per CPU regfile) input features 2901, network architecture, and / or network parameters. The outputs of the hardware block may include (in per CPU regfile) DNN outputs 2951. In some embodiments, the hardware block may use inputs to perform or execute QT-DNN operations (or a QT-DNN engine), said inputs including (in per CPU regfile) a QT Histogram 2902 and additional inputs 2903, such as a set of rows (row set), loop range (optional), and / or temperature during programming and / or reading. Using the inputs, the QT-DNN engine may output (in per CPU regfile) a QT read threshold 2911.
[0176] Figure 30 This is a diagram illustrating an example hardware implementation of R2R target row to reference row threshold estimation using a LUT according to some embodiments.Figure 30 The diagram illustrates a flow or hardware block for implementing or activating a target row to reference row operation (or target row to reference row engine 3000). In some embodiments, Figure 30 The process shown can activate LUT engines 2620 and 2650. Inputs to the hardware block may include (in a per-CPU regfile) a target row threshold 3001 and / or a target row 3002. Outputs to the hardware block may include (in a per-CPU regfile) a reference row threshold 3051. In some embodiments, the offset from the target row to the reference row can be read or obtained from the R2R estimator 2620 based on the target row. In some embodiments, the R2R transformation can be performed based on the target row threshold and the offset.
[0177] Figure 31 This is a diagram illustrating an example hardware implementation of R2R reference row to target row threshold estimation using a LUT according to some embodiments. Figure 31 This illustrates a flow or hardware block that implements or activates a reference line for a target line operation (or reference line to target line engine 3100). The inputs to the hardware block may include (in per CPU regfile) a reference line threshold 3101, a reference line, and / or a target line 3102. The outputs of the hardware block may include (in per CPU regfile) a read threshold 3151 for the target line 3102.
[0178] Figure 32 This is a diagram illustrating an example hardware implementation of an HT-Set that uses K-means search to compute the CB index given an input threshold, according to some embodiments. Figure 32 A flow or hardware block for implementing or activating a K-means search operation (or K-means search engine 3200) is illustrated. The input to the hardware block may include a reference row threshold 3201 (in per CPU regfile). The output of the hardware block may include a CB index 3251 (in the per CPU regfile). In some embodiments, the K-means engine 2630 may compare the reference row threshold with all clusters in the codebook 2640 and find the CB index 3251 associated with the best matching centroid entry.
[0179] Figure 33 This is a block diagram illustrating an example flash memory system based on some arrangement.
[0180] refer to Figure 33 The flash memory system 3300 may include a computing device 20 and a solid-state drive (SSD) 10, which is a storage device and can be used as the main memory of an information processing device (e.g., a host computer). The SSD 10 may be incorporated into the information processing device or may be connected to the information processing device via a cable or network.
[0181] The computing device 20 may be an information processing device (computing device). In some arrangements, the computing device 20, configured to process or manipulate data for training and executing a neural network (e.g., DNN 300) and the training data, may be collected from multiple SSDs by multiple computing devices. The data collected from multiple SSDs may be recorded and processed / manipulated by different computing devices, which are not necessarily connected to any SSD and perform training based on the collected data. The computing device 20 includes a processor 21 and / or a database system 26. The database system 26 may store read thresholds including the training set or training results.
[0182] SSD 10 includes, for example, a controller 3320 and flash memory 3380 as non-volatile memory (e.g., NAND flash memory). SSD 10 may include random access memory (RAM), which is volatile memory, such as DRAM (Dynamic Random Access Memory) 3310 and / or SRAM (Static Random Access Memory) 3315. The RAM may have, for example, a read buffer that serves as a buffer for temporarily storing data read from flash memory 3380, a write buffer that serves as a buffer for temporarily storing data written to flash memory 3380, and a buffer for garbage collection. In some arrangements, controller 3320 may include DRAM or SRAM.
[0183] In some arrangements, flash memory 3380 may include an array of memory cells comprising multiple flash memory blocks (e.g., NAND blocks) 3382-1 to 3382-m. Each of blocks 3382-1 to 3382-m can serve as an erase unit. Each of blocks 3382-1 to 3382-m contains multiple physical pages. In some arrangements, in flash memory 3380, data reads and writes are performed on a page-by-page basis, and data erasure is performed on a block-by-block basis.
[0184] In some arrangements, controller 3320 may be a memory controller configured to control flash memory 3380. Controller 3320 includes, for example, one or more processors (e.g., CPUs) 3326, flash memory interface 3328, memory interface 3322, and network interface 3324, all of which may be interconnected via bus 3328. Memory interface 3322 may include a DRAM controller configured to control access to DRAM 3310 and an SRAM controller configured to control access to SRAM 3315. Flash memory interface 3328 may act as flash memory control circuitry (e.g., NAND control circuitry) configured to control flash memory 3380 (e.g., NAND-type flash memory). Network interface 3324 may be used as circuitry for receiving and transmitting various types of data from computing device 20. Data may include multiple sets of read thresholds or other data collected from flash memory 3380 or multiple SSDs used for training neural networks (e.g., DNN 300).
[0185] The controller 3320 may include a read circuit 3330, a programming circuit (e.g., a programming DSP) 3340, and / or a programming parameter adapter 3350. For example... Figure 33 As shown, adapter 3350 can adjust programming parameters 3344 used by programming circuitry 3340 as described above. In this example, adapter 3350 may include a program / erase (P / E) cycle counter 3352. Although shown separately for ease of illustration, some or all of adapter 3350 may be incorporated into programming circuitry 3340. In some arrangements, read circuitry 230 may include ECC decoder 3332 and read hardware engine 3334 (e.g., [TBD] system 500, system 600, RdDSP HW engine 820, DNN-based R2R estimator 900, hardware engine 1000). In some arrangements, programming circuitry 3340 may include ECC encoder 3342. The arrangement of memory controller 3320 may include additional or fewer components, such as... Figure 33 The components shown in the document.
[0186] In some embodiments, a flash memory system (e.g., SSD 10) may include non-volatile memory (e.g., flash memory 3380), a controller (e.g., controller 3320) for performing operations on the non-volatile memory, and a plurality of processors (e.g., processors 2510-0, 2510-1, 2510-2) including a first processor (e.g., processor 2510-0) and a second processor (e.g., processor 2510-1). The non-volatile memory (e.g., flash memory 3380) may include one or more blocks (e.g., flash memory blocks 3382-1, 3382-2, ..., 3382-m), each block including multiple rows of cells. The first processor (e.g., processor 2510-0) may generate a first configuration (e.g., configuration 2551) including pointers to a first set of predefined configurations (e.g., regfile CFG set 2151 in SRAM 2150) of a plurality of predefined configurations for performing read operations on the non-volatile memory. In response to generating a first configuration (e.g., configuration 2551), the controller may generate a first set of predefined configurations in memory (e.g., mem-regfile load 2561). The controller may perform a first operation (e.g., task execution 2571) based on the first set of predefined configurations generated in memory. A second processor (e.g., processor 2510-1) may generate a second configuration (e.g., configuration 2552) including pointers to a second set of predefined configurations (e.g., regfile CFG set 2159 in SRAM 2150) among a plurality of predefined configurations. In response to generating a second configuration (e.g., configuration 2552), the controller may generate a second set of predefined configurations in memory (e.g., mem-regfile load 2562). The controller may perform a second operation (e.g., task execution 2572) based on the second set of predefined configurations generated in memory.
[0187] In some embodiments, the first set of predefined configurations may be the same as the second set of predefined configurations. In some embodiments, the multiple predefined configurations may correspond to multiple circuits for performing read operations on non-volatile memory. A first operation (e.g., an R2R operation) may be performed using a first set of circuits (e.g., circuits in R2R engine 2126). A second operation (e.g., a DNN operation) may be performed using a second set of circuits (e.g., circuits in DNN engine 2127). The second set of circuits may include at least one circuit not included in the first set of circuits.
[0188] In some embodiments, when performing the first operation, the controller may obtain a row identifier (e.g., target row 2702) that identifies the row of the target page among multiple rows. A machine learning model (e.g., DNN 2610) may generate one or more voltage thresholds (e.g., read threshold 2713) for the read operation, at least based on the row identifier. The controller may perform a read operation on the target page of the non-volatile memory using one or more voltage thresholds. The controller may obtain a shift index corresponding to a subset of one or more stress conditions and defined as a shift to a default voltage threshold. The controller may generate a lookup table storing multiple voltage thresholds for each row. The controller may generate one or more voltage thresholds using the lookup table based on the shift index and the row identifier (e.g., using Equations 1 and 2).
[0189] In some embodiments, when generating one or more voltage thresholds, the controller may receive a shift index and a row identifier as input features to a machine learning model. In response to receiving the shift index and row identifier, the machine learning model may output one or more voltage thresholds (e.g., using Equation 3).
[0190] In some embodiments, when generating one or more voltage thresholds, the controller may receive a shift index, a row identifier, and one or more voltage thresholds extracted from a history table as input features to a machine learning model. The history table may store multiple voltage thresholds for each block used historically and lead to successful decoding, and the shift index is an index into the history table. In response to receiving the shift index, row identifier, and one or more voltage thresholds, the machine learning model may output one or more voltage thresholds (e.g., using Equation 3).
[0191] In some embodiments, when performing the first operation, the controller may perform multiple read operations using fixed voltage thresholds. The controller may generate a histogram (e.g., QT histogram 2902) based on the results of the multiple read operations. The controller may generate a set of voltage thresholds for read operations on target pages of non-volatile memory based on the histogram.
[0192] In some embodiments, when performing the first operation, the controller may generate voltage thresholds representing the set of voltage thresholds. The controller may store the voltage thresholds in a lookup table that stores multiple voltage thresholds.
[0193] In some embodiments, during a first operation, the controller may perform multiple read operations having one or more voltage thresholds. The controller may generate a histogram (e.g., QT histogram 2902) based on the results of the multiple read operations. The controller may generate a set of voltage thresholds for read operations on target pages of non-volatile memory based on the histogram.
[0194] Figure 34This is a flowchart illustrating an example method for providing configurable hardware blocks to perform read operations on flash memory according to some embodiments. In some arrangements, the example method involves a process 3400 for performing operations on non-volatile memory (e.g., flash memory 3380) comprising one or more blocks (e.g., flash memory blocks 3382-1, 3382-2, ..., 3382-m), each block containing multiple rows of cells. This process may be performed by one or more controllers (e.g., controller 3320) and / or one or more processors (e.g., processor 3326) of a flash memory system (e.g., a NAND flash memory device, SSD 10).
[0195] In this example, process 3400 begins in step S3402 by generating a first configuration (e.g., configuration 2551) by a first processor (e.g., processor 2510-0) that contains pointers to a first set of predefined configurations (e.g., regfile CFG set 2151) among a plurality of predefined configurations for performing read operations on nonvolatile memory.
[0196] In step S3404, in some embodiments, in response to generating a first configuration (e.g., processor 2510-0), a first set of predefined configurations (e.g., mem-regfile load 2561) can be generated in memory (e.g., DRAM 3310).
[0197] In step S3406, in some embodiments, a controller (e.g., controller 3320) can perform a first operation (e.g., task execution 2571) based on a first set of predefined configurations generated in memory. In some embodiments, the multiple predefined configurations may correspond to multiple circuits for performing read operations on non-volatile memory. The first operation (e.g., R2R operation) can be performed using a first set of circuits from the multiple circuits (e.g., circuits in R2R engine 2126).
[0198] In step S3408, in some embodiments, a second configuration (e.g., configuration 2552) may be generated by a second processor (e.g., processor 2510-1) including pointers to a second set of predefined configurations (e.g., regfile CFG set 2159 in SRAM 2150) among a plurality of predefined configurations.
[0199] In step S3410, in some embodiments, in response to generating a second configuration (e.g., configuration 2552), a second set of predefined configurations can be generated in memory (e.g., mem-regfile load 2562). In some embodiments, the first set of predefined configurations may be the same as the second set of predefined configurations.
[0200] In step S3412, in some embodiments, the controller may perform a second operation (e.g., task execution 2572) based on a second set of predefined configurations generated in memory. A second set of circuits from a plurality of circuits may be used to perform the second operation (e.g., DNN operation). The second set of circuits (e.g., circuits in DNN engine 2127) may include at least one circuit not included in the first set of circuits.
[0201] In some embodiments, the first operation may include obtaining a row identifier (e.g., target row 2702) that identifies the row of a target page among a plurality of rows. A machine learning model (e.g., DNN 2610) may generate one or more voltage thresholds (e.g., read threshold 2713) for a read operation, at least based on the row identifier. The read operation on the target page of non-volatile memory may be performed using one or more voltage thresholds. A shift index corresponding to a subset of one or more stress conditions and defined to a default voltage threshold may be obtained. The controller may generate a lookup table storing the plurality of voltage thresholds for each row. The lookup table may be used to generate one or more voltage thresholds based on the shift index and the row identifier (e.g., using Equations 1 and 2).
[0202] In some embodiments, when generating one or more voltage thresholds, the shift index and row identifier can be received as input features to a machine learning model. In response to receiving the shift index and row identifier, the machine learning model can output one or more voltage thresholds (e.g., using Equation 3).
[0203] In some embodiments, when generating one or more voltage thresholds, a shift index, a row identifier, and one or more voltage thresholds extracted from a history table may be received as input features to a machine learning model. The history table may store multiple voltage thresholds for each block used historically and lead to successful decoding. The shift index may be an index of the history table. In response to receiving the shift index, row identifier, and one or more voltage thresholds, the machine learning model may output one or more voltage thresholds (e.g., using Equation 3).
[0204] In some embodiments, the first operation may include performing a plurality of read operations using fixed voltage thresholds. A histogram (e.g., QT histogram 2902) may be generated based on the results of the plurality of read operations. Based on the histogram, a set of voltage thresholds may be generated for read operations on target pages of non-volatile memory.
[0205] In some embodiments, the first operation may include generating voltage thresholds representing a set of voltage thresholds. The voltage thresholds may be stored in a lookup table that stores multiple voltage thresholds.
[0206] In some embodiments, the first operation may include performing multiple read operations using one or more voltage thresholds. A histogram (e.g., QT histogram 2902) may be generated based on the results of the multiple read operations. Based on the histogram, a set of voltage thresholds may be generated for read operations on target pages of non-volatile memory.
[0207] The foregoing description is provided to enable those skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects. Therefore, the claims are not intended to be limited to the aspects shown herein, but are to be given the full scope consistent with the language of the claims, wherein references to elements in the singular form are not intended to mean “one” and “only one,” unless specifically stated otherwise, but rather “one or more.” Unless otherwise specifically stated, the term “some” means one or more. All structural and functional equivalents of the elements throughout the various aspects described in the foregoing description are known or will later be known to those skilled in the art and are intended to be covered by the claims. Furthermore, nothing disclosed herein is intended to be exclusive to the public, regardless of whether such disclosure is expressly recited in the claims. Unless an element is explicitly described using the phrase “means for…”, a claim element is not to be interpreted as means plus function.
[0208] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of illustrative method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged while remaining within the scope of the previously described method. The appended method asserts the essential elements of various steps in a sample order and is not intended to be limited to the specific order or hierarchy presented.
[0209] The prior description of the disclosed embodiments is provided to enable those skilled in the art to make or use the disclosed subject matter. Various modifications to these implementations will be apparent to those skilled in the art, and the general principles defined herein can be applied to other implementations without departing from the spirit or scope of the prior description. Therefore, the prior description is not intended to limit itself to the embodiments shown herein, but is accorded the widest scope consistent with the principles and novel features disclosed herein.
[0210] The various examples illustrated and described are provided only as examples to illustrate the various features of the claims. However, the features shown and described with respect to any given example are not necessarily limited to the associated example and may be used or combined with other examples shown and described. Furthermore, the claims are not intended to be limited to any one example.
[0211] The foregoing method descriptions and process flowcharts are provided as illustrative examples only and are not intended to require or imply that the steps of the various examples must be performed in the presented order. As those skilled in the art will understand, the order of steps in the foregoing examples can be performed in any order. Words such as “afterward,” “then,” “next,” etc., are not intended to limit the order of steps; these words are merely used to guide the reader through the description of the method. Furthermore, any reference to singular claim elements, such as the use of the articles “a,” “an,” or “the,” should not be construed as limiting the element to the singular.
[0212] The various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and design constraints imposed on the system as a whole. Those skilled in the art can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this disclosure.
[0213] Hardware used to implement the various illustrative logics, logic blocks, modules, and circuits described in conjunction with the examples disclosed herein can be implemented or executed using a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but alternatively, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors incorporating a DSP core, or any other such configuration. Alternatively, some steps or methods can be performed by circuitry specific to a given function.
[0214] In some exemplary examples, the described functionality can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or code on a non-transitory computer-readable storage medium or a non-transitory processor-readable storage medium. The steps of the methods or algorithms disclosed herein can be embodied in a processor-executable software module that may reside on a non-transitory computer-readable or processor-readable storage medium. A non-transitory computer-readable or processor-readable storage medium can be any storage medium accessible by a computer or processor. By way of example and not limitation, such non-transitory computer-readable or processor-readable storage media may include RAM, ROM, EEPROM, flash memory, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer. As used herein, disks and optical discs include optical discs (CDs), laser discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically reproduce data, while optical discs optically reproduce data using lasers. Combinations of the above are also included within the scope of non-transitory computer-readable and processor-readable media. Additionally, the operation of a method or algorithm may exist as one or any combination or set of code and / or instructions on a non-transitory processor-readable storage medium and / or computer-readable storage medium that can be incorporated into a computer program product.
[0215] The foregoing description of the disclosed examples is provided to enable those skilled in the art to make or use this disclosure. Various modifications to these examples will be apparent to those skilled in the art, and the general principles defined herein may be applied to some examples without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not intended to be limited to the examples shown herein, but is accorded the widest scope consistent with the appended claims and the principles and novel features disclosed herein.
Claims
1. A method for performing operations on a non-volatile memory comprising one or more blocks, each block comprising a plurality of rows of cells, the method comprising: generating, by a first processor, a first configuration comprising a pointer to a first set of pre-defined configurations among a plurality of pre-defined configurations for performing a read operation on the non-volatile memory; in response to generating the first configuration, generating the first set of pre-defined configurations in a memory; controlling, by a controller, a first operation in accordance with the first set of pre-defined configurations generated in the memory; generating, by a second processor, a second configuration comprising a pointer to a second set of pre-defined configurations among the plurality of pre-defined configurations; in response to generating the second configuration, generating the second set of pre-defined configurations in the memory; and controlling, by the controller, a second operation in accordance with the second set of pre-defined configurations generated in the memory.
2. The method of claim 1, wherein: the first set of pre-defined configurations is the same as the second set of pre-defined configurations.
3. The method of claim 1, wherein: the plurality of pre-defined configurations correspond to a plurality of circuits for performing the read operation on the non-volatile memory, the first operation is performed using a first set of circuits among the plurality of circuits, the second operation is performed using a second set of circuits among the plurality of circuits, and the second set of circuits comprises at least one circuit not included in the first set of circuits. the first operation comprises:
4. The method of claim 1, wherein, obtaining a row identifier identifying a target row among the plurality of rows; generating, by a machine learning model, one or more voltage thresholds for the read operation based at least on the row identifier; and performing a read operation on the target row of the non-volatile memory with the one or more voltage thresholds.
5. The method of claim 4, further comprising: obtaining a subset corresponding to one or more stress conditions and defining a shift index to a shift to a default voltage threshold; generating, by the machine learning model, a lookup table storing a plurality of voltage thresholds for each row; and generating the one or more voltage thresholds based on the shift index and the row identifier using the lookup table. generating the one or more voltage thresholds comprises:
6. The method of claim 4, wherein, receiving the shift index and the row identifier as input features to the machine learning model; in response to receiving the shift index and the row identifier, outputting the one or more voltage thresholds by the machine learning model. generating the one or more voltage thresholds comprises:
7. The method of claim 4, wherein, receiving the shift index, the row identifier, and one or more voltage thresholds extracted from a history table as input features to the machine learning model, wherein, the history table stores a plurality of voltage thresholds for each block used historically and resulting in a successful decode, and the shift index is an index to the history table; in response to receiving the shift index, the row identifier, and the one or more voltage thresholds, outputting the one or more voltage thresholds by the machine learning model. the first operation comprises:
8. The method of claim 4, wherein, performing a plurality of read operations with a fixed voltage threshold; generating a histogram based on results of the plurality of read operations; and generating a set of voltage thresholds for read operations on the target page of the non-volatile memory based on the histogram.
9. The method of claim 8, wherein the first operation comprises: generating a voltage threshold representing the set of voltage thresholds; and storing the voltage threshold in a lookup table storing a plurality of voltage thresholds.
10. The method of claim 4, wherein, the first operation comprises: performing a plurality of read operations with the one or more voltage thresholds; generating a histogram based on results of the plurality of read operations; and generating a set of voltage thresholds for read operations on the target page of the non-volatile memory based on the histogram.
11. A flash system comprising: a non-volatile memory comprising one or more blocks, each block comprising a plurality of rows of cells; a controller to perform operations on the non-volatile memory; and a plurality of processors comprising a first processor and a second processor, wherein: the first processor generates a first configuration comprising a pointer to a first set of pre-defined configurations of a plurality of pre-defined configurations for performing read operations on the non-volatile memory, in response to generating the first configuration, the controller generates a first set of pre-defined configurations in a memory, the controller performs a first operation according to the first set of pre-defined configurations generated in the memory, the second processor generates a second configuration comprising a pointer to a second set of pre-defined configurations of the plurality of pre-defined configurations, in response to generating the second configuration, the controller generates a second set of pre-defined configurations in the memory; and the controller performs a second operation according to the second set of pre-defined configurations generated in the memory.
12. The system of claim 11, wherein: the first set of pre-defined configurations is the same as the second set of pre-defined configurations.
13. The system of claim 11, wherein: the plurality of pre-defined configurations correspond to a plurality of circuits for performing the read operations on the non-volatile memory, the first operation is performed using a first set of circuits of the plurality of circuits, the second operation is performed using a second set of circuits of the plurality of circuits, and the second set of circuits comprises at least one circuit not included in the first set of circuits.
14. The system of claim 11, wherein the first operation comprises: obtaining a row identifier identifying a row of the target page of the plurality of rows; generating, by a machine learning model, one or more voltage thresholds for read operations based at least on the row identifier; and performing a read operation on the target page of the non-volatile memory with the one or more voltage thresholds.
15. The system of claim 14, further comprising: obtaining a shift index corresponding to a subset of one or more stress conditions and defining a shift to a default voltage threshold; generating, by the controller, a lookup table storing a plurality of voltage thresholds for each row; and generating the one or more voltage thresholds based on the shift index and the row identifier using the lookup table.
16. The system of claim 14, wherein generating the one or more voltage thresholds comprises: receiving the shift index and the row identifier as input features for the machine learning model; in response to receiving the shift index and the row identifier, outputting, by the machine learning model, the one or more voltage thresholds.
17. The system of claim 14, wherein generating the one or more voltage thresholds comprises: receiving the shift index, the row identifier, and one or more voltage thresholds extracted from a history table as input features for the machine learning model, wherein the history table stores a plurality of voltage thresholds for each block that were historically used and resulted in a successful decode, and the shift index is an index into the history table; in response to receiving the shift index, the row identifier, and the one or more voltage thresholds, outputting, by the machine learning model, the one or more voltage thresholds.
18. The system of claim 14, wherein the first operation comprises: performing a plurality of read operations with a fixed voltage threshold; generating a histogram based on results of the plurality of read operations; and generating a set of voltage thresholds for read operations on the target page of the non-volatile memory based on the histogram.
19. The system of claim 18, wherein the first operation comprises: generating a voltage threshold representing the set of voltage thresholds; and storing the voltage threshold in a lookup table that stores a plurality of voltage thresholds.
20. The system of claim 14, wherein the first operation comprises: performing a plurality of read operations with the one or more voltage thresholds; generating a histogram based on results of the plurality of read operations; and generating a set of voltage thresholds for read operations on the target page of the non-volatile memory based on the histogram.