Memory system and method for circuit boundary array-based time and temperature tagging management and inference of read thresholds

A machine learning-based system adjusts read thresholds in memory systems using a circuit boundary array to address inconsistencies caused by NAND process scaling and varying conditions, enhancing performance and reliability.

JP2025520224AActive Publication Date: 2025-07-02SANDISK TECHNOLOGIES LLC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2024568124
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-07-11
Filing Date
2023-10-03
Publication Date
2025-07-02
Estimated Expiration
2043-10-03

AI Technical Summary

Technical Problem

The challenge of maintaining process uniformity in NAND process scaling and 3D stacking, coupled with varying operating conditions, leads to inconsistent read thresholds in memory systems, resulting in higher bit error rates and degraded performance due to decoding failures.

Method used

A memory system utilizing a machine learning-based inference engine to dynamically adjust read thresholds based on multiple parameters, including time and temperature groups, program/erase counts, and physical location, optimizing the read threshold for each memory die using a circuit boundary array (CBA) for parallel processing.

Benefits of technology

This approach reduces bit error rates, improves latency and throughput, and enhances quality of service by providing a consistent and nearly optimal read threshold under varying conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520224000001_ABST
    Figure 2025520224000001_ABST
Patent Text Reader

Abstract

The memory system has an inference engine that can infer a read threshold based on multiple parameters of the memory. The read threshold can be used when reading word lines in the memory during normal read operations or as part of an error handling process. Using a machine learning-based approach to infer the read threshold can provide a significant improvement in read threshold accuracy, which can reduce the bit error rate and improve latency, throughput, power consumption, and quality of service. In another embodiment, the circuit boundary array manages updates to time and temperature tag information and is used to infer the read threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of priority to U.S. Patent Application No. 18 / 220,363, filed on July 11, 2023, entitled "Storage System and Method for Circuit - Bounded - Array - Based Time and Temperature Tag Management and Inference of Read Thresholds", the entire content of which is incorporated herein by reference for all purposes. The said patent application is (a) a continuation - in - part of U.S. Patent Application No. 17 / 838,481, filed on June 13, 2022, (b) a continuation - in - part of U.S. Patent Application No. 17 / 838,481, filed on June 13, 2022, and also a continuation - in - part of U.S. Patent Application No. 17 / 899,073, filed on August 30, 2022, and (c) claims priority to U.S. Provisional Patent Application No. 63 / 421,647, filed on November 2, 2022, all of which are incorporated herein by reference.

Background Art

[0002] One of the main challenges posed by negative logic AND (NAND) process scaling and three-dimensional stacking is maintaining process uniformity. In addition, memory products need to support a wide range of operating conditions such as different program / erase cycles, retention times, and temperatures, which leads to increased variability between memory dies, blocks, and pages across different operating conditions. Due to these variations, the read threshold used to read a memory page is not fixed and can vary significantly as a function of physical location and operating conditions, especially for newer, less mature memory nodes. Reading with an inaccurate read threshold can lead to a higher bit error rate, which can degrade performance and quality of service due to decoding failures, which may require invoking a high-latency recovery flow, causing a temporary degradation in latency and performance.

Brief Description of the Drawings

[0003]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5A

Figure 5B

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11A

Figure 11B

Figure 11C

Figure 12

Figure 13

Figure 14

Figure 15

DETAILED DESCRIPTION OF THE INVENTION

[0004] The following embodiments generally relate to a memory system and method for inferring an optimal read threshold based on memory parameters and conditions. In one embodiment, a memory system is provided that includes a memory and a controller. The controller is configured to use an inference engine to infer a read threshold based on a plurality of parameters of the memory and to use the read threshold when reading word lines in the memory. In another embodiment, a method is provided that is executed within a memory system that includes a memory. The method includes generating an inference of a read threshold based on a plurality of parameters of the memory and using the read threshold when reading word lines in the memory. In yet another embodiment, a memory system is provided that includes a memory, an inference engine configured to provide an inference of a read threshold based on a plurality of parameters of the memory, and means for retraining the inference engine based on the quality of the inference.

[0005] In another embodiment, a memory system is provided that includes a memory that includes a plurality of memory dies, each memory die including a respective circuit boundary array, and a memory controller coupled to the memory. The memory controller is configured to read word lines in each of the plurality of memory dies, determine a read threshold based on the read word lines, and transmit the read threshold to the circuit boundary array within the memory die. Each circuit boundary array is configured to apply a machine learning-based adjustment to the read threshold.

[0006] In another embodiment, a method is provided that is executed in a memory system that includes a memory that includes a plurality of circuit boundary arrays (CBAs), each CBA including a memory that includes a memory die. In each CBA, reading word lines within a time and temperature group within the memory die of the CBA, determining a read threshold based on the read word lines, storing the read threshold in the CBA, and applying a machine learning-based adjustment to the read threshold are executed.

[0007] In yet another embodiment, a memory comprising a circuit border array including memory dies, and means located within the circuit border array for reading word lines within a memory die, determining a read threshold based on the read word lines, and applying on-the-fly a machine learning-based adjustment to the read threshold before a target word line within the memory die is read, is provided. Other embodiments are provided and may be used alone or in combination.

[0008] Referring now to the drawings, a memory system suitable for use in implementing aspects of these embodiments is shown in FIGS. 1A - 1C. FIG. 1A is a block diagram illustrating a non-volatile memory system 100 (which may be referred to herein as a memory device or simply a device) according to one embodiment of the subject matter described herein. Referring to FIG. 1A, the non-volatile memory system 100 includes a controller 102 and a non-volatile memory that may be composed of one or more non-volatile memory dies 104. As used herein, the term die refers to an aggregate of non-volatile memory cells formed on a single semiconductor substrate and associated circuitry for managing the physical operation of these non-volatile memory cells. The controller 102 interfaces with a host system and transmits command sequences for read, program, and erase operations to the non-volatile memory die 104.

[0009] The controller 102 (the controller 102 may be a non-volatile memory controller (e.g., a flash, resistive random-access memory (ReRAM), phase-change memory (PCM), or magneto-resistive random-access memory (MRAM) controller)) can take the form of a processing circuit, a microprocessor or a processor, and a computer-readable medium. The computer-readable medium stores computer-readable program code (e.g., firmware) executable by, for example, a (micro)processor, logic gates, switches, application specific integrated circuit (ASIC), programmable logic controller, and embedded microcontroller. The controller 102 can be composed of hardware and / or firmware for performing various functions described below and shown in the flow diagrams. Also, some of the components shown as being inside the controller may also be stored outside the controller, and other components may be used. Additionally, the phrase "operably communicate with" can mean communicate directly with, or communicate indirectly (wired or wireless) through one or more components, which may or may not be illustrated and described herein.

[0010] As used herein, a non-volatile memory controller is a device that manages data stored in a non-volatile memory and communicates with a host such as a computer or an electronic device. The non-volatile memory controller can have various functions in addition to the specific functions described herein. For example, the non-volatile memory controller can format the memory to ensure that the non-volatile memory is operating properly, map out defective non-volatile memory cells, and allocate spare cells to be replaced with future failing cells. Some of the spare cells can be used to hold firmware for operating the non-volatile memory controller and implementing other features. In operation, when a host needs to read data from or write data to the non-volatile memory, the host can communicate with the non-volatile memory controller. If the host provides a logical address to which the data is to be read / written, the non-volatile memory controller can translate the logical address received from the host into a physical address within the non-volatile memory. (Alternatively, the host can provide a physical address.) The non-volatile memory controller can also perform various memory management functions including, but not limited to, wear leveling (to avoid wearing out particular memory cell blocks that would otherwise be repeatedly written to) and garbage collection (to move only valid data pages to a new block after a block becomes full, so that full blocks can be erased and reused). Also, the structure for the "means" recited in the claims can include, for example, some or all of the structures of the controller described herein that are programmed or manufactured as necessary to cause the controller to operate to perform the recited functions.

[0011] The non-volatile memory die 104 may include any suitable non-volatile storage medium including ReRAM, MRAM, PCM, NAND flash memory cells and / or NOR flash memory cells. The memory cells can be in the form of solid state (e.g., flash) memory cells and can be once programmable, multi-programmable, or many times programmable. The memory cells can also be single-level (1 bit per cell) cells (single-level cell, SLC), or multi-level cells (multiple-level cell, MLC) such as two-level cells, triple-level cells (TLC), quad-level cells (QLC), or other memory cell level technologies known currently or developed in the future can be used. Also, the memory cells can be fabricated two-dimensionally or three-dimensionally.

[0012] The interface between the controller 102 and the non-volatile memory die 104 may be any suitable flash interface such as toggle mode 200, 400, or 800. In one embodiment, the memory system 100 may be a card-based system such as a Secure Digital (SD) or Micro Secure Digital (Micro SD) card (or USB, SSD, etc.). In an alternative embodiment, the memory system 100 may be part of an embedded memory system.

[0013] In the example illustrated in FIG. 1A, a non-volatile memory system 100 (which may be referred to herein as a memory module) includes a single channel between a controller 102 and a non-volatile memory die 104, and the subject matter described herein is not limited to having a single memory channel. For example, in some memory system architectures (such as those shown in FIGS. 1B and 1C), two, four, eight or more memory channels may exist between the controller and the memory device, depending on the capabilities of the controller. In any of the embodiments described herein, even if a single channel is shown in the figure, more than one channel may exist between the controller and the memory die.

[0014] FIG. 1B illustrates a memory module 200 that includes a plurality of non-volatile memory systems 100. Thus, the memory module 200 may include a memory controller 202, and the memory controller 202 interfaces with a host and a memory system 204, and the memory system 204 includes a plurality of non-volatile memory systems 100. The interface between the memory controller 202 and the non-volatile memory system 100 may be a bus interface such as a Serial Advanced Technology Attachment (SATA), a Peripheral Component Interconnect express (PCIe) interface, or a Double-Data-Rate (DDR) interface. In one embodiment, the memory module 200 may be a Solid State Drive (SSD), or a Non-Volatile Dual In-line Memory Module (NVDIMM), such as found in a server PC or a portable computing device such as a laptop computer and a tablet computer.

[0015] FIG. 1C is a block diagram illustrating a hierarchical memory system. The hierarchical memory system 250 includes a plurality of memory controllers 202, and each of the plurality of memory controllers 202 controls a respective memory system 204. The host system 252 can access the memory in the memory system via a bus interface. In one embodiment, the bus interface may be a Non-Volatile Memory express (NVMe) or a Fiber Channel over Ethernet (FCoE) interface. In one embodiment, the system illustrated in FIG. 1C may be a rack-mountable mass storage system accessible by a plurality of host computers, such as found in a data center or other locations where a mass storage device is required.

[0016] FIG. 2A is a block diagram illustrating the components of the controller 102 in more detail. The controller 102 includes a front-end module 108 that interfaces with a host, a back-end module 110 that interfaces with one or more non-volatile memory dies 104, and various other modules that perform the functions described in detail herein. The modules can take the form of, for example, a packaged functional hardware unit designed for use with other components, a portion of program code (e.g., software or firmware) executable by a (micro)processor or processing circuit that normally performs a specific function of the associated function, or a self-contained hardware or software component that interfaces with a larger system. The controller 102 may be referred to herein as a NAND controller or a flash controller, although it should be understood that the controller 102 can be used with any suitable memory technology, some examples of which are provided below.

[0017] Referring again to the modules of the controller 102, the buffer manager / bus controller 114 manages buffers in the Random Access Memory (RAM) 116 and controls the internal bus arbitration of the controller 102. The Read Only Memory (ROM) 118 stores the system startup code. Although illustrated in FIG. 2A as being located separately from the controller 102, in other embodiments, one or both of the RAM 116 and the ROM 118 may be located within the controller. In still other embodiments, portions of the RAM and the ROM may be located both within and outside the controller 102.

[0018] The front-end module 108 includes a host interface 120 and a Physical Layer Interface (PHY) 122, and the host interface 120 and the PHY 122 provide an electrical interface with a host or the next-level storage controller. The choice of the type of the host interface 120 may depend on the type of memory being used. Examples of the host interface 120 include, but are not limited to, SATA, SATA Express, Serially Attached Small Computer System Interface (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 typically facilitates the transfer of data, control signals, and timing signals.

[0019] The back-end module 110 includes an Error Correction Code (ECC) engine 124 that encodes data bytes received from a host and decodes and error corrects data bytes read from the non-volatile memory. A command sequencer 126 generates command sequences such as program command sequences and erase command sequences to be sent to the non-volatile memory die 104. A Redundant Array of Independent Drive (RAID) module 128 manages the generation of RAID parity and the recovery of failed data. The RAID parity can be used as an additional level of integrity protection for the data written within the memory device 104. In some cases, the RAID module 128 may be part of the ECC engine 124. A memory interface 130 provides command sequences to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104. In one embodiment, the memory interface 130 can be a double data rate (DDR) interface such as a toggle mode 200, 400, or 800 interface. A flash control layer 132 controls the overall operation of the back-end module 110.

[0020] The storage system 100 also includes other discrete components 140 such as an external electrical interface, an external RAM, resistors, capacitors, or other components that can interface with the controller 102. In an alternative embodiment, one or more of the physical layer interface 122, the RAID module 128, the media management layer 138, and the buffer management / bus controller 114 are optional components not necessary within the controller 102.

[0021] FIG. 2B is a block diagram illustrating in more detail the components of the non-volatile memory die 104. The non-volatile memory die 104 includes a peripheral circuit 141 and a non-volatile memory array 142. The non-volatile memory array 142 includes non-volatile memory cells used to store data. The non-volatile memory cells may be any suitable non-volatile memory cells including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in two-dimensional and / or three-dimensional configurations. The non-volatile memory die 104 further includes a data cache 156 that caches data. The peripheral circuit 141 includes a state machine 152 that provides status information to the controller 102.

[0022] Referring back to FIG. 2A, the flash control layer 132 (referred to herein as the Flash Translation Layer (FTL), or more generally, when the memory may not be flash, the "media management layer") processes flash errors and interfaces with the host. In particular, the FTL, which can be an algorithm within the firmware, is involved in the internal memory management and converts writes from the host into writes into the memory 104. The memory 104 may have limited durability, may only be written into within a plurality of pages, and / or may not be written into unless the memory 104 is erased as a block of memory cells, so the FTL may be required. The FTL understands these potential limitations of the memory 104 that may not be visible to the host. Thus, the FTL attempts to convert writes from the host into writes into the memory 104.

[0023] The FTL may include a logical-to-physical address (L2P) map (which may be referred to herein as a table or data structure) and an allocated cache memory. In this way, the FTL converts a logical block address (“LBA”) from the host to a physical address within the memory 104. The FTL can include other features including, but not limited to, power-off recovery (for which the data structure of the FTL can be recovered in the event of a sudden power loss), and wear leveling (for which the wear across the memory blocks is uniform to prevent excessive wear on certain blocks that could present a greater chance of failure).

[0024] Referring back to the drawings, FIG. 3 is a block diagram of a host 300 and a memory system 100 (which may be referred to herein as a device) in one embodiment. The host 300 can take any suitable form including, but not limited to, a computer, a mobile phone, a digital camera, a tablet, a wearable device, a digital video recorder, a surveillance system, etc. The host 300 includes a processor 330 configured to send data (initially stored, for example, in the host's memory 340 (e.g., DRAM)) to the memory system 100 for storage within the memory 104 of the memory system (e.g., a non-volatile memory die). Note that although the host 300 and the memory system 100 are shown as separate boxes in FIG. 3, the memory system 100 may be integrated within the host 300, the memory system 100 may be removably connected to the host 300, and the memory system 100 and the host 300 can communicate via a network. Also note that the memory 104 may be integrated within the memory system 100 or may be removably connected to the memory system 100.

[0025] As described above, one of the main challenges posed by NAND process shrink and 3D stacking is to maintain process uniformity. In addition, memory products need to support a wide range of operating conditions such as different program / erase cycles, retention times, and temperatures, which leads to an increase in variability between memory dies, blocks, and pages across different operating conditions. Due to these variations, the read threshold (RT) used to read a memory page is not fixed and varies significantly as a function of physical location and operating conditions, especially for newly emerging and less mature memory nodes. Reading with an inaccurate read threshold can lead to a higher bit error rate (BER), which can degrade performance and quality of service (QoS) due to decoding failures, which requires invoking a high latency recovery flow, causing a temporary degradation in latency and performance.

[0026] The problem of maintaining an optimal read threshold is particularly important for enterprise memory systems with very strict quality of service requirements, as well as for mobile, Internet of Things (IoT), and automotive memory systems where the range of required operating conditions is wide and the frequency of condition changes (e.g., temperature) can be high. This problem is even more difficult during the transition to new and less mature memory nodes.

[0027] Current solutions for read threshold calibration, such as bit error rate (BER) estimation scan (BES) and valley search (VS), are high-latency operations aimed at optimizing the read threshold for specific word lines. This is good for the rare read recovery flow when data decoding fails, but not so suitable for frequent operations in the case of frequent read threshold changes. Therefore, to address this issue, a flash memory system can implement a read threshold management scheme that attempts to track changes in the read threshold in the background through a maintenance process to ensure that an appropriate read threshold is used when the host issues a read command.

[0028] One approach is to track the read threshold for each group of blocks that share the same conditions. More specifically, blocks written at approximately the same time and temperature are grouped into time and temperature (TT) groups. The read threshold is tracked for each TT group and is typically obtained on a representative word line from the blocks within the group. When the host performs a read operation, the read threshold associated with the TT group corresponding to the read block is used, and additional adaptation to the read threshold may be performed based on a pre-calibrated word line zoning table according to the specific word line read.

[0029] Unfortunately, existing read threshold management schemes may be sub-optimal and may not properly track read thresholds under frequently changing conditions and high variability between memory pages. For example, as described above, blocks can be grouped according to programming time and temperature, and the maintenance process can track appropriate read thresholds for each group of blocks by finding the optimal read threshold for a representative word line from the block (e.g., via a BER estimation scan or valley search). An example of this technique is shown in flowchart 400 of FIG. 4. As shown in FIG. 4, in this method, the memory system receives a bit error rate that exceeds the threshold indication from the indicated word line (operation 410). Next, the memory system 100 obtains the read threshold for the representative word line from the same time and temperature group (e.g., via a BER estimation scan or cell voltage distribution valley search) (operation 420). Next, the memory system 100 associates the obtained read threshold with the time and temperature group (to be used for subsequent read operations from the group) (operation 430) and adjusts the read threshold based on the word line indexing zone when reading other word lines (operation 440).

[0030] Accordingly, in this technique, a predefined correction can be applied to the read threshold of the representative word line based on the word line number being read (using a word line zoning table). If a particular word line shows an increase in bit error rate or if the decoding of the data on the word line fails, a BER estimation scan or valley search can be applied in the foreground to calibrate the read threshold of the word line as part of the read error handling (REH) flow. The indicated word line is typically selected on the edge of the block so that a BER increase is quickly captured. However, this approach is sub-optimal and can lead to a temporary degradation in performance and a violation of service quality under stress conditions (such as rapid temperature changes).

[0031] As a function of various memory parameters (such as Program / Erase Count (PEC), WL#, etc.), other table-based methods can be used to set the read threshold based on a predefined table. However, due to limitations in the actual table size, such methods may consider only a limited number of parameters, or alternatively, may assume a simplified model, where each factor (e.g., word line number, program erase count, temperature, die-dependency, etc.) affects the read threshold in an independent additional form. In reality, interactions may be further involved and can be more complex and non-linear.

[0032] Using the following embodiments, an optimal read threshold can be inferred from all available information, including TT group information, temperature information, BER information, Program / Erase Count (PEC) information, and physical page location. In one embodiment, a machine learning methodology is used to train a low-complexity inference model under all relevant conditions to learn the intricate non-linear dependencies of the read threshold on each of the available features. Using this approach, the memory system can fine-tune the TT group read threshold based on additional information sources to provide a consistent and nearly optimal read threshold. This, in turn, can reduce the BER level of the read data, which improves the performance and quality of the service, reduces power consumption, and reduces the decoder failure event rate.

[0033] In one embodiment, the controller 102 of the memory system 100 infers an optimal read threshold based on a non-linear function of a plurality of inputs that reflect the current memory and data conditions. Using a machine learning (ML) methodology, an inference function for the read threshold can be derived that utilizes all available information sources, including the latest TT information, BER information, temperature information (program temperature / TT acquisition temperature / current read temperature), PEC information, and physical location information (WL#, string#, plane#, edge block, die X / Y information, etc.) of the block. In this way, an improved read threshold is used to reduce the BER level of the read data. In one implementation, the controller 102 uses a low-complexity hardware and firmware implementation of the inference function that selects appropriate operated features and an appropriate machine learning model. Of course, other implementations are possible.

[0034] As described above, some of the methods for setting the read threshold are sub-optimal and do not use all of the information sources (e.g., TT group information, NAND conditions, temperature, physical address, etc.) available for inferring the read threshold in an optimal and holistic manner. More specifically, the optimal read threshold for a particular page under particular memory conditions is information regarding the time and temperature group of the block to which the read block belongs, the read threshold obtained on the representative WLx, the BER information of the representative WLx (SW / BER / BER1→0 / BER0→1), the temperature at which the WLx read threshold was obtained, the time at which the WLx read threshold was obtained, the read threshold obtained on the representative WLy, the BER information of the representative WLy (SW / BER / BER1→0 / BER0→1), the temperature at which the WLy read threshold was obtained, the time at which the WLx read threshold was obtained, the program temperature of the block being read ("Prog-Temp"), the current read temperature ("Read-Temp"), the difference between the Prog-Temp and the current Read-Temp (also referred to as "X-Temp"), the PEC of the block being read, the data retention level of the block being read (i.e., a function of the time elapsed since the block was programmed, normalized by temperature, the TimePool index of the block), the BER information of the previous WL / page for the page being read (which may be available under a sequential read scenario), the default read threshold of the die, the physical address information of the read page, WL / page#, string#, plane#, block position (e.g., edge / non-edge block), and die information (e.g., X / Y position on the wafer), etc., but not limited to these, and may be correlated to a plurality of parameters that may be available to the controller 102 during operation.

[0035] In theory, an optimal read level for each case could be memorized using a large multi-dimensional table indexed by all of these parameters. However, this is not feasible as it requires an exponentially large table. Instead, one embodiment applies a machine learning-based methodology to learn a low-complexity inference model for the optimal read threshold based on all available parameters (or the most beneficial and / or easily available parameters).

[0036] FIG. 5A is a diagram of an inference engine 500 in one embodiment that illustrates this approach. Model training can be done either offline or online, or in combination (e.g., basic offline training with online fine-tuning and retraining). With respect to offline training, one of the main problems in any machine learning project is obtaining a large labeled dataset (i.e., obtaining “big data” to train the model). Fortunately, for nearby problems, this task is fairly straightforward. A large number of state-by-state cell-voltage-distribution (SbS CVD) measurements can be obtained under different NAND conditions, and based on those measurements, both features (such as those listed above) and the optimal read threshold for each measurement that will act as a label can be obtained. Then, a machine learning model can be trained to non-linearly couple input features (including target page parameters and TT group parameters) with the optimal read threshold. During normal operation of the device, the machine learning model can be applied to normal read operations, or read error handling (REH) can be used. As demonstrated below, a compact and very simple machine learning model can provide a significant reduction in the resulting failed bit count (FBC) compared to the default read threshold. The relevant model parameters may be held in ROM or SRAM, or in non-volatile memory 104 (e.g., in an SLC block). However, a very simple model that requires very little RAM / ROM can provide excellent results.

[0037] For example, when performing data collection and constructing a training test, the cell voltage distribution for each state (SbS CVD) can be collected for various conditions (PEC, memory bake time, program / read temperature, etc.). FIG. 5B is a graph showing the SbS CVD from which the optimal read threshold and other related information (such as BER information for any set of read thresholds (optimal / BES / VS / default)) can be derived. From each SbS-CVD, the optimal read threshold and the BER statistics when reading at any read level can be determined. This enables the construction of a training dataset. For example, for any WL j within block k, the optimal read level label can be the BES / VS read level on a representative word line taken from block m having similar conditions (potentially excluding its read temperature that emulates the TT management scheme used by the system), the TT read threshold acquisition temperature (= the read temperature of block m), the elapsed time since acquisition (= the additional bake time of block m), the PEC of block k, the elapsed time normalized by temperature, the read temperature of block k, the program temperature of block k, and the BER information of various pages (which may include the target page within WL j of block k, the previous page within block k, the selected representative page within block m), and can be associated with several features including but not limited to the physical address of the target WL j (which may include WL index j, plane #, string #, edge block indication, die X / Y information).

[0038] Regarding the online training approach, the approach can include continuous data collection during the lifetime of the device (from similar SbS-CVD data or other available data), at which time the machine learning model can be trained or modified based on this dynamic database. Online training can continue throughout the lifetime of the device and model adjustment can be performed.

[0039] The method in this embodiment can be much more sensitive and scalable, and additional conditions and modifications are not required. This method is illustrated in the flowchart 600 of FIG. 6. As shown in FIG. 6, first data is collected (operation 610), and a machine learning model is trained (operation 620). These operations can be performed using offline analysis. Then, the controller 102 can apply the inference model (operation 630), execute a read operation within an optimized read threshold (threshold, TH) (operation 640), monitor the quality of the read threshold (operation 650), and retrain / fine-tune the model (operation 660).

[0040] These embodiments can be generalized for the optimization and inference of one or more of the following parameters, namely, read threshold, program / verify threshold, log-likelihood ratio (LLR) table, soft-bit read threshold, or soft-bit delta value, based on the same or similar characteristics that affect the performance of the memory system.

[0041] One embodiment is based on a hardware implementation of an inference engine such that inference is performed as part of the main stream read operation from the host 300. In this case, the inference engine can directly access the memory that holds the features related to the current read operation (e.g., TT table, PEC table, temperature sensor, physical address, etc.). The low-level RISC can prepare a descriptor having the features related to the current read operation. In this way, the inference engine can provide an optimized read level for each read operation.

[0042] In another embodiment, selective use of the inference engine may be applied when the inference is based on a firmware implementation or when the latency to access all relevant features is prohibitive for mainstream use. FIG. 7 is a flowchart 700 illustrating selective use. As shown in FIG. 7, when a read is issued, the controller 102 checks whether the estimated BER (e.g., via ECC syndrome weight (SW) calculation) is greater than a threshold (operation 710), which indicates whether the condition under which the TT read threshold was obtained is the same as the current one. If the SW is below the threshold, the current read threshold is used for the read (operation 720). However, if the SW is greater than the threshold, the read threshold is adjusted (operation 730) and used for the read (operation 740).

[0043] In another alternative, these embodiments can be used as part of a read error handling (REH) flow. For example, if the decoder fails to decode after a normal hard bit (HB) read, the read threshold inference module can be applied to the failed word line, followed by another HB read. The conventional REH flow performs a long read threshold calibration operation (e.g., via BES or VS) immediately after the HB decode failure, and the proposed additional step may significantly reduce the overall read latency. This alternative is shown in the flowchart 800 of FIG. 8. As shown in FIG. 8, in this embodiment, after a read command is issued to a word line, the controller 102 detects an ECC HB decode failure (operation 810) and determines whether the SW is greater than the read threshold (operation 820). If the SW is greater than the read threshold, BES / VS is executed on the failed word line (operation 830). If the SW is less than the read threshold, the read threshold of the word line is adjusted based on the inference function / engine (operation 840), and an HB read is performed using the adjusted read threshold (operation 850).

[0044] Figure 9 is a sigma plot graph of the average FBC for different test conditions. The three curves represent the reference value, the optimal read threshold, and the results generated using a machine learning-based inference engine. The results show a significant gain when using the machine learning-based inference engine compared to the reference value. The large gain is particularly observed for samples where the TT read threshold acquisition temperature is different from the current read temperature. In that situation, the TT read threshold becomes significantly sub-optimal, but the machine learning-based inference engine provides effective compensation for the temperature difference. This information is provided for illustrative purposes and should be understood not to be intended as a limitation in any way to the claims.

[0045] There are several advantages associated with these embodiments. For example, using a machine learning-based approach can provide a significant improvement in read threshold accuracy over the reference method, as shown above. The improved read threshold can result in a reduced BER, which can improve NAND latency and throughput, improve power consumption, reduce the error rate, and improve the quality of service.

[0046] Many alternative forms can be used with these embodiments. For example, some of the embodiments described above rely on a controller placement inference module based on a non-linear function of multiple inputs that reflect the current memory and data conditions. In one alternative embodiment, a circuit boundary array (CBA) is used because it has properties that can better accommodate a distributed read threshold calculation locally for each die by an inference module. The following paragraphs provide a brief overview of the CBA, followed by examples of how the CBA can be used for time and temperature tag management and inference of read thresholds.

[0047] The Circuit Border Array (CBA) architecture is a technology under development that separates the complementary metal oxide semiconductor (CMOS) logic typically executed under the memory array in the "Circuit Under the Array" (CUA) architecture and implements it in a separate CMOS chip, thereby enabling faster operation. FIG. 10 is an example of the CBA architecture 1000 of one embodiment. As shown in FIG. 10, the CBA architecture 1000 of this embodiment includes one or more CMOS chips coupled to one or more arrays via one or more connection units, and the array corresponds to one or more memory locations of the non-volatile memory 104, such as the first die among the plurality of dies of the memory 104. The CMOS device and the related architecture may be referred to as the CMOS work function (WF) 1002. Further, the array and the related architecture may be referred to as the array WF 1004.

[0048] In one embodiment, the CMOS WF 1002 is a "Circuit Above the Array" (CAA) device. Since the CMOS device is separated from the array WF 1004, the CMOS logic may be executed faster than the CuA device. Each CAA device of the plurality of CAA devices may include an error correction code (ECC) module configured to encode and decode an error correction code with each of the associated memory dies.

[0049] As described above, in the CBA and CAA architectures, the CMOS and array portions are placed in separate wafers using separate fabrication processes. This enables better overall input-output (IO) performance, power, and cost, as well as a new approach that increases three-dimensional stacking. For example, in a conventional architecture 1100 (see FIG. 11A) where the CMOS and memory array are on the same wafer, the wafers can be stacked in an offset pattern to allow bond wires to connect to the CMOS portion of each wafer and then to the substrate. As shown in the stacking configuration 1102 of FIG. 11B, offset stacking can also be used (each array can include multiple array wafers). Alternatively, as shown in the stacking configuration 1104 of FIG. 11C, each CMOS / array pair can be stacked directly on top of each other. Since each memory die is coupled to a CMOS device and each CMOS device includes an ECC unit, the architectures 1102, 1104 of FIGS. 11B and 11C include an equal number of memory dies, CMOS devices, and ECC units.

[0050] Also, the CMOS of each CMOS / array pair can manage the programming and reading of data with the paired memory. Thus, when the controller 102 receives a write command from the host 300, the controller 102 can transfer the data associated with the write command to the associated CMOS / array pair. The CMOS device within the pair can use its ECC module to encode the data with error correction codes (e.g., low-density parity check (LDPC) codes and / or parity data), and then program the encoded data into the paired memory array. The reverse operation can be performed when the controller 102 receives a read command from the host 300. The CMOS device can include a sense amplifier that senses a low-power signal representing the bits of the memory cells at the word lines in the memory 104 and amplifies it to a recognizable logic level. The CMOS device can also include a latch that latches the data and error correction code parity read from or written to the word lines in the memory.

[0051] Another advantage of CBA is that each die can have its own system for read threshold calibration, and as a result, the process of read threshold calibration can be offloaded to CBA for faster parallel operation. Additionally, the individual die characteristics can be used to improve the read threshold calibration and better tune it for the individual dies using CBA. For example, CBA can be used to infer the optimal read threshold based on a non-linear function of multiple inputs that reflects the current memory and data conditions. In this way, CBA can use an efficient mechanism to adapt the solution to the corresponding die.

[0052] Also, the CBA can be used to improve the above-mentioned machine learning (ML) methodology to derive a read threshold inference function that utilizes all available information sources. FIG. 12 is a block diagram illustrating an architecture that uses the above-described machine learning methodology. As shown in FIG. 12, in this architecture, the controller 1002 has a read threshold update module 1200 that applies a machine learning-based methodology to learn a low-complexity inference model for optimal read thresholds in a plurality of memory dies within the memory 104 based on all various available parameters. As shown in FIG. 13, in this embodiment, the controller 102 has a read threshold update management module 1300 that applies a CBA-based calculation to provide an optimized read threshold for each die within the memory 104. In this embodiment, each of the dies has an attached CBA read threshold update module 1304 that serves to update the read threshold. These modules 104 can include a machine learning module that uses all available parameters to optimize the read threshold. In FIG. 13, the time and temperature (TT) group is managed by the read threshold update management module 1300 within the controller 102.

[0053] FIG. 14 is a flowchart of a method of an embodiment in which time and temperature group information is managed by the controller 102. As shown in FIG. 14, in this embodiment, the controller 102 triggers a time tag update (operation 1402) (e.g., based on a BER indication, Xtemp, elapsed time, etc.). Next, the controller 102 reads a representative word line and performs a read threshold calibration thereon (operation 1404). Next, the controller 102 passes the updated read threshold to the corresponding CBA (operation 1406). Next, that CBA applies an on-the-fly machine learning read threshold adjustment when the corresponding time and temperature group on the die is read (operation 1408).

[0054] In this embodiment, the module on the CBA can perform a machine learning read threshold operation, and the machine learning read threshold operation may involve decision trees with different optimization characteristics. (An example of a symmetric tree model that can be used for read threshold calibration is described in U.S. Patent Application No. 17 / 899,073, filed on August 30, 2022, which is incorporated herein by reference.) Since the CBA can be executed in parallel, the read threshold adjustment may be performed "on the fly" simultaneously immediately before the read of the target WL word line is performed on the corresponding die.

[0055] In another embodiment, the time and temperature group information is managed at the CBA level, and only the trigger for updating the time and temperature group information is passed from the controller 102 to the CBA of the corresponding die. This is shown in the flowchart 1500 of FIG. 15. As shown in FIG. 15, the controller 102 triggers a time tag update (operation 1502) (e.g., based on a BER indication, Xtemp, elapsed time, etc.). Next, a command to update the time and temperature group information is passed from the controller 102 to the CBA of the corresponding die (operation 1504). Next, the CBA reads the representative word line and performs a read threshold calibration on it (operation 1506). Then, the read threshold determined for the time and temperature group information is updated and stored in the corresponding CBA (operation 1508). Finally, the CBA applies the machine learning read threshold adjustment on the fly when the corresponding time and temperature group on the die is read (operation 1510). Note that the time tag and temperature tag can be defined at the per-die level, which facilitates implementation as each die can update the time and temperature tags without changing the currently operating mode.

[0056] Both the update of time and temperature (which can be applied to representative word lines using a read threshold calibration algorithm such as BES) and the inference for individual word lines based on the time and temperature read thresholds can be performed in parallel across all different dies. Autonomous CBA time and temperature updates triggered by elapsed time or idle periods can also be triggered in this embodiment.

[0057] In one embodiment, hardware that supports read threshold calibration such as BES or valley search (VS) can be present on the CBA (VS is a NAND function, and thus the CBA can operate VS on NAND for this purpose). Also, the update of time and temperature can be performed during the idle time of the die or when a signal from the controller 102 notifies the CBA to update the time and temperature. The trigger for this signal from the controller 102 can be based on any suitable factor, including but not limited to a high BER indication, elapsed time, or a change in conditions.

[0058] In another embodiment, each of the CBA-based machine learning read threshold adjustment modules can be modified to better fit the individual die. The controller-based system can be based on the same implementation for all dies. A die-based system can be implemented in the controller 102 itself, but much of the information such as die temperature or the position of the die within the die stack may not be present within the controller 102. Instead of having the controller 102 work on all dies and perform all calculations for all dies, it may be desirable to offload the effort to each of the CBAs (each of which processes only its corresponding die in parallel).

[0059] One embodiment includes both an "online" approach for machine learning read thresholds where the training of the model is done during the lifetime of device 100, and an "offline" approach where most of the training is done in the lab over a large dataset of different dies and the fine-tuning is done on the CBA of the die itself. For example, if one of the dies is an outlier for the lab dataset, the CBA can make adjustments based on feedback from the read results of the die itself. Also, each of the dies' CBAs can have a "model retraining / fine-tuning" block for adapting to the corresponding die. Additionally, to shorten the retraining time, die settings can be selected from a given set of classes.

[0060] There are several advantages associated with these embodiments. For example, implementing a machine language-based approach in the CBA has significant advantages due to the parallel processing capabilities provided by the CBA-based system. Additionally, the model for each die can be specifically adapted to the die on which the model operates without considering other dies and without consuming additional space. The improved read threshold can result in a reduction in BER, which can improve NAND latency and throughput, power consumption, and quality of service (QoS) while reducing the correctable ECC (CECC) rate.

[0061] Finally, as described above, any suitable type of memory may be used. Semiconductor memory devices include volatile memory devices such as Dynamic Random Access Memory ("DRAM"), Static Random Access Memory ("SRAM") devices, ReRAM, Electrically Erasable Programmable Read Only Memory ("EEPROM"), flash memory (which may also be considered a subset of EEPROM), Ferroelectric Random Access Memory ("FRAM"), and MRAM, as well as other semiconductor elements capable of storing information. Each of these types of memory devices may have different configurations. For example, flash memory devices may be configured in a NAND or NOR configuration.

[0062] Memory devices may be formed from passive and / or active elements in any combination. By way of non-limiting example, passive semiconductor memory elements include ReRAM device elements, which in some embodiments include resistive switching memory elements such as anti-fuses, phase change materials, and optionally, steering elements such as diodes. Further by way of non-limiting example, active semiconductor memory elements include EEPROM and flash memory device elements, which in some embodiments include elements having charge storage regions such as floating gates, conductive nanoparticles, or charge storage dielectric materials.

[0063] The plurality of memory elements can be configured such that the plurality of memory elements are connected in series or such that each element is individually accessible. As a non-limiting example, flash memory devices within a NAND configuration (NAND memory) typically include memory elements connected in series. A NAND memory array can be configured such that the array is composed of a plurality of memory strings, and in the plurality of memory strings, a string is composed of a plurality of memory elements that share a single bit line and are accessed as a group. Alternatively, the memory elements can be configured such that each element is individually accessible, for example, configured as a NOR memory array. The NAND and NOR memory configurations are examples, and the memory elements may be configured otherwise.

[0064] Semiconductor memory elements located within and / or on a substrate may be arranged two-dimensionally (2D) or three-dimensionally (3D), such as in a 2D memory structure or a 3D memory structure.

[0065] In a 2D memory structure, the semiconductor memory elements are arranged in a single plane or at a single memory device level. Typically, in a 2D memory structure, the memory elements are arranged in a plane (e.g., an x-z plane) that extends substantially parallel to the main surface of the substrate that supports the memory elements. The substrate may be a wafer, and a layer of memory elements is formed on or within the wafer, or the substrate may be a carrier substrate that is attached to the memory elements after the memory elements are formed. As a non-limiting example, the substrate may include a semiconductor such as silicon.

[0066] The memory elements can be arranged at a single memory device level in an aligned array such as a plurality of rows and / or columns. However, the memory elements can be arranged in an irregular or non-orthogonal configuration. Each of the memory elements can have two or more electrodes, or contact lines such as bit lines and word lines.

[0067] A 3D memory array is arranged such that memory elements occupy multiple planes or multiple memory device levels, thereby forming a three-dimensional (i.e., the x, y, and z directions, where the y direction is substantially perpendicular to the main surface of the substrate, and the x and z directions are substantially parallel to the main surface of the substrate) structure.

[0068] As a non-limiting example, a 3D memory structure can be arranged vertically as a stack of multiple 2D memory device levels. As another non-limiting example, a 3D memory array can be arranged as a plurality of vertical columns (e.g., columns extending substantially perpendicular to the main surface of the substrate, i.e., in the y direction), where each column has multiple memory elements in each column. The columns may be arranged in a 2D configuration, e.g., in the x-z plane, resulting in a 3D arrangement of memory elements on multiple vertically stacked memory planes. Other configurations of three-dimensional memory elements can also build a 3D memory array.

[0069] As a non-limiting example, in a 3D NAND memory array, the memory elements can be coupled together to form NAND strings within a single horizontal (e.g., x-z) memory device level. Alternatively, the memory elements can be coupled together to form vertical NAND strings that traverse multiple horizontal memory device levels. Other 3D configurations can be contemplated, and in other 3D configurations, some NAND strings contain memory elements at a single memory level, while other strings contain memory elements spanning multiple memory levels. A 3D memory array may also be designed in a NOR configuration and a ReRAM configuration.

[0070] Typically, in a monolithic 3D memory array, one or more memory device levels are formed on a single substrate. Optionally, the monolithic 3D memory array may also have at least partially within a single substrate one or more memory layers. As a non-limiting example, the substrate may include a semiconductor such as silicon. In a monolithic 3D array, the layers that make up each of the memory device levels of the array are typically formed on top of the layer of the memory device level that is beneath the array. However, the layers of adjacent memory device levels of the monolithic 3D memory array may be shared or may have intervening layers between the memory device levels.

[0071] Similarly, a two-dimensional array may be formed separately and then packaged together to form a non-monolithic memory device having multiple memory layers. For example, a non-monolithic stacked memory can be constructed by forming memory levels on separate substrates and then stacking the memory levels on top of each other. The substrate may be thin or may be removed from the memory device levels prior to stacking, but since the memory device levels are initially formed on separate substrates, the resulting memory array is not a monolithic 3D memory array. Further, multiple 2D memory arrays or 3D memory arrays (monolithic or non-monolithic) may be formed on separate chips and then packaged together to form a stacked chip memory device.

[0072] Associated circuitry is typically required for the operation of the memory elements and communication with the memory elements. As a non-limiting example, a memory device may have circuitry for controlling and driving the memory elements to achieve functions such as programming and reading. This associated circuitry may be on the same substrate as the memory elements and / or on a separate substrate. For example, a controller for memory read and write operations may be located on a separate controller chip and / or on the same substrate as the memory elements.

[0073] The present invention is not limited to the 2D and 3D structures described, and it will be understood by those skilled in the art that it encompasses all relevant memory structures within the spirit and scope of the present invention, as described herein and as understood by those skilled in the art.

[0074] The above detailed description is intended to be understood as an illustration of selected forms the invention can take and is not intended to be understood as a definition of the invention. Only the following claims, which include all equivalents, are intended to define the scope of the claimed invention. Finally, note that any aspect of any of the embodiments described herein may be used alone or in combination with each other.

Claims

1. A memory system, comprising: a memory having a plurality of memory dies, each memory die including a respective circuit boundary array; a memory controller coupled to the memory and configured to, for each of the plurality of memory dies: read a word line within the memory die; determine a read threshold based on the read word line; and send the read threshold to the circuit boundary array within the memory die; wherein each circuit boundary array is configured to apply machine learning-based adjustment to the read threshold. A memory system, wherein each circuit boundary array is configured to apply machine learning-based adjustment to the read threshold.

2. The memory system according to claim 1, wherein the word lines read by the memory controller are part of a time and temperature group, and the circuit boundary array is further configured to apply the machine learning-based adjustment to the read threshold in response to the time and temperature group being read.

3. The memory system according to claim 2, wherein the memory controller is further configured to manage all time and temperature groups of the plurality of memory dies.

4. The memory system according to claim 3, wherein the memory controller is further configured to determine the read threshold in response to a time and temperature group update.

5. The memory system according to claim 1, wherein each circuit boundary array is further configured to apply the machine learning-based adjustment to the read threshold on-the-fly before reading a target word line.

6. The memory system according to claim 1, wherein at least one circuit boundary array uses a different machine learning algorithm than another one of the circuit boundary arrays.

7. The memory system according to claim 1, wherein at least one circuit boundary array uses a machine learning algorithm that is at least partially trained offline.

8. The memory system according to claim 1, wherein each circuit boundary array includes a respective memory chip and a respective separate complementary metal oxide semiconductor (CMOS) chip.

9. The memory system according to claim 1, wherein at least one of the circuit boundary arrays is a circuit array on (CAO) device.

10. The memory system according to claim 1, wherein at least one of the circuit boundary arrays is a circuit array under (CAU) device.

11. The memory includes a three-dimensional memory, and the storage system according to claim 1.

12. A storage system comprising a memory including a plurality of circuit boundary arrays (CBAs), each CBA including a memory die, the method comprising: In each CBA, Reading word lines within a time and temperature group in the memory die of the CBA; Determining a read threshold based on the read word lines; Storing the read threshold in the CBA; Applying machine learning-based adjustment to the read threshold, the method comprising.

13. The method according to claim 12, wherein each CBA applies the machine learning-based adjustment to the read threshold on the fly before reading a target word line.

14. The method according to claim 12, wherein each CBA comprises a respective memory die and a respective separate complementary metal oxide semiconductor (CMOS) chip coupled to the memory die.

15. The method according to claim 12, wherein each CBA determines the read threshold based on the read word lines in parallel with at least one other CBA.

16. The method according to claim 12, wherein at least one of the CBAs reads the word lines within the time and temperature group during an idle time of the memory die of the CBA.

17. The method according to claim 12, wherein at least one of the CBAs reads the word lines within the time and temperature group in response to receiving a command from a memory controller of the storage system.

18. The method according to claim 12, wherein at least one CBA uses a different machine learning algorithm than another one of the CBAs.

19. The method according to claim 12, wherein at least one CBA uses a machine learning algorithm that is trained at least partially offline.

20. A storage system, A memory comprising a circuit boundary array including a memory die, Means located within the circuit boundary array, Reading word lines in the memory die, Determining a read threshold based on the read word lines, Means for applying machine learning-based adjustment to the read threshold on the fly before a target word line in the memory die is read.

Citation Information

Patent Citations

  • Memory system and method

    JP2020144958A

  • Read integration time calibration for non-volatile storage device

    JP2022045317A

  • Read threshold management and calibration

    JP2022063210A

  • Memory system and method

    US20200285419A1

  • Non-volatile memory array driven from both sides for performance improvement

    US20200402587A1