Memory system and method for pre-die-rite due to detection of critical word line leak

The fatal wordline leak detection mode in memory systems addresses critical word line leaks by independently biasing word lines, preventing fatal failures and maintaining system performance by proactive block management.

JP7706028B2Active Publication Date: 2025-07-10SANDISK TECHNOLOGIES LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024568048
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-07-01
Filing Date
2023-05-26
Publication Date
2025-07-10
Estimated Expiration
2043-05-26

AI Technical Summary

Technical Problem

Existing memory systems fail to accurately detect and address critical word line leaks near the global control gate interface, leading to premature die retirement and performance degradation due to undetected fatal plane/die-level failures.

Method used

Implementing a fatal wordline leak detection (F-WLLD) mode that independently biases word lines to identify and proactively mark or transfer data from affected blocks, avoiding unnecessary die retirement and maintaining system performance.

Benefits of technology

Prevents catastrophic data loss by accurately identifying and managing critical word line leaks, reducing unnecessary die retirement, and minimizing performance degradation in high-capacity memory systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007706028000001
    Figure 0007706028000001
  • Figure 0007706028000002
    Figure 0007706028000002
  • Figure 0007706028000003
    Figure 0007706028000003
Patent Text Reader

Abstract

In some situations, leakage on a word line can be a local problem that causes data loss within the block containing the word line. In other situations, such as when leakage occurs near the peripheral word line routing area, the leakage can potentially affect the entire memory die. The memory system provided herein has a detector for critical word line leakage that determines the type of leakage and thus determines whether only the block should be retired or whether the associated blocks should be retired.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross - Reference to Related Applications) This application claims the benefit of U.S. Non - Provisional Application No. 17 / 856,073, entitled "Storage System and Method for Proactive Die Retirement by Fatal Wordline Leakage Detection", filed on July 1, 2022, the entire contents of which are incorporated herein by reference for all purposes.

Background Art

[0002] A single or multiple word - line short - circuits in a NAND memory array usually cause only data loss of several pages. During the factory test of the memory, the built - in self - test (BIST) leakage detection mode can be used to screen out leaky blocks and mark them as factory bad blocks (FBB). If any data loss occurs due to a word - line short - circuit in the field, the memory system can attempt to recover the user data. If it fails, the block can be retired as a grown bad block (GBB) to prevent future use. Some GBBs may deteriorate into global failures in later use, causing preemptive die retirement (PDR), which may affect performance.

Brief Description of the Drawings

[0003]

Figure 1A

Figure 1B

Figure 1C

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0004] The following embodiments generally relate to a memory system and method for proactive die retirement by detecting critical word line leaks. In one embodiment, a method implemented within a memory system comprising memory dies is provided. The method includes detecting a short circuit of word lines in a block within a memory die, determining whether the short circuit in the word lines affects only the block or affects the memory die, retiring the block without retiring the memory die in response to determining that the short circuit in the word lines affects only the block, and retiring the memory die in response to determining that the short circuit in the word lines affects the memory die.

[0005] In another embodiment, a memory system is provided that includes a memory die and a detector for detecting a critical word line leak. The detector for the critical word line leak determines whether a leak detected in a word line in a block within the memory die affects only the block or the entire memory die. In response to determining that the leak affects only the block, only the block is marked as defective and the other blocks within the memory die are made available for use. In response to determining that the leak affects the entire memory die, the entire die is marked as defective.

[0006] In yet another embodiment, a memory system is provided that includes a memory, means for detecting a short circuit in a word line in a block within the memory die, means for determining whether the short circuit in the word line affects only the block or the memory die, means for retiring the block without retiring the memory die in response to determining that the short circuit in the word line affects only the block, and means for retiring the memory die in response to determining that the short circuit in the word line affects the memory die.

[0007] Other embodiments are provided and may be used alone or in combination.

[0008] Referring now to the drawings, a memory system suitable for use in implementing aspects of these embodiments is shown in FIGS. 1A-1C. FIG. 1A is a block diagram showing a non-volatile memory system 100 (which may be referred to herein as a memory device or simply a device) according to one embodiment of the subject matter described herein. Referring to FIG. 1A, the non-volatile memory system 100 includes a controller 102 and a non-volatile memory that may be composed of one or more non-volatile memory dies 104. As used herein, the term die refers to an aggregate of non-volatile memory cells formed on a single semiconductor substrate and associated circuitry for managing the physical operation of these non-volatile memory cells. The controller 102 interfaces with a host system and transmits command sequences for read, program, and erase operations to the non-volatile memory die 104.

[0009] The controller 102 (the controller 102 may be a non-volatile memory controller (e.g., a flash, resistive random-access memory (ReRAM), phase-change memory (PCM), or magneto-resistive random-access memory (MRAM) controller)) can take the form of a processing circuit, a microprocessor or a processor, and a computer-readable medium. The computer-readable medium stores computer-readable program code (e.g., firmware) executable by, for example, a (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. The controller 102 can be composed of hardware and / or firmware for performing various functions described below and shown in the flow diagrams. Also, some of the components shown in the figure as being inside the controller may also be stored outside the controller, and other components may be used. Also, the phrase "operably communicate with" can mean "directly communicate with" or "indirectly (wired or wirelessly) communicate with via one or more components". However, the above components may or may not be illustrated and described in this specification.

[0010] As used herein, a non-volatile memory controller is a device that manages data stored in a non-volatile memory and communicates with a host such as a computer or an electronic device. In addition to the specific functions described herein, a non-volatile memory controller can have various functions. For example, a non-volatile memory controller can format the memory to ensure that the non-volatile memory is operating properly, map out defective non-volatile memory cells, and allocate spare cells to be replaced with future failing cells. A portion of the spare cells can be used to hold firmware for operating the non-volatile memory controller and implementing other features. In operation, when the host needs to read data from or write data to the non-volatile memory, the host can communicate with the non-volatile memory controller. If the host provides a logical address to which the data is to be read / written, the non-volatile memory controller can translate the logical address received from the host into a physical address within the non-volatile memory. (Alternatively, the host can provide a physical address.) The non-volatile memory controller can also perform various memory management functions including, but not limited to, wear leveling (to avoid wearing out specific memory cell blocks that would otherwise be repeatedly written to) and garbage collection (after a block is full, moving only valid data pages to a new block so that full blocks can be erased and reused). Also, the structure for the "means" recited in the claims can include, for example, some or all of the structures of the controllers described herein that are programmed or manufactured as necessary to cause the controller to operate to perform the recited functions.

[0011] The non-volatile memory die 104 can include any suitable non-volatile storage medium, including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells. The memory cells can be in the form of solid state (e.g., flash) memory cells and can be one-time programmable, multi-time programmable, or many-time programmable. The memory cells can also be single-level (1 bit per cell) cells (Single-Level Cell, SLC), or multi-level cells (Multiple-Level Cell, MLC), e.g., two-level cells, triple-level cells (Triple-Level Cell, TLC), quad-level cells (Quad-Level Cell, QLC). Alternatively, the memory cells can use other memory cell level technologies that are currently known or developed in the future. Also, the memory cells can be fabricated two-dimensionally or three-dimensionally.

[0012] The interface between the controller 102 and the non-volatile memory die 104 can be any suitable flash interface, such as toggle mode 200, 400, or 800. In one embodiment, the memory system 100 can be a card-based system, such as a Secure Digital (SD) or Micro Secure Digital (Micro SD) card (or USB, SSD, etc.). In an alternative embodiment, the memory system 100 can be part of an embedded memory system.

[0013] In the example shown in FIG. 1A, a non-volatile memory system 100 (which may be referred to herein as a memory module) includes a single channel between a controller 102 and a non-volatile memory die 104, but the subject matter described herein is not limited to having a single memory channel. For example, in some memory system architectures (such as those shown in FIGS. 1B and 1C), two, four, eight, or more memory channels may exist between the controller and the memory device, depending on the capabilities of the controller. In any of the embodiments described herein, even if a single channel is shown in the figure, more than one channel may exist between the controller and the memory die.

[0014] FIG. 1B shows a memory module 200 that includes a plurality of non-volatile memory systems 100. As shown in the figure, the memory module 200 may include a memory controller 202, and the memory controller 202 interfaces with a host and a memory system 204, and the memory system 204 includes a plurality of non-volatile memory systems 100. The interface between the memory controller 202 and the non-volatile memory system 100 may be a bus interface, for example, a Serial Advanced Technology Attachment (SATA), a Peripheral Component Interconnect express (PCIe) interface, or a Double-Data-Rate (DDR) interface. In one embodiment, the memory module 200 may be a Solid State Drive (SSD) or a Non-Volatile Dual In-line Memory Module (NVDIMM) as found in server PCs or portable computing devices such as laptop computers and tablet computers.

[0015] Figure 1C is a block diagram showing a hierarchical memory system. The hierarchical memory system 250 includes a plurality of memory controllers 202, and each of the plurality of memory controllers 202 controls an individual memory system 204. The host system 252 can access the memory in the memory system via a bus interface. In one embodiment, the bus interface may be a Non-Volatile Memory express (NVMe) or Fiber Channel over Ethernet (FCoE) interface. In one embodiment, the system shown in Figure 1C may be a rack-mountable mass storage system that can be accessed by a plurality of host computers, such as those found in a data center or other locations where a mass storage device is required.

[0016] Figure 2A is a block diagram showing the components of the controller 102 in more detail. The controller 102 includes a front-end module 108 that interfaces with the host, a back-end module 110 that interfaces with one or more non-volatile memory dies 104, and various other modules that perform the functions described in detail herein. The modules can take the form of, for example, a packaged functional hardware unit designed for use with other components, a portion of program code (e.g., software or firmware) executable by a (micro)processor or processing circuit that normally performs a specific function of the associated function, or a self-contained hardware or software component that interfaces with a larger system. The controller 102 may be referred to herein as a NAND controller or a flash controller, although it should be understood that the controller 102 can be used with any suitable memory technology, some examples of which are provided below.

[0017] Referring back to the modules of the controller 102, the buffer manager / bus controller 114 manages a buffer in the Random Access Memory (RAM) 116 and controls the internal bus arbitration of the controller 102. The Read Only Memory (ROM) 118 stores system startup code. Although shown in FIG. 2A as being located separately from the controller 102, in other embodiments, one or both of the RAM 116 and the ROM 118 may be located within the controller. In still other embodiments, portions of the RAM and the ROM may be located both within and outside of the controller 102.

[0018] The front-end module 108 includes a host interface 120 and a Physical Layer Interface (PHY) 122, and the host interface 120 and the PHY 122 provide an electrical interface with a host or the next-level storage controller. The choice of the type of the host interface 120 may depend on the type of memory being used. Examples of the host interface 120 include, but are not limited to, SATA, SATA Express, Serially Attached Small Computer System Interface (SAS), Fibre Channel, Universal Serial Bus (USB), PCIe, and NVMe. The host interface 120 typically facilitates transfers for data, control signals, and timing signals.

[0019] The back-end module 110 includes an Error Correction Code (ECC) engine 124, which encodes data bytes received from the host and decodes and corrects error data bytes read from the non-volatile memory. The command sequencer 126 generates command sequences such as a program command sequence and an erase command sequence to be sent to the non-volatile memory die 104. The Redundant Array of Independent Drive (RAID) module 128 manages the generation of RAID parity and the recovery of failed data. The RAID parity can be used as an additional level of integrity protection for the data written in the memory device 104. In some cases, the RAID module 128 may be a part of the ECC engine 124. The memory interface 130 provides the command sequence to the non-volatile memory die 104 and receives status information from the non-volatile memory die 104. In one embodiment, the memory interface 130 can be a double data rate (DDR) interface such as a toggle mode 200, 400, or 800 interface. The flash control layer 132 controls the overall operation of the back-end module 110.

[0020] The storage system 100 also includes other discrete components 140 such as an external electrical interface, an external RAM, a resistor, a capacitor, or other components that can interface with the controller 102. In an alternative embodiment, one or more of the physical layer interface 122, the RAID module 128, the media management layer 138, and the buffer management / bus controller 114 are optional components that are not necessary within the controller 102.

[0021] FIG. 2B is a block diagram showing the components of the non-volatile memory die 104 in more detail. The non-volatile memory die 104 includes a peripheral circuit 141 and a non-volatile memory array 142. The non-volatile memory array 142 includes non-volatile memory cells used to store data. The non-volatile memory cells may be any suitable non-volatile memory cells including ReRAM, MRAM, PCM, NAND flash memory cells, and / or NOR flash memory cells in two-dimensional and / or three-dimensional configurations. The non-volatile memory die 104 further includes a data cache 156 that caches data. The peripheral circuit 141 includes a state machine 152 that provides state information to the controller 102.

[0022] Referring again to FIG. 2A, the flash control layer 132 (referred to herein as the Flash Translation Layer (FTL), or more generally, as the “media management layer” when the memory may not be flash) processes flash errors and interfaces with the host. In particular, the FTL, which can be an algorithm within the firmware, is involved in the internal memory management and converts writes from the host into writes into the memory 104. The memory 104 may have limited durability, may only be written within a plurality of pages, and / or may not be written unless the memory 104 is erased as a block of memory cells, so the FTL may be required. The FTL understands these potential limitations of the memory 104 and may not be visible to the host. Thus, the FTL attempts to convert writes from the host into writes into the memory 104.

[0023] The FTL may include a logical-to-physical (L2P) map (which may be referred to herein as a table or data structure) and an allocated cache memory. In this way, the FTL converts a logical block address ("LBA") from the host to a physical address within the memory 104. The FTL can include other features such as power-off recovery (for which the FTL's data structure can be recovered in the event of a sudden power loss), and wear leveling (for which the wear across the memory blocks is uniform to prevent excessive wear of a certain block that can lead to a greater chance of failure), but is not limited thereto.

[0024] Referring to the drawings again, FIG. 3 is a block diagram of a host 300 and a memory system 100 (which may be referred to herein as a device) according to one embodiment. The host 300 can take any suitable form including, but not limited to, a computer, a mobile phone, a digital camera, a tablet, a wearable device, a digital video recorder, a surveillance system, etc. The host 300 includes a processor 330, which is configured to send data (initially stored, for example, in the host's memory 340 (such as DRAM)) to the memory system 100 for storage in the memory 104 (such as a non-volatile memory die) of the memory system. Note that although the host 300 and the memory system 100 are shown as separate boxes in FIG. 3, the memory system 100 may be integrated within the host 300, the memory system 100 may be removably connected to the host 300, and the memory system 100 and the host 300 can communicate via a network. Note also that the memory 104 may be integrated within the memory system 100 or may be removably connected to the memory system 100.

[0025] As described above, a single or multiple word line shorts in a NAND memory array usually only cause data loss of several pages. During the factory test of the memory, the built-in self-test (BIST) leak detection mode can be used to screen for leaky blocks and mark them as factory bad blocks (FBBs). If any data loss occurs due to a word line short in the field, the memory system can attempt to recover the user data. If this fails, the block can be retired as a growth bad block (GBB) to prevent future use.

[0026] However, unlike word line shorts inside the array, word line shorts in the peripheral word line routing area can cause a fatal plane / die-level failure during the life of the memory. More specifically, if the word line short is close to the global control gate interface (CGI) area, a fatal plane / die-level data loss can occur even if the shorted block was marked as an FBB during the factory test or was retired as a GBB by the memory system in the field. This is because the CGI is a global signal to the transfer gates of the local word lines and is biased during normal block operation, so the defective area may still be stressed during user operation.

[0027] Figures 4 and 5 show a scenario in which stress is applied to the local word line and CGI during user erasure of a good block. In this scenario, neither block is selected, but they share the same CGI as the selected block. The FBB had an SGS-WL short circuit at M1. In this case, the CGI is biased as an isolation voltage (VISO) (e.g., a very low voltage such as about 0.5V), and the local word line of the bad block is coupled to a verification voltage (VERA) (e.g., a very high voltage such as about 18V) from the memory hole. Therefore, over time, the short circuit may grow towards the CGI side and ultimately lead to a global CGI short circuit.

[0028] In the case of high-capacity products such as enterprise memory systems, a die with a large number of GBBs is preemptively retired if the number of GBBs exceeds a criterion, typically 30 for each type of failure (erase / program / read failure, GBB due to aging retirement, etc.). Generally, a die with a high number of GBBs means a high defect density. Therefore, in order to reduce the risk of drive failure, a die with strong defects is preemptively retired. Thus, when a global CGI short circuit occurs on an enterprise product, XOR can be triggered first for data recovery, and then the die can be retired to avoid future use according to the preemptive die retirement (PDR) function. However, this can cause some problems. For example, since all valid data from the retired die needs to be moved to a good die, there may be a temporary performance degradation due to the strong system overhead from XOR recovery and due to data rearrangement. As another example, a permanent overprovisioning (OP) loss may occur, which can reduce the random write (RW) performance and increase the end-of-life (EOL) program / erase (P / E) requirements for other dies. Therefore, unnecessary PDR in such cases can bring undesirable system impacts as described above.

[0029] Common test modes for array leak detection do not specify which word lines should be measured, as leakage is measured for either all word lines or any of the word lines within a group (e.g., by even / odd word lines or by different drivers). This leads to a dilemma where, although it can be quickly determined that a block is leaking by measuring all word lines or all grouped word lines, there is no opportunity to know which word line is leaking. This is particularly true for the M1 defect described above, where instead of retiring a single block or marking it as FBB, the die may be retired during customer use or rejected during factory testing.

[0030] Some test modes can specify the bias for each CG driver. However, the defective modes described here still have unresolved problems: (1) leakage detected between different CG groups cannot be avoided, and (2) it is difficult to determine a short circuit between two specific CG groups. Moreover, such modes may become invalid when two shorted word lines are from the same CG group. Therefore, there is no good way to detect a specific word line short circuit and retire / reject the die when the short circuit is located in a critical area such as near the global CGI contact.

[0031] The following embodiments can be used to address this problem. In one embodiment, a new algorithm is provided that uses fatal wordline leak detection (F-WLLD). The F-WLLD mode can be used to accurately detect potential CGI shorts, such that instead of retiring an entire die, the associated blocks can be proactively marked as bad (since one global CGI short only affects one common CGI block). After a GBB event, to avoid memory performance penalties, F-WLLD can be run on the GBB during system background time. If F-WLLD fails, all valid data from one common CGI block is transferred to other good blocks, which are then retired to avoid future use. Thus, even a GBB with a fatal wordline leak does not degrade into a large-scale GBB event. Thus, unnecessary PDR is avoided. This feature can be used in any suitable memory and may be particularly desirable for high-capacity memory products such as enterprise storage systems. For example, the loss of one or more dies may be acceptable for a high-capacity (e.g., 32 / 64 / 128 die) drive still within its lifespan, depending on the system performance degradation specification. Thus, proactively retiring a die can effectively avoid catastrophic plane / die-level data loss for the system if a fatal wordline leak is detected on the GBB.

[0032] FIG. 6 is a flowchart 600 of a proactive die retirement method according to an embodiment. As shown in FIG. 6, a memory system 100 (here, an enterprise storage system (ESS), although any type of memory system can be used) enables its proactive die retirement function (operation 610). After a GBB event (operation 620), a critical word line leak detection F-WLLD mechanism is executed against the GBB (e.g., during system background time to avoid NAND performance penalties) (operations 630, 640, 650). Next, a determination is made to check whether the F-WLLD has failed (operation 660). If the F-WLLD has failed, data is transferred from the defective die to another good die, and the die is retired from future use (operation 670). In this way, a GBB with a critical word line leak does not deteriorate to a plane / die-level failure. However, if the F-WLLD has succeeded, the die is retained and the GBB is retired (operation 680).

[0033] FIG. 7 is a schematic diagram showing a high voltage switch (HVSW) according to an embodiment. As shown in FIG. 7, the HVSW of this embodiment includes an SG decoding module, an XY decoding module, a zone decoding module, a chunk decoding module, a tier decoding module, and an edge / dummy decoding module. In the normal WLLD mode, the voltage is passed to the SG and the data and dummy word lines through different CG drivers. However, as described above, this may not provide the flexibility of independent bias for each word line. Therefore, in this embodiment, in the F-WLLD mode, the XY decoding circuit is enabled. Here, CGX and CGY are CG drivers for providing different biases to the word lines WL1-110 in the F-WLLD mode. In this mode, all HVSW gate signals for zone, chunk, and tier decoding are off, and only XY decoding, SG decoding, and edge / dummy word line decoding are operating. In XY decoding, G_CGX_SW and G_CGY_SW on each word line can be independently turned on or off based on the input word line address for high and low biases.

[0034] FIG. 8 is a flowchart 800 showing this process. As shown in FIG. 8, in response to an F-WLLD mode trigger (operation 805), all HVSW gates are closed (operation 810), and WLm is input for high bias (operation 815). Next, a determination is made as to whether it is an SG, dummy word line, or edge word line (operation 820). If so, the corresponding SG, dummy word line, or edge word line HVSW gate is turned on (operation 825). If not, G_CGX_SW <m>is turned on (operation 830), and the input WLn is set to a low bias (operation 835). Next, a determination is made as to whether it is an SG, a dummy word line, or an edge word line (operation 840). If so, the HVSW gate of the corresponding SG / dummy / edge WL is turned on (operation 855), and the CGX, CGY, SG / dummy / edge WL bias is set directly to the SIN mode (operation 850). If not, G_CGX_SW <n>is turned on (operation 845), and the CGX, CGY, SG / dummy / edge WL biases are directly set to the SIN mode (operation 850). Thus, in the F-WLLD mode, only two given word lines including the SG word line and the edge / dummy word line receive high and low biases from the XY decoder and the SG decoder and the edge / dummy decoder, respectively. All other unselected word lines are in a floating state.

[0035] Considering the variations in metal routing for different generations, the F-WLLD mode is also layout-adaptive. For example, the M1 word line pair running near the CGI contact is a high-risk word line, and thus any die having those shorted word line pairs can be retired / rejected. In some layouts, only four word line pairs are adjacent to the global CGI. By looping all four word line pairs in the F-WLLD mode, the risk of catastrophic plane / die-level data loss due to global CGI short circuit is significantly suppressed. In the case of other layouts, the word line routing near the CGI may be different, but the word line pairs for the F-WLLD mode can be pre-stored in the firmware after checking the specific product layout. This is shown in the flowchart 900 of FIG. 9.

[0036] As shown in FIG. 9, after the product layout is checked (operation 905), N pairs (any positive integer) of dangerous word line pairs can be pre-stored in the firmware (operation 910). After a GBB event occurs (operation 915), and after data recovery and system background time (operations 920, 925), an F-WLLD occurs on the nth word line pair on the GBB (operation 930). Next, a determination is made as to whether the F-WLLD has failed (operation 935). If the F-WLLD has failed, data is transferred and the die is retired (operation 940). If the F-WLLD has not failed, a determination is made as to whether n = N (operation 945). If so, the die is retained and the GBB is retired (operation 950). If not, n is incremented by 1 only (operation 955), and the method loops back to operation 930.

[0037] There are several advantages associated with these embodiments. For example, these embodiments present a new feature of proactive die retirement that uses critical word line leak detection, which can avoid critical plane / die level data loss. Also, this new GBB management algorithm that uses critical word line leak detection can avoid unnecessary PDR for enterprise memory systems, which benefits the system. For example, these embodiments can limit the impact on performance by proactively retiring related blocks without data loss. Thus, time-consuming heavy XOR recovery is not required. Further, only a portion of the block data needs to be reallocated, which has a much lower granularity compared to current PDR designs. Additionally, in some situations, these embodiments can provide savings of over 80% in overprovisioning loss. These embodiments can retire only one common CGI block, so most array blocks can still be retained for overprovisioning or user capacity. This significant savings in overprovisioning loss can help mitigate read / write performance degradation and reduce the extra erase / program cycle requirements for other good dies. Also, these advantages can be realized without sacrificing NAND performance (since it can operate during the system's background time) and with a negligible increase in die size.

[0038] Finally, as described above, any suitable type of memory may be used. Semiconductor memory devices include volatile memory devices such as Dynamic Random Access Memory ("DRAM"), Static Random Access Memory ("SRAM") devices, ReRAM, Electrically Erasable Programmable Read Only Memory ("EEPROM"), flash memory (which may also be considered a subset of EEPROM), Ferroelectric Random Access Memory ("FRAM"), and MRAM, as well as other semiconductor elements capable of storing information. Each of these types of memory devices may have different configurations. For example, flash memory devices may be configured in a NAND or NOR configuration.

[0039] Memory devices may be formed from passive and / or active elements in any combination. As non-limiting examples, passive semiconductor memory elements include ReRAM device elements, which in some embodiments include resistive switching memory elements such as anti-fuses, phase change materials, and optionally, steering elements such as diodes. Further non-limiting examples of active semiconductor memory elements include EEPROM and flash memory device elements, which in some embodiments include elements having charge storage regions such as floating gates, conductive nanoparticles, or charge storage dielectric materials.

[0040] The plurality of memory elements may be configured such that the plurality of memory elements are connected in series, or such that each of the plurality of memory elements is individually accessible. As a non-limiting example, flash memory devices within a NAND configuration (NAND memory) typically include memory elements connected in series. A NAND memory array may be configured such that the array is composed of a plurality of memory strings, where in the plurality of memory strings, a string is composed of a plurality of memory elements that share a single bit line and are accessed as a group. Alternatively, the memory elements may be configured such that each of the elements is individually accessible, for example, may be configured as a NOR memory array. The NAND and NOR memory configurations are examples, and the memory elements may be configured otherwise.

[0041] Semiconductor memory elements located within and / or on a substrate may be arranged two-dimensionally (2D) or three-dimensionally (3D), such as in a two-dimensional (two Dimensional, 2D) memory structure, a three-dimensional (three Dimensional, 3D) memory structure, and the like.

[0042] In a 2D memory structure, the semiconductor memory elements are arranged in a single plane or at a single memory device level. Typically, in a 2D memory structure, the memory elements are arranged in a plane (e.g., an xz-direction plane) that extends substantially parallel to the main surface of the substrate that supports the memory elements. The substrate may be a wafer, and a layer of memory elements is formed on or within the wafer, or the substrate may be a carrier substrate that is attached to the memory elements after the memory elements are formed. As a non-limiting example, the substrate may include a semiconductor such as silicon.

[0043] The memory elements may be arranged at a single memory device level in an aligned array such as a plurality of rows and / or columns. However, the memory elements may be arranged in an irregular or non-orthogonal configuration. Each of the memory elements may have two or more electrodes, or contact lines such as bit lines and word lines.

[0044] The 3D memory array is arranged such that memory elements occupy multiple planes or multiple memory device levels, thereby forming a three-dimensional (i.e., in the x, y, and z directions, where the y direction is substantially perpendicular to the main surface of the substrate and the x and z directions are substantially parallel to the main surface of the substrate) structure.

[0045] As a non-limiting example, the 3D memory structure can be arranged vertically as a stack of multiple 2D memory device levels. As another non-limiting example, the 3D memory array can be arranged as a plurality of vertical columns (e.g., columns extending substantially perpendicular to the main surface of the substrate, i.e., in the y direction), each column having a plurality of memory elements in each of the columns. The columns may be arranged in a 2D configuration, e.g., in the xz plane, resulting in a 3D arrangement of memory elements on a plurality of vertically stacked memory planes. Other configurations of three-dimensional memory elements can also build a 3D memory array.

[0046] As a non-limiting example, in a 3D NAND memory array, the memory elements can be coupled together to form NAND strings within a single horizontal (e.g., xz) memory device level. Alternatively, the memory elements can be coupled together to form vertical NAND strings that traverse across multiple horizontal memory device levels. Other 3D configurations can be contemplated, and in other 3D configurations, some NAND strings contain memory elements at a single memory level and other strings contain memory elements across multiple memory levels. The 3D memory array may also be designed in a NOR configuration and a ReRAM configuration.

[0047] Typically, in a monolithic 3D memory array, one or more memory device levels are formed on a single substrate. Optionally, the monolithic 3D memory array may also have at least partially within a single substrate one or more memory layers. By way of non-limiting example, the substrate may include a semiconductor such as silicon. In a monolithic 3D array, the layers that make up each of the memory device levels of the array are typically formed on top of the layer of the memory device level beneath the array. However, the layers of adjacent memory device levels of a monolithic 3D memory array may be shared or may have intervening layers between the memory device levels.

[0048] Similarly, a two-dimensional array may be formed separately and then packaged together to form a non-monolithic memory device having multiple memory layers. For example, a non-monolithic stacked memory can be constructed by forming memory levels on separate substrates and then stacking the memory levels on top of each other. The substrate may be thin or may be removed from the memory device levels prior to stacking, but since the memory device levels are initially formed on separate substrates, the resulting memory array is not a monolithic 3D memory array. Further, multiple 2D memory arrays or 3D memory arrays (monolithic or non-monolithic) may be formed on separate chips and then packaged together to form a stacked chip memory device.

[0049] Associated circuitry is typically required for the operation of the memory elements and communication with the memory elements. By way of non-limiting example, a memory device may have circuitry for controlling and driving the memory elements to achieve functions such as programming and reading. This associated circuitry may be on the same substrate as the memory elements and / or on a separate substrate. For example, a controller for memory read and write operations may be located on a separate controller chip and / or on the same substrate as the memory elements.

[0050] The present invention is not limited to the 2D and 3D structures described, and it will be understood by those skilled in the art that it encompasses all relevant memory structures within the spirit and scope of the present invention as described herein and as understood by those skilled in the art.

[0051] The foregoing detailed description is to be understood as illustrative of selected forms the invention can take and is not intended to be construed as a definition of the invention. Only the following claims, including all equivalents, are intended to define the scope of the claimed invention. Finally, note that any aspect of any of the embodiments described herein may be used alone or in combination with each other.< / n> < / m>

Claims

1. A method in a memory system comprising a memory die, comprising: detecting a growth bad block (GBB) in the memory die in response to detecting a word line short circuit; performing critical word line leak detection on the detected growth bad block by independently biasing each word line in the detected growth bad block to identify the location of the word line short circuit, wherein in performing the critical word line leak detection, at any given time, only two word lines receive low and high biases through the XY decode circuit of the high voltage switch, and all other word lines are in a floating state; identifying that the location of the word line short circuit is within the memory array of the memory die, and in response to thus identifying a successful critical word line leak detection, transferring data from the detected growth bad block; retiring the detected growth bad block without retiring the memory die; identifying that the location of the word line short circuit is within the peripheral word line routing area of the memory die that is susceptible to control gate interface (CGI) short circuits, and in response to thus identifying a failed critical word line leak detection, transferring data from the detected growth bad block and other blocks within the memory die; retiring the memory die; A method comprising the above steps.

2. The method according to claim 1, wherein the critical word line leak detection is performed during system background time.

3. The method according to claim 1, wherein the critical word line leak detection is performed as part of an embedded self-test.

4. The method according to claim 1, further comprising attempting to recover data lost in the detected growth bad block.

5. The method according to claim 1, wherein the memory die comprises a three-dimensional memory.

6. A memory system, comprising: a memory die; a detector for critical word line leaks, the detector being configured to: detect a growth bad block (GBB) in the memory die in response to detecting a word line short circuit; To identify the location of the word line short circuit, by independently biasing each word line in the detected defective growth block, perform a critical word line leak detection on the detected defective growth block. In the implementation of the critical word line leak detection, at any given time, only two word lines receive low bias and high bias through the XY decode circuit of the high voltage switch, and all other word lines are in a floating state. Identify that the location of the word line short circuit is within the memory array of the memory die, and in response to identifying the success of the critical word line leak detection thereby, Transfer data from the detected defective growth block, Retire the detected defective growth block without retiring the memory die, Identify that the location of the word line short circuit is within the peripheral word line routing area of the memory die that is susceptible to the influence of a control gate interface (CGI) short circuit, and in response to identifying the failure of the critical word line leak detection thereby, Transfer data from the detected defective growth block and other blocks within the memory die, A memory system configured to retire the memory die.

7. The memory die includes a 3D memory, the memory system according to claim 6.

8. The memory system includes an enterprise memory system comprising a plurality of memory dies, the memory system according to claim 6.

9. A memory system, comprising: A memory die; and Means, wherein the means: Detect a defective growth block (GBB) within the memory die in response to detecting a word line short circuit; To identify the location of the word line short circuit, by independently biasing each word line in the detected defective growth block, perform a critical word line leak detection on the detected defective growth block. In the implementation of the critical word line leak detection, at any given time, only two word lines receive low bias and high bias through the XY decode circuit of the high voltage switch, and all other word lines are in a floating state; Identify that the location of the word line short circuit is within the memory array of the memory die, and in response to identifying the success of the critical word line leak detection thereby, Transfer data from the detected defective growth block, Retire the detected underperforming block without retiring the memory die, Identify that the word line short circuit location is within the peripheral word line routing area of the memory die that is susceptible to the influence of a control gate interface (CGI) short circuit, and in response to identifying this as a fatal word line leak detection failure, Transfer data from the detected underperforming block and other blocks within the memory die, A memory system that retires the memory die.

Citation Information

Patent Citations

  • Semiconductor memory device

    JP1987262162A

  • Nonvolatile semiconductor memory device

    JP2012146369A

  • Defective block management

    US20140321202A1

  • Behavior-driven die management on solid-state drives

    US20210182188A1