Error detection of activated pages in memory device

By introducing lightweight error detection circuits and counter mechanisms into memory devices, the problem of low probability of error detection and mitigation in emerging memory technologies is solved, and the reliability and security of the memory is improved.

CN120407245APending Publication Date: 2025-08-01MICRON TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510106201.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-07-24
Filing Date
2025-01-23
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Emerging memory technology Due to the increase in the number of pages and the decrease in hammer threshold, the probability of error detection and mitigation is reduced. It is difficult for existing methods to effectively detect and correct errors in memory.

Method used

The lightweight error detection circuit system and error correction circuit system are used to detect errors in the memory array through parity and correct errors during memory management operations, and track memory management commands using counters to trigger a clear operation.

Benefits of technology

The detection probability of hammering pages is improved, the reliability and security of the memory device are enhanced, the time delay for error relief is reduced, and the reliability and life of the memory device are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407245A_ABST
    Figure CN120407245A_ABST
Patent Text Reader

Abstract

The invention relates to error detection of activated pages in a memory device. Systems, methods, and devices for detecting errors in data accessed in a memory array. In one method, a memory device includes error detection circuitry, error correction circuitry, and a controller. The controller accesses a portion of the memory array. The error detection circuitry determines whether there is an error in the accessed portion. An error is detected by comparing the parity stored in each portion with a parity calculated for all data stored in the portion when accessed. If an error is detected for a portion, an address of the portion is stored in a clear queue for later correction using the error correction circuitry.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 548,589, filed Feb. 1, 2024, the entire disclosure of which is hereby incorporated by reference herein. Technical Field

[0003] At least some embodiments disclosed herein generally relate to memory devices, and more particularly but not limited to, error detection when accessing data stored in a memory. Background Art

[0004] A memory device may include semiconductor circuitry for electronic storage that provides data to a host system, such as a server or other computing device. The memory device may be volatile or non-volatile. Volatile memory requires power to maintain data and includes devices such as random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), or synchronous dynamic random access memory (SDRAM). Non-volatile memory may retain stored data when not powered and includes devices such as flash memory, read-only memory (ROM), electrically erasable programmable ROM (EEPROM), erasable programmable ROM (EPROM), resistive change memory (such as phase change random access memory (PCRAM)), resistive random access memory (RRAM), and magnetoresistive random access memory (MRAM).

[0005] A host system (i.e., a host device) may include a host processor, a first quantity of host memory for supporting the host processor (such as main memory, typically volatile memory, such as DRAM), and one or more storage systems that provide additional storage devices to hold data in addition to or independent of the main memory (e.g., non-volatile memory, such as flash memory).

[0006] For example, a storage system such as a solid state drive (SSD) may include a memory controller and one or more memory devices, including several (e.g., multiple) dies or logical units (LUNs). In a particular instance, each die may include several memory arrays and peripheral circuitry thereon, such as die logic or a die processor. The memory controller may include interface circuitry configured to communicate with a host device (e.g., a host processor or interface circuitry) via a communication interface (e.g., a bidirectional parallel or serial communication interface). The memory controller may receive, for example, commands or operations associated with memory operations or instructions from the host, such as read or write operations to transfer data (e.g., user data and associated integrity data, such as error data or address data, etc.) between the memory device and the host device, an erase operation to erase data from the memory device, perform drive management operations (e.g., data migration, garbage collection, block retirement), etc.

[0007] Many memory devices (specifically non-volatile memory devices, such as NAND flash devices, etc.) frequently relocate data or otherwise manage data in the memory device (e.g., garbage collection, wear leveling, drive management, etc.). NAND flash memory is a type of flash memory constructed using NAND logic gates. Alternatively, NOR flash memory is a type of flash memory constructed using NOR logic gates.

[0008] Volatile memory devices such as DRAM typically refresh the stored data. For example, a refresh is to activate a row and then precharge the row. At the activation time, the data in the cell is sensed (implicit read), and at the precharge time, the data is written back to the cell (implicit write).

[0009] A storage device has a controller that receives data access requests from a host computer and performs programmed computational tasks to implement the requests in a manner that may be specific to the media and architecture configured in the storage device. In one instance, a flash memory controller manages data stored in a flash memory and communicates with a computing device. In some cases, flash memory controllers are used in solid state drives for mobile devices or in SD cards or similar media for digital cameras.

[0010] Firmware may be used to operate the flash memory controller of a particular storage device. In one instance, when a computer system or device reads data from or writes data to a flash memory device, it communicates with the flash memory controller. SUMMARY OF THE INVENTION

[0011] In one aspect, the present disclosure provides an apparatus comprising: error detection circuitry; and at least one controller configured to: access a portion of a memory array; use the error detection circuitry to detect whether an error exists in data stored in each accessed portion; and in response to detecting a first error in a first portion, record a first location of the first portion.

[0012] In another aspect, the present disclosure provides an apparatus comprising: a counter; and at least one controller configured to: maintain a record of rows in which an error has been detected in stored data when a row of a memory array is activated; use the counter to count a number of memory management operations of the memory array; and correct the detected error of the recorded rows based on the number.

[0013] In another aspect, the present disclosure provides a method comprising: in response to receiving an activation command, sensing a first data page and a parity stored in the first page; calculating a first parity of the first page; and comparing the first parity with the stored parity. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings, in which like reference numerals indicate similar elements.

[0015] Figure 1 Illustrates various mathematical relationships associated with hammer detection and mitigation.

[0016] Figure 2 Illustrates an exemplary graph of hammer detection and mitigation probability versus a hammer threshold.

[0017] Figure 3 Illustrates a memory device having error detection circuitry for detecting errors in an accessed portion of a memory array, according to some embodiments.

[0018] Figure 4 Illustrates circuitry for detecting errors in a page accessed in a memory array, according to some embodiments.

[0019] Figure 5 Illustrates a sense amplifier latch for holding data associated with memory cells of a memory array, according to some embodiments.

[0020] Figure 6 Illustrates a flowchart for error detection and recording of an address associated with a detected error for future scrubbing, according to some embodiments.

[0021] Figure 7An example is shown of comparing a calculated parity with a stored parity to determine if an error has been detected.

[0022] Figure 8 A exemplary method for calculating parity is shown.

[0023] Figure 9 A method for detecting errors in an accessed page and recording the errors for later correction according to some embodiments is shown. Detailed Description

[0024] The following disclosure describes various embodiments for performing error detection when accessing data stored in a memory. Errors detected in the data are recorded for future correction (e.g., during memory management). In one example, the accessed data is a page, and the detected errors are stored in a queue for cleaning. At least some embodiments herein relate to a non-volatile memory device that includes a cleaning queue for storing physical addresses of pages with detected errors. The memory device may store, for example, data used by a host device (such as a computing device of an autonomous vehicle or another computing device accessing data stored in the memory device). In one example, the memory device is a solid state drive installed in an electric vehicle.

[0025] Memory devices may suffer from soft errors caused by activities during memory device operation. For example, DRAM and certain emerging memory devices (e.g., ferroelectric RAM devices) may exhibit row hammer, where repeated activation of rows causes one or more soft errors to occur. This may result in an incorrect data value not being read from a row when the row is accessed during a read operation. An error occurs when the number of activations exceeds a hammering threshold.

[0026] In some cases, to mitigate errors, the memory device may perform an operation such as ECC cleaning after a page has caused a large number of activations equal to the hammering threshold. In existing methods, the physical page addresses associated with the activations may be randomly sampled in an attempt to capture (i.e., detect) the hammered page addresses.

[0027] Figure 1 Various mathematical relationships associated with hammering detection and mitigation are shown. As can be seen in relationships 102 and 106, the detection and mitigation probabilities are proportional to the hammering threshold and inversely proportional to the number of pages. As can be seen in equation 104, the number of pages hammered in parallel is equal to the number of pages divided by the hammering threshold.

[0028] Figure 2Exemplary graph showing the hammering detection and mitigation probability versus the hammering threshold. For a satisfactory probability 202, the hammered page address is cleared so that errors from row hammering are corrected. To meet the reliability requirements, a satisfactory probability is met to ensure that it is highly likely to detect all hammered rows and appropriate mitigation operations occur (e.g., to avoid failure to read data). This results in a hammering threshold associated with the satisfactory detection and mitigation probability 202. A hammering threshold less than this value will not result in a satisfactory detection and mitigation probability. The result could be a failure of the correct memory device operation.

[0029] It has been observed that emerging memory technologies (e.g., emerging non-volatile memory devices) sometimes require more activation power than DRAM. To reduce the power associated with activation, the page size can be reduced. However, the reduced page size results in larger row address entries and fewer column entries, which results in an increase in the number of addressable pages. The number of addressable pages determines the number of pages that can be hammered in parallel to reach the hammering threshold. This is illustrated by Figure 1 Equation 104.

[0030] Emerging memory technologies can suffer from particularly low detection and mitigation probabilities due to both a reduced hammering threshold and an increased number of pages. To address the above technical problems, despite the increased number of pages and the reduced hammering threshold, it is still necessary to increase the probability of detecting hammered pages.

[0031] Various embodiments of the present disclosure provide technical solutions to one or more of the above technical problems. In one embodiment, a memory device includes an error detection circuitry (e.g., an even / odd parity circuitry) and an error correction circuitry (e.g., an ECC engine). A controller of the memory device accesses a portion of a memory array in response to a command from a host device. The controller uses the error detection circuitry to detect whether an error exists in the data stored in each accessed portion (e.g., an entire data page on an activated row of the memory array). The error is detected by comparing the parity stored in the accessed portion with the parity calculated at the time of access.

[0032] In response to detecting an error in any accessed portion, the location of the accessed portion is stored in a flush queue for future correction. The error correction circuitry corrects the error detected for the location (e.g., the physical address of a page) in the flush queue by using the error correction circuitry. The error is corrected at a later time after the detection (e.g., during a memory management operation).

[0033] In one embodiment, the memory device implements an on-die ECC scheme to detect and correct errors on a given codeword. Multiple codewords may reside on a page. The latency of the codeword ECC engine is typically too large such that it is not practical to run every codeword through the ECC to determine if there are errors across the page.

[0034] When a page is hammered, activity-induced soft errors are expected. The memory device implements a lightweight error detection scheme to detect if the page has errors when it is activated.

[0035] As used herein, "lightweight" indicates that the functionality of the error detection scheme is less than that of the ECC engine. The lightweight functionality is used to reduce the time required to perform error detection on all data stored in a page when accessing the page.

[0036] In one example, the lightweight error detection scheme consists of an even / odd parity scheme for detecting single-bit errors. An even / odd parity is stored for each page, which requires storing one bit per page.

[0037] The lightweight error detection latency is less than the codeword ECC engine latency. The number of bits that the codeword ECC engine can correct is greater than or equal to the number of bit errors that the lightweight error detection scheme can detect.

[0038] If an error is detected by the lightweight error detection scheme during activation, then the physical page address associated with the activation is recorded (e.g., stored in a flush queue) for future activity-based ECC flushing. During the ECC flushing of a page, each codeword residing on the page is passed through the ECC engine. Any correctable errors are corrected and written back to the same page after the ECC flush.

[0039] In one embodiment, the activity-based memory management operations of the controller will trigger a wear leveling operation or a flush operation. In one example, 95% of the memory management (MM) operations are for wear leveling, while 5% of the memory management operations are for flushing the addresses in the flush queue.

[0040] In one embodiment, the memory device uses a counter to count memory management commands. The counter keeps track of the number of activity-based memory management (MM) commands issued. In one example, a flush is performed when the counter reaches a threshold, and then the counter is reset. In one embodiment, the threshold for performing a flush may be randomized. This can help improve the security of the memory device. In one embodiment, other mitigation operations may be performed in addition to and / or instead of flushing.

[0041] In one example, the scavenge steal ratio is based on the number of wear leveling operations divided by the number of scavenge operations. The scavenge steal ratio determines how often memory management commands are stolen to perform scavenging. The controller steals commands by performing scavenging instead of the operation originally indicated to the controller by the memory management command.

[0042] The scavenge queue stores physical page addresses that have been marked as containing errors by a lightweight error detection scheme. When the memory device is powered on, the counter and the scavenge queue are reset.

[0043] In one embodiment, the controller receives various commands during operation. The commands include, for example, activation and precharge commands.

[0044] When an activation is issued, the controller senses the page, which causes the data in the page to reside in the sense amplifier latch. The controller then calculates an error detection parity. If the parity stored in the page does not match the calculated parity, the physical page address is added to the scavenge queue.

[0045] When a precharge command is issued, the controller calculates an error detection parity. The page data including the calculated parity is written to the memory cells of the memory array.

[0046] When an activity-based memory management command (e.g., a wear leveling command or a refresh command) is issued, the controller determines whether the counter has reached a value corresponding to the scavenge steal ratio. If the memory management (MM) count of the counter is equal to the scavenge steal ratio, the physical page addresses from the scavenge queue are scavenged. In one example of such scavenging, each codeword of the page is passed through the codeword ECC engine, and the corrected data is written back to the physical page address. The MM count is then reset to zero.

[0047] If the memory management (MM) count is not equal to the scavenge steal ratio, the controller performs a wear leveling operation and increments the MM count.

[0048] In one embodiment, the memory device implements a page error detection scheme to detect whether a page has an error when it is activated, which will cause the address associated with the activation to be recorded in a queue for future scavenging by the codeword ECC engine. A page consists of a plurality of codewords representing the entire page and a parity. The queue can have any depth. The page error detection scheme calculates the parity associated with a set of codewords or for the entire page, where if the calculated parity does not equal the stored parity, an error can be identified.

[0049] In one embodiment, the codeword ECC engine is used to detect and correct errors on a given codeword. A codeword consists of the data to be processed by the codeword ECC engine and a parity. The scavenging performed by the codeword ECC engine is triggered by an activity-based memory management operation.

[0050] In one embodiment, a page error detection method identifies whether an error exists on a page when an activate command is issued. The page error detection method calculates a parity check representing the entire page and writes this parity check to the memory array when a precharge command is issued.

[0051] At least some embodiments described herein may provide various advantages. For example, the probability of detecting a hammered page may increase, even though the page count is high and the hammering threshold is low. For example, greater protection may be provided against targeted row hammer attacks. Error mitigation occurs when the error manifests. For example, hammer detection does not depend on the read pattern, which can increase device security and can increase the reliability and lifespan of the memory device.

[0052] Figure 3 A memory device 302 is shown having error detection circuitry 312 for detecting errors in an accessed portion of one or more memory arrays 306, according to some embodiments. A controller 304 accesses portions of the memory array 306 in response to commands received from a host device 301 via a communication interface 316. The controller 304 uses the error detection circuitry 312 to detect whether an error exists in data stored in each accessed portion (e.g., an entire data page on an activated row of the memory array 306).

[0053] In response to detecting an error in any accessed portion, the location of the accessed portion is stored in a scrub queue 318 for future correction. An error correction circuitry 310 (e.g., a codeword ECC engine such as described above) corrects the error detected for the location (e.g., the physical address of the page) in the scrub queue 318 by using the error correction circuitry 310. The error is corrected at a later time after detection (e.g., during a memory management operation when the MM count reaches a scrub threshold).

[0054] In one embodiment, the controller 304 uses a counter 320 to count memory management commands (e.g., wear leveling commands). When the counter reaches a defined threshold, one or more addresses in the scrub queue 318 are cleared.

[0055] A sense amplifier 308 senses data stored in memory cells of the memory array 306. The controller 304 accesses the stored data by activating one or more rows of the memory array 306. In one instance, the activated rows correspond to pages of the stored data. When one or more rows are activated, the error detection circuitry 312 determines whether an error is detected in the stored data. In one instance, the error detection circuitry 312 determines whether an error exists in any data stored in the entire accessed page.

[0056] In one embodiment, the error detection circuitry 312 calculates the parity of the data accessed by the activated row. The previously determined parity is stored together with the data page. The error detection circuitry 312 compares the calculated parity with the stored parity. If they are not equal, an error is detected. In response to detecting an error, the controller 304 or the error detection circuitry 312 adds the physical address of the page to the cleaning queue 318.

[0057] The cleaning queue 318 is used to record the addresses of the pages in which errors have been detected. The addresses are cleared in future operations. When the addresses are cleared, the error correction circuitry 310 is used to correct the errors.

[0058] When a row of the memory array 306 is activated, data may be read from the row as part of a read or other operation (e.g., wear leveling). The error correction circuitry 310 is used to detect and correct any errors identified in the accessed data on the row. The corrected read data is provided for output by the I / O circuitry 314 on the communication interface 316.

[0059] In one embodiment, the communication interface (I / F) 316 is a bidirectional parallel or serial communication interface. The host device 301 may include a host processor (e.g., a host central processing unit (CPU) or other processor or processing circuitry such as a memory management unit (MMU), interface circuitry, etc.).

[0060] In one embodiment, the memory array 306 may be configured in several non-volatile memory devices (e.g., dies or LUNs) (e.g., one or more stacked flash memory devices each including non-volatile memory (NVM), having one or more groups of non-volatile memory cells and a local device controller or other peripheral circuitry (e.g., device logic, etc.)), and is controlled by the controller 304 through an internal storage system communication interface (e.g., an Open NAND Flash Interface (ONFI) bus, etc.) separate from the communication interface 316.

[0061] In one embodiment, each memory cell in the NOR, NAND, 3D cross-point, MRAM, or one or more other architecture semiconductor memory arrays 306 may be individually or jointly programmed to one or several programmed states. A single-level cell (SLC) may represent one data bit having one of two programmed states (e.g., 1 or 0) per cell. A multi-level cell (MLC) may represent several programmed states per cell (e.g., 2 n, where n is the number of data bits) of two or more data bits. In a particular instance, MLC may refer to a memory cell that can store two data bits in one of four programmed states. A triple-level cell (TLC) may represent three data bits in one of eight programmed states per cell. A quad-level cell (QLC) may represent four data bits in one of sixteen programmed states per cell. In other instances, MLC may refer to any memory cell that can store more than one data bit per cell, including TLCs and QLCs, etc.

[0062] The controller 304 may receive instructions from the host device 301 and may transfer (e.g., write or erase) data to or from one or more of the memory cells of the memory array 306. The controller 304 may particularly include circuitry or firmware, such as several components or integrated circuits. For example, the controller 304 may include one or more memory control units, circuits, or components configured to control access across the memory array and to provide a translation layer between the host device 301 and the storage system (e.g., a memory manager, one or more memory management tables, etc.).

[0063] In one embodiment, the controller 304 may include circuitry or firmware, such as several components or integrated circuits associated with various memory management functions, which particularly include wear leveling (e.g., garbage collection or recycling), error detection or correction, bank or block retirement, or one or more other memory management functions.

[0064] In one embodiment, the controller 304 may include a set of management tables configured to maintain various information associated with one or more components of the memory device 302 (e.g., various information associated with the memory array or one or more memory cells coupled to the controller 304). For example, the management tables may include information about the bank or block age, block erase count, error history, or one or more error counts (e.g., write operation error count, read bit error count, read operation error count, erase error count, etc.) of one or more banks or blocks of memory cells coupled to the controller 304. In certain instances, a bit error may be referred to as an error of uncorrectable bits if the number of detected errors of one or more of the error counts is higher than a threshold. The management tables may particularly maintain a count of correctable or uncorrectable bit errors.

[0065] In one embodiment, the memory device 302 may include one or more 3D NAND architecture semiconductor memory arrays 306. The memory array 306 may include a plurality of memory cells arranged as, for example, banks, a number of devices, planes, blocks, physical pages, super blocks, or super pages. As an example, a TLC memory device may include 18,592 data bytes (B) per page, 1536 pages per block, 548 blocks per plane, and 4 planes per device.

[0066] In one embodiment, data may be written to the memory device 302 or read from the memory device 302 in pages and erased in blocks. However, one or more memory operations (such as read, write, erase, etc.) may be performed on larger or smaller groups of memory cells as needed. For example, during data migration or garbage collection, partial updates of tagged data from offloaded units may be collected to ensure it is rewritten efficiently.

[0067] In one example, a data page includes several bytes of user data (e.g., data payload) and its corresponding metadata. As an example, a data page may include 4kB of user data and several bytes (e.g., 32B, 54B, 224B, etc.) of auxiliary or metadata corresponding to the user data, such as integrity data (e.g., error detection or correction code data), address data (e.g., logical address data, etc.), or other metadata associated with the user data. Different types of memory cells or memory arrays may provide different page sizes, or may require different amounts of metadata associated therewith.

[0068] Figure 4 Circuitry for detecting errors in a page accessed in a memory array is shown in accordance with some embodiments. In one example, a page is accessed by activating a row in the memory array 306. An error detection circuitry 408 is used to detect errors in the accessed page. In one embodiment, the error detection circuitry 408 uses one or more parity bits to check for errors in the page when the page is accessed. Other error detection schemes may be used in other embodiments.

[0069] In one example, page 402 contains a plurality of codewords 0, 1, …, 2 n -1. Page 402 also contains a page parity 410 stored together with page 402. The parity 410 is calculated during the precharge time of page 402. During the precharge time, the parity 410 may be calculated by the error detection circuitry 408 or by other logic circuitry.

[0070] In one embodiment, all of the data stored in the codewords of page 402 is used to calculate parity 410. This includes the user data and parity data stored for each codeword. In one instance, an exclusive OR (XOR) logic operation may be used to combine all such data to calculate parity 410. Other parity schemes may be used in alternative embodiments.

[0071] In one embodiment, page parity 410 may be stored outside of page 402. For example, parity 410 may be stored in the memory of memory device 302 outside of memory array 306. This may be done for the parity of all pages or only a subset of the pages.

[0072] Each page 402 in the memory array has a plurality of columns [n:0]. Data read from or written to page 402 is addressed by a row address and a column address. The row address corresponds to a word line that is activated to access the data stored in page 402. The column address is used by column decoder 404 to select the column of memory cells containing the data to be accessed.

[0073] During a read operation, the data read from page 402 is processed by codeword ECC engine 406 to detect and correct errors. For example, the corrected data is communicated to the host device via a data path to an input / output pin (e.g., a DQ pin).

[0074] In one embodiment, each codeword of page 402 includes user data (e.g., data 0) and parity previously calculated for the data (e.g., parity 0). In one instance, the parity is an error correction code that provides the ability to correct one or more bits of the codeword. When storing a codeword, ECC engine 406 may calculate the parity stored for each codeword. ECC engine 406 may use the parity stored for each codeword to detect and correct one or more bit errors in the codeword when reading the codeword.

[0075] In one embodiment, a lightweight error detection scheme is used to assist in selecting pages (e.g., page 402) to be erased. This error detection scheme has reduced error detection capabilities and no error correction capabilities relative to codeword ECC engine 406. For example, a simple even / odd parity scheme may be implemented that only requires storing one additional bit per page. This error detection scheme can detect any number of odd errors, but cannot detect an even number of errors. When a page is hammered, soft errors caused by the activity will manifest. Thus, this scheme detects single-bit errors that occur on the page.

[0076] In one embodiment, this scheme assumes that single-bit errors are more likely to occur than double-bit errors. Or, in other words, single-bit errors will occur before double-bit errors. However, in the case where activity-induced double-bit errors may occur from one activation cycle to another, the lightweight error detection scheme is paired with a randomly selected scrubbing scheme. In any case, the lightweight error detection scheme is used to identify single-bit errors present on a page in much less time than running all codewords through the ECC engine 406.

[0077] In one instance, when an activation command is issued, the sense page 402 is sensed and the data of the page is stored in the sense amplifier latch. After the activation operation is complete, all page data except the new parity bits 410 (e.g., the data associated with all codewords) is input to an XOR tree that calculates even / odd parity. If the calculated parity is the same as the stored parity 410, then no error is detected. Conversely, if the calculated parity is opposite to the stored parity 410, then an error is identified and the physical page address associated with the activation is recorded in a queue (e.g., the scrub queue 318). For example, the time required for overall page error detection can be appropriately considered and disposed of within the row address strobe to column address strobe (tRCD) timing budget.

[0078] After a precharge command is issued, the even / odd parity of the data residing in the sense amplifier latch is calculated and the appropriate parity bit polarity is written to the array along with the data (e.g., all codewords of the page). Subsequent activation operations may cause bit flips to occur, which will then be detected by the lightweight ECC engine.

[0079] Later, a memory management operation is allocated to scrub the page whose physical address was previously recorded in the queue. Scrubbing requires the use of the standard codeword ECC engine 406, where each codeword is scrubbed one at a time (e.g., a read-modify-write operation occurs such that the corrected data is written back to the array one codeword at a time). Thus, the lightweight error detection scheme can be used to quickly identify page errors, and the codeword ECC engine is used to scrub the identified page errors. Data transfer from the DQ (input / output) pins of the memory device involves the use of the codeword ECC engine 406 and does not involve the lightweight error detection scheme.

[0080] One advantage of this exemplary lightweight error detection scheme is that aggressor detection does not depend on the activation or read mode. In some cases, if the queue depth is too small, there is a risk that not all identified aggressors can be scrubbed. Thus, the queue depth can be determined based on the processing capacity.

[0081] In various embodiments, when lightweight error detection is used (e.g., by circuitry 408), the data state of the memory array is considered (e.g., the initial state of page data and page parity 410). In the absence of proper initialization of the data state throughout the memory array, there may be random page data and random page parity. Random page data and random page parity may lead to many errors. For example, in the absence of proper initialization, the erase queue may be flooded with false errors, compromising the ability to accurately erase rows with soft errors.

[0082] In one embodiment, for non-volatile memory devices, the initial page data and parity are initialized at the factory during manufacturing to avoid the possibility of flagging false errors on first use. In one embodiment, for volatile memory devices (or if a non-volatile memory device emulates a volatile memory device), the initial page data and parity are initialized during each power-up cycle.

[0083] In one example, the initialization state of a simple even / odd parity state is an all-zero state. That is, for all page data bits and all page parity bits in the memory device, all pages in the memory device are initialized to an all-zero state.

[0084] Figure 5 Sense amplifier latches 520, 521, 522 for holding data associated with memory cells 510, 511, 512, 513 of a memory array are shown in accordance with some embodiments. In one example, the memory cells are located in memory array 306. The memory cells can be of various memory types, including volatile and / or non-volatile memory cells.

[0085] The memory cells are accessed using word lines (e.g., WL0) and digit lines (e.g., DL0) or bit lines. Individual memory cells are accessed by activating the word line selected by row decoder 530 and selecting the digit line or bit line selected by column decoder 540. When the word line is activated, the data from each memory cell on the row resides in the corresponding sense amplifier latch for each digit line or bit line.

[0086] The data residing in the sense amplifier latches can be used as inputs to logic circuitry 550, 551 for various calculations. These calculations can include detecting and / or correcting errors in the data retrieved from the memory cells using parity or other metadata stored using the memory cells. In one embodiment, logic circuitry 550 includes error detection circuitry 312. In one example, logic circuitry 550 is any logic that operates on data at the page level (e.g., lightweight error detection as described for Figure 4 ).

[0087] The logic circuitry 551 is coupled to the column decoder 540. In one embodiment, the logic circuitry includes error correction circuitry 310. In one instance, the logic circuitry 551 is any logic (e.g., ECC engine 406) that operates on data at the column (e.g., word - line) level.

[0088] In one embodiment, a memory device that includes a memory array has a plurality of memory cells 510, 511, 512, 513, etc., and one or more circuits or components for providing communication with the memory array or performing one or more memory operations on the memory array. A single memory array or additional memory arrays, dies, or LUNs may be used. The memory device may include a row decoder 530, a column decoder 540, sense amplifiers, page buffers, selectors, input / output (I / O) circuitry, and a controller.

[0089] In some non - volatile memory devices (e.g., NAND flash), the memory cells of the memory array may be arranged in blocks. Each block may include sub - blocks. Each sub - block may include a number of physical pages, and each page includes a number of memory cells. In some instances, the memory cells may be arranged in rows, columns, pages, sub - blocks, blocks, etc., and are accessed using, for example, access lines, data lines, or one or more select gates, source lines, etc.

[0090] In volatile memory devices (e.g., DRAM) and some emerging non - volatile memory technologies, the memory cells of the memory array may be arranged in banks or other forms of partitions. In one instance, when an activation of a row address is issued, the bits on the activation command may be addressed by using a bank address (to specify which bank within the memory device) and a row address (to specify which row within the specified bank). The word line associated with the row address is made high.

[0091] A controller (e.g., controller 304) may control the memory operations of the memory device based on one or more signals or instructions received on control lines (e.g., from a host device 301) (e.g., including one or more clock signals or control signals indicating a desired operation (e.g., write, read, erase, etc.)) or address signals (A0 to AX) received on one or more address lines. One or more devices external to the memory device may control the values of the control signals on the control lines or the address signals on the address lines. Examples of devices external to the memory device may include, but are not limited to, a host, a memory controller, a processor, or one or more circuits or components.

[0092] A memory device may use access lines and data lines to transfer data (e.g., write or erase) to or from one or more of the memory cells, or transfer (e.g., read) data from one or more of the memory cells. A row decoder and a column decoder may receive and decode an address signal (A0 to AX) from an address line, determine which memory cells to access, and provide signals to one or more of the access lines (e.g., one or more of a plurality of word lines (e.g., WL0 to WLm)) or one or more of the data lines (e.g., one or more of a plurality of bit lines (BL0 to BLn)).

[0093] The memory device may include sensing circuitry (e.g., sense amplifier 308) configured to determine the value of data on a memory cell (e.g., read) or determine the value of data written to a memory cell. In one example, the sense amplifier is used to sense a voltage (e.g., in the case of charge sharing in a DRAM). In one example, in a selected string of memory cells, one or more of the sense amplifiers may read a logic level in a selected memory cell in response to a read current flowing through the selected string in the memory array to the data line.

[0094] One or more devices external to the memory device may communicate with the memory device using I / O lines (e.g., DQ0 to DQN), address lines (e.g., A0 to AX), or control lines. I / O circuitry (e.g., 314) may transfer data values into or out of the memory device using the I / O lines according to, for example, control lines and address lines, e.g., transfer into or out of a page buffer or a memory array. The page buffer may store data received from one or more devices external to the memory device before programming the data into a relevant portion of the memory array, or may store data read from the memory array before transferring the data to one or more devices external to the memory device.

[0095] A column decoder 540 may receive an address signal (e.g., A0 to AX) and decode it into one or more column select signals (e.g., CSEL1 to CSELn). A selector (e.g., a selection circuit) may receive the column select signals (CSEL1 to CSELn) and select data values in the page buffer representing data to be read from or programmed into the memory cells. The selected data may be transferred between the page buffer and the I / O circuitry.

[0096] Figure 6A flowchart is shown for error detection and recording of an address associated with a detected error for future clearing, according to some embodiments. When a memory device (e.g., memory device 302) powers on from a power-off state (transitioning from block 602 to block 604), a memory management counter (e.g., counter 320) is reset to zero, and a clearing queue (e.g., clearing queue 318) is reset. Then, the memory device performs various operations at block 606.

[0097] In response to receiving an activate command (e.g., from host device 301 by controller 304), a page is sensed. The data from the page now resides in a sense amplifier latch. This data is used to calculate an error detection parity. In one example, the error detection parity is calculated by error detection circuitry 312.

[0098] At block 610, the stored error detection parity from the page is compared with the calculated error detection parity. In one example, the parity stored in the page is page parity 410. If the stored parity does not equal the calculated parity, then at block 612, the physical address of the page is added to the clearing queue. In one example, the physical address is a row address. If the stored parity equals the calculated parity, then normal memory device operation continues at block 606.

[0099] In response to receiving a precharge command (e.g., from host device 301 by controller 304), at block 614, the data in the sense amplifier latch is used to calculate an error detection parity to be stored with the data written to the page. In one example, the calculated error detection parity is page parity 410. The page data is written to the memory cells corresponding to the page, including the write of the error detection parity. The error detection parity is calculated at block 614 because there is a possibility that some data has been written to the page, which causes the codeword data to change.

[0100] Another advantage of calculating the lightweight error detection parity at this time is that the new lightweight error detection parity will account for any new soft errors. This ensures that a page is not added to the cleaning queue multiple times due to a single soft error. In this case, the first activate cycle after the soft error occurs will mark the error, which will cause the page to be added to the clearing queue.

[0101] Note that soft errors occurring in memory cells are sensed and the associated data now resides in the sense amplifier latch. During the precharge time, the calculated lightweight error detection scheme will now be updated considering soft errors based on the data residing in the sense amplifier latch. If there are no further soft errors before a subsequent activation, then the calculated lightweight error detection parity will be equivalent to the stored lightweight error detection parity. Thus, this page will not be added to the flush queue again until a new soft error occurs.

[0102] Additionally, the problem of hard failure flooding address collection when using a codeword ECC engine to detect invaders can be largely avoided by the above scheme (e.g., the lightweight error detection scheme). For example, when the codeword ECC engine detects an error on a read codeword, an error flag may occur. Not all codewords in a page are guaranteed to be read in every activation cycle. Running every codeword of a page through the codeword ECC engine in every activation cycle would take too much time. Hard errors can flood address collection. This is attributed to parity generation occurring on the written data. However, this can be avoided by the lightweight error detection scheme because parity is generated based on the data in the sense amplifier (SA) latch. Hard data failures are considered during parity calculation so that addresses with hard errors are not repeatedly added to the flush queue. In some cases, hard parity failures may still cause an address to be mis-identified. However, additional parity bits and logic can be added to reduce the likelihood of this occurring.

[0103] Thus, in this way, the lightweight error detection is calculated at precharge time to update the lightweight error detection parity based on any data present in the sense amplifier latch. Thus, the lightweight error detection parity calculation at this time can adapt to new written data, new soft errors, or new hard errors.

[0104] In some embodiments, to save power, the calculation of the lightweight error detection parity is selected (e.g., optimized) to occur only when needed. For example, the calculation of the lightweight error detection parity at precharge time is performed only after:

[0105] - a write to the page has occurred after the row has been activated

[0106] - a lightweight error detection parity mismatch (the calculated parity is not equal to the stored parity) at activation time

[0107] Memory device operations at block 606 can include various memory management operations (e.g., wear leveling or refreshing). These operations can be initiated, for example, by commands received by the controller. In one instance, the controller 304 receives memory management commands from the host device 301. In other embodiments, the controller 304 can initiate memory management operations (e.g., self-management functions) without communication from an external device.

[0108] Memory management commands and / or operations are counted by a counter (e.g., counter 320). The counter 320 can be incremented by the controller for each command or operation. The controller and / or an external device such as a host can define a threshold. In one instance, the threshold is the scrub steal ratio.

[0109] In one embodiment, the scrub ratio is fusibly programmed at the factory during manufacturing. The process capabilities are characterized in the factory. Based on the characterized process capabilities, the scrub ratio value is determined and fusibly programmed at the factory (e.g., before being sent to the customer).

[0110] If the customer (e.g., the host) has the ability to adjust the scrub ratio, then a security risk can occur. For example, assume the scrub ratio is 10. This means that 9 / 10 memory management commands will be used for wear leveling, while 1 / 10 memory management commands will be used to scrub detected hammered rows. If the customer has the ability to adjust the scrub ratio, then the customer can, for example, adjust the scrub ratio to 2. In this case, only 1 / 2 of the commands will be used for wear leveling. This will indicate that less wear leveling is occurring and the wear is likely not being mitigated efficiently. Therefore, the scrub ratio should be selected by the memory manufacturer considering the process requirements. For example, this scrub ratio is fusibly programmed by the memory manufacturer at the factory and should not be adjusted by the customer later.

[0111] At block 616, the current value of the memory management count is compared with the threshold. If the count is less than the threshold, then at block 618, wear leveling and / or another operation is performed. The count is incremented by 1 (or some other defined value). Then, the memory device operations continue at block 606.

[0112] If the count reaches the threshold, then at block 620, one or more physical addresses are retrieved from the scrub queue. The pages at each physical address are scrubbed to correct errors in the pages. In one instance, each codeword in the page is passed through a codeword ECC engine (e.g., 406). The corrected data for the codewords is written back to the memory cells in the memory array. The memory management count is reset to zero. Then, the memory device operations continue at block 606 until the memory device is powered down to the powered-off state at block 602.

[0113] At block 620, in the case where the purge queue is empty, the controller may purge the last physical address accessed in the memory or do nothing.

[0114] In one embodiment, instead of using the memory management count as described above to initiate the purge operation, a timer may be used to trigger the purge of addresses in the queue at regular or other defined time intervals. In some embodiments, a combination of memory management activity and the timer may be used to trigger the purge. In one instance, as an alternative to issuing memory management commands in an activity-based manner, memory management commands are issued at some defined or selected time interval (e.g., periodic memory management commands).

[0115] Figure 7 An example of comparing the calculated parity with the stored parity to determine if an error is detected is shown. For illustrative purposes only, a simplified example where eight-bit data is stored in a page is used. The stored parity of the page is 1, as Figure 7 shown by the original data in. After writing to the page, a single-bit error occurs on the page. Thus, the calculated parity is 0. Since the calculated parity does not equal the stored parity, the controller detects an error in the page. The location of the page or other indication is added to the purge queue or otherwise recorded.

[0116] An example where there is a single-bit error in the stored parity is also illustrated. In this case, the stored parity does not equal the stored parity, and the controller detects an error. The page is added to the purge queue or otherwise recorded.

[0117] Figure 8 An exemplary method for calculating parity is shown. The codeword is stored in the page as bit data [0] to data [n]. As illustrated, the parity to be stored in the page is calculated by performing an exclusive OR (XOR) on all bits of all codewords in the page.

[0118] In one instance, memory device operations include issuing various commands. Non-limiting details regarding certain exemplary commands are provided below:

[0119] ■ Memory device operations

[0120] - When an activation of row address x is issued

[0121] ■ Row address x can be addressed by addressing the bits on the activation command in a pressing manner

[0122] ■ Bank address (to specify which bank within the memory device)

[0123] ■ Row address (to specify the row within the specified bank)

[0124] ■ Set the word line associated with row address x to high level

[0125] ■ Sense the page that causes the data to reside in the sense amplifier latch

[0126] ■ Calculate the lightweight error detection parity

[0127] ■ If the stored parity does not match the calculated parity, then add the physical page address to the purge queue

[0128] - When a read command for column y of the activated row address x is issued (the read command must occur when the row is activated)

[0129] ■ Summary: Read the corrected data of the data residing in the sense amplifier (SA) latch associated with column y

[0130] ■ Note: The read command contains the following addressing bits

[0131] ■ Bank address (used to specify which bank within the memory device)

[0132] ■ Column address (used to specify which column in the activated row within the specified bank)

[0133] ■ Details:

[0134] ■ Note: The column y data is represented by the state of the SA latch associated with column y

[0135] ■ In addition, column y is associated with a codeword consisting of data bits and parity bits

[0136] ■ The data and parity from the SA latch of the appropriate column fed through the codeword ECC engine

[0137] ■ Error detection

[0138] ■ Syndrome generated based on the calculated parity (parity calculated from the data residing in the "data" SA latch) and the stored parity (parity directly stored in the SA latch)

[0139] ■ Correct any correctable errors

[0140] ■ The SA latch data remains unchanged

[0141] ■ Corrected data output

[0142] ■ The possibly corrected data sent out from the memory device on the DQ (IO pin) of the memory device

[0143] - When a write command for column y of the activated row address x is issued (the write command must occur when the row is activated)

[0144] ■Summary: Change the SA latch associated with column y to contain the new write data and associated parity

[0145] ■Note: The write command contains the following addressing bits

[0146] ■Bank address (to specify which bank within the memory device)

[0147] ■Column address (to specify which column within the activated row in the specified bank)

[0148] ■Details

[0149] ■Data input from the DQ (IO pin) of the memory device

[0150] ■Data fed through the codeword ECC engine

[0151] ■Parity generated based on the input data

[0152] ■Input data and generated parity (codeword) written to the appropriate column's SA latch

[0153] -When a precharge command is issued

[0154] ■Summary: The precharge command will cause an implicit write to the cell and then make the word line go low

[0155] ■The cell will now contain

[0156] ■Data residing in the SA latch (all data and all parity of all codewords)

[0157] ■Calculated (at precharge time) lightweight error detection parity

[0158] ■Note: The precharge command contains the following addressing bits

[0159] ■Bank address (precharge any row activated in the bank)

[0160] ■Details:

[0161] ■Calculate lightweight error detection parity

[0162] ■Change one or more SA latches associated with the lightweight error detection parity to reflect the state of this calculation

[0163] ■Write the page data containing the calculated lightweight error detection parity to the memory cell

[0164] ■Data written to the memory cell residing in all SA latches

[0165] ■After writing data into a memory cell, the word line associated with the row address x is made low.

[0166] -When an activity-based MM command is issued

[0167] ■If the MM count is equal to the erase steal ratio, then the physical page address is cleared from the erase queue. In this case, each codeword of the page is passed through the codeword ECC engine, and the corrected data is written back to the physical page address. Thereafter, the MM count is reset to zero.

[0168] ■Otherwise, a wear leveling operation is performed and the MM count is incremented.

[0169] A non-limiting example of a memory device is now described to illustrate the technical problems associated with a method of detecting a hammered row that may use only a codeword ECC engine.

[0170] ● Various details on how using only a codeword ECC engine poses problems for detecting a hammered row are presented below:

[0171] ○ Various details on the operation and behavior of the memory device are provided below:

[0172] ■ Memory array behavior

[0173] ● Regarding the activate command (ACT): Sense the page of the memory array that makes the data reside in the sense amplifier (or simply "sense amp") latch.

[0174] ● Regarding the precharge command (PRE): Write the data in the sense amplifier latch to the page of the memory array.

[0175] ■ Codeword ECC engine behavior

[0176] ● During writing

[0177] ○ The data input from the DQ (IO pin) of the memory device

[0178] ○ The data fed through the codeword ECC engine

[0179] ■ The parity generated based on the input data

[0180] ○ The input data and the generated parity (codeword) written into the sense amplifier (SA) latch of the appropriate column

[0181] ● During reading

[0182] ○ The data from the SA latch of the appropriate column fed through the codeword ECC engine

[0183] ○ Syndrome (error detection) generated based on the calculated parity and the stored parity

[0184] ○ Correct any correctable errors

[0185] ■ SA latch data remains unchanged

[0186] ■ Corrected data output

[0187] ○ Possibly corrected data sent out from the memory device on DQ

[0188] ○ Technical problems of the method of using only the codeword ECC engine for hammer row detection:

[0189] ■ Error detection only occurs when reading with the codeword ECC engine. This specification does not require reading all codewords of the accessed page, so errors may not be noticed.

[0190] ● To solve this problem, when activated, all codewords can be cycled through the codeword ECC engine. However, the latency of the codeword ECC engine is too large to determine whether there are errors on the entire page by running each codeword through ECC.

[0191] ■ Hard failures can overwhelm address collection and cause correctable failures to be missed.

[0192] ○ At least some embodiments described herein provide a technical solution to the above technical problems by performing error detection when accessing a page:

[0193] ■ Error detection occurs on the entire page during each activation cycle (e.g., in contrast to error detection only for a portion of the codewords stored on the page). This allows a lightweight error detection scheme's aggressor detection ability to be, for example, independent of the read pattern.

[0194] ● A lightweight error detection scheme can identify errors much faster than running each codeword through the codeword ECC engine.

[0195] ■ Hard data failures will be considered during parity calculation so that addresses with hard errors will not be repeatedly added to the queue. However, hard parity failures may still cause an address to be misidentified as an aggressor.

[0196] · A lightweight error detection scheme can add, for example, hard errors to the queue when there are hard errors on the parity bits. However, it should be noted that there are fewer lightweight error detection parity bits compared to the total lightweight error detection data bits. Therefore, this probability is very low. Additionally, extra parity bits and logic can be added to further reduce the likelihood of this happening.

[0197] ● The depth or size of the flush queue depends on the process capabilities that govern the hammering threshold. The flush queue depth is inversely proportional to the hammering threshold.

[0198] ● If a simple even / odd parity scheme is chosen for the lightweight error detection scheme:

[0199] ○ This error detection scheme can detect any number of odd errors, but cannot detect an even number of errors.

[0200] ○ When a page is hammered, soft errors caused by activity become apparent. Therefore, this scheme can detect single-bit errors that occur on the page.

[0201] ○ This scheme generally assumes that single-bit errors are more likely to occur than double-bit errors. Or, in other words, single-bit errors will occur before double-bit errors.

[0202] ○ However, if activity-induced double-bit errors can occur from one activation cycle to another, then the lightweight error detection scheme can be paired with a random address sampling scheme.

[0203] ■ This lightweight error detection scheme will not be able to detect double-bit errors (even-bit errors) that suddenly appear from one activation cycle to another.

[0204] ● It should be noted that if the lightweight error detection scheme uses an error detection scheme that has a greater error detection capability at the cost of a larger die size and latency than the simple even / odd parity scheme, then this limitation can be resolved. To save die size and reduce latency, the simple even / odd parity scheme can be paired with a random address sampling scheme.

[0205] ■ When the lightweight error detection scheme is paired with the random address sampling scheme

[0206] ● A certain percentage of aggressors (e.g., hammered rows) for activity-based flushing will be identified by the lightweight error detection scheme

[0207] ● A certain percentage of aggressors for activity-based flushing will be identified by random address sampling

[0208] ● As mentioned above, emerging memory technologies may suffer from technical problems with particularly low detection and mitigation probabilities due to small hammering thresholds and increased page counts. At least some of the embodiments described herein for whole-page error detection provide solutions that increase the probability of detecting hammered pages despite the increased page count and decreased hammering threshold.

[0209] ○ It should be noted that some methods attempt to address this problem by using a per-row activation counter (PRAC).

[0210] ■ For PRAC, additional bits are added to each page to act as a row activation counter

[0211] ● These bits are incremented in each activation cycle

[0212] ■ The number of PRAC counter bits is proportional to log2(hammer threshold)

[0213] ○ At least some embodiments of the per-page error detection for hammer detection as described herein provide a solution to this problem with less die area increase than using PRAC. Fewer bits need to be added per row. At least, for per-page error detection, only one additional bit per page is needed for hammer detection.

[0214] Figure 9 Disclosed is a method for detecting errors in an accessed page and recording the errors for later correction (e.g., clearing queue 318) according to some embodiments. For example, Figure 9 the method may be implemented in Figure 3 memory device 302. In one example, error detection circuitry 312 is used to detect errors in a page. In one example, error correction circuitry 310 is used to correct the errors later.

[0215] Figure 9 the method may be executed by processing logic, which may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, Figure 9 the method is at least partially executed by one or more processing devices (e.g., Figure 3 controller 304) and / or logic circuitry.

[0216] Although shown in a particular sequence or order, the order of the process may be modified unless otherwise specified. Accordingly, the illustrated embodiments should be understood only as examples, and the illustrated process may be performed in a different order and some processes may be performed in parallel. Additionally, one or more processes may be omitted in various embodiments. Thus, all processes are not required in every embodiment. Other process flows are possible.

[0217] At block 901, sense a first data page and the parity stored in the first page. In one example, in response to an activation command, sense the stored codeword of page 402 and page parity 410.

[0218] At block 903, calculate a first parity of the first page. In one example, the first parity is calculated by error detection circuitry 408.

[0219] At block 905, the calculated first parity is compared with the stored parity. In one example, the controller 304 compares the calculated first parity with the page parity 410.

[0220] At block 907, if an error is detected, the physical address of the first page is added to the purge queue. In one example, the physical address of page 402 is added to the purge queue 318.

[0221] At block 909, when new data (e.g., in response to a precharge command) is written to the first page, a second parity of the first data page is calculated. In one example, a lightweight error detection parity is calculated at block 614.

[0222] At block 911, the first data page and the calculated second parity are written to the first page. For example, the second parity can be used for error detection when a future activate command is issued after writing to the first page. In one example, the first data page including the calculated second parity is written to the memory cells of the memory array 306.

[0223] In some aspects, the techniques described herein relate to an apparatus comprising: error detection circuitry (e.g., 312); and at least one controller (e.g., 304) configured to: access a portion of a memory array (e.g., 306); use the error detection circuitry to detect whether an error exists in data stored in each accessed portion; and in response to detecting a first error in a first portion, record a first location of the first portion.

[0224] In some aspects, the techniques described herein relate to an apparatus further comprising error correction circuitry (e.g., 310), wherein the controller is further configured to use the error correction circuitry to correct the detected first error.

[0225] In some aspects, the techniques described herein relate to an apparatus, wherein the number of bit errors correctable by the error correction circuitry is greater than or equal to the number of bit errors detectable by the error detection circuitry.

[0226] In some aspects, the techniques described herein relate to an apparatus, wherein: the controller is further configured to perform memory management on a memory device including the memory array; and the memory management includes using the error correction circuitry to correct the first error.

[0227] In some aspects, the techniques described herein relate to an apparatus that further includes a queue (e.g., flush queue 318) to store locations in a memory array where errors in the stored data are detected, where the queue includes a first location, and where the controller is further configured to correct errors at the stored locations during a flush operation.

[0228] In some aspects, the techniques described herein relate to an apparatus where error detection circuitry is configured to determine a parity of data stored in each accessed portion.

[0229] In some aspects, the techniques described herein relate to an apparatus where a first portion is a row of a memory array, the row storing a first parity (e.g., page parity 410) of data written to the row, and when data is read from the row, the first parity is compared to a second parity determined by the error detection circuitry.

[0230] In some aspects, the techniques described herein relate to an apparatus where accessing a portion of a memory array includes activating a row of the memory array to read or write data stored on each activated row.

[0231] In some aspects, the techniques described herein relate to an apparatus where the accessed portion is a page storing a codeword (e.g., the page includes page 402), and each page stores a parity determined using the codeword stored in the page.

[0232] In some aspects, the techniques described herein relate to an apparatus where a first error is detected by comparing a parity stored with a codeword in a first page to a parity calculated for the codeword when the first page is activated.

[0233] In some aspects, the techniques described herein relate to an apparatus that includes: a counter (e.g., 320); and at least one controller configured to: maintain a record of rows in which errors are detected in stored data when a row of a memory array is activated; count a number of memory management operations (e.g., Figure 6 MM count) of the memory array using the counter; and correct detected errors in the recorded rows based on the number.

[0234] In some aspects, the techniques described herein relate to an apparatus where detected errors are corrected when the number reaches a threshold.

[0235] In some aspects, the techniques described herein relate to an apparatus where errors are detected in stored data in response to an activation command.

[0236] In some aspects, the techniques described herein relate to an apparatus that corrects detected errors when a count reaches a purge threshold, and the controller is further configured to randomize the purge threshold.

[0237] In one example, a random purge threshold is determined when the MM count is reset (within a certain range). For example:

[0238] Reset the MM count to 0;

[0239] Set the purge threshold to a first random value within a certain range;

[0240] Increment the MM count for each memory management command (during which the memory management commands are used for wear leveling);

[0241] Once the first random value is reached, a purge can occur and the MM count is reset;

[0242] Set the purge threshold to a second random value within a certain range;

[0243] Increment the MM count for each memory management command (during which the memory management commands are used for wear leveling);

[0244] Once the second random value is reached, a purge can occur and the MM count is reset; . . .

[0248] In some aspects, the techniques described herein relate to an apparatus, wherein: the controller is further configured to operate in a first mode and a second mode, perform a wear leveling operation and increment a counter in the first mode (e.g., block 618), and in the second mode, correct at least one of the detected errors and reset the counter (e.g., block 620); and select the first mode or the second mode based on the quantity.

[0249] In some aspects, the techniques described herein relate to an apparatus, wherein correcting the detected errors includes passing a codeword stored in a memory cell through an ECC engine (e.g., the codeword ECC engine 406) and writing the corrected data back to the memory cell.

[0250] In some aspects, the techniques described herein relate to an apparatus, wherein a record is a queue of physical addresses of each row in which an error is detected.

[0251] In some aspects, the techniques described herein relate to a method that includes: in response to receiving an activation command, sensing a first data page and a parity stored in the first page (e.g., block 608); calculating a first parity of the first page; and comparing the first parity with the stored parity.

[0252] In some aspects, the techniques described herein relate to a method that further includes, in response to determining that the first parity does not match the stored parity, adding a physical address of the first page to a purge queue.

[0253] In some aspects, the techniques described herein relate to a method that further includes, in response to receiving a precharge command: calculating a second parity of the first data page; and writing the first data page and the second parity (e.g., block 614).

[0254] The present disclosure includes various apparatuses that perform the methods described above and implement the systems described above, including a data processing system that executes these methods and a computer-readable medium that contains instructions that, when executed on the data processing system, cause the system to execute these methods.

[0255] The description and drawings are illustrative and should not be construed as restrictive. Many specific details are described to provide a thorough understanding. However, in specific instances, well-known or conventional details are not described to avoid obscuring the description. References in this disclosure to one or an embodiment are not necessarily references to the same embodiment; and such references mean at least one.

[0256] As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without an intermediate component), whether wired or wireless, including connections such as electrical, optical, magnetic, etc.

[0257] References in this specification to "an embodiment" or "one embodiment" mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present disclosure. The phrase "in an embodiment" appearing in various places in the specification does not necessarily refer to the same embodiment, and separate or alternative embodiments are not mutually exclusive of other embodiments. Additionally, various features are described that may be exhibited by some embodiments and not by others. Similarly, various requirements are described that may be requirements of some embodiments and not of others.

[0258] In this description, various functions and operations may be described as being performed or caused by software code to simplify the description. However, those skilled in the art will recognize that such a statement means that the functions and / or operations are produced by executing code by one or more processing devices such as a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit, and / or a field programmable gate array (FPGA). Alternatively or in combination, the functions and operations may be implemented using special circuit systems (such as logic circuit systems) with or without software instructions. Embodiments may be implemented using hardwired circuit systems without software instructions or in combination with software instructions. Thus, the technology is not limited to any particular combination of hardware circuit systems and software, nor to any particular source of instructions executed by a computing device.

[0259] Although some embodiments may be implemented in a fully functional computer and computer system, various embodiments can be distributed in a variety of forms as a computing product and can be applied regardless of the particular type of computer-readable medium used to actually implement the distribution.

[0260] At least some aspects of the disclosure may be embodied, at least in part, in software. That is, the technology may be practiced in a computing device or other system in response to its processing device (such as a microprocessor) executing a sequence of instructions contained in a memory (such as ROM, volatile RAM, non-volatile memory, cache, or remote storage device).

[0261] The routines executed to implement the embodiments may be implemented as part of an operating system, middleware, service delivery platform, SDK (software development kit) component, web service, or other specific application, component, program, object, module, or sequence of instructions (sometimes referred to as a computer program). The call interface to these routines may be open to the software development community as an API (application programming interface). Computer programs typically include one or more sets of instructions in various memories and storage devices in a computer at various times, and when the sets of instructions are read and executed by one or more processors in the computer, the computer performs the operations necessary to execute the elements involved in various aspects.

[0262] A computer-readable medium can be used to store software and data that, when executed by a computing device, cause the device to perform various methods. The executable software and data can be stored in various places, including, for example, ROM, volatile RAM, non-volatile memory, and / or cache. Portions of this software and / or data can be stored in any of these storage devices. Additionally, data and instructions can be obtained from a centralized server or a peer-to-peer network. Different portions of the data and instructions can be obtained at different times and in different communication sessions or in the same communication session from different centralized servers and / or peer-to-peer networks. The complete data and instructions can be obtained before an application is executed. Alternatively, portions of the data and instructions can be obtained dynamically and in a timely manner as needed during execution. Thus, in a particular instance of time, it is not required that all of the data and instructions be on the computer-readable medium.

[0263] Examples of computer-readable media include, but are not limited to, recordable and non-recordable types of media such as volatile and non-volatile memory devices, read-only memory (ROM), random access memory (RAM), flash memory devices, solid-state drive storage media, removable disks, magnetic disk storage media, optical storage media (e.g., compact disc read-only memory (CDROM), digital versatile disc (DVD), etc.), and the like. A computer-readable medium can store instructions. Other examples of computer-readable media include, but are not limited to, non-volatile embedded devices using NOR flash memory or NAND flash memory architectures. The media used in these architectures can include unmanaged NAND devices and / or managed NAND devices, including, for example, eMMC, SD, CF, UFS, and SSD.

[0264] In general, a non-transitory computer-readable medium includes any mechanism that provides (e.g., stores) information in a form accessible by a computing device (e.g., a computer, a mobile device, a network device, a personal digital assistant, a manufacturing tool with a controller, any device having a set of one or more processors, etc.). As used herein, "computer-readable media" can include a single medium or multiple media (e.g., that store one or more sets of instructions).

[0265] In various embodiments, hardwired circuitry can be combined with software and firmware instructions to implement techniques. Thus, the techniques are not limited to any particular combination of hardware circuitry and software, nor to any particular source of the instructions executed by a computing device.

[0266] The various embodiments described herein can be implemented using a variety of different types of computing devices. As used herein, examples of "computing devices" include, but are not limited to, servers, centralized computing platforms, systems of multiple computing processors and / or components, mobile devices, user terminals, vehicles, personal communication devices, wearable digital devices, electronic kiosks, general-purpose computers, electronic file readers, tablet computers, laptop computers, smartphones, digital cameras, residential household appliances, televisions, or digital music players. Additional examples of computing devices include devices that are part of the so-called "Internet of Things" (IoT). Such "things" may occasionally interact with their owners or administrators who can monitor the things or modify the settings of these things. In some cases, these owners or administrators act in the role of users with respect to the "thing" devices. In some instances, the user's primary mobile device (e.g., an Apple iPhone) can be an administrator server with respect to a paired "thing" device (e.g., an Apple Watch) worn by the user.

[0267] In some embodiments, the computing device can be a computer or a host system, which can be implemented as, for example, a desktop computer, a laptop computer, a network server, a mobile device, or other computing devices that include a memory and a processing device. The host system can include or be coupled to a memory subsystem such that the host system can read data from or write data to the memory subsystem. The host system can be coupled to the memory subsystem via a physical host interface. Generally, the host system can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0268] In some embodiments, the computing device is a system that includes one or more processing devices. Examples of processing devices can include microcontrollers, central processing units (CPUs), dedicated logic circuitry (e.g., field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc.), system-on-chips (SoCs), or another suitable processor.

[0269] In one example, the computing device is a controller of a memory system. The controller includes a processing device and a memory that contains instructions executed by the processing device to control various operations of the memory system.

[0270] Although some figures illustrate several operations in a particular order, non-sequential related operations can be reordered, and other operations can be combined or decomposed. Although some reordering or other grouping is explicitly mentioned, other groupings will be obvious to those of ordinary skill in the art and thus an exhaustive list of alternatives is not presented. Additionally, it should be recognized that the stages can be implemented in hardware, firmware, software, or any combination thereof.

[0271] Unless otherwise specifically stated, disjunctive language such as the phrase "at least one of X, Y, or Z" will be understood within the context of the general use to present items, terms, etc. as being either X, Y, or Z, or any combination thereof (e.g., X, Y, and / or Z). Thus, such disjunctive language is not generally intended and should not be construed to imply that certain embodiments require the presence of at least one of each of X, at least one of each of Y, or at least one of each of Z.

[0272] In the foregoing specification, the disclosure has been described with reference to specific exemplary embodiments thereof. Obviously, various modifications can be made thereto without departing from the broader spirit and scope set forth in the appended claims. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.

Claims

1. An apparatus, comprising: An error detection circuitry; And At least one controller configured to: Access a portion of a memory array; Use the error detection circuitry to detect whether an error exists in data stored in each accessed portion; And In response to detecting a first error in a first portion, record a first location of the first portion.

2. The apparatus according to claim 1, further comprising an error correction circuitry, wherein the controller is further configured to use the error correction circuitry to correct the detected first error, wherein: The number of bit errors correctable by the error correction circuitry is greater than or equal to the number of bit errors detectable by the error detection circuitry; The controller is further configured to perform memory management on a memory device including the memory array; The memory management includes using the error correction circuitry to correct the first error; And Detect the first error by comparing a parity check stored in a first page together with a code word with a parity check calculated for the code word when the first page is activated.

3. The apparatus according to claim 1, further comprising a queue for storing locations of detected errors in the stored data of the memory array, wherein the queue includes the first location, wherein the controller is further configured to correct an error at the stored location during a clear operation, and wherein the error detection circuitry is configured to determine a parity check of the data stored in each accessed portion.

4. The apparatus according to claim 1, wherein: The first portion is a row of the memory array, the row stores a first parity check of data written to the row, and when reading data from the row, the first parity check is compared with a second parity check determined by the error detection circuitry; Accessing a portion of the memory array includes activating a row of the memory array to read or write data stored on each activated row; And The accessed portion is a page storing a code word, and each page stores a parity check determined using the code word stored in the page.

5. An apparatus, comprising: A counter; And At least one controller configured to: When activating a row of a memory array, maintain a record of the row in which an error is detected in the stored data; Use the counter to count the number of memory management operations of the memory array; And Based on the number, correct the detected error of the recorded row.

6. The apparatus according to claim 5, wherein when the number reaches a threshold, correct the detected error, and detect the error in the stored data in response to an activation command.

7. The apparatus according to claim 5, wherein when the number reaches a clear threshold, correct the detected error, and the controller is further configured to randomize the clear threshold.

8. The apparatus according to claim 5, wherein: The controller is further configured to operate in a first mode and a second mode, in the first mode, performing a wear leveling operation and incrementing the counter, and in the second mode, correcting at least one of the detected errors and resetting the counter; selecting the first or second mode based on the quantity; and correcting the detected error includes passing a codeword stored in a memory cell through an ECC engine and writing the corrected data back to the memory cell; and the record is a queue including physical addresses of each row where an error is detected.

9. A method, comprising: sensing a first data page and parity stored in the first page in response to receiving an activation command; calculating a first parity of the first page; and comparing the first parity with the stored parity.

10. The method according to claim 9, further comprising: adding a physical address of the first page to a purge queue in response to determining that the first parity does not match the stored parity; and in response to receiving a precharge command: calculating a second parity of the first data page, and writing the first data page and the second parity.