Error management for memory devices

By using logical to physical (L2P) data structure management errors in the memory subsystem, locking uncorrectable and unrepairable management units is solved, and the reliability and life of the memory device is improved.

CN120371593APending Publication Date: 2025-07-25MICRON TECHNOLOGY INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510114423.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-01-16
Filing Date
2025-01-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing memory subsystem cannot effectively manage uncorrectable and unrepairable management units during loss equalization operations, resulting in a reduced reliability and life of the memory device.

Method used

By using the logic to physical (L2P) data structure to manage errors, locking uncorrectable and irreparable management units, maintaining the mapping of logical addresses to physical addresses, and providing location information of unavailable units.

Benefits of technology

Effectively remove uncorrectable and unrepairable management units, improve the reliability and life of the memory subsystem, and provide users with location information of unavailable units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371593A_ABST
    Figure CN120371593A_ABST
Patent Text Reader

Abstract

The invention relates to error management of a memory device. Media management operations are initiated on a plurality of management units of one or more memory devices managed by a controller. A first error state and a second error state associated with a read phase of the media management operation performed on a first management unit of the plurality of management units are received. An error correction operation is performed on the first management unit in response to determining that the first error state and the second error state indicate correctable errors in the first management unit. An entry that maps a logical address to a physical address associated with the first management unit is locked in response to determining that no spare management unit is available.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to memory subsystems, and more particularly, to error management of CXL devices. Background Art

[0002] A memory subsystem may include one or more memory devices that store data. For example, the memory devices may be non-volatile memory devices and volatile memory devices. Generally, a host system may utilize the memory subsystem to store data at and retrieve data from the memory devices. Summary of the Invention

[0003] Aspects of the present disclosure relate to a method that includes: initiating a media management operation on a plurality of management units of one or more memory devices managed by a controller by a processing device of the controller; receiving, from the controller, a first error status associated with a read phase of the media management operation performed on a first management unit of the plurality of management units, wherein the read phase reads data from the first management unit for movement to a second management unit of the plurality of management units; receiving, from the one or more memory devices, a second error status associated with the read phase of the media management operation on the first management unit; in response to determining that the first error status and the second error status indicate a correctable error in the first management unit, performing an error correction operation on the first management unit; in response to an unsuccessful error correction of the correctable error in the first management unit, determining whether a spare management unit is available, wherein the spare management unit is located on the one or more memory devices but is not one of the plurality of management units; and in response to determining that no spare management unit is available, locking an entry in a logical-to-physical (L2P) data structure that maps a logical address to a physical address associated with the first management unit.

[0004] Another aspect of the present disclosure relates to a system, comprising: one or more memory devices; and a processing device coupled to the one or more memory devices, the processing device configured to perform operations including: in response to receiving a request from a host system, performing, by the processing device, a read operation on a management unit of the one or more memory devices; storing data from the management unit into a second memory device among the one or more memory devices; transmitting the data from the second memory device to the host system; receiving, from a controller managing the one or more memory devices, a first error status associated with the read operation; receiving, from a first memory device, a second error status associated with the read operation; in response to determining that the first error status and the second error status indicate a correctable error in the management unit, performing an error correction operation on the management unit; in response to an unsuccessful error correction of the correctable error in the management unit, determining whether a spare management unit of the first memory device is available; and in response to determining that no spare management unit is available, locking an entry in a logical-to-physical (L2P) data structure that maps a logical address to a physical address associated with the management unit.

[0005] Another aspect of the present disclosure relates to a non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations including: initiating, by a processing device of a controller, a media management operation on a plurality of management units of one or more memory devices managed by the controller; receiving, from the controller, a first error status associated with a read phase of the media management operation performed on a first management unit among the plurality of management units, wherein the read phase reads data from the first management unit for transfer to a second management unit among the plurality of management units; receiving, from the one or more memory devices, a second error status associated with the read phase of the media management operation performed on the first management unit; in response to determining that the first error status and the second error status indicate a correctable error in the first management unit, performing an error correction operation on the first management unit; in response to an unsuccessful error correction of the correctable error in the first management unit, determining whether a spare management unit is available, wherein the spare management unit is located on the one or more memory devices but is not one of the plurality of management units; and in response to determining that no spare management unit is available, locking an entry in a logical-to-physical (L2P) data structure that maps a logical address to a physical address associated with the first management unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments of the present disclosure. However, the drawings should not be regarded as limiting the present disclosure to a particular embodiment, but are for explanation and understanding only.

[0007] Figure 1 Illustrate an example computing system that includes a memory subsystem in accordance with some embodiments of the present disclosure.

[0008] Figure 2A Is an illustrative example of a logical-to-physical (L2P) data structure managed by an error management unit based on wear leveling operations in accordance with some embodiments of the present disclosure.

[0009] Figure 2B Is an illustrative example of an L2P data structure managed by an error management unit based on wear leveling operations in accordance with some embodiments of the present disclosure.

[0010] Figure 2C Is an illustrative example of an L2P data structure managed by an error management unit based on wear leveling operations in accordance with some embodiments of the present disclosure.

[0011] Figure 3 Is an illustrative error condition table used by an error management unit to modify an L2P data structure in accordance with some embodiments of the present disclosure.

[0012] Figure 4 Is a flowchart of an example method for error management of a CXL device in accordance with some embodiments of the present disclosure.

[0013] Figure 5 Is a flowchart of an example method for error management of a CXL device in accordance with some embodiments of the present disclosure.

[0014] Figure 6 Is a block diagram of an example computer system in which embodiments of the present disclosure may operate. Detailed Description

[0015] Aspects of the present disclosure relate to error management of a CXL device. The memory subsystem can be a storage device, a memory module, or a combination of a storage device and a memory module. Examples of storage devices and memory modules are described below in connection with Figure 1 Generally, a host system can utilize a memory subsystem that includes one or more memory components, such as a memory device that stores data. The host system can provide data stored at the memory subsystem and can request data retrieved from the memory subsystem.

[0016] The memory subsystem can include high-density non-volatile memory devices where it is desirable to retain data without power being supplied to the memory device. An example of a non-volatile memory device is a NAND memory device. Examples of storage devices and memory modules are described below in connection with Figure 1Describe other examples of non-volatile memory devices. A non-volatile memory device is a package of one or more dies. Each die may include one or more planes. For some types of non-volatile memory devices (e.g., NAND devices), each plane includes a set of physical blocks. Each block includes a set of pages. Each page includes a set of memory cells ("cells"). A cell is an electronic circuit that stores information. Depending on the cell type, a cell may store one or more bits of binary information and have various logical states related to the number of bits stored. The logical states may be represented by binary values such as "0" and "1" or combinations of such values.

[0017] The memory device may include a plurality of memory cells arranged in a two-dimensional or three-dimensional grid. The memory cells may be formed on a silicon wafer by an array of columns connected by conductive lines (hereinafter also referred to as bit lines or BLs) and rows connected by conductive lines (hereinafter also referred to as word lines or WLs). A word line may have a row of associated memory cells in the memory device, which together with one or more bit lines are used to generate an address for each of the memory cells. The intersection of a bit line and a word line constitutes the address of a memory cell. A block, hereinafter, refers to a unit of the memory device for storing data and may include a group of memory cells, a group of word lines, a word line, or an individual memory cell. One or more blocks may be grouped together to form a separate partition (e.g., a plane) of the memory device to allow concurrent operations to occur on each plane. The memory device may include circuitry for performing concurrent memory page accesses of two or more memory planes. For example, the memory device may include a plurality of access line driver circuits and power circuits that may be shared by the planes of the memory device to facilitate concurrent access to pages of different page types in two or more memory planes. For ease of description, these circuits may generally be referred to as independent plane driver circuits. Depending on the storage architecture employed, data may be stored across memory planes (i.e., in stripes). Thus, a single request to read a data segment (e.g., corresponding to one or more data addresses) may result in read operations being performed on two or more of the memory planes of the memory device.

[0018] Some memory devices (e.g., non-volatile memory devices) may have limited durability. For example, some memory devices may be written to, read from, or erased a limited number of times before the memory device begins to physically degrade or wear out and ultimately fails.

[0019] The local controller of a memory component (or memory device) can perform media management operations to mitigate the impact of physical wear on the memory device and extend the overall lifespan of the memory subsystem. For example, the local controller can perform wear leveling operations to distribute physical wear across the management units of the memory device. A management unit refers to a specific amount of memory of the memory device, such as a page or a block. To perform a wear leveling operation, the local controller can identify management units at the memory device that are experiencing a large amount of physical wear and can move the data stored at the management unit to another management unit that is experiencing a lesser amount of physical wear. In some examples, a management unit may experience a large amount of physical wear if a large number of memory access operations (e.g., write operations (i.e., programming operations) or read operations) are performed at the management unit. Thus, in some embodiments, the local controller can identify management units experiencing a large amount of physical wear based on, for example, the write count of each management unit. The write count refers to the number of write operations the local controller has performed at a management unit during the lifespan of the specific management unit. Data from a management unit with a high write count can be exchanged with data from a management unit with a low write count to attempt to evenly distribute wear across the management units of the memory component.

[0020] In some embodiments, the local controller can initiate gap wear leveling, which uses the mapping between logical addresses and physical addresses. The controller can then perform a wear leveling operation by periodically moving each management unit to its adjacent location, regardless of the write traffic to the management unit. Initiating the gap wear leveling operation employs two registers (a start register and a gap register) and additional memory management units to facilitate data movement. The controller can use the gap register to keep track of the number of management units that have been moved. When all the management units in the pool of management units designated for the wear leveling operation have been moved, the controller can increment the start register, thus keeping track of the number of times all the management units have been moved. The mapping of management units from logical addresses to physical addresses is done by operating on the gap and start registers using logical addresses, as explained below.

[0021] Specifically, the memory system designates multiple management units in the management unit pool for wear leveling operations, and each management unit corresponds to a physical address (e.g., 16 management units correspond to physical addresses 0 to 15). To implement start-gap wear leveling, an additional management unit is added at a location adjacent to the multiple management units (e.g., the location corresponding to physical address 16). All the management units including the additional management unit (e.g., 16 management units + 1 additional management unit) can form a circular pool. The additional management unit is a memory location that does not contain useful data. The start register and the gap register are initialized such that the start register initially points to a location corresponding to a physical address (e.g., physical address 0), and the gap register that always points to the location of the additional management unit initially points to a location corresponding to a physical address (e.g., physical address 16). After performing a predefined number of memory write operations, the content of the location decremented by 1 (e.g., the location corresponding to physical address 15) referenced by the gap register is copied to the location of the additional management unit (e.g., the location corresponding to physical address 16). That is, the content is moved to an adjacent location. Then, the value stored in the gap register is decremented such that it will reference the adjacent location, which will be the location identified by the physical address immediately preceding the physical address in the address range (e.g., physical address 15) after moving the content. After the number of moves of the gap register reaches the number of management units in the management unit pool for wear leveling operations (e.g., 16 moves), the gap register wraps around the address range (e.g., points to physical address 0). For the next data move, the gap register is reset to the initial address of the range (e.g., points to physical address 16), and since the content of all (e.g., 16) management units in the pool has been moved once, the start register is incremented by 1. Thus, each move of the gap register provides a remapping of the content (specific to a logical address) to its adjacent location (corresponding to a physical address).

[0022] During wear leveling operations, some management units may experience read errors. Specifically, data cannot be successfully read from the original location due to read interference, inability to accurately read the original data, or memory cells that cannot be reliably read, preventing it from being relocated to a new location. Due to the continuous movement and remapping of the incorrect locations, the memory subsystem controller cannot consistently and accurately obtain information corresponding to the incorrect locations from the local controller.

[0023] Aspects of the present disclosure address the above and other disadvantages by using a pool of management units to a physical (L2P) data structure based on error management logic associated with media management operations (such as wear leveling operations). The media management operation performs a read operation on a management unit using a logical address mapped to a physical address of a management unit in the pool of management units, where the management unit is adjacent to a gap management unit of the pool of management units (e.g., the next management unit after the gap management unit). One or more errors may be identified in response to the read operation performed on the management unit. For example, an error status may be received from a controller of the memory subsystem and another error status may be received from a memory device of the memory subsystem that contains the pool of management units. Based on the one or more error statuses, the read operation may result in no error, a correctable error, or an uncorrectable error at the physical address associated with the logical address. Specifically, based on an error condition table, which combinations of error results may result in no error, a correctable error, or an uncorrectable error are indicated.

[0024] If the read operation results in no error, then the read data is written to the gap management unit and an entry in the L2P data structure may be updated to reflect the mapping of the logical address to the physical address associated with the gap management unit. If the read operation results in a correctable error, then an error correction operation may be performed on the management unit to correct the error. Based on successful error correction, the read data is written to the gap management unit and the L2P data structure may be updated to reflect the mapping of the logical address to the physical address associated with the gap management unit. If the error correction is unsuccessful and a spare management unit (e.g., a management unit in another pool of management units) is available, then the read data is written to the spare management unit and an entry in the L2P data structure may be updated to reflect the mapping of the logical address to the physical address associated with the spare management unit. If the error correction is unsuccessful and no spare management unit is available, then the entry that maps the logical address to the physical address associated with the management unit, also referred to as a locked entry (e.g., locked), is maintained, e.g., locked in the L2P data structure. Similarly, if the read operation results in an uncorrectable error, then the entry that maps the logical address to the physical address associated with the management unit is maintained (e.g., locked) in the L2P data structure.

[0025] Advantages of the present disclosure include (but are not limited to) removing uncorrectable and / or irreparable management units from the pool of management units by maintaining (e.g., locking) the mapping of logical addresses to uncorrectable and / or irreparable management units, thereby providing the user with information about the location of uncorrectable and / or irreparable management units.

[0026] Figure 1Describe an example computing system 100 that includes a memory subsystem 110 according to some embodiments of the present disclosure. The memory subsystem 110 may include media such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., memory device 130), or a combination thereof.

[0027] The memory subsystem 110 may be a storage device, a memory module, or a combination of a storage device and a memory module. Examples of storage devices include solid state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controllers (eMMCs), universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).

[0028] The computing system 100 may be a computing device such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (e.g., an airplane, a drone, a train, an automobile, or other transportation means), an Internet of Things (IoT) enabled device, an embedded computer (e.g., an embedded computer included in a vehicle, an industrial device, or a networked commercial device), or such a computing device that includes a memory and a processing device.

[0029] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to multiple memory subsystems 110 of different types. Figure 1 Describe an example of a host system 120 coupled to one memory subsystem 110. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which may be an indirect communication connection or a direct communication connection (e.g., without an intermediate component), whether wired or wireless, including connections such as electrical, optical, magnetic, etc.

[0030] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 writes data to and reads data from the memory subsystem 110 using the memory subsystem 110, for example.

[0031] The host system 120 can be coupled to the memory subsystem 110 via a physical host interface. Examples of physical host interfaces include, but are not limited to, Serial Advanced Technology Attachment (SATA) interfaces, Peripheral Component Interconnect Express (PCIe) interfaces, Universal Serial Bus (USB) interfaces, Fibre Channel, Serial Attached SCSI (SAS), Double Data Rate (DDR) memory buses, Small Computer System Interface (SCSI), Dual In-line Memory Module (DIMM) interfaces (such as DIMM slot interfaces that support Double Data Rate (DDR)), and the like. The physical host interface can be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 via a physical host interface (such as a PCIe bus), the host system 120 can further utilize the Non-Volatile Memory Express (NVMe) interface to access components (such as the memory device 130). The physical host interface can provide an interface for transferring control, address, data, and other signals between the memory subsystem 110 and the host system 120. Figure 1 The memory subsystem 110 is illustrated as an example. In general, the host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0032] The memory devices 130, 140 can include any combination of different types of non-volatile memory devices and / or volatile memory devices. The volatile memory device (such as the memory device 140) can be, but is not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).

[0033] Some examples of non-volatile memory devices (such as the memory device 130) include NAND-type flash memory and write-in-place memory, such as three-dimensional cross-point (“3D cross-point”) memory devices, which are cross-point arrays of non-volatile memory cells. The cross-point array of non-volatile memory cells can perform bit storage based on changes in bulk resistance in combination with a stackable cross-grid format data access array. Additionally, compared to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, where non-volatile memory cells can be programmed without first erasing the non-volatile memory cells. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).

[0034] Each of the memory devices 130 may include one or more memory cell arrays. One type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), triple-level cells (TLC), quad-level cells (QLC), and penta-level cells (PLC), may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more memory cell arrays, such as SLC, MLC, TLC, QLC, PLC, or any combination thereof. In some embodiments, a particular memory device may include an SLC portion and an MLC portion, a TLC portion, a QLC portion, or a PLC portion of memory cells. The memory cells of the memory devices 130 may be grouped into pages, which may refer to logical units of the memory device for storing data. For some types of memory, such as NAND, pages may be grouped to form blocks. Some types of memory, such as 3D cross-point, may group pages across dies and channels to form management units (MUs).

[0035] Although non-volatile memory components such as 3D cross-point arrays of non-volatile memory cells and NAND-type flash memories (such as 2D NAND, 3D NAND) are described, the memory devices 130 may be based on any other type of non-volatile memory, such as read-only memory (ROM), phase change memory (PCM), self-selecting memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive-bridge RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), nor flash memory, or electrically erasable programmable read-only memory (EEPROM).

[0036] The memory subsystem controller 115 (or simply referred to as the controller 115) may communicate with the memory devices 130 to perform operations such as reading data, writing data, or erasing data and other such operations at the memory devices 130. The memory subsystem controller 115 may include hardware such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry having dedicated (i.e., hard-coded) logic for performing the operations described herein. The memory subsystem controller 115 may be a microcontroller, dedicated logic circuitry (such as a field programmable gate array (FPGA), application specific integrated circuit (ASIC), etc.), or other suitable processor.

[0037] The memory subsystem controller 115 may include processing circuitry configured to execute instructions stored in local memory 119, which includes one or more processors (e.g., processor 117). In the illustrated example, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines for controlling the operation of the memory subsystem 110, including handling communication between the memory subsystem 110 and the host system 120.

[0038] In some embodiments, the local memory 119 may include memory registers for storing memory pointers, fetching data, etc. The local memory 119 may also include a read-only memory (ROM) for storing microcode. Although Figure 1 the illustrated memory subsystem 110 has been shown to include a memory subsystem controller 115, in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115 and instead may rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).

[0039] Generally, the memory subsystem controller 115 may receive commands or operations from the host system 120 and may translate the commands or operations into instructions or appropriate commands to achieve a desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block address (LBA), namespace) associated with the memory device 130 and physical addresses (e.g., physical MU address, physical block address). The memory subsystem controller 115 may further include host interface circuitry to communicate with the host system 120 via a physical host interface. The host interface circuitry may translate commands received from the host system into command instructions to access the memory device 130 and translate responses associated with the memory device 130 into information for the host system 120.

[0040] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.

[0041] In some embodiments, the memory device 130 includes a local media controller 135 that operates in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller, such as the memory subsystem controller 115, may manage the memory device 130 externally (e.g., perform media management operations on the memory device 130). In some embodiments, the memory subsystem 110 is a managed memory device that is an original memory device 130 having control logic on the die (e.g., the local media controller 135) and a controller (e.g., the memory subsystem controller 115) for media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.

[0042] The memory subsystem 110 includes an error management component 113 that can remove uncorrectable and / or irreparable management units from the pool of management units used by the wear-leveling operation by maintaining (e.g., locking) a mapping of logical addresses to uncorrectable and / or irreparable management units, thereby providing information about the location of the uncorrectable and / or irreparable management units. In some embodiments, the memory subsystem controller 115 includes at least a portion of the error management component 113. In some embodiments, the error management component 113 is part of the host system 120, an application, or an operating system. In other embodiments, the local media controller 135 includes at least a portion of the error management component 113 and is configured to perform the functionality described herein.

[0043] The error management component 113 receives requests to perform memory management operations, such as wear-leveling operations on the memory devices 130 and / or 140 that are interconnected and synchronized, executing the same instructions and storing / retrieving data simultaneously (e.g., locking independent disk redundant array (LRAID)). This redundancy enables error detection and correction to ensure that the data stored and retrieved by each memory device matches. If a mismatch or error is detected between the redundant memory devices, it indicates a potential memory error and appropriate measures can be taken to address it.

[0044] The error management component 113 includes a wear leveling pool. The wear leveling pool (also referred to as a circular pool) includes a plurality of management units. For example, each virtual storage body has a wear leveling pool. A virtual storage body is a group of homogeneous storage bodies in a locked die (i.e., dies accessed in parallel through the same channel). At the beginning, all physical locations in the virtual storage body belong to the wear leveling pool. During the usage period, a certain location may be affected by a fault. When a fault is identified, the corresponding affected location will determine a vulnerability and it will be removed from the WL pool. The management units of the wear leveling pool are used as additional management units, also referred to as gap management units that do not contain useful data. Each management unit corresponds to a physical address. A logical to physical (L2P) data structure is used to maintain the mapping between the logical address and the physical address of the wear leveling pool.

[0045] The error management component 113 initiates a wear leveling operation. For each step of the wear leveling operation, the error management component 113 performs a read operation on the management units. The read operation is performed on the management unit adjacent to the gap management unit (e.g., the next management unit after the gap management unit).

[0046] The error management component 113 may receive one or more error correction codes (ECCs) based on the results of the read operations of the wear leveling operation. The error management component 113 may receive the DECTED ECC from all memory devices of the LRAID. For example, each memory device may use a double error correction, triple error detection (DEC-TED) scheme, which refers to the ability of the error correction code to correct up to 2 errors within a data word and detect up to 3 errors within a data word. Thus, if the DEC-TED ECCs of all memory devices do not detect an error during the read operation, the DECTED ECC is "0E", if at least one DEC-TED ECC of all memory devices detects and corrects 1 error, the DECTED ECC is "1E", if at least one DEC-TED ECC of all memory devices detects and corrects 2 errors, the DECTED ECC is "2E", and if at least one DEC-TED ECC of all memory devices detects 3 errors, the DECTED ECC is "3E". In some embodiments, the ECC mechanism may perform a single error correction (SEC) scheme, which refers to the ability of the error correction code to detect and correct a single error within a data word.

[0047] The error management component 113 may further receive LRAID ECC from the memory subsystem controller 115. For example, if no error is detected at LRAID, then the LRAID ECC is "NE", if at least one error is detected and corrected at LRAID, then the LRAID ECC is "CE", and if at least two errors are detected at LRAID, then the LRAID ECC is "UE".

[0048] The error management component 113 may determine the error condition of the read data based on the DEC-TED ECC and the LRAID ECC. For example, the error condition may be: "correct", which indicates that the data is correct and does not contain errors; "correctable", which indicates that the data contains correctable errors; and "uncorrectable", which indicates that the data contains uncorrectable errors. Specifically, the error management component 113 may manage an error status table that indicates which combinations of the DEC-TED ECC and the LRAID ECC can result in no error, correctable error, or uncorrectable error. The error status table may be defined by a logic circuit.

[0049] If the DEC-TED ECC is "0E" or "1E" and the LRAID ECC is "NE", then the error management component 113 may determine that the error condition is "correct". Accordingly, the error management component 113 performs a write operation to write the data read during the read operation to the gap management unit using the physical address of the gap management unit stored in the gap register of the error management component 113. The error management component 113 updates the entry of the L2P data structure that maps the logical address to the physical address of the management unit to map the logical address to the physical address of the gap management unit. The error management component 113 updates the gap register to point to the physical address of the management unit that becomes the new gap management unit. If the error condition is "correct" caused by the DEC-TED ECC being "1E", then the error management component 113 may record the DEC-TED ECC and the LRAID ECC in the error log of the error management component 113.

[0050] If the DEC-TED ECC is "2E" and the LRAID ECC is "NE" or if the LRAID ECC is "CE" regardless of the DEC-TED ECC, then the error management component 113 may determine that the error condition is "correctable". The error management component 113 may perform an error correction operation (such as a memory clear operation) to correct the error in the data and perform a memory access operation (such as a read operation) to access the corrected data. The memory clear operation corrects the error in the data and writes the corrected data back.

[0051] In some embodiments, the error correction operation can be successful and the error management component 113 may not receive an ECC (e.g., DEC-TED ECC or LRAID ECC). Accordingly, the error management component 113 performs a write operation to write the correction data to the gap management unit using the physical address of the gap management unit stored in the gap register of the error management component 113. The error management component 113 updates the entry of the L2P data structure that maps the logical address to the physical address of the management unit to map the logical address to the physical address of the gap management unit. The error management component 113 updates the gap register to point to the physical address of the management unit that becomes the new gap management unit. The error management component 113 may record the DEC-TED ECC and LRAID ECC to the error log of the error management component 113.

[0052] In some embodiments, the error management component 113 may receive at least one ECC (e.g., DEC-TED ECC and / or LRAID ECC) based on the result of a read operation after the error correction operation.

[0053] The error management component 113 may include a spare wear-leveling pool different from the wear-leveling pool. The spare wear-leveling pool may include a plurality of management units different from the plurality of management units of the wear-leveling pool. The error management component 113 may identify the management units of the available spare wear-leveling pool. The error management component 113 performs a write operation to write the correction data to the gap management unit using the physical address of the gap management unit stored in the gap register of the error management component 113. The error management component 113 updates the entry of the L2P data structure that maps the logical address to the physical address of the management unit to map the logical address to the physical address of the gap management unit. The error management component 113 updates the gap register to point to the identified management unit in the spare wear-leveling pool that becomes the new gap management unit. The error management component 113 may record the DEC-TED ECC and LRAID ECC to the error log of the error management component 113.

[0054] In some embodiments, the error management component 113 may not include a spare wear-leveling pool or may not identify any available management units of the spare wear-leveling pool. Accordingly, the error management component 113 does not update (e.g., maintain) the entries of the L2P data structure that map logical addresses to the physical addresses of the management units. In other words, the error management component 113 creates a hole in the L2P data structure by locking the entries of the L2P data structure that map logical addresses to the physical addresses of the management units. The error management component 113 does not update (e.g., maintain) the physical addresses in the gap register. The error management component 113 may log the DEC-TEDECC and the LRAID ECC to the error log of the error management component 113. In some embodiments, the error management component 113 may generate event record log information associated with the LRAID ECC that includes the physical address of the management unit recorded in the event log of the memory subsystem 110.

[0055] If the DEC-TED ECC is "3E" and the LRAID ECC is "NE" or if the LRAID ECC is "UE" and regardless of what the DEC-TED ECC is, then the error management component 113 may determine that the error condition is "uncorrectable". Accordingly, the error management component 113 does not update (e.g., maintain) the entries of the L2P data structure that map logical addresses to the physical addresses of the management units. In other words, the error management component 113 creates a hole in the L2P data structure by locking the entries of the L2P data structure that map logical addresses to the physical addresses of the management units. The error management component 113 does not update (e.g., maintain) the physical addresses in the gap register. The error management component 113 may log the DEC-TED ECC and the LRAID ECC to the error log of the error management component 113. In some embodiments, the error management component 113 may generate event record log information associated with the LRAID ECC that includes the physical address of the management unit recorded in the event log of the memory subsystem 110. In some embodiments, the memory subsystem may include a list that facilitates the discovery of faults in persistent memory (e.g., the memory devices of the LRAID). In some embodiments, the error management component 113 may update the list with the physical address of the management unit.

[0056] Depending on the embodiment, the error management component 113 may receive a request to perform a memory access operation (e.g., a read operation). The error management component 113 may receive one or more error correction codes (ECCs) based on the result of the read operation.

[0057] If the DEC-TED ECC is "0E" or "1E" and the LRAID ECC is "NE", then the error management component 113 may determine that the error condition is "correct". Accordingly, the error management component 113 performs a write operation to write the data read during the read operation to the cache of the storage subsystem 110. The error management component 113 may transfer the data from the cache to the host system 120. If the error condition is "correct" and the DEC-TED ECC is "1E", then the error management component 113 may record the DEC-TED ECC and the LRAID ECC. The error management component 113 may further perform an error correction operation (e.g., a memory scrubbing operation) to correct the error in the data.

[0058] If the DEC-TED ECC is "2E" and the LRAID ECC is "NE" or if the LRAID ECC is "CE" regardless of the DEC-TED ECC, then the error management component 113 may determine that the error condition is "correctable". The error management component 113 performs a write operation to write the data read during the read operation to the cache of the storage subsystem 110. The error management component 113 may transfer the data from the cache to the host system 120. The error management component 113 may perform an error correction operation (e.g., a memory scrubbing operation) to correct the error in the data and perform a memory access operation (e.g., a read operation) to access the corrected data. The memory scrubbing operation corrects the error in the data and writes the corrected data back.

[0059] In some embodiments, the error management component 113 may receive at least one ECC (e.g., DEC-TED ECC and / or LRAID ECC) based on the result of a read operation after an error correction operation. The error management component 113 may include a spare wear leveling pool. The spare wear leveling pool may include a plurality of management units. The error management component 113 may identify the management units of the available spare wear leveling pool. Accordingly, the error management component 113 performs a write operation to write the corrected data to the identified management units of the spare wear leveling pool using the physical addresses of the identified management units of the spare wear leveling pool.

[0060] In some embodiments, the error management component 113 may not include a spare wear leveling pool or may not identify any available management units of the spare wear leveling pool. Accordingly, the error management component 113 does not update (e.g., maintain) entries of the L2P data structure that map logical addresses to physical addresses of management units. In other words, the error management component 113 creates a hole in the L2P data structure by locking entries of the L2P data structure that map logical addresses to physical addresses of management units. The error management component 113 does not update (e.g., maintain) the physical addresses in the gap registers. The error management component 113 may log DEC-TEDECC and LRAID ECC to an error log of the error management component 113. In some embodiments, the error management component 113 may generate event record log information associated with the LRAID ECC that includes the physical address of the management unit recorded in the event log of the memory subsystem 110.

[0061] If the DEC-TED ECC is "3E" and the LRAID ECC is "NE" or if the LRAID ECC is "UE" and regardless of the DEC-TED ECC, then the error management component 113 may determine that the error condition is "uncorrectable". The error management component 113 performs a write operation to write the data read during a read operation to a cache of the storage subsystem 110. The error management component 113 may transfer the data to the host system 120. The error management component 113 does not update (e.g., maintain) entries of the L2P data structure that map logical addresses to physical addresses of management units. In other words, the error management component 113 creates a hole in the L2P data structure by locking entries of the L2P data structure that map logical addresses to physical addresses of management units. The error management component 113 does not update (e.g., maintain) the physical addresses in the gap registers. The error management component 113 may log DEC-TEDECC and LRAID ECC to an error log of the error management component 113. In some embodiments, the error management component 113 may generate event record log information associated with the LRAID ECC that includes the physical address of the management unit recorded in the event log of the memory subsystem 110. In some embodiments, the memory subsystem may include a list that facilitates discovery of faults in persistent memory (e.g., memory devices of LRAID). In some embodiments, the error management component 113 may update the list with the physical address of the management unit. More details regarding the operation of the error management component 113 are described below.

[0062] Figure 2A is an error management unit of a CXL device according to some embodiments of the present disclosure (e.g., Figure 1Illustrative example of the error management component 113 managing the L2P data structure based on the wear leveling operation 200. During the wear leveling operation 200, the error management component 113 may or may not modify the L2P data structure at each of the times 222, 224, 226, and / or 228.

[0063] First, at time 222 of the wear leveling operation 200, the L2P data structure shows that the logical address 210B maps to the physical address 212C corresponding to the management unit of the wear leveling pool, the logical address 210C maps to the physical address 212D corresponding to the management unit of the wear leveling pool, and the logical address 210D maps to the physical address 212E corresponding to the management unit of the wear leveling pool. The gap register of the error management component 113 may indicate that the physical address 212F corresponding to the management unit of the wear leveling pool is the gap management unit in the wear leveling pool.

[0064] The error management component 113 may perform a read operation on the management unit associated with the logical address 210D (e.g., the management unit at the physical address 212E) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 may perform a write operation to write the data from the management unit (e.g., the management unit at the physical address 212E) to the gap management unit (e.g., the management unit at the physical address 212F).

[0065] Thus, at time 224 of the wear leveling operation 200, the L2P data structure may indicate that the error management component 113 has updated the L2P mapping such that the logical address 210D points to the physical address 212F and the gap register is updated with the physical address 212E representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 212C, the logical address 210C maps to the physical address 212D, and the logical address 210D maps to the physical address 212F.

[0066] The error management component 113 may perform a read operation on the management unit associated with the logical address 210C (e.g., the management unit at the physical address 212D) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 may perform a write operation to write the data from the management unit (e.g., the management unit at the physical address 212D) to the gap management unit (e.g., the management unit at the physical address 212E).

[0067] Thus, at time 226 of the wear leveling operation 200, the L2P data structure may indicate that the error management component 113 has updated the L2P mapping such that the logical address 210C points to the physical address 212E and the gap register is updated with the physical address 212D representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 212C, the logical address 210C maps to the physical address 212E, and the logical address 210D maps to the physical address 212F. The gap register of the error management component 113 may indicate that the physical address 212D is the gap management unit.

[0068] The error management component 113 may perform a read operation on the management unit associated with the logical address 210B (e.g., the management unit at the physical address 212C) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 may perform a write operation to write data from the management unit (e.g., the management unit at the physical address 212B) to the gap management unit (e.g., the management unit at the physical address 212D).

[0069] Thus, at time 228 of the wear leveling operation 200, the L2P data structure may indicate that the error management component 113 has updated the L2P mapping such that the logical address 210B points to the physical address 212D and the gap register is updated with the physical address 212C representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 212D, the logical address 210C maps to the physical address 212E, and the logical address 210D maps to the physical address 212F. The gap register of the error management component 113 may indicate that the physical address 212C is the gap management unit.

[0070] Figure 2B This is an illustrative example where the error management unit of the CXL device (e.g., Figure 1 the error management component 113) manages the L2P data structure based on the wear leveling operation 230. The error management component 113 may or may not modify the L2P data structure at each of the times 232, 234, 236, and / or 238.

[0071] First, at time 232 of the wear leveling operation 230, the L2P data structure shows that the logical address 210B maps to the physical address 212C of the management unit corresponding to the wear leveling pool, the logical address 210C maps to the physical address 212D of the management unit corresponding to the wear leveling pool, and the logical address 210D maps to the physical address 212E of the management unit corresponding to the wear leveling pool. The gap register of the error management component 113 may indicate that the physical address 212F of the management unit corresponding to the wear leveling pool is the gap management unit in the wear leveling pool.

[0072] The error management component 113 may perform a read operation on a management unit associated with the logical address 210D (e.g., the management unit at physical address 212E) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 may perform a write operation to write data from the management unit (e.g., the management unit at physical address 212E) to a gap management unit (e.g., the management unit at physical address 212F).

[0073] Therefore, at time 234 of the wear leveling operation 230, the L2P data structure may indicate that the error management component 113 has updated the L2P mapping such that the logical address 210D points to the physical address 212F and the gap register is updated with the physical address 212E representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 212C, the logical address 210C maps to the physical address 212D, and the logical address 210D maps to the physical address 212F.

[0074] The error management component 113 may perform a read operation on a management unit associated with the logical address 210C (e.g., the management unit at physical address 212D) to read the data of the management unit. Due to the read operation, the error management component 113 may receive one or more ECC codes. If the error condition caused by the one or more ECC codes is "correctable", then the error management component 113 may perform an error correction operation (e.g., a memory clear operation) to correct the error in the data (e.g., the corrected data) and perform a memory access operation (e.g., a read operation) to access the corrected data.

[0075] Due to the read operation after the error correction operation, the error management component 113 may continue to receive one or more ECC codes. The error management component 113 may determine and identify the management unit of an available spare wear leveling pool (e.g., the management unit at physical address 216D). The error management component 113 may perform a write operation to write the corrected data from the management unit (e.g., the management unit at physical address 212D) to a gap management unit (e.g., the management unit at physical address 212E).

[0076] Therefore, at time 236 of the wear leveling operation 230, the L2P data structure may indicate that the error management component 113 has updated the L2P mapping such that the logical address 210C points to the physical address 212E and the gap register is updated with the physical address 216D representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 212C, the logical address 210C maps to the physical address 212E, and the logical address 210D maps to the physical address 212F.

[0077] The error management component 113 may perform a read operation on the management unit associated with the logical address 210B (e.g., the management unit at physical address 212C) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 may perform a write operation to write data from the management unit (e.g., the management unit at physical address 212B) to the gap management unit (e.g., the management unit at physical address 216D).

[0078] Thus, at time 238 of the wear leveling operation 230, the L2P data structure may indicate that the error management component 113 has updated the L2P mapping such that the logical address 210B points to the physical address 216D and the gap register is updated with the physical address 212C representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 216D, the logical address 210C maps to the physical address 212E, and the logical address 210D maps to the physical address 212F. The gap register of the error management component 113 may indicate that the physical address 212C is the gap management unit.

[0079] Figure 2C This is an illustrative example in which the error management unit of the CXL device (e.g., Figure 1 the error management component 113) manages the L2P data structure based on the wear leveling operation 240, and the error management component 113 may or may not modify the L2P data structure at each of the times 242, 244, 246, and / or 248.

[0080] First, at time 242 of the wear leveling operation 240, the L2P data structure may indicate that the logical address 210B maps to the physical address 212C of the management unit corresponding to the wear leveling pool, the logical address 210C maps to the physical address 212D of the management unit corresponding to the wear leveling pool, and the logical address 210D maps to the physical address 212E of the management unit corresponding to the wear leveling pool. The gap register of the error management component 113 may indicate that the physical address 212F is the gap management unit of the wear leveling pool.

[0081] The error management component 113 may perform a read operation on the management unit associated with the logical address 210D (e.g., the management unit at physical address 212E) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 may perform a write operation to write data from the management unit (e.g., the management unit at physical address 212E) to the gap management unit (e.g., the management unit at physical address 212F).

[0082] Thus, at time 244 of wear leveling operation 240, the L2P data structure may indicate that error management component 113 has updated the L2P mapping such that logical address 210D points to physical address 212F and the gap register is updated with physical address 212E representing the new gap management unit. Thus, the L2P data structure shows that logical address 210B maps to physical address 212C, logical address 210C maps to physical address 212D, and logical address 210D maps to physical address 212F.

[0083] Error management component 113 may perform a read operation on the management unit associated with logical address 210C (e.g., the management unit at physical address 212D) to read the data of the management unit. Due to the read operation, error management component 113 may receive one or more ECC codes.

[0084] In an embodiment, if the error condition caused by one or more ECC codes is "correctable", then error management component 113 may perform an error correction operation (e.g., a memory scrubbing operation) to correct the error in the data (e.g., the corrected data) and perform a memory access operation (e.g., a read operation) to access the corrected data. Due to the read operation after the error correction operation, error management component 113 may continue to receive one or more ECC codes. Error management component 113 may not identify the management unit of the available spare wear leveling pool (e.g., Figure 2B the management unit at physical address 216D). Error management component 113 may create a hole in the L2P data structure due to maintaining (e.g., locking) the mapping of logical address 210C to physical address 212D. In another embodiment, if the error condition caused by one or more ECC codes is "uncorrectable", then error management component 113 may create a hole in the L2P data structure due to maintaining (e.g., locking) the mapping of logical address 210C to physical address 212D.

[0085] Thus, at time 246 of wear leveling operation 240, the L2P data structure may indicate that error management component 113 maintains (e.g., does not update) the L2P mapping and thus does not update the gap register. Thus, the L2P data structure shows that logical address 210B maps to physical address 212C, logical address 210C still maps to physical address 212D, and logical address 210D maps to physical address 212F.

[0086] The error management component 113 can perform a read operation on a management unit associated with the logical address 210B (e.g., the management unit at physical address 212C) to read the data of the management unit. Due to the read operation, the error management component 113 may not receive an ECC code. The error management component 113 can perform a write operation to write data from a management unit (e.g., the management unit at physical address 212B) to a gap management unit (e.g., the management unit at physical address 212E).

[0087] Thus, at time 248 of the wear leveling operation 240, the L2P data structure can indicate that the error management component 113 has updated the L2P mapping such that the logical address 210B points to the physical address 212E and the gap register is updated with the physical address 212C representing the new gap management unit. Thus, the L2P data structure shows that the logical address 210B maps to the physical address 212E, the logical address 210C maps to the physical address 212D, and the logical address 210D maps to the physical address 212F. The gap register of the error management component 113 can indicate that the physical address 212C is the gap management unit.

[0088] Figure 3 is an illustrative error condition table 300 for a CXL device's error management unit (e.g., Figure 1 the error management component 113) to modify the L2P data structure according to some embodiments of the present disclosure.

[0089] The error condition table 300 can be stored in the local memory 119 of the memory subsystem 110. The error condition table 300 contains multiple rows (e.g., "NE", "CE", and "UE") identified by LRAID ECC. The LRAID ECC is received by the error management component. "NE" indicates that no error is detected at the LRAID. "CE" indicates that at least one error is detected and corrected at the LRAID. "UE" indicates that at least two errors are detected at the LRAID.

[0090] The error condition table 300 contains multiple columns (e.g., "0E", "1E", "2E", and "3E") identified by DEC-TED ECC. The DEC-TED ECC is received by the error management component. "0E" indicates that no error of DEC-TED ECC is detected in all memory devices of the LRAID. "1E" indicates that at least one DEC-TED ECC of all memory devices of the LRAID detects and corrects an error. "2E" indicates that at least one DEC-TED ECC of all memory devices of the LRAID detects and corrects two errors. "3E" indicates that at least one DEC-TED ECC of all memory devices of the LRAID detects three errors.

[0091] Each intersection point between DEC-TED ECCs (such as "0E", "1E", "2E", and "3E") and LRAID ECCs (such as "NE", "CE", and "UE") represents an error condition (such as "correct", "correctable", or "uncorrectable"). For example, the DEC-TED ECC of "0E" and the LRAID ECC of "NE" result in a "correct" error condition, the DEC-TED ECCs of "1E" or "2E" and the LRAID ECC of "NE" result in a "correctable" error condition, the DEC-TED ECCs of "0E", "1E", "2E", or "3E" and the LRAID ECC of "CE" result in a "correctable" error condition, the DEC-TED ECC of "3E" and the LRAID ECC of "NE" result in an "uncorrectable" error condition, and the DEC-TED ECCs of "0E", "1E", "2E", or "3E" and the LRAID ECC of "UE" result in an "uncorrectable" error condition.

[0092] Due to a read operation, the error management component 113 can use the received ECCs (such as DEC-TED ECC and LRAID ECC) to query the error condition table 300 to determine the error condition. The read operation can be initiated by a wear leveling operation or a host read request. Based on the determined error condition, the error management component 113 can update the L2P data structure or create a hole in the L2P data structure, as previously described.

[0093] Figure 4 is a flowchart of an example method 400 for error management of a CXL device according to some embodiments of the present disclosure. Method 400 can be executed by processing logic, which can include hardware (such as a processing device, circuitry, dedicated logic, programmable logic, microcode, the hardware of the device, an integrated circuit, etc.), software (such as instructions running or executing on a processing device), or a combination thereof. In some embodiments, method 400 is executed by Figure 1 the error management component 113. Although shown in a specific sequence or order, the order of the process can be modified unless otherwise specified. Therefore, the illustrated embodiments should only be understood as examples, and the illustrated process can be executed in a different order, and some processes can be executed in parallel. Additionally, one or more processes can be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0094] At operation 410, processing logic, via a processing device of the controller, initiates a media management operation on a plurality of management units of one or more memory devices managed by the controller. As previously described, the plurality of management units of the one or more memory devices can be a wear-leveling pool that includes additional management units (e.g., gap management units) that do not contain useful data. Each management unit corresponds to a physical address. A logical-to-physical (L2P) data structure is used to maintain a mapping between logical addresses and physical addresses of the wear-leveling pool. As previously described, the media management operation can be a wear-leveling operation. Performing the wear-leveling operation includes performing a read operation (e.g., a read phase). The read phase reads data from a first management unit to move to a second management unit among the plurality of management units. The first management unit refers to the management unit corresponding to the logical address that is typically adjacent to the second management unit. The second management unit refers to a gap management unit.

[0095] At operation 420, processing logic receives a first error status associated with a read phase of a media management operation performed on a first management unit among the plurality of management units from the controller. At operation 430, processing logic receives a second error status associated with the read phase of the media management operation on the first management unit from the one or more memory devices. As previously described, the one or more memory devices can include a double error correction, triple error detection (DEC-TED) scheme, which refers to the ability of an error correction code to correct up to two errors within a data word and detect up to three errors within the data word.

[0096] At operation 440, in response to determining that the first error status and the second error status indicate a correctable error in the first management unit, processing logic performs an error correction operation (e.g., a memory clear operation) on the first management unit. As previously described, the first and second errors can be used to determine an error condition of the read data. The error condition can be no error, a correctable error, and an uncorrectable error.

[0097] At operation 450, in response to an unsuccessful error correction of a correctable error in the first management unit, processing logic determines whether a spare management unit is available. The spare management unit is located on the one or more memory devices but is not one of the plurality of management units. As previously described, the spare management unit is from a spare wear-leveling pool that is different from the wear-leveling pool.

[0098] At operation 460, in response to determining that no spare management unit is available, the processing logic locks an entry in the logical-to-physical (L2P) data structure that maps a logical address to a physical address associated with the first management unit. Locking the entry in the L2P data structure that maps a logical address to a physical address associated with the first management unit causes the processing logic to generate an event record based on the first error state and the second error state, which includes the logical address associated with the locked mapping entry and the physical address associated with the locked mapping entry.

[0099] Depending on the embodiment, in response to determining that a spare management unit is available, the processing logic performs a write phase of a media management operation on the spare management unit. Performing a wear leveling operation includes performing a write operation (e.g., the write phase). The write phase writes data read from the first management unit to the spare management unit. The processing logic then updates the entry in the L2P data structure that maps the logical address to the physical address associated with the spare management unit.

[0100] Depending on the embodiment, in response to successful error correction of a correctable error in the first management unit, the processing logic performs a write phase of a media management operation on a second management unit among the plurality of management units. The write phase writes data read from the first management unit to the second management unit. The processing logic then updates the entry in the L2P data structure that maps the logical address to the physical address associated with the second management unit.

[0101] Depending on the embodiment, in response to determining that the first error state and the second error state indicate no error in the first management unit, the processing logic performs a write phase of a media management operation on a second management unit among the plurality of management units. The write phase writes data read from the first management unit to the second management unit. The processing logic then updates the entry in the L2P data structure that maps the logical address to the physical address associated with the second management unit.

[0102] Depending on the embodiment, in response to determining that the first error state and the second error state indicate an uncorrectable error in the first management unit, the processing logic locks the entry in the L2P data structure that maps the logical address to the physical address associated with the first management unit. The processing logic then records the logical address associated with the first management unit to a list.

[0103] Depending on the embodiment, the processing logic records the physical address associated with the first management unit to an error log based on the first error state and the second error state associated with the first management unit.

[0104] Figure 5FIG. 500 is a flow chart of an example method for error management of a CXL device in accordance with some embodiments of the present disclosure. Method 500 may be executed by processing logic that may include hardware (e.g., a processing device, circuitry, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executing on a processing device), or a combination thereof. In some embodiments, method 500 is executed by the error management component 113 of Figure 1 Although shown in a particular sequence or order, the order of the processes may be modified unless otherwise specified. Accordingly, the illustrated embodiments should be understood only as examples, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. Additionally, one or more processes may be omitted in various embodiments. Accordingly, not all processes are required in every embodiment. Other process flows are possible.

[0105] At operation 510, in response to receiving a request from a host system, the processing logic performs a read operation on a management unit of a first memory device among one or more memory devices via a processing device. The first memory device is a non-volatile memory device. At operation 520, the processing logic stores data from the management unit into a second memory device among one or more memory devices. The second memory device is a volatile memory device. At operation 530, the processing logic transfers the data from the second memory device to the host system.

[0106] At operation 540, the processing logic receives a first error status associated with the read operation from a controller that manages one or more memory devices. At operation 550, the processing logic receives a second error status associated with the read operation from the first memory device. As previously described, one or more memory devices may include a double error correction, triple error detection (DEC-TED) scheme, which refers to the ability to correct up to two errors within a data word and detect up to three errors within a data word with an error correction code.

[0107] At operation 560, in response to determining that the first error status and the second error status indicate a correctable error in the management unit, the processing logic performs an error correction operation (e.g., a memory scrubbing operation) on the management unit. As previously described, the first and second errors may be used to determine an error condition of the read data. The error condition may be no error, a correctable error, and an uncorrectable error.

[0108] At operation 570, in response to an unsuccessful error correction of a correctable error in the management unit, the processing logic determines whether a spare management unit of the first memory device is available. The spare management unit is located on one or more memory devices but is not one of the multiple management units. As previously described, the spare management unit is from a spare wear leveling pool different from the wear leveling pool.

[0109] At operation 580, in response to determining that no spare management unit is available, the processing logic locks the entry in the logical-to-physical (L2P) data structure that maps a logical address to the physical address associated with the first management unit. Locking the entry in the L2P data structure that maps a logical address to the physical address associated with the first management unit causes the processing logic to generate an event record based on the first error state and the second error state, which includes the logical address associated with the locked mapping entry and the physical address associated with the locked mapping entry.

[0110] Depending on the embodiment, in response to determining that a spare management unit is available, the processing logic performs the write phase of the media management operation on the spare management unit. Performing the wear leveling operation includes performing a write operation (e.g., the write phase). The write phase writes the data read from the first management unit to the spare management unit. The processing logic then updates the entry in the L2P data structure that maps the logical address to the physical address associated with the spare management unit.

[0111] Depending on the embodiment, in response to determining that the first error state and the second error state indicate an uncorrectable error in the first management unit, the processing logic locks the entry in the L2P data structure that maps a logical address to the physical address associated with the first management unit. The processing logic then records the logical address associated with the first management unit to a list.

[0112] Depending on the embodiment, the processing logic records the physical address associated with the first management unit to an error log based on the first error state and the second error state associated with the first management unit.

[0113] Figure 6 An example machine of computer system 600 is illustrated, within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein. In some embodiments, computer system 500 can correspond to a host system (e.g., Figure 1 host system 120) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 memory subsystem 110) or can be used to perform the operations of a controller (e.g., for executing an operating system to perform the operations corresponding to Figure 1 error management component 113). In alternative embodiments, the machine can be connected (e.g., networked) to other machines in a LAN, intranet, extranet, and / or the Internet. The machine can operate as a server or client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.

[0114] The machine can be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network device, a server, a network router, a switch, or a bridge, or any machine capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by the machine. Additionally, although a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methods discussed herein.

[0115] Example computer system 600 includes a processing device 602, a main memory 604 (such as read only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM) or RDRAM), etc.), a static memory 606 (such as flash memory, static random access memory (SRAM), etc.), and a data storage system 618, which communicate with each other via a bus 630.

[0116] The processing device 602 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or several processors implementing a combination of instruction sets. The processing device 602 can also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device 602 is configured to execute instructions 626 for performing the operations and steps discussed herein. The computer system 600 can further include a network interface device 608 for communicating via a network 620.

[0117] The data storage system 618 can include a machine-readable storage medium 624 (also referred to as a computer-readable medium) having stored thereon one or more sets of instructions 626 or software embodying any one or more of the methods or functions described herein. The instructions 626 can also reside, completely or at least partially, within the main memory 604 and / or within the processing device 602 during execution by the computer system 600, and the main memory 604 and the processing device 602 also constitute machine-readable storage media. The machine-readable storage medium 624, the data storage system 618, and / or the main memory 604 can correspond to Figure 1 the memory subsystem 110.

[0118] In one embodiment, the instructions 626 include those for implementing a corresponding error management component (such as Figure 1instructions for the functionality of the error management component 113). Although the machine-readable storage medium 624 is shown as a single medium in the example embodiments, the term "machine-readable storage medium" should be regarded as including a single medium or multiple media that store one or more sets of instructions. The term "machine-readable storage medium" should also be regarded as including any medium that is capable of storing or encoding a set of instructions for execution by a machine and that causes the machine to perform any one or more of the methods of the present disclosure. Thus, the term "machine-readable storage medium" should be regarded as including (but not limited to) solid-state memories, optical media, and magnetic media.

[0119] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived as a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0120] However, it should be borne in mind that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may relate to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the memories or registers or other such information storage systems of the computer system.

[0121] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purposes, or it may comprise a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as (but not limited to) any type of disk (including floppy disks, optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to the computer system bus.

[0122] The algorithms and displays presented herein are not inherently related to any particular computer or other device. A variety of general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized devices to perform the methods. The structure of various such systems will appear as set forth in the appended claims. Additionally, the present disclosure has been described without reference to any particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the present disclosure described herein.

[0123] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, the instructions being usable to program a computer system (or other electronic device) to perform a process in accordance with the present disclosure. The machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, the machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium such as read only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory components, etc.

[0124] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments of the present disclosure. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure set forth in the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A method, comprising: initiating, by a processing device of a controller, a media management operation on a plurality of management units of one or more memory devices managed by the controller; receiving, from the controller, a first error status associated with a read phase of the media management operation performed on a first management unit of the plurality of management units, wherein the read phase reads data from the first management unit for transfer to a second management unit of the plurality of management units; receiving, from the one or more memory devices, a second error status associated with the read phase of the media management operation on the first management unit; performing, in response to determining that the first error status and the second error status indicate a correctable error in the first management unit, an error correction operation on the first management unit; determining, in response to an unsuccessful error correction of the correctable error in the first management unit, whether a spare management unit is available, wherein the spare management unit is located on the one or more memory devices but is not one of the plurality of management units; and locking, in response to determining that no spare management unit is available, an entry in a logical-to-physical L2P data structure that maps a logical address to a physical address associated with the first management unit.

2. The method according to claim 1, further comprising: performing, in response to determining that the spare management unit is available, a write phase of the media management operation on the spare management unit, wherein the write phase writes the data read from the first management unit to the spare management unit; and updating, in the L2P data structure, an entry that maps the logical address to a physical address associated with the spare management unit.

3. The method according to claim 1, further comprising: performing, in response to a successful error correction of the correctable error in the first management unit, a write phase of the media management operation on the second management unit of the plurality of management units, wherein the write phase writes the data read from the first management unit to the second management unit; and updating, in the L2P data structure, an entry that maps the logical address to a physical address associated with the second management unit.

4. The method according to claim 1, further comprising: performing, in response to determining that the first error status and the second error status indicate no error in the first management unit, a write phase of the media management operation on the second management unit of the plurality of management units, wherein the write phase writes the data read from the first management unit to the second management unit; and updating, in the L2P data structure, an entry that maps the logical address to a physical address associated with the second management unit.

5. The method according to claim 1, further comprising: locking, in response to determining that the first error status and the second error status indicate an uncorrectable error in the first management unit, an entry in the L2P data structure that maps a logical address to a physical address associated with the first management unit.

6. The method according to claim 4, wherein determining that the first error state and the second error state indicate an uncorrectable error in the first management unit includes: Recording the logical address associated with the first management unit into a list.

7. The method according to claim 1, further comprising: Based on the first error state and the second error state associated with the first management unit, recording the physical address associated with the first management unit into an error log.

8. The method according to claim 1, wherein locking the entry in the L2P data structure that maps the logical address to the physical address associated with the first management unit includes: Generating an event record based on the first error state and the second error state, wherein the event record includes the logical address associated with the locked mapping entry and the physical address associated with the locked mapping entry.

9. The method according to claim 1, wherein the error correction operation is a memory clear operation.

10. A system, comprising: One or more memory devices; And A processing device coupled to the one or more memory devices, the processing device being configured to perform operations including: In response to receiving a request from a host system, performing, by the processing device, a read operation on a management unit of the one or more memory devices; Storing data from the management unit into a second memory device among the one or more memory devices; Transmitting data from the second memory device to the host system; Receiving, from a controller managing the one or more memory devices, a first error state associated with the read operation; Receiving, from a first memory device, a second error state associated with the read operation; In response to determining that the first error state and the second error state indicate a correctable error in the management unit, performing an error correction operation on the management unit; In response to an unsuccessful error correction of the correctable error in the management unit, determining whether a spare management unit of the first memory device is available; And In response to determining that no spare management unit is available, locking an entry in a logical-to-physical L2P data structure that maps a logical address to a physical address associated with the management unit.

11. The system according to claim 10, wherein the operations that the processing device is configured to perform further include: In response to determining that a spare management unit is available, performing a write operation on the spare management unit, wherein the write operation writes the data read from the management unit into the spare management unit; and Updating an entry in the L2P data structure that maps the logical address to the physical address associated with the spare management unit.

12. The system according to claim 10, wherein the operations that the processing device is configured to perform further include: In response to determining that the first error state and the second error state indicate an uncorrectable error in the management unit, lock an entry in the L2P data structure that maps the logical address to the physical address associated with the management unit.

13. The system according to claim 12, wherein determining that the first error state and the second error state indicate an uncorrectable error in the management unit includes: Recording a logical address associated with the management unit into a list.

14. The system according to claim 10, wherein the processing device is configured to perform further operations including: based on the first error state and the second error state associated with the management unit, recording the physical address associated with the management unit into an error log.

15. The system according to claim 10, wherein locking the entry in the L2P data structure that maps the logical address to the physical address associated with the management unit includes: Generating an event record based on the first error state and the second error state, wherein the event record includes the logical address associated with the locked mapping entry and the physical address associated with the locked mapping entry.

16. The system according to claim 10, wherein the error correction operation is a memory clear operation.

17. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device, cause the processing device to perform operations including: Initiating a media management operation on a plurality of management units of one or more memory devices managed by a controller by a processing device of the controller; Receiving, from the controller, a first error state associated with a read phase of the media management operation performed on a first management unit of the plurality of management units, wherein the read phase reads data from the first management unit for transfer to a second management unit of the plurality of management units; Receiving, from the one or more memory devices, a second error state associated with the read phase of the media management operation on the first management unit; In response to determining that the first error state and the second error state indicate a correctable error in the first management unit, performing an error correction operation on the first management unit; In response to an unsuccessful error correction of the correctable error in the first management unit, determining whether a spare management unit is available, wherein the spare management unit is located on the one or more memory devices but is not one of the plurality of management units; and In response to determining that no spare management unit is available, locking an entry in a logical-to-physical L2P data structure that maps a logical address to a physical address associated with the first management unit.

18. The non-transitory computer-readable storage medium according to claim 17, further comprising: In response to determining that the spare management unit is available, performing a write phase of the media management operation on the spare management unit, wherein the write phase writes data read from the first management unit to the spare management unit; and Update the entry in the L2P data structure that maps the logical address to the physical address associated with the spare management unit.

19. The non-transitory computer-readable storage medium according to claim 17, further comprising: In response to successful error correction of the correctable error in the first management unit, perform a write phase of the media management operation on the second management unit among the plurality of management units, wherein the write phase writes the data read from the first management unit to the second management unit; And Update the entry in the L2P data structure that maps the logical address to the physical address associated with the second management unit.

20. The non-transitory computer-readable storage medium according to claim 17, further comprising: In response to determining that the first error state and the second error state indicate no error in the first management unit, perform a write operation of the media management operation on the second management unit among the plurality of management units, wherein the write operation writes the data read from the first management unit to the second management unit; And Update the entry in the L2P data structure that maps the logical address to the physical address associated with the second management unit.