Error recovery method, non-transitory computer storage medium, and memory subsystem

By extracting and configuring the error recovery parameters of the namespace in the controller of the memory subsystem, the problem of data access speed in the error recovery process in the prior art is solved, and the flexible demand for error recovery for different data applications is realized, and the trade-off between data access performance and accuracy is improved.

CN111831469BActive Publication Date: 2025-05-16MICRON TECHNOLOGY INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202010269299.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-04-23
Filing Date
2020-04-08
Publication Date
2025-05-16
Estimated Expiration
2041-05-16

AI Technical Summary

Technical Problem

Existing memory systems may reduce data access speed during error recovery, and different data applications have different requirements for error recovery, so it is difficult for the prior art to achieve an effective trade-off between data access performance and accuracy.

Method used

By extracting the identification and error recovery parameters of the namespace in the controller of the memory subsystem, the namespace is configured on the nonvolatile media according to these parameters, and the error recovery parameters are stored in association with the namespace to control the error recovery operation of data access.

Benefits of technology

It realizes dynamic adjustment of error recovery settings according to the needs of different data applications, improves the trade-off between data access performance and accuracy, and enhances the flexibility and efficiency of the memory system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111831469B_ABST
    Figure CN111831469B_ABST
Patent Text Reader

Abstract

The present application relates to an error recovery method, a non-transitory computer storage medium, and a memory subsystem. A memory subsystem having a non-volatile medium on which a plurality of namespaces are allocated. A command from a host system has an identification of a namespace and at least one error recovery parameter. A controller of the memory subsystem configures the namespace on the non-volatile medium according to the at least one error recovery parameter, stores the at least one error recovery parameter in association with the namespace, and controls an error recovery operation for data access in the namespace according to the at least one error recovery parameter stored in association with the namespace.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least some embodiments disclosed herein relate generally to customization of error recovery in memory systems, and more particularly, but not limited to, data storage devices. Background Art

[0002] The memory subsystem may include one or more memory components that store data. The memory subsystem may be a data storage system, such as a solid-state drive (SSD) or a hard disk drive (HDD). The memory subsystem may be a memory module, such as a dual in-line memory module (DIMM), a small outline DIMM (SO-DIMM), or a non-volatile dual in-line memory module (NVDIMM). The memory components may be, for example, non-volatile memory components and volatile memory components. Examples of memory components include memory integrated circuits. Some memory integrated circuits are volatile and require power to maintain the stored data. Some memory integrated circuits are non-volatile and can retain the stored data even when power is not applied. Examples of non-volatile memory include fast flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), and electrically erasable programmable read-only memory (EEPROM), etc. Examples of volatile memory include dynamic random access memory (DRAM) and static random access memory (SRAM). Generally speaking, a host system may utilize a memory subsystem to store data at and retrieve data from memory components.

[0003] A computer may include a host system and one or more memory subsystems attached to the host system. The host system may have a central processing unit (CPU) that communicates with the one or more memory subsystems to store and / or retrieve data and instructions. Instructions for a computer may include an operating system, device drivers, and application programs. The operating system manages resources in the computer and provides common services to application programs, such as memory allocation and time sharing of resources. Device drivers may operate or control a specific type of device in the computer; and the operating system uses the device drivers to provide resources and / or services provided by the device of that type. The central processing unit (CPU) of the computer system may run the operating system and device drivers to provide services and / or resources to the application program. The central processing unit (CPU) may run an application program that uses the services and / or resources. For example, an application program that implements a type of application of the computer system may instruct the central processing unit (CPU) to store data in a memory component of the memory subsystem and retrieve data from the memory component.

[0004] The host system may communicate with the memory subsystem according to a predefined communication protocol, such as the Non-Volatile Memory Host Controller Interface Specification (NVMHCI), also known as NVM Express (NVMe), which specifies a logical device interface protocol for accessing non-volatile storage devices via a Peripheral Component Interconnect Express (PCI Express or PCIe) bus. According to the communication protocol, the host system may send different types of commands to the memory subsystem; and the memory subsystem may execute the commands and provide responses to the commands. Some commands instruct the memory subsystem to store a data item at an address specified in the command, or to retrieve a data item from an address specified in the command, such as read commands and write commands. Some commands manage infrastructure and / or management tasks in the memory subsystem, such as commands to manage namespaces, commands to attach namespaces, commands to create input / output submission or completion queues, commands to delete input / output submission or completion queues, commands for firmware management, and the like.

[0005] The date retrieved from the memory component in a read operation may contain an error. The controller of the data storage device / system may retry the read operation to obtain error-free data. For example, a "NAND" type flash memory can be used to store one or more data bits in a memory cell by programming the voltage threshold level of the memory cell. When the voltage threshold level is within the voltage window, the memory cell is in a state associated with the voltage window and thus stores one or more value bits pre-associated with the state. However, various factors (e.g., temperature, time, charge leakage) may cause the voltage threshold of the memory cell to shift outside the voltage window, resulting in errors in determining the data stored in the memory cell. The data storage device / system can retry the read operation at a different voltage threshold to reduce and / or eliminate errors. Summary of the invention

[0006] In an embodiment of the present invention, a method is provided, the method comprising: receiving a command from a host system in a controller of a memory subsystem, the memory subsystem having a non-volatile medium; extracting, by the controller, an identification of a namespace and at least one error recovery parameter of the namespace; configuring, by the controller, the namespace on the non-volatile medium according to the at least one error recovery parameter; storing the at least one error recovery parameter in association with the namespace; and controlling, by the controller, an error recovery operation for data access in the namespace according to the at least one error recovery parameter stored in association with the namespace.

[0007] In an embodiment of the present invention, a non-transitory computer storage medium is provided. The non-transitory computer storage medium stores instructions, which, when executed by a controller of a memory subsystem, cause the controller to execute a method, the method comprising: receiving a command from a host system in the controller of the memory subsystem, the memory subsystem having a non-volatile medium; extracting, by the controller, an identification of a namespace and at least one error recovery parameter of the namespace; configuring, by the controller, the namespace on the non-volatile medium according to the at least one error recovery parameter; storing the at least one error recovery parameter in association with the namespace; and controlling, by the controller, an error recovery operation for data access in the namespace according to the at least one error recovery parameter stored in association with the namespace.

[0008] In an embodiment of the present invention, a memory subsystem is provided. The memory subsystem includes: a non-volatile medium; a processing device coupled to a buffer memory and the non-volatile medium; and an error recovery manager configured to: store at least one error recovery parameter in the data storage system in association with a namespace allocated on a portion of the non-volatile medium, wherein the namespace is one of a plurality of namespaces allocated on the non-volatile medium; and in response to determining that an input / output operation in the namespace has an error, retrieve the error recovery parameter and control an error recovery operation for the input / output operation according to the error recovery parameter. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Embodiments are illustrated by way of example and not limitation in the figures of the accompanying drawings in which like references indicate similar elements.

[0010] Figure 1 An example computing system having a memory subsystem according to some embodiments of the invention is described.

[0011] Figure 2 Describes customization of namespaces for different error recovery operations.

[0012] Figure 3 Demonstrates how to configure namespaces in a storage device.

[0013] Figure 4 Shows how to perform error recovery in a namespace.

[0014] Figure 5 is a block diagram of an example computer system in which embodiments of the invention may operate. DETAILED DESCRIPTION

[0015] At least some aspects of the present invention relate to techniques for customizing error recovery of selected portions of the storage capacity of a data storage device.

[0016] Error recovery operations in data storage devices slow down data access. It is advantageous to customize error recovery for different data applications to access performance and accuracy. For example, some data applications (e.g., video, audio, image data) can tolerate a certain degree of errors in the retrieved data without seriously affecting the performance of the application. However, some data applications (e.g., instruction code and file system data) must be as error-free as possible. Therefore, it may be desirable to prolong read operations due to repeated retries to retrieve error-free data for this data application. In general, a computing device may have different types of data applications.

[0017] In at least some embodiments disclosed herein, error recovery settings are configured based on namespaces. Namespaces can be dynamically mapped to different portions and / or locations in a data storage device. When different namespaces can have different error recovery settings, a computing device can use different namespaces to store different categories of data, so that a trade-off between data access performance and accuracy for different data applications can be achieved by storing the corresponding data in namespaces customized in different ways.

[0018] A namespace configured on a storage device may be considered to be a logical storage device dynamically allocated on a portion of the storage device. Multiple namespaces configured on the storage device correspond to different logical storage devices implemented using different portions of the storage capacity of the storage device. A computer system may access storage units in a logical storage device represented by a namespace by identifying a namespace and a logical block addressing (LBA) address defined within the namespace. Different namespaces may have separate LBA address spaces. For example, a first namespace allocated on a first portion of a storage device having n memory units may have an LBA address ranging from 0 to n-1; and a second namespace allocated on a second portion of a storage device having m memory units may have an LBA address ranging from 0 to m-1. The same LBA address may be used in different namespaces to identify different memory units in different portions of the storage device.

[0019] A namespace mapping may be used to map a logical address space defined in a namespace to an address space in a storage device. For example, an LBA address space defined in a namespace allocated on a portion of a storage device may be mapped in blocks to an LBA address space of a fictitious namespace that may be defined on the entire capacity of the storage device according to a customizable block size. The block size of the namespace mapping may be configured to balance the flexibility of defining a namespace according to blocks in a fictitious namespace and the overhead of performing address translation using a namespace mapping. The addresses in the fictitious namespace defined on the entire capacity may be further translated into physical addresses (e.g., via a flash translation layer (FTL) of a storage device) in a manner independent of the namespace. Thus, the complexity of the namespace is isolated from the operation of the flash translation layer (FTL). Further details of such techniques for namespace mapping may be found in U.S. Patent No. 10,223,254, issued on March 5, 2019 and entitled “Namespace Change Propagation in Non-Volatile Memory Devices,” the entire disclosure of which is hereby incorporated herein by reference.

[0020] Generally speaking, a host computer of a storage device may send a request to the storage device to create, delete, or reserve a namespace. After a portion of the storage capacity of the storage device is allocated to the namespace, an LBA address in the corresponding namespace logically represents a specific memory unit in the data storage medium of the storage device, but the specific memory unit logically represented by the LBA address in the namespace may physically correspond to a different storage unit (e.g., as in an SSD) at different time instances (e.g., as determined by a flash translation layer (FTL) of the storage device).

[0021] In general, a memory subsystem may also be referred to as a “memory device.” An example of a memory subsystem is a memory module connected to a central processing unit (CPU) via a memory bus. Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), non-volatile dual in-line memory modules (NVDIMMs), and the like.

[0022] Another example of a memory subsystem is a data storage device / system connected to a central processing unit (CPU) via a peripheral interconnect (e.g., input / output bus, storage area network). Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, and hard disk drives (HDDs).

[0023] In some embodiments, the memory subsystem is a hybrid memory / storage subsystem that provides both memory and storage functions. In general, a host system may utilize a memory subsystem that includes one or more memory components. The host system may provide data to be stored at the memory subsystem and may request to retrieve data from the memory subsystem.

[0024] Figure 1 An example computing system having a memory subsystem (110) according to some embodiments of the invention is illustrated.

[0025] The memory subsystem (110) may include non-volatile media (109) including memory components. In general, the memory components may be volatile memory components, non-volatile memory components, or a combination thereof. In some embodiments, the memory subsystem (110) is a data storage system. An example of a data storage system is an SSD. In other embodiments, the memory subsystem (110) is a memory module. Examples of memory modules include DIMMs, NVDIMMs, and NVDIMM-Ps. In some embodiments, the memory subsystem (110) is a hybrid memory / storage subsystem.

[0026] Generally speaking, a computing environment may include a host system (120) that uses a memory subsystem (110). For example, the host system (120) may write data to the memory subsystem (110) and read data from the memory subsystem (110).

[0027] The host system (120) may be part of a computing device, such as a desktop computer, a laptop computer, a network server, a mobile device, or such computing device that includes a memory and a processing device. The host system (120) may include or be coupled to a memory subsystem (110) so that the host system (120) can read data from the memory subsystem (110) or write data to the memory subsystem (110). The host system (120) may be coupled to the memory subsystem (110) via a physical host interface. As used herein, "coupled to" generally refers to a connection between components, which may be an indirect communication connection or a direct communication connection (e.g., without intermediate components), whether wired or wireless, including, for example, electrical, optical, magnetic, etc. Examples of physical host interfaces include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, a Fibre Channel, a Serial Attached SCSI (SAS), a Double Data Rate (DDR) memory bus, etc. The physical host interface may be used to transfer data and / or commands between the host system (120) and the memory subsystem (110). When the memory subsystem (110) is coupled to the host system (120) via a PCIe interface, the host system (120) may further utilize an NVM Express (NVMe) interface to access the non-volatile media (109). The physical host interface may provide an interface for transferring control, address, data, and other signals between the memory subsystem (110) and the host system (120). Figure 1 The memory subsystem (110) is illustrated as an example. In general, the host system (120) can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0028] The host system (120) includes a processing device (118) and a controller (116). The processing device (118) of the host system (120) may be, for example, a microprocessor, a central processing unit (CPU), a processing core of a processor, an execution unit, etc. In some cases, the controller (116) may be referred to as a memory controller, a memory management unit, and / or an initiator. In one example, the controller (116) controls communication through a bus coupled between the host system (120) and the memory subsystem (110).

[0029] In general, the controller (116) may send commands or requests to the memory subsystem (110) to obtain the desired access to the non-volatile media (109). The controller (116) may further include interface circuitry for communicating with the memory subsystem (110). The interface circuitry may convert responses received from the memory subsystem (110) into information for the host system (120).

[0030] The controller (116) of the host system (120) can communicate with the controller (115) of the memory subsystem (110) to perform operations such as reading data, writing data, or erasing data in the non-volatile media (109) and other such operations. In some cases, the controller (116) is integrated into the same package of the processing device (118). In other cases, the controller (116) is separate from the package of the processing device (118). The controller (116) and / or the processing device (118) may include hardware such as one or more integrated circuits and / or discrete components, buffer memory, cache memory, or a combination thereof. The controller (116) and / or the processing device (118) may be a microcontroller, a dedicated logic circuit (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or another suitable processor.

[0031] The non-volatile media (109) may include any combination of different types of non-volatile memory components. In some cases, volatile memory components may also be used. Examples of non-volatile memory components include "NAND" (NAND) type flash memory. The memory components in the media (109) may include one or more memory cell arrays, such as single-level cells (SLC) or multi-level cells (MLC) (e.g., triple-level cells (TLC) or quad-level cells (QLC)). In some embodiments, a particular memory component may include both an SLC portion and an MLC portion of a memory cell. Each of the memory cells may store one or more data bits (e.g., data blocks) used by the host system (120). Although non-volatile memory components such as NAND-type flash memory are described, the memory components used in the non-volatile media (109) may be based on any other type of memory. In addition, volatile memory may be used. In some embodiments, the memory components in the medium (109) may include, but are not limited to, random access memory (RAM), read-only memory (ROM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), phase change memory (PCM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, ferroelectric random access memory (FeTRAM), ferroelectric RAM (FeRAM), conductive bridging RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), "NOR" (NOR) flash memory, electrically erasable programmable read-only memory (EEPROM), nanowire-based non-volatile memory, memory incorporating memristor technology, or a cross-point array of non-volatile memory cells, or any combination thereof. The cross-point array of the non-volatile memory may perform bit storage based on changes in bulk resistance in conjunction with a stackable cross-grid data access array. In addition, in contrast to many flash-based memories, cross-point non-volatile memories can perform write-in-place operations, where non-volatile memory cells can be programmed without having to erase them in advance. In addition, memory cells of a memory component in the medium (109) can be grouped into memory pages or data blocks, which can refer to a unit of a memory component for storing data.

[0032] The controller (115) of the memory subsystem (110) can communicate with the memory components in the media (109) to perform operations such as reading data, writing data, or erasing data at the memory components, as well as other such operations (e.g., in response to commands dispatched by the controller (116) on a command bus). The controller (115) may include hardware such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The controller (115) may be a microcontroller, a dedicated logic circuit (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or another suitable processor. The controller (115) may include a processing device (117) (e.g., a processor) configured to execute instructions stored in a local memory (119). In the illustrated example, the buffer memory (119) of the controller (115) includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines for controlling the operation of the memory subsystem (110), including handling communications between the memory subsystem (110) and the host system (120). In some embodiments, the controller (115) may include memory registers to store memory pointers, fetched data, etc. The controller (115) may also include a read-only memory (ROM) for storing microcode. Figure 1 It has been described as including a controller (115), but in another embodiment of the present invention, the memory subsystem (110) may not include a controller (115), but may rely on external control (e.g., provided by an external host, or by a processor or controller separate from the memory subsystem).

[0033] In general, the controller (115) may receive commands or operations from the host system (120) and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory components in the medium (109). The controller (115) may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical block addresses and physical block addresses associated with the memory components in the medium (109). The controller (115) may further include host interface circuitry to communicate with the host system (120) via a physical host interface. The host interface circuitry may convert commands received from the host system into command instructions to access the memory components in the medium (109), and convert responses associated with the memory components into information for the host system (120).

[0034] The memory subsystem (110) may also include additional circuits or components not illustrated. In some embodiments, the memory subsystem (110) may include a cache or buffer (e.g., DRAM) and address circuits (e.g., row decoders and column decoders) that can receive addresses from the controller (115) and decode the addresses to access memory components of the medium (109).

[0035] The computing system includes an error recovery manager (113) in the memory subsystem (110) that applies different error recovery preferences / settings for different portions of the non-volatile media (109) based on parameters set for the namespace.

[0036] In some embodiments, a controller (115) in the memory subsystem (110) includes at least a portion of the error recovery manager (113). In other embodiments or combinations, a controller (116) and / or a processing device (118) in the host system (120) includes at least a portion of the error recovery manager (113). For example, the controller (115), the controller (116), and / or the processing device (118) may include logic circuitry that implements the error recovery manager (113). For example, the controller (115) or the processing device (118) (processor) of the host system (120) may be configured to execute instructions stored in a memory for performing the operations of the error recovery manager (113) described herein. In some embodiments, the error recovery manager (113) is implemented in an integrated circuit chip disposed in the memory subsystem (110). In other embodiments, the error recovery manager (113) is part of an operating system, a device driver, or an application of the host system (120).

[0037] The memory subsystem (110) may have one or more queues (e.g., 123) for receiving commands from the host system (120). For example, the queue (123) may be configured for typical input / output commands, such as read commands and write commands. For example, the queue (125) may be configured for management commands that are not typical input / output commands. The memory subsystem (110) may include one or more completion queues (121) for reporting the execution results of the commands in the command queue (123) to the host system (120).

[0038] The error recovery manager (113) is configured to control operations related to error handling in the memory subsystem (110). For example, data retrieved from the non-volatile media (109) is detected as containing errors (e.g., in error detection and error correction code (ECC) operations), and the error recovery manager (113) is configured to determine the namespace in which the error occurred. Error recovery settings for the corresponding namespace are retrieved to control the error recovery process. For example, when an error is detected in a particular namespace, a threshold level for retrying read operations associated with the particular namespace is retrieved; and the error recovery manager (113) configures the controller (115) to perform read operation retries that do not exceed the threshold level in an attempt to read error-free data. If the controller (115) fails to retrieve error-free data with the retry read operations at the threshold level, the controller (115) is configured to provide a version of the retrieved data (e.g., with fewer errors) in the completion queue (121) for reporting to the host system (120). Optionally, a flag may be set in the response provided in the completion queue (121) to indicate that the retrieved data contains errors. In addition, the error recovery manager (113) may configure the namespace to store data using a RAID (Redundant Array of Independent Disks) configuration to increase fault tolerance for data stored in the namespace. Alternatively or in combination, the error recovery manager (113) may configure the namespace so that the memory cells in the namespace store data in a mode that has a desired level of reliability in data retrieval. For example, the error recovery manager (113) may configure the namespace to operate memory cells assigned to the non-volatile media (109) in SLC mode (e.g., rather than in MLC, TLC, or QLC mode) to improve the life of the memory cells and / or data access speed using reduced storage capacity, as will be discussed further below.

[0039] Figure 2 The customization of the namespace (131, ..., 133) for different error recovery operations is described.

[0040] exist Figure 2 In the example, a plurality of namespaces (131, ..., 133) are allocated on a non-volatile medium (109) (e.g., Figure 1 ).

[0041] For example, namespace A (131) has a namespace mapping (137) that defines a mapping between logical address regions (141, ..., 143) in the namespace (131) and physical address regions (142, ..., 144) of corresponding memory regions (151, ..., 153) in the non-volatile medium (109). Using the namespace mapping (137), a read or write operation requested for a logical address defined in the namespace (131) can be translated into a corresponding read or write operation in a memory region (151, ..., or 153) in the non-volatile medium (109).

[0042] Figure 2 A simplified example of a namespace mapping (137) that maps logical address regions (141, ..., 143) to physical address regions (142, ..., 144) is illustrated. In an improved mapping technique, logical address regions (141, ..., 143) in a namespace (131) may be mapped to logical address regions (141, ..., 143) in a fictitious namespace configured throughout a non-volatile medium (109); and a flash translation layer may be configured to further map the logical address regions (141, ..., 143) in the fictitious namespace to memory cells in the non-volatile medium (109), as discussed in U.S. Pat. No. 10,223,254, issued on Mar. 5, 2019, and entitled "Namespace Change Propagation in a Non-Volatile Memory Device," the entire disclosure of which is hereby incorporated herein by reference. Thus, the error recovery techniques of the present invention are not limited to a specific technique of namespace mapping.

[0043] exist Figure 2 In the example, namespace A (131) may have multiple configuration parameters related to error recovery operations, such as RAID settings (161), SLC mode settings (163), and / or threshold levels (165) for error recovery.

[0044] For example, different namespaces (131, ..., 133) may have different RAID settings (eg, 161). For example, namespace A (131) may be configured to have RAID operation; and namespace B (133) may be configured not to have RAID operation.

[0045] Optionally, the error recovery manager (113) can be configured to support multiple levels of RAID operations, such as mirroring, parity, byte-level striping with parity, block-level striping with parity, block-level striping with distributed parity.

[0046] For example, when the namespace (131) is configured to enable RAID operations, the error recovery manager (113) configures the namespace mapping (137) to map from the logical address areas (141, ..., 143) in the namespace (131) to multiple sets of physical address areas (e.g., 142, ..., 144). This mapping may be performed by mapping the logical address areas (141, ..., 143) defined in the namespace (131) to a separate set of logical address areas defined in a fictitious namespace that is allocated across the entire capacity of the non-volatile medium (109), wherein the flash translation layer is configured to map the logical address areas defined in the fictitious namespace to physical addresses in the non-volatile medium (109) in a manner that is independent of the namespace and / or RAID operations of the namespace. Thus, the operation of the flash translation layer may be independent of namespace and / or RAID considerations. The error recovery manager (113) can translate RAID operations in the namespace (131) into separate operations for different sets of logical address regions defined in the fictitious namespace; and therefore, RAID operations can be enabled / disabled without changing the operation / functionality of the flash translation layer.

[0047] Different namespaces (131, ..., 133) may have different SLC mode settings (e.g., 163). For example, namespace A (131) may be configured to operate in SLC mode; and namespace B (133) may be configured to operate in non-SLC mode (e.g., in MLC / TLC / QLC mode).

[0048] When the namespace (131) is configured to be in a mode with reduced storage capacity to increase access speed, reduce error rate, and / or increase lifespan with respect to erase program cycles, the error recovery manager (113) can configure the namespace map (137) to map logical address areas (e.g., 141, ..., 143) defined in the namespace (131) to multiple sets of logical address areas defined in the virtual namespace, and translate operations in the namespace (131) into corresponding operations in the virtual namespace in a manner similar to RAID operations. In general, this mode with reduced storage capacity to increase access speed, reduce error rate, and / or increase lifespan with respect to erase program cycles can be implemented via RAID and / or reducing the number of bits stored per memory cell (e.g., from QLC mode to MLC mode or SLC mode).

[0049] Different namespaces (131, ..., 133) may have different error recovery levels (e.g., 165). For example, namespace A (131) may be configured to limit retries of read operations at a first threshold that is greater than a threshold for namespace B (133), such that when namespace A (131) and namespace B (133) have the same other settings, namespace A (131) is configured for fewer errors, while namespace B (133) is configured to increase access speed at a tolerable error level. In some embodiments, the recovery level (165) identifies a fixed limit for retries of input / output operations. In other embodiments, the recovery level (165) identifies an error rate, such that further retries may be skipped when the retrieved data has an error rate below that specified by the recovery level (165).

[0050] Error recovery settings (e.g., 161, 163, and 165) allow namespaces (131) to be customized with specific tradeoffs in performance, capacity, and data errors. Thus, different namespaces (e.g., 131, 133) can be customized as logical storage devices with different error recovery characteristics using the same underlying non-volatile media (109).

[0051] For example, some data is used in error-intolerant situations, such as instructions for operating systems and / or applications, metadata for file systems. This data may be stored in a namespace (e.g., 131) that is configured to perform exhaustive error recovery, use memory cells in SLC mode, and / or perform RAID operations for data redundancy and recovery.

[0052] For example, some data may tolerate a certain degree of error in the retrieved data, such as video files, audio files, and / or image files. Such data may be stored in a namespace (e.g., 133) that is configured to perform less data error recovery operations in the event of an error.

[0053] Settings (eg, 161, 163, 165) related to error recovery operations of a namespace may be specified by the host system (120) using commands for recreating and / or managing the namespace (131).

[0054] For example, the host system (120) may submit a command for creating a namespace (131) using a command queue (e.g., 123). The command includes parameters for settings (e.g., 161, 163, 165). In some cases, data specifying portions of a setting is optional; and the command may use a linked list to specify portions of a setting in any order reflected in the linked list. Thus, the settings (e.g., 161, 163, 165) do not have to be specified in a particular order.

[0055] Furthermore, after creating the namespace (131), the host system (120) may change portions of the settings (eg, 161, 163, 165) using commands sent to the memory subsystem (110) using a command queue (eg, 123).

[0056] For example, the level (165) of error recovery operations for a namespace (131) may be changed on the fly using commands in a command queue (eg, 123).

[0057] For example, after the namespace (131) is created, a command from the host system (120) may change the SLC mode setting (163) or the RAID setting (161). The error recovery manager (113) may implement the change by adjusting the namespace map (137) and / or restore the data of the namespace (131) according to the updated setting.

[0058] In one embodiment, the error recovery manager (113) stores settings (e.g., 161, 163, 165) for namespaces (e.g., 131, ..., 133) allocated in the memory subsystem (110) in a centralized error recovery table. When an input / output error occurs in a namespace (e.g., 131), the error recovery manager (113) retrieves the recovery level (e.g., 165) of the namespace (e.g., 131) from the error recovery table and controls the error recovery operation based on the recovery level (e.g., 165) of the namespace (e.g., 131).

[0059] Figure 3 A method for configuring a namespace (e.g., 131) in a storage device (e.g., 110) is shown. For example, a method combining Figure 2 The technology discussed in Figure 1 Implementation in a computer system Figure 3 method.

[0060] At block 171, the host system (120) transmits a command to the memory subsystem (110).

[0061] For example, commands may be communicated from the host system (120) to the memory subsystem (110) via the command queue (123) using a predefined communication protocol such as the Non-Volatile Memory Host Controller Interface Specification (NVMHCI), also known as NVM Express (NVMe).

[0062] The command may be a command for creating a namespace, a command for configuring a namespace, a command for changing an attribute of a namespace, or a command for adjusting an error recovery parameter in a namespace. The command includes an identification of a namespace (e.g., 131) and at least one error recovery parameter (e.g., 161, 163, 165) of the namespace (e.g., 131).

[0063] The memory subsystem (110) may be configured to support customization of a predefined set of error recovery parameters (e.g., 161, 163, 165). Each parameter in the predefined set has a default value. The command does not have to explicitly identify the default value of the corresponding parameter of the namespace. Therefore, it is sufficient to specify a customized value for the parameter that is different from the default value.

[0064] At block 173, the controller (115) of the memory subsystem (110) extracts from the command an identification of the namespace (e.g., 131) and at least one error recovery parameter (e.g., 161, 163, and / or 165) for the namespace (e.g., 131).

[0065] At block 175, the controller (115) configures the namespace (131) according to at least one error recovery parameter (eg, 161, 163, and / or 165).

[0066] For example, when a RAID setup (161) requires RAID operations, the controller (115) may configure a namespace mapping (137) to groups of logical address areas (141, ..., 143) defined in the namespace (131) as multiple groups of memory units in the non-volatile media (109), where each group of memory units may be considered as an independent disk for RAID operations.

[0067] At block 177, the controller (115) stores at least one error recovery parameter (eg, 161, 163, and / or 165) in association with the namespace (eg, 131).

[0068] For example, a centralized error recovery table may be used to store values ​​of parameters (eg, 161, 163, ..., 165) for each namespace (eg, 131) configured in the memory subsystem (110).

[0069] At block 179, the controller (115) controls error recovery operations for data access in the namespace (eg, 131) according to at least one error recovery parameter stored in association with the namespace (eg, 131).

[0070] For example, when the namespace (131) has redundant data stored according to a RAID setting (eg, 161), the error recovery manager (113) may perform calculations to recover error-free data from the redundant data.

[0071] For example, when an input / output operation has an error that needs to be resolved via a retry, the error recovery manager (113) limits the number of retry operations according to the recovery level (165). For example, a predetermined function or a lookup table may be used to determine the maximum number of retry operations allowed by the recovery level (165). Once the maximum number of retry operations is reached, the controller (115) may report the result of the input / output operation with the error to the host system (120) via the completion queue (121), such as Figure 4 As described in .

[0072] Figure 4 A method for performing error recovery in a namespace (e.g., 131) is shown. For example, a combination of Figure 2 The technology discussed in Figure 1 Implementation in a computer system Figure 4 method.

[0073] At block 181 , the controller ( 115 ) receives a command to read data at a logical block address in the namespace ( 131 ).

[0074] At block 183 , the controller ( 115 ) retrieves data from the non-volatile media ( 109 ) using the namespace map ( 137 ) of the namespace ( 131 ).

[0075] At block 185, the controller (115) determines whether the retrieved data is error-free.

[0076] If the retrieved data is not error-free, the controller (115) retrieves the error recovery level (165) associated with the namespace (137) at block 187. Otherwise, the controller reports the retrieved data at block 191.

[0077] At block 189, the controller (115) determines whether the retried read operation has reached an error recovery level associated with the namespace (137); if so, the controller (115) reports the retrieved data at block 191. Otherwise, the controller (115) retrieves data from the non-volatile media (109) again using the namespace map (137) of the namespace (131) at block 183. For example, the controller (115) may adjust parameters (e.g., reference voltages) used in the read operation in an attempt to obtain error-free data.

[0078] If the controller (115) determines at block (189) that the retried read operation has reached an error recovery level associated with the namespace (137), the controller (115) may report an error to the host system (120). For example, the error may be reported via a flag set in the response containing the retrieved data or via a separate message.

[0079] In some embodiments, the communication channel between the processing device (118) and the memory subsystem (110) includes a computer network, such as a local area network, a wireless local area network, a wireless personal area network, a cellular communication network, a broadband high-speed always-connected wireless communication connection (e.g., a contemporary or next-generation mobile network link); and the processing device (118) and the memory subsystem may be configured to communicate with each other using data storage management and usage commands similar to those in the NVMe protocol.

[0080] The memory subsystem (110) may generally have a non-volatile storage medium. Examples of non-volatile storage media include memory cells formed in integrated circuits and magnetic materials coated on hard disks. Non-volatile storage media can maintain data / information stored therein without consuming power. Memory cells can be implemented using various memory / storage technologies, such as NAND logic gates, NOR logic gates, phase change memory (PCM), magnetic memory (MRAM), resistive random access memory, cross-point storage, and memory devices (e.g., 3D XPoint memory). Cross-point memory devices use transistor-free memory elements, each of which has memory cells and selectors stacked together as a column. The memory element columns are connected via two vertical layers of wires, one of which is located above the memory element columns and the other is located below the memory element columns. Each memory element can be individually selected at the intersection of a wire on each of the two layers. Cross-point memory devices are fast and non-volatile, and can be used as a unified memory pool for processing and storage.

[0081] A controller (eg, 115) of a memory subsystem (eg, 110) may run firmware to perform operations in response to communications from a processing device (118). In general, firmware is a type of computer program that provides control, monitoring, and data manipulation of an engineering computing device.

[0082] Some embodiments involving the operation of the controller (115) and / or the error recovery manager (113) may be implemented using computer instructions executed by the controller (115), such as firmware for the controller (115). In some cases, hardware circuits may be used to implement at least some of the functionality. The firmware may be initially stored in a non-volatile storage medium or another non-volatile device and loaded into volatile DRAM and / or intra-processor cache memory for execution by the controller (115).

[0083] The non-transitory computer storage medium may be used to store instructions of the firmware of the memory subsystem (e.g., 110). When the instructions are executed by the controller (115) and / or the processing device (117), the instructions cause the controller (115) and / or the processing device (117) to perform the methods discussed above.

[0084] Figure 5 An example machine of a computer system (200) is illustrated within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system (200) may correspond to a host system (e.g., Figure 1 A host system (120) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 The memory subsystem (110) may be used to perform operations of the error recovery manager (113) (e.g., execute instructions to perform operations corresponding to the reference Figures 1 to 4 In some embodiments, the machine may be connected (e.g., using a network) to other machines. The machine may operate in the capacity of a server or a client user machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client user machine in a cloud computing infrastructure or environment.

[0085] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine. Further, while a single machine is described, the term "machine" shall also be taken to include any collection of machines that individually or collectively execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0086] The example computer system (200) includes a processing device (202), a main memory (204) (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM) (e.g., synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), static random access memory (SRAM), etc.), and a data storage system (218), which communicate with each other via a bus (230) (which may include multiple buses).

[0087] The processing device (202) represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device (202) may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device (202) is configured to execute instructions (226) for performing the operations and steps discussed herein. The computer system (200) may further include a network interface device (208) for communicating over a network (220).

[0088] The data storage system (218) may include a machine-readable storage medium (224) (also referred to as a computer-readable medium) on which is stored one or more sets of instructions (226) or software embodying any one or more of the methodologies or functions described herein. During execution of the instructions (226) by the computer system (200), the instructions (226) may also reside, in whole or in part, within the main memory (204) and / or within the processing device (202), the main memory (204) and the processing device (202) also constituting machine-readable storage media. The machine-readable storage medium (224), the data storage system (218), and / or the main memory (204) may correspond to Figure 1 A memory subsystem (110) is provided.

[0089] In one embodiment, the instructions (226) include instructions for implementing a corresponding error recovery manager (113) (e.g., referring to Figures 1 to 4The machine-readable storage medium (224) is a storage medium that is used to store or encode a set of instructions for execution by a machine and causes the machine to perform any one or more of the methods of the present invention. Thus, the term "machine-readable storage medium" should be taken to include, but not be limited to, solid-state memory, optical media, and magnetic media.

[0090] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is generally considered here to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulations of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common sense reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0091] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) numbers within the computer system's registers and memories into physical quantities similarly represented as within the computer system's memories or registers or other such information storage systems.

[0092] The present invention also relates to an apparatus for performing the operations herein. This apparatus may be specifically constructed for the intended purposes, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. This computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.

[0093] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used with programs according to the teachings herein, or it may prove convenient to construct more specialized equipment to perform the methods. The structures of various these systems will appear as described in the following description. In addition, the present invention is not described with reference to any particular programming language. It will be appreciated that various programming languages ​​may be used to implement the teachings of the present invention as described herein.

[0094] The present invention may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, which instructions may be used to program a computer system (or other electronic device) to perform a process according to the present invention. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.

[0095] In this description, various functions and operations are described as being performed by or caused by computer instructions to simplify the description. However, those skilled in the art will recognize that such expressions mean that the functions are generated by the execution of computer instructions by one or more controllers or processors (such as microprocessors). Alternatively or in combination, dedicated circuits with or without software instructions may be used to implement the functions and operations, such as using application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). Embodiments may be implemented using hard-wired circuits without software instructions or in combination with software instructions. Therefore, the technology is neither limited to any specific combination of hardware circuits and software, nor to any specific source of instructions executed by the data processing system.

[0096] In the foregoing specification, embodiments of the present invention have been described with reference to specific example embodiments of the invention. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of embodiments of the invention as set forth in the appended claims. Accordingly, the description and drawings are to be regarded in an illustrative rather than a restrictive sense.

Claims

1. A method performed in a memory subsystem, comprising: In a controller of the memory subsystem, receiving a command from a host system, the command including at least one error recovery parameter to specify a preference related to error recovery operations of a namespace, the memory subsystem having a non-volatile medium; extracting, by the controller, from the command, an identification of the namespace and the at least one error recovery parameter of the namespace; configuring, by the controller, the namespace on the non-volatile medium based on the at least one error recovery parameter; storing the at least one error recovery parameter associated with the namespace; as well as The error recovery operation for data access in the namespace is controlled by the controller according to the at least one error recovery parameter stored associated with the namespace.

2. The method of claim 1, wherein the memory subsystem is a data storage device; and the non-volatile media comprises NAND flash memory.

3. The method of claim 2, wherein the non-volatile medium of the data storage device has a plurality of namespaces; and the method further comprises: A plurality of sets of error recovery parameters are stored in association with the plurality of namespaces, respectively.

4. The method of claim 3, wherein configuring the namespace on the non-volatile medium comprises: A namespace mapping is generated for the namespace based at least in part on the at least one error recovery parameter.

5. The method of claim 4, wherein the namespace mapping defines a mapping from a first logical block address space defined in the namespace to a portion of a second logical block address space defined in a fictitious namespace allocated over the entire capacity of the non-volatile media.

6. The method of claim 4, wherein the at least one error recovery parameter comprises a Redundant Array of Independent Disks (RAID) configuration.

7. The method of claim 6, wherein the RAID setup is implemented at least in part based on the namespace mapping, the namespace mapping mapping a first local block address space defined in the namespace to a redundant portion of a second logical block address space defined in a fictitious namespace allocated over the entire capacity of the non-volatile media. 8 . The method of claim 6 , wherein the at least one error resilience parameter further comprises a single-level cell (SLC) mode setting.

9. The method of claim 8, wherein the at least one error recovery parameter further comprises a recovery level; and the controlling of the error recovery operation for data access in the namespace comprises: Retries of input / output operations performed on commands from the host system are limited according to the recovery level.

10. The method of claim 1, wherein the memory subsystem is a solid state drive; the non-volatile media comprises flash memory; and the method further comprises: The namespace is created on a portion of the non-volatile media in response to the command.

11. The method of claim 1 , wherein the command is a second command; and the method further comprises: receiving, in the controller, a first command for creating the namespace on a portion of the non-volatile media; as well as The namespace is created in response to the first command before receiving the second command.

12. The method of claim 11, wherein the second command changes a recovery level of the namespace while the namespace is in use.

13. A non-transitory computer storage medium storing instructions which, when executed by a controller of a memory subsystem, cause the controller to perform a method comprising: In a controller of a memory subsystem, receiving a command from a host system, the command including at least one error recovery parameter to specify preferences related to error recovery operations of a namespace, the memory subsystem having non-volatile media; extracting, by the controller, from the command, an identification of the namespace and the at least one error recovery parameter of the namespace; configuring, by the controller, the namespace on the non-volatile medium based on the at least one error recovery parameter; storing the at least one error recovery parameter associated with the namespace; as well as The error recovery operation for data access in the namespace is controlled by the controller according to the at least one error recovery parameter stored associated with the namespace.

14. The non-transitory computer storage medium of claim 13, wherein the method further comprises: In the controller, a command for changing the error recovery parameter of the namespace among a plurality of namespaces allocated on the nonvolatile medium is received.

15. A memory subsystem comprising: Non-volatile media; a processing device coupled to the buffer memory and the non-volatile medium; and An error recovery manager configured to: storing, in the memory subsystem, at least one error recovery parameter to specify a preference associated with error recovery operations for a namespace allocated on a portion of the non-volatile medium, wherein the namespace is one of a plurality of namespaces allocated on the non-volatile medium; as well as In response to determining that an input / output operation in the namespace has an error, the error recovery parameter is retrieved and the error recovery operation for the input / output operation is controlled according to the error recovery parameter.

16. The memory subsystem of claim 15, wherein the error recovery parameters identify a recovery level; and the error recovery manager is configured to limit retries of the input / output operations according to the recovery level.

17. The memory subsystem of claim 16, wherein the error recovery manager is configured to change the recovery level in response to a command from a host system.

18. The memory subsystem of claim 15, wherein the error recovery parameters identify a redundant array of independent disks (RAID) setup; and the error recovery manager is configured to perform error recovery based on the RAID setup.

19. The memory subsystem of claim 18, wherein the nonvolatile media comprises flash memory; and the memory subsystem is a solid-state drive configured to receive commands from a host system over a Peripheral Component Interconnect Express (PCI Express or PCIe) bus.

20. The memory subsystem of claim 17, wherein the processing device is configured to receive a command from the host system; and in response to the command, create the namespace and associate it with the at least one error recovery parameter.

Citation Information

Patent Citations

  • Namespace change propagation in non-volatile memory devices

    US10223254B1

  • Adjustable Error Correction Based on Memory Health in a Storage Unit

    US20160041870A1

  • Memory system and method of controlling nonvolatile memory

    US20170024276A1