Storage Class Memory

The memory system addresses SCM's finite write endurance by using a logical to virtual translation table and drift buffer to manage wear levels, ensuring data is written to healthy memory areas, thereby extending device lifespan and maintaining data integrity.

JP7702195B2Active Publication Date: 2025-07-03INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2022539008
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-01-07
Filing Date
2020-12-09
Publication Date
2025-07-03
Estimated Expiration
2040-12-09

AI Technical Summary

Technical Problem

Storage class memory (SCM) devices have finite write endurance, leading to potential data loss due to defects if not managed properly, necessitating effective strategies to identify and manage defective memory bytes.

Method used

A memory system with a logical to virtual translation table (LVT) that tracks write and read operations, moving data to new locations when thresholds are reached, using a drift buffer for temporary storage and a VBA free list to manage wear levels, ensuring data is written to healthy memory areas.

Benefits of technology

Effectively extends the lifespan of SCM devices by preventing excessive writes to defective memory cells, maintaining data integrity and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007702195000004
    Figure 0007702195000004
  • Figure 0007702195000005
    Figure 0007702195000005
  • Figure 0007702195000006
    Figure 0007702195000006
Patent Text Reader

Abstract

A memory system and method for storing data on one or more storage chips includes one or more memory cards, each having a plurality of storage chips, each chip having a plurality of dies with a plurality of memory cells; a memory controller with a translation module, the translation module further including a logical-to-virtual translation table (LVT) having a plurality of entries, each entry in the LVT configured to map a logical address to a virtual block address (VBA), the VBA corresponding to a group of memory cells on the one or more memory cards; and each entry in the LVT further includes a write wear level count for tracking the number of write operations to the VBA and a read wear level count for tracking the number of read operations to the VBA mapped to that LVT entry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosure herein generally relates to memory translation from logical to virtual and from virtual to physical for storage class memory (SCM).

Background Art

[0002] Storage class memory (SCM) is a type of persistent memory that combines the low latency and byte addressability of dynamic read access memory (DRAM) with the non-volatility, areal density, and economic characteristics of conventional storage media. Furthermore, considering the byte addressability and low latency of SCM technology, a central processing unit (CPU) can access data stored in SCM without buffering the data in DRAM. As a result, SCM technology blurs the distinction between computer memory and conventional storage media and enables a single-level architecture without DRAM. Different from conventional main memory and disk storage configurations, SCM provides a single-level architecture.

[0003] Typically, an SCM is implemented as a group of solid-state devices connected to a computing system via several input / output (I / O) adapters that are used to map the technology of I / O devices to the memory bus of a central processing unit (CPU). However, care must be taken in the details of SCM technology to write data to an SCM. That is, an SCM media card is organized as a collection of multiple packages, each containing "N" dies with millions of byte-addressable memory elements. One common characteristic of SCMs is that these memory devices have finite write endurance. What a memory device with finite write endurance means is that it cannot be written to indefinitely before parts of the SCM begin to develop defects. Identifying which memory is likely to be defective or error-prone helps reduce the risk of losing stored data. For example, memory bytes (or bit arrays) identified as defective may be completely avoided, while memory bytes not identified as defective may be used without restriction. Additionally, defective memory bytes in multiple embodiments can be replaced with spare bytes. SUMMARY OF THE INVENTION

[0004] The summary of the present disclosure is provided to assist in understanding computer systems, computer architecture, processors, storage class memory (SCM), and media management methods, and is not intended to limit the present disclosure or the invention. The present disclosure is directed to those skilled in the art. It should be understood that the various aspects and features of the present disclosure may be advantageously used separately in some cases or in combination with other aspects and features of the present disclosure in other cases. Accordingly, changes and modifications may be made to computer systems, architectures, processors, SCMs, and the methods of their operation to achieve different effects.

[0005] A memory system for storing data in one or more storage chips is disclosed. The memory system in one embodiment includes one or more memory cards, each card having a plurality of storage chips, and each chip having a plurality of dies with a plurality of memory cells; a memory controller including a translation module, the translation module further including a logical to virtual translation table (LVT) having a plurality of entries, each entry in the LVT being configured to map a logical address to a virtual block address (VBA), the VBA corresponding to a group of memory cells on one or more memory cards, e.g., corresponding to a logical block address (LBA), and each entry in the LVT further including a write wear level count for tracking the number of write operations to the VBA mapped to that LVT entry, and a read wear level count for tracking the number of read operations to the VBA mapped to that LVT entry. In one or more embodiments, the write wear level count in the LVT is programmable to have a write level threshold corresponding to the maximum number of write operations to the VBA, and in response to a write operation that exceeds (or equals) the write level threshold in the LVT entry, data in the memory card corresponding to the LTV entry that exceeds (or equals) the write level threshold is moved to a new location on a memory card having a different VBA.

[0006] A memory system in one embodiment includes a VBA free list that identifies VBAs available for writing data, and in response to a write operation, a new VBA is obtained from the VBA free list. The system in one or more embodiments is configured to obtain a new VBA from the VBA free list based on a wear level count. In one aspect, the read wear level count in the LVT is programmable to have a read level threshold corresponding to the maximum number of read operations of the VBA, and in response to a read operation that exceeds the read level threshold in the LVT entry, data in the memory card corresponding to the LVT entry that exceeds (or equals) the read level threshold is written to a new location on that memory card with a different VBA. The system is configured to obtain a new VBA from the VBA free list in response to a read operation that exceeds (or equals) the read level threshold for the LVT entry, such that data in the memory card corresponding to the LVT entry that exceeds (or equals) the read level threshold is written to a new location on that memory card with a different VBA.

[0007] In one or more embodiments, a memory system includes a drift buffer having a plurality of entries for temporarily storing data, and a drift table having a plurality of entries, wherein each drift table entry is configured to index to one of the plurality of entries in the drift buffer, each entry of the drift table maps a drift buffer index to a VBA, and the system is configured to write data to an entry in the drift buffer in response to writing the data to a memory card, and further to write the VBA and a corresponding logical address, e.g., a logical block address (LBA), to an entry in the drift table indexed to the corresponding entry in the drift buffer. A memory system in an embodiment is configured to read from the drift buffer if data corresponding to the VBA exists in the drift buffer. In one aspect, the drift buffer is a cyclic FIFO buffer included in a memory card. Each LVT entry includes a field for indicating whether the drift buffer contains data corresponding to that LVT entry, and the system is configured in response to a hit on an LVT entry indicated by the LVT field that the data is in the drift buffer, and the LVT entry points to an entry in the drift table.In one embodiment, each LVT entry has a drift buffer index valid field for indicating whether the drift buffer contains data corresponding to each respective LVT entry. The system is configured to read data from the memory card in response to a request. The system is configured to look up a logical address in the LVT and, in response to finding the LVT entry corresponding to the logical address, the system checks the drift buffer index valid field. In response to the drift buffer index valid field indicating that the requested data is not in the drift buffer, the system utilizes the VBA from that LVT entry. In response to the drift buffer index valid field indicating that the requested data is in the drift buffer, the system reads the requested data from the drift buffer. In one aspect, in response to the drift buffer valid field indicating that the requested data is in the drift buffer, the LVT is configured to point to an entry in the drift table and the system utilizes the information in the drift table to obtain the requested data from the corresponding entry in the drift buffer. According to one embodiment, in response to data being removed from a drift buffer entry, the LVT entry corresponding to the drift buffer entry from which the data is removed is configured to be updated to include the VBA of the entry being removed from the drift buffer.

[0008] A method according to one or more embodiments for reading data from one or more memory cards is disclosed, where each memory card has a plurality of storage chips, and each storage chip has a plurality of dies having a plurality of memory cells. The method, in one aspect, issues a request for data located on one or more memory cards and looks up a logical address for the requested data in a logical-to-virtual translation table (LVT) having a plurality of entries, where each entry in the LVT maps a logical address to a virtual block address (VBA), and the VBA corresponds to a group of memory cells in one or more memory cards; checks the LVT entry to determine whether the data is located in a drift buffer in response to identifying the location of the logical address of the requested data in the LVT entry, e.g., a logical block address (LBA); reads the requested data from the drift buffer in response to determining that the data is located within the drift buffer; and obtains the VBA from the LVT entry corresponding to the logical address of the requested data, e.g., the logical block address (LBA), and reads the requested data in the memory card corresponding to the VBA in response to determining that the data is not located in the drift buffer.

[0009] The method in one embodiment further includes updating the read level count field in the LVT in response to reading requested data from the memory card. The method according to one aspect includes comparing the read level count in the LVT entry with the read level threshold field in the LVT entry, and in response to the read level count being equal to or exceeding the read level threshold, writing the data to be read to a new location on one or more memory cards having different VBAs. In response to writing the data to be read to a new location on one or more memory cards having different VBAs, the method includes, in one embodiment, updating the LVT with a different VBA. The method includes, in one aspect, further updating the VBA in the corresponding LVT entry in response to the data being removed from the drift buffer. The method according to another aspect includes moving an entry to the head of the drift buffer in response to reading data from the drift buffer.

[0010] A further method of writing data to one or more memory cards is disclosed, where each memory card has a plurality of storage chips, and each storage chip has a plurality of dies having a plurality of memory cells. The method in one or more embodiments includes issuing a request to write data to one or more memory cards, obtaining available VBAs from a VBA free list, writing the data to a memory card location corresponding to the available VBA obtained from the VBA free list, writing the data to an entry in a drift buffer, and writing the VBA of the available VBA and its corresponding logical address to the available VBA to an entry in a drift table corresponding to the entry in the drift buffer. The method in one aspect further includes writing a corresponding LVT entry having the available VBA. In one embodiment, the method further includes writing a drift table index identifying a drift table entry corresponding to the drift buffer entry to which the data is written into an LVT entry corresponding to the VBA corresponding to the location on the memory card to which the data is written, and setting a 1-bit to identify that the data is in the drift buffer.

[0011] The foregoing and other objects, features, and advantages of the present invention will be apparent from the following more detailed description of exemplary embodiments of the present invention as shown in the accompanying drawings, in which like reference numerals generally represent like parts of the exemplary embodiments of the present invention.

[0012] Aspects, features, and embodiments of computer systems, computer architecture structures, processors, memory systems, and their methods of operation will be better understood when read in conjunction with the provided figures. Embodiments are provided in the figures for the purpose of illustrating aspects, features, and / or various embodiments of computer systems, computer architecture structures, processors, SCM, and their methods of operation, but the claims should not be limited to the arrangements, structures, features, aspects, assemblies, sub-assemblies, systems, circuits, embodiments, methods, processes, techniques, and / or devices as shown, and the arrangements, structures, systems, assemblies, sub-assemblies, features, aspects, methods, processes, techniques, circuits, embodiments, and devices shown may be used alone or in combination with other arrangements, structures, assemblies, sub-assemblies, systems, features, aspects, circuits, embodiments, methods, techniques, processes, and / or devices.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2A

Figure 2B

Figure 3

Figure 4

Figure 5

Figure 6A

Figure 6B

Figure 6C

Figure 7

Figure 8

Figure 9A

Figure 9B

Figure 9C

DETAILED DESCRIPTION OF THE INVENTION

[0014] The following description is made to illustrate the general principles of the present invention and is not intended to limit the inventive concept claimed in the claims. In the following detailed description, many details are presented to provide an understanding of computer systems, computer architecture structures, processors, caches, memory systems, and their operating methods. However, many different embodiments of computer systems, computer architecture structures, processors, caches, memory systems, and their operating methods may be implemented without those specific details, and it will be understood by those skilled in the art that the claims and the disclosure should not be limited to the arrangements, structures, systems, assemblies, sub-assemblies, circuits, features, aspects, processes, methods, techniques, embodiments, and / or details specifically described and shown herein. Further, the specific features, aspects, arrangements, systems, embodiments, techniques, etc. described herein can be combined and used in each of various possible combinations and arrangements with other described features, aspects, arrangements, systems, embodiments, techniques, etc.

[0015] Unless specifically defined otherwise herein, all terms should be given their broadest possible interpretation, including the meanings suggested by the specification and understood by those skilled in the art and / or as defined in dictionaries, treatises, etc. As used in the specification and the appended claims, the singular forms "a", "an", and "the" include plural referents unless otherwise specified, and the terms "comprises" and / or "comprising", when used in the specification and the claims, specify the presence of the recited features, integers, aspects, arrangements, embodiments, structures, systems, assemblies, sub-assemblies, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, aspects, arrangements, embodiments, structures, systems, assemblies, sub-assemblies, steps, operations, elements, components, and / or groups thereof.

[0016] The following discussion omits or simply describes in brief conventional features of information processing systems, including processors, microprocessor systems and architectures, and address translation techniques and systems, which should be apparent to those skilled in the art. It is assumed that those skilled in the art are familiar with the general architecture of processors, and in particular address translation techniques and systems, and their operation. It will be noted that the numbered elements are numbered according to the figure in which the element is introduced and are typically referred to by that number throughout the subsequent figures.

[0017] FIG. 1 depicts a high-level block diagram representation of a computer 100-A connected via a network 130 to another computer 100-B, according to an embodiment of the present invention. The term "computer" is used herein merely for convenience and in various embodiments is a more general data handling system such as a cellular phone, tablet, server computer, etc. The mechanisms and apparatus of embodiments of the present invention equally apply to any suitable data handling device.

[0018] The main components of computer 100 may include one or more processors 101, main memory system 102, terminal interface 111, storage interface 112, I / O (input / output) device interface 113, and network adapter or interface 114, all of which are communicatively coupled, either directly or indirectly, for component - to - component communication via memory bus 103, I / O bus 104, and I / O bus interface device 105. Computer 100 includes one or more general - purpose programmable central processing units (CPUs) 101A, 101B, 101C, and 101D, which are generally referred to herein as processor 101. In certain embodiments, computer 100 includes multiple processors typical of a relatively large system. However, in another embodiment, computer 100 may alternatively be a single - CPU system. Each processor 101 executes instructions stored in main memory system 102 and may include one or more levels of on - board cache.

[0019] In one embodiment, the main memory system 102 may comprise a random access semiconductor memory (e.g., DRAM, SCM, or both), a storage device, or a storage medium for storing or encoding data and programs. In another embodiment, the main memory system 102 represents the entire virtual memory of the computer 100 and may also include the virtual memory of other computer systems coupled to the computer 100 or connected via the network 130. Although conceptually a single monolithic entity, in other embodiments, the main memory system 102 is a more complex arrangement, such as a hierarchy of caches and other memory devices. For example, the memory may exist in multiple levels of cache, and these caches may be further divided by the functions used by one or more processors such that one cache holds instructions while the other holds non-instruction data. The memory may be further distributed and associated with different CPUs or sets of CPUs, as is known in any of various so-called non-uniform memory access (NUMA) computer architectures.

[0020] The main memory system 102 stores, or encodes, an operating system (OS) 150, applications 160, and / or other program instructions. Although the operating system (OS) 150, applications 160, etc. are shown as being included within the main memory system 102 in computer 100, in other embodiments some or all of them may be on different computer systems and may, for example, be accessed remotely via network 130. Computer 100 may use a virtual addressing mechanism, which allows a program of computer 100 to behave as if it only has access to a single, large storage entity, instead of access to multiple, smaller storage entities. Thus, although the operating system 150, applications 160, or other program instructions are shown as being included within the main memory system 102, not all of these elements necessarily need to be fully included in the same storage device at the same time. Further, although the operating system 150, applications 160, other program instructions, etc. are shown as separate entities, in other embodiments some of them, some portions of them, or all of them may be packaged together.

[0021] In one embodiment, the operating system 150, application 160, and / or other program instructions include instructions or statements that execute on the processor 101, or instructions or statements interpreted by instructions or statements that execute on the processor 101, to perform functions as further described below. When such program instructions can be run by the processor 101, the computer 100 becomes a specific machine configured to execute such instructions. For example, instructions for a memory mirroring application 160A that mirrors the main memory system 102 to a first portion and a redundant second portion may be loaded onto one or more computers 100A. In another example, the main memory system 102 may be mirrored by the operating system 150. In another example, the main memory system 102 may be mirrored by a virtualizer application 170, such as a hypervisor.

[0022] One or more processors 101 may function as a general-purpose programmable graphics processing unit (GPU) that constructs an image (e.g., a GUI) for output to a display. The GPU works with one or more applications 160 to create a display image or user interface and determines how pixels should be manipulated, for example, on a display, touch screen, etc. Ultimately, an image (e.g., a GUI, etc.) is displayed to the user. The processor 101 and the GPU may be separate components or integrated into a single component.

[0023] Memory bus 103 provides a data communication path for transferring data among processor 101, main memory system 102, and I / O bus interface device 105. I / O bus interface device 105 is further coupled to system I / O bus 104 for transferring data to and from various I / O devices. I / O bus interface device 105 communicates through system I / O bus 104 with a plurality of I / O interface devices 111, 112, 113, and 114, also known as I / O processors (IOPs) or I / O adapters (IOAs). The I / O interface devices support communication with various storage and I / O devices. For example, terminal interface device 111 may support connection of one or more user I / O devices 121, which may include user output devices (e.g., video display devices, speakers, and / or television sets) and user input devices (e.g., keyboards, mice, keypads, touch pads, trackballs, buttons, light pens, or other pointing devices). A user may operate a user input device using a user interface to provide input data and commands to user I / O device 121 and computer 100, and may receive output data via a user output device. For example, the user interface may be presented via user I / O device 121, such as being displayed on a display device, played through speakers, or printed via a printer. The user interface may be a user interface that provides content to the user visually (e.g., via a screen), audibly (e.g., via speakers), and / or via touch (e.g., vibration, etc.). In some embodiments, computer 100 may itself act as a user interface, as the user may interact with computer application 160 data, functions, etc., inputting or operating them in a way that controls computer 100.

[0024] The storage interface device 112 supports the connection of one or more local disk drives or secondary storage devices 125. In certain embodiments, the secondary storage devices 125 are rotating magnetic disk drive storage devices, but in other embodiments, they are an array of disk drives configured to appear to the host computer as a single large storage device, or other types of storage devices. The contents of the main memory system 102, or any portion thereof, may be stored to and then retrieved from the secondary storage device 125 as needed. The local secondary storage device 125 typically has an access time slower than that of the main memory system 102, that is, the time required to read and / or write data from / to the main memory system 102 is less than the time required to read and / or write data from / to the local secondary storage device 125.

[0025] The I / O device interface 113 provides an interface to various other input / output devices, or to other types of devices such as printers or facsimile machines. The network adapter 114 provides one or more communication paths from the computer 100 to other data handling devices such as many other computers, and such paths may comprise, for example, one or more networks 130. The memory bus 103 is shown in FIG. 2 as a relatively simple single bus structure providing a direct communication path between the processor 101, the main memory system 102, and the I / O bus interface 105, but in practice, the memory bus 103 may be arranged in various forms, for example, point-to-point links in hierarchical, star or web configurations, multiple hierarchical buses, any of parallel and redundant paths, or other suitable types of configurations, and may comprise a plurality of different buses or communication paths. Furthermore, although the I / O bus interface 105 and the I / O bus 104 are shown as single respective devices, the computer 100 may in fact include a plurality of I / O bus interface devices 105 and / or a plurality of I / O buses 104. A plurality of I / O interface devices are shown separating the system I / O bus 104 from the various communication paths running to the various I / O devices, but in other embodiments, some or all of the I / O devices are directly connected to one or more system buses.

[0026] The I / O interface 113 may include electronic components and logic for adapting or converting data of one protocol on the I / O bus 104 to another protocol on another bus. Therefore, the I / O interface 113 may use one or more protocols including, but not limited to, Token Ring, Gigabit Ethernet, Ethernet, Fiber Channel, SSA, Fiber Channel Arbitrated Loop (FCAL), Serial SCSI, ULTRA3 SCSI, InfiniBand, FDDI, ATM, 1394, ESCON, Wireless Relay, Twinax, LAN connection, WAN connection, High Performance Graphics, etc. to connect a variety of devices to the computer 100 and to each other, such as, but not limited to, tape drives, optical drives, printers, disk controllers, other bus adapters, PCI adapters, workstations. Although shown as separate entities, the functions of the plurality of I / O interface devices 111, 112, 113, and 114, or the I / O interface devices 111, 112, 113, and 114 may be integrated into devices with similar functionality.

[0027] In various embodiments, the computer 100 is a multi-user mainframe computer system, a single-user system, a server computer, a storage system, or a similar device that has little or no direct user interface but receives requests from other computer systems (clients). In other embodiments, the computer 100 is implemented as a desktop computer, a portable computer, a laptop or notebook computer, a tablet computer, a pocket computer, a telephone, a smartphone, a pager, an automobile, a teleconference system, an appliance, or any other suitable type of electronic device.

[0028] Network 130 may be any suitable network or combination of networks, and may support any suitable protocol suitable for communicating data and / or code to / from computer 100A and at least computer 100B. In various embodiments, network 130 may represent a data handling device or combination of data handling devices directly or indirectly connected to computer 100. In another embodiment, network 130 may support wireless communication. Alternatively and / or in addition, network 130 may support hardwired communication, such as over telephone lines or cables. In one embodiment, network 130 may be the Internet and may support the Internet Protocol (IP). In one embodiment, network 130 is implemented as a local area network (LAN) or wide area network (WAN). In one embodiment, network 130 is implemented as a hotspot service provider network. In another embodiment, network 130 is implemented on an intranet. In one embodiment, network 130 is implemented as any suitable cellular data network, cell-based wireless network technology, or wireless network. In one embodiment, network 130 is implemented as any suitable network or combination of networks. Although one network 130 is shown, in other embodiments, there may be multiple networks (of the same or different types).

[0029] FIG. 1 is intended to depict representative major components of computer 100. However, individual components may be more complex than those represented in FIG. 1, and there may be components in addition to or instead of those shown in FIG. 1, and the number, type, and configuration of such components may vary. Some specific examples of such additional complexity or additional variation are disclosed herein, and these are for illustrative purposes only and are not necessarily the only such variations. For example, the various program instructions implemented on computer system 100 according to various embodiments of the present invention may be implemented in several ways, including using various computer applications, routines, components, programs, objects, modules, data structures, and the like.

[0030] Referring next to FIG. 2A, a schematic block diagram of an example main memory system 102 that communicates with processor 101 via memory controller 200 is shown. As shown in FIG. 2A, memory module(s) or card(s) 102 (e.g., SCM media card) are configured to store data in a plurality of “K” packages (i.e., chips) 252a-k (e.g., K = 24), where each package includes a plurality of “N” dies 251a-n (e.g., N = 16). Each package in some embodiments can include the same number “N” of dies (e.g., N = 8, 16, etc.). Each of dies 251a-n includes “M” memory cells, particularly memory cells 250a-m. Some of the memory cells 250 in each die may be grouped into “X” media replacement unit (MRU) groups 253a-p, and each die in some embodiments can include a fixed number of MRU groups 253. For example, there could be 16 (X = 16) MRU groups in die 251. Memory module / card 102 in one or more embodiments also includes a drift buffer 260 for storing data, e.g., temporarily storing data as described later. Memory module / card 102 in some aspects also includes a drift table 230 for mapping entries to drift buffer 260 as described later.

[0031] Furthermore, as shown in FIG. 2B, each MRU group 253a - p may include a plurality of MRUs 254a - n, each MRU including a plurality of "B" bit arrays, and each bit in the bit array including one memory cell. That is, as shown in FIG. 2B, in one embodiment, the MRU group is divided into 128 columns called bit arrays. The MRU group is divided horizontally into rows called MRUs. Each MRU has 128 bit arrays. Each box in FIG. 2B can be a 1 million - bit bit array. Additionally, each MRU includes "P" pages, where a page is a unit of data that can be written or read in the SCM (e.g., 16 bytes or 128 bits). In one embodiment, an MRU can include 1 million pages. Pages use bits from the active bit arrays during read / write operations. The number of active bit arrays from which each page takes memory cells may depend on the redundancy of memory cells required in the memory module / card 102 and / or the scrub process (as described below). For example, in each MRU, if 4 bit arrays are reserved as spares for swapping operations with the bit arrays that are likely to cause failures / errors, each page will include bits from 124 of the 128 active bit arrays during write / read operations (128 bit arrays in the MRU - 4 reserved bit arrays = 124 active bit arrays). This means that one page can store 15.5 bytes of actual data. As described below, all MRU groups and spare MRUs from dies in a single package can replace the failed MRUs in that package.

[0032] 1 million bits per bit array (1024 *It should be noted that using 1024 is a design choice for the example embodiments disclosed in the present disclosure. However, the present disclosure is not so limited, and several bits per bit array (e.g., 500,000, 2 million, 3 million, etc.) may be used. The number of bits per bit array is thus used to determine the number of MRUs and, accordingly, the size of tables (e.g., Chip Select Table (CST), Media Repair Table (MRT), Bit Array Repair Table (BART), etc.).

[0033] The total capacity of the memory module / card 102, K (measured in bytes), * can be determined according to C, which is the capacity of each package. Of the K packages of the memory module / card 102, some packages may be used to store data, and the other or remaining packages may be used for error-correcting code (ECC) and metadata used for data management. The error-correcting code is used to correct errors contained in the data stored in the data area. Each memory module / card 102 (e.g., SCM media card) has I / O data with a data width of z bits and address bits of an appropriate size depending on the capacity. SCM may be, for example, Phase Change Memory (PCM), Resistive RAM (RRAM), or any suitable non-volatile storage.

[0034] FIG. 2A shows the memory controller 200 located outside the memory module / card 102, but the present disclosure is not so limited, and the controller 200 may be part of the memory module / card 102. As shown in FIG. 2A, the controller 200 may include at least one of an address translation module 202, a scrub module 204, and a media repair module 206. Modules 202, 204, 206 may be implemented in software, firmware, hardware, or a combination of two or more of software, firmware, and hardware. In other examples, the controller 200 may include additional modules or hardware devices, or may include fewer modules or hardware devices. The controller 200 may include a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other digital logic circuits.

[0035] The address translation module 202 may associate a logical block address (LBA) and / or a virtual block address (VBA) used by the processor(s) 101 (as discussed below) with the physical block address (PBA) of the memory module / card 102. For example, in response to receiving an LBA from the processor as part of a read or write command, the address translation module 202 may look up the VBA through a logical-to-virtual table (LVT) 210, and the address translation module 202 may then use the VBA to determine the PBA of the memory module / card 102 that matches the received LBA. In some examples, the address translation module 202 may use a hierarchical virtual-to-physical table structure (V2P) to perform the translation from VBA to PBA (as described below). For example, the address translation module 202 may include a chip select table (CST) 220, a media repair table (MRT) 222, and a bit array repair table (BART) 224.

[0036] The scrub module 204 may be configured to detect and correct faults or errors in the memory module / card 102. Errors in the memory may occur due to alpha or other particles, or due to physical defects in the memory cells. As used herein, the term "scrubbing" generally refers to the process of detecting errors in a memory system and correcting correctable errors. Errors can include soft (or transient) errors as well as, in certain environments, hard errors. In various embodiments, memory scrubbing may use a process of detecting and correcting bit errors in the memory as described below. To avoid disturbing normal memory requests from the CPU and thus prevent performance degradation, scrubbing may be done by taking specific portions of the memory module / card 102 out of service for the scrubbing process (described below). Since scrubbing may consist of normal read and / or write operations, it may increase the power consumption by the memory compared to operations without scrubbing. Therefore, according to various embodiments, scrubbing is performed periodically rather than continuously. In many servers, the timing or period for scrubbing may be configured in the BIOS setup program.

[0037] In certain embodiments, the scrub module 204 may include an error rate table (ERT). The error rate table may be configured to store information regarding memory defects or errors in the memory module / card 102. In certain embodiments, an ERT may be allocated for each of the packages included in the memory module / card 102 (e.g., 24 ERTs for 24 packages of a media card).

[0038] The media repair module 206 may be configured to repair errors or defects in the memory module / card 102 by analyzing information stored in the ERT(s). For example, the media repair module 206 may create an ERT summary table and cause the memory controller to perform appropriate corrective actions, such as replacing a portion that caused one or more malfunctions of the memory module / card 102 with a properly functioning spare (e.g., replacing a defective bit array, MRU, MRU group, die, or package with an appropriate spare), modifying various translation tables, searching for spare memory portions / locations, creating spare memory portions / locations, rebalancing spare memory portions / locations (to increase the likelihood of identifying the location of a spare for future replacement), but is not limited thereto. In some embodiments, the media repair module 206 can be configured to avoid replacements that, while considering functional accuracy, would likely have an adverse effect on the performance of the memory module / card 102.

[0039] Referring now to FIG. 3, there is shown a block diagram of an embodiment of the hardware in a memory system and the process flow that can translate a logical address to a virtual address (logical-to-virtual (L2V) translation) and translate the virtual address to an actual or physical address on media card 390 (virtual-to-physical (V2P) translation). The host addresses the memory / media card 390 via a logical unit number (LUN) and a logical unit number offset (LUN offset). The translation table 360 translates these LUNs and LUN offsets to a logical block address (LBA). As shown in the LBA table 360 in FIG. 3, multiple LUNs are supported. In one embodiment, there is an LBA for each 4KB block accessible by the host. The LBA is fed to a logical-to-virtual translation table (LVT) 370 which converts or translates each LBA to a virtual entry or address, e.g., a virtual block address or VBA. The VBA is fed to a media table 380 where, as will be described in detail in a later part of this disclosure, the virtual address, e.g., the VBA, is translated to a physical or actual address on media card 390. The virtual-to-physical translation in one or more embodiments includes a chip select table (CST), a media repair table (MRT), and a bit array repair table (BART) for repairing the media.

[0040] In one or more embodiments, the LVT 370 is arranged as a table structure and has a plurality of entries 372. In certain embodiments, each entry 372 in the LVT 370 includes a virtual block address (VBA) or a drift table index, write wear level count data, and read wear level count data. For each LBA, there is a 1:1 mapping to the VBA (assuming the LBA is not trimmed). Other metadata can be included in the LVT. The LVT 370 in certain embodiments is stored in DRAM and, in one aspect, has a small field size for each entry, e.g., 8 bytes per entry. The LVT 370 in certain embodiments is located within the memory controller 200 and, in one aspect, can be included in the address translation module 202 shown in FIG. 2A (LVT 210 in FIG. 2A). Table 1 below shows an example of an LVT entry 372.

[0041]

Table 1

[0042] The LVT entries in Table 1 are merely examples, and other storage media, fields, and field sizes are contemplated for the LVT 370 and the LVT entries 372. The media on the media card 390 wears out at several locations over time, and thus, in one or more embodiments, the design will over-provision. The over-provision that should be made to account for media wear should be small, and thus, according to one aspect, the number of LBAs will typically be slightly less than the number of VBAs, e.g., 10% less. Thus, for example, if the media card 390 supports 614M VBAs, the LVT 370 will have entries close to 614M. In certain embodiments, the host is instructed that the total size of the memory is smaller than its actual size.

[0043] Media card 390 permits write operations to the same address / location, but since writes can disturb adjacent areas on media card 390, in certain embodiments, new data is written to different locations on media card 390. In certain aspects, no location on media card 390 is written more than a threshold number, e.g., more than 10,000 times, more than other locations on media card 390. For example, in certain embodiments, a location on media card 390, e.g., a block of memory cells, is not written more than “Y” times the number of times another location on the media card is written. The write wear level field (bits 39:32 in Table 1) in LVT 370 is used to control this aspect of operation. Since reads can also disturb the media, after “N” reads of a location, the data should be rewritten to a new location / address. LVT 370 has a read wear level count (bits 53:40 in Table 1) to track and record the number of reads that have occurred to that block since that block was last written. In response to a read count exceeding a certain value, e.g., a threshold, that block will be moved and / or copied to a new location on the memory card that will have a new VBA. For example, after 10,000 reads, the hardware will perform a wear level migration where the data is copied / written to a new location on the memory card with a corresponding new VBA.

[0044] After a new location on the memory card has been written, it should not be read for a period of time, typically milliseconds, e.g., 10 milliseconds. To overcome the potential latency of waiting to read the new location on the memory card, in one embodiment a drift buffer is provided. The drift buffer 260 in one embodiment is included on the media card 102 shown in FIG. 2A. The drift buffer is preferably a FIFO buffer used to hold a copy of the newly written data, such that the drift buffer can be read instead of the media card 390 during the time frame when the newly written media should not be read. The drift buffer has a plurality of entities and in one embodiment has a table or index, e.g., a drift buffer index or drift table 330, for correlating the VBA with an entry in the drift buffer. When data is newly written to the media, e.g., the memory card, the data is also written to a drift buffer entry and the VBA is stored in a drift table entry 332 corresponding to the drift buffer entry. The drift table 330 stores the VBA and the corresponding LBA. For example, the drift table entry indexes to the drift buffer entry to map the VBA to the corresponding LBA. When data is written to the drift buffer, the VBA field in the LVT 370 indicates the drift table entry 372 using the drift table index instead of the VBA. When a read operation is processed and the entry in the LVT corresponding to the LBA has a drift buffer index valid bit set indicating that the data is in the drift buffer, the system identifies the drift buffer entry and uses the drift table to read the data from the corresponding drift buffer. In that context where the drift buffer index valid bit is set, in one embodiment the LVT does not return the VBA.In one embodiment, the drift table 330 can be smaller than the LVT 370 and can be on the controller 200 and / or the DRAM on the media card 102 / 390. An example of a drift table entry 332 is shown in Table 2 below.

[0045]

Table 2

[0046] When data is aged out of the drift buffer, the VBA field from the drift table 330 is copied to the corresponding entry 372 in the LVT 370. The table structures of the drift table 330 and the LVT 370 facilitate utilization such that the drift buffer is also a read cache. When utilized as a read cache, the drift buffer will have writes and reads or prefetching taking place. Also, the time that frequently used data remains in the drift buffer can be extended using drift table 330 updates, and the data being used will remain in the drift buffer. For example, read hit data could be moved to the front of the drift buffer. In this way, the time that frequently read data remains in the drift buffer can be extended.

[0047] FIG. 3 also shows a VBA free list 340. The VBA free list 340 has several entries 342. VBAs that do not actively store data for the associated LBA are maintained on the VBA free list 340 and, in some embodiments, are organized or prioritized based on wear level. When a write command requires a free VBA to store data, the free VBA is obtained from the VBA free list 340. If the write command targets a previously written LBA, the previously used VBA (corresponding to the previously used LBA) is added to the VBA free list 340. The VBA free list 340 may be present on the controller 200 and / or the media card 102 / 390. Table 3 below shows the free list entries in the VBA free list 340, and the entries are indexed by VBA.

[0048]

Table 3

[0049] Referring now to FIG. 4, there is disclosed a flowchart illustrating an example method 400 for processing a host read command, including, for example, converting and / or translating a logical address, such as an LBA, to a virtual block address or VBA. The flowcharts and block diagrams in the figures illustrate the exemplary architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur in an order different than that noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified function or acts, or combinations of dedicated hardware and computer instructions.

[0050] At 405, a host read command is issued by the host. The host addresses the media card via the logical unit number (LUN) and LUN offset as described above. The LUN and LUN offset are converted to a logical block address or LBA at 410. At 415, a conversion table or LVT entry from logical to virtual address corresponding to the LBA is read. That is, in one embodiment, the LVT is searched using the LBA converted from the LUN and LUN offset, a comparison is made, and the entry in the LVT corresponding to that LBA is read. In one or more embodiments, it is determined at 420 whether the entry read from the LVT has been trimmed. A trimmed entry in the LVT means that the entry is not available on the media and / or should not be read. In the embodiment of Table 1, bit 54 indicates whether entry 372 in the LVT has been trimmed. If the entry read at 420 has been trimmed (420: Yes), a zero is returned at 425 and a good completion is notified to the host at 430.

[0051] At 420, if the LVT entry to be read is not trimmed (420: No), the processor 400 continues to 435 where it is determined whether the LBA is in the drift buffer (indexed in the drift table). Thus, for example, for an LVT entry as in the example of Table 1, bit 31, the drift buffer index valid bit is read, and if bit 31 is set to 1, the LBA is in the drift buffer. If the LBA is in the drift buffer (435: Yes), at 440 an entry in the drift table 330 having the corresponding LBA is read. From the drift table, an entry to the drift buffer is obtained and at 445 the corresponding drift buffer is read. In the example of Table 1, if the LBA is in the drift buffer, bits 20:0 in the LVT are the index to the drift table 330, and the corresponding entry in the drift table indicates which entry in the drift buffer should be read. After reading the drift buffer at 445, the process continues to 450 where it is determined whether there is any read error in one or more embodiments. If there is a read error (450: Yes), at 460 the host is notified of a read failure. If there is no read error (450: No), at 455 the host is notified of a successful completion, e.g., a successful read completion.

[0052] If it is determined at 435 that the LBA is not in the drift buffer (435: No), the LVT read level count is updated at 470. In the LVT entry example of Table 1, if the drift buffer index valid bit 31 in the LVT is set to 0, bits 29:0 in the LVT refer to the VBA that selects a 5120-byte block on the storage class memory (SCM). In that example, the LBA is not in the drift buffer, and thus, the process 400 continues at 475 where it is checked whether the read level count for that entry (media address) exceeds a certain threshold. For example, it can be determined whether the read count level exceeds a threshold, for example, 10,000 reads. The threshold level can be set to different values that can be pre-programmed or changed during operation based on several factors. If the read level exceeds the read level count (475: Yes), at 480, for example, a wear level migration operation where data is moved to a new location is scheduled as considered in the present disclosure. After the wear level migration at 480, the process 400 continues at 485 where the VBA and metadata are read from the LVT. On the other hand, if it is determined at 475 that the read count threshold has not been exceeded (475: No), the process continues to 485 where the VBA and metadata are read from the LVT. The process continues after 485 as shown in FIG. 4, and at 450, it is determined whether there is any read error, and depending on whether there is any error, a good completion is notified to the host at 455 or a read failure is notified at 460.

[0053] Referring now to FIG. 5, there is disclosed a flowchart illustrating an example method 500 for processing a host write command that includes, for example, converting and / or translating a logical address, such as an LBA, to a virtual block address or VBA. At 505, a host write command executable indication is issued by a host. An address for the write command is issued as a LUN and a LUN offset. At 510, the LUN and LUN offset are converted to a logical block address or LBA. At 512, in one or more embodiments, it is determined whether all of the host write data is zero. If all of the host write data is zero (512: Yes), the process skips to step 540 and no data is written to the media (or drift buffer). If all of the host write data is not zero (512: No), the process continues to 515 where an entry available on the VBA free list is requested. In one or more embodiments, the request is targeted to the VBA with the lowest wear level and / or the VBA free list is configured to provide the VBA with the lowest wear level.

[0054] At 520, data is written to a memory, such as a media card. Additionally, at 525 data is written to a drift buffer and at 530 a new drift table entry is written. That is, at 525 and 530, data is written to an entry in the drift buffer and the corresponding entry in the drift table that indexes to that entry in the drift buffer is written with the VBA, LBA, and optionally metadata shown in Table 2, or other data. In one aspect, CRC and other metadata are provided in response to a request from the host. At 535, it is determined whether the write is complete. The process remains at 535 until the write is complete (535: Yes).

[0055] When the writing of the media, the drift buffer, and the drift table is complete, process 500 continues to 540 where the old LVT entry is read. The LVT corresponding to the LBA is read at 540. In one or more embodiments, the old LVT entry is replaced with a new LVT entry. At 545, it is determined whether the old LVT entry is trimmed. If the LVT entry is trimmed, it indicates that the media should not be read or is not available. In one or more embodiments, an entry in the LVT is marked as trimmed when the LBA has never been written or was last written using zeroed data. If the old LVT entry is marked as trimmed (545: Yes), the process continues to 565 where a new LVT is written. If the old LVT entry is not trimmed at 545 (545: No), the process continues to 550 where it is determined whether the LBA corresponding to the written VBA is in the drift buffer. If it is determined at 550 that the LBA is in the drift buffer (550: Yes), the old drift table entry is read at 555 to obtain the VBA. Note that the entry in the drift table obtained at 555 is different from the entry written to the drift table at 530 because the write process overwrote the old entry in the drift table and the read entry was not updated. Subsequently, at 560, the VBA is returned to the VBA free list along with the wear level indicated by the VBA. If it is determined at 550 that the LBA is not in the drift buffer (550: No), in one embodiment, the process proceeds directly to 560 where the VBA is returned to the free list along with its indicated wear level data so that available VBAs can be prioritized in the VBA free list. After 560, the process proceeds to 565 where a new LVT is written.

[0056] In one embodiment, after 560, the process proceeds to 565 where a new LVT is written. The content of the LVT to be written will depend on the previous processing. If the new LVT entry is to be marked as trimmed because all host write data is zero (512: Yes), the new LVT entry will include a Trimmed indicator, invalid VBA, and a cleared (set to zero) drift buffer index valid field. If the new LVT entry is not to be marked as trimmed because all host write data is not zero (512: No), the new LVT entry will include a drift buffer index valid field set to indicate that the drift buffer index is valid, along with a valid drift buffer index. When a new LVT entry is written, the read count field in the LVT entry is always cleared (set to zero). In one or more aspects, other metadata, including wear level data, is written to the LVT entry at 565. The write wear level in one aspect comes from the VBA free list 340. After 565, the process continues to 570 where the host is notified of the write completion.

[0057] In one or more embodiments, it may be recognized that a write operation can trigger a drift table cast out 575. At the start, the drift buffer has room for the next write operation. However, after the drift buffer becomes full, the oldest entry should have its allocation canceled to create room for the next write operation. In one or more embodiments, the drift buffer may need to cast out one or more entries in the drift buffer or cancel an allocation, and may cast out some entries in front of the write pointer so that there is room for the next write to the drift buffer. In certain embodiments of casting out drift buffer entries, the drift table is read to obtain cast out information, including the LBA of one or more drift table entries to be removed from the drift buffer. Next, LVT entries for one or more LBAs to be removed from the drift buffer are read to determine whether the LVT entry corresponding to one or more LBAs to be cast out from the drift buffer points to the drift table. If the LVT entry points to the drift table (e.g., the drift buffer index valid bit is set), the drift buffer index valid bit is cleared and the LVT entry is written to change the DBI in the LVT entry to VBA. If the LVT entry does not point to the drift table, e.g., if the drift buffer index valid bit is not set, the LVT has already been updated and no update of the LVT entry is needed.

[0058] Referring next to FIG. 6A, a block diagram overview of the hardware and process involved in translating a virtual block address (VBA) to a physical address on a media card in the memory module / card 102 of FIGS. 1, 2A, and 2B is shown, and referring to FIG. 6B, a flowchart showing an example of the translation of a virtual block address (VBA) to a physical address indicating the appropriate bits for performing read / write operations (byte addressable) in the memory module / card 102 of FIGS. 1, 2A, and 2B is shown. The flowcharts and block diagrams in the figures illustrate the exemplary architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur in an order different than that noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially simultaneously, or depending on the functionality involved, the blocks may sometimes be executed in the reverse order. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by a dedicated hardware-based system that performs the specified function or acts, or combinations of dedicated hardware and computer instructions.

[0059] Those skilled in the art will understand that VBA may be derived from the logical block address (LBA) received from the host using any currently or hereafter known method in the above methods and structures, and in one or more embodiments. The flowchart of FIG. 6B is based on the assumption that the lowest addressable subdivision of the memory module / card 102 of FIG. 1 is 16 bits, and that the memory module / card 102 includes 24 packages, 8 or 16 dies per package, 16 MRU groups per die, 64 MRUs per MRU group, 128-bit arrays per MRU (4 of which are spare), 1 million bits in each bit array, and 1 million pages per MRU (each page including 128 bits, i.e., 16 bytes). It will be recognized that the process of FIG. 6B can be utilized in other memory system configurations.

[0060] The flowchart further assumes that the read / write operations are performed on data that includes 4K-byte sized data blocks. In so doing, each read / write operation is performed on a total of 5120 bytes of data that includes 4K bytes of data received from the host, ECC bytes, and metadata. Considering that the lowest addressable granularity of the memory module / card 102 is 16 bytes, the 5120 bytes of data may be stored across the memory module / card 102 in 320 separate pages (each 16 bytes, with 15.5 bytes used to hold the actual data). It will be appreciated that for wear leveling, the 320 separate pages should be evenly distributed across the memory module / card 102. Therefore, the translation process translates the VBA to 320 separate pages of physical addresses in the memory module / card 102, and the VBA represents such 320 separate pages. The VBA is a 30-bit address, and the lower 20 bits of the VBA (i.e., the least significant bits) are ignored during translation because they represent a collection of one million 4K-byte sized blocks (collectively, the one million 4K-byte blocks are referred to as a virtual repair unit (VRU) represented by the VBA), all of which have the same upper 10 bits (most significant bits) of the VBA. It should be noted that the embodiments disclosed herein use a 30-bit VBA to address a 4TB media card. However, the present disclosure is not so limited and may be used in a similar manner for other sized addresses for different media cards. The upper 10 bits may be used to identify the VRU using a chip select table, as discussed below. The number of VRUs is configured based on the actual media card (e.g., if the actual storage capacity of a 4TB media card is approximately 2.5TB, the number of VRUs is 614). Different upper bits may be used to identify the VRU using a chip select table (CST) for different media cards.

[0061] As shown in Figure 6C, the translation of the VBA physical address is a multi-layer lookup that includes at least three lookup tables - a chip selection table (CST) (e.g., the same for all 24 packages), 24 media repair tables (MRTs) (e.g., one per package), and 24 bit array repair tables (BARTs) (e.g., one per package). At 602, the VBA is indexed into the CST. As previously discussed, in one or more embodiments, the lower 20 bits of the VBA representing the VRU are ignored during translation, and only the upper 10 bits of the VBA are indexed into the CST at step 602.

[0062] As shown in Figure 6C, the CST 610 includes columns indexed by package number, e.g., Pkg0 - 23, and rows indexed by VRU number. The VRU number is identified using the upper (most significant) 10 bits of the VBA. Each entry 611 in the CST 610 contains a total of 16 bits. That is, the least significant 9 bits are for determining the intermediate repair unit (IRU) number (i.e., 9 IRU bits), the next 3 bits are reserved for future use, the next 1 bit indicates whether the VRU should be scrubbed (i.e., the scrub bit), the next bit is a flag bit indicating whether the package is included in the VRU for VBA read / write operations (i.e., the included flag bit), the next bit indicates whether the package contains an IRU for which some or all of the MRUs have failed (i.e., the partial failure bit), and the last most significant bit (of the 16 bits total) indicates whether the package contains a full spare IRU (i.e., the spare bit). Similarly, in certain embodiments, IRUs that are partially marked as defective and / or spare cannot be used for read / write operations.

[0063] As described above, the flag bits included in the CST entry 611 indicate whether the package will contain data for VBA (i.e., read / write operations). For example, if the included flag bit value for a certain package is "0", it will not contain data for VBA read / write operations, and if the included flag bit value for a certain package is "1", it will contain data for VBA read / write operations (and vice versa). In one embodiment, out of the 24 packages on the media card, only 20 packages are used for read / write operations, and 4 packages are spare packages for use in replacing defective packages by the media repair module 206. Since 320 pages for storing data of 5120-byte-sized blocks are evenly distributed across the memory module / card 102, during translation, the flag bits included in the 20 packages will indicate that the corresponding packages will contain data for VBA read / write operations. Furthermore, each of the 20 packages will contain 16 separate pages (or beats) that will contain data to achieve an even distribution of data. In this specification, the term "beat" is used to describe a page in a package that will contain data for performing read / write operations corresponding to the received VBA. These 16 beats per package may be identified using MRT and BART as discussed below. In that way, during translation, the upper 10 bits of VBA are used to index VBA to the CST at 602, identify the VRU number, and identify the corresponding included flag bits and IRU bits per package using the CST.

[0064] Referring back to FIG. 6B, in step 604, for each package, the 9 IRU bits retrieved from the corresponding CST entry may be used to index the MRT for that package. Specifically, for each of the 20 valid packages, the 9 IRU bits from CST610 are used to identify the IRU number (0 - 511), and then that number is used to index the MRT620.

[0065] As shown in FIG. 6C, the MRT620 includes rows indexed by IRU number and columns indexed by beat number. Each entry 621 in the MRT620 contains a total of 16 bits and will contain data for performing VBA read / write operations (determining the physical address of each beat on the (all 16) packages). The physical address is represented as a combination of the die index in the package (the least significant 4 bits of entry 621), the MRU group index in the die (the next 4 bits of entry 621), and the MRU index in the MRU group containing the beat (the next 6 bits of entry 621) (D / MG / M). Entry 621 also includes a fault bit (the bit next to the most significant bit) indicating whether the page represented by D / MG / M has been previously declared as faulty, and a spare bit indicating whether the page represented by D / MG / M has been previously declared as an individual spare (if the entire IRU is a spare, the MRU will not be marked as a spare in entry 621). Faulty and / or spare pages cannot be used for read / write operations. The MRT620 for each package may therefore return 16 physical addresses for the beats in each package that may contain data for performing read / write operations for the VBA.

[0066] As previously discussed, one MRU contains a 128-bit array, four of which are reserved as spare bit arrays, and each of the one million pages in the MRU takes one bit from each of the remaining 124-bit arrays. Therefore, in step 606, BART 630 may be used to determine the index corresponding to the 124-bit array in each MRU of the physical address, and the corresponding beats will take bits from those bit arrays for the VBA read / write operation. As shown in FIG. 6C, each row of BART 630 may be indexed using the D / MG / M identification from MRT 620, and each column of BART 630 is indexed by 0 to 3 since each MRU contains four unused bit arrays. Each entry 631 of BART 630 contains 8 bits - the 7 least significant bits indicate one of the four unused bit arrays for the MRU of that row (from among the 128-bit array), and the most significant bit is reserved for future use. The 124-bit array from which bits will be taken by the beats may be determined by excluding the unused bit arrays from the MRU.

[0067] In step 608, the system may use the 20 least significant bits of the VBA, along with the physical address (D / MG / M) and the unused bit array index, to perform the read / write operation in the appropriate memory cell, and the actual read / write operation may be performed according to the interface specifications of each particular SCM memory technology.

[0068] Note that the unused bit array index of BART330 may be determined using an error rate table (ERT) 640 created during the memory scrub process (described below with respect to FIG. 7). As shown in FIG. 6C, the ERT 640 for one package includes the rows corresponding to the 128-bit arrays and the columns corresponding to the MRUs in that package. Since one VRU is scrubbed at a time during the scrub process, the MRU of the ERT 640 corresponds to the VRU being scrubbed (as discussed below). Further still, the scrub process described below targets 16 bits (i.e., MRUs) per package and updates one ERT per package such that there are 16 MRUs in the ERT. Each entry 641 in the ERT includes the total number of stuck bits observed during the scrub process of one VRU. Since one VRU includes 1 million consecutive VBAs, the total number of stuck bits can be 2 million (not common to both, 1 million stuck-at-0 and 1 million stuck-at-1).

[0069] Referring now to FIG. 7, a flowchart depicting an example scrub process for memory module / card 102 is described. The scrub process may be performed to periodically rewrite data within a threshold time to ensure data readability and accuracy, to condition each memory cell for future reprogrammability, and / or to test the memory system (which may then be used to initiate media repair procedures) by periodically inverting the bits of the memory system to collect statistical data corresponding to stuck fault failures. In certain embodiments, the scrub process may be performed on one VRU at a time (in a round-robin fashion and / or upon the occurrence of a trigger event) by taking 1 million consecutive VBAs corresponding to the VRU being scrubbed out of service by disabling any further read / write operations from the VRU.

[0070] During the scrub process, at 702, the system identifies the VRUs to be scrubbed. In one or more embodiments, the scrub process for each VRU may be performed periodically in a round-robin fashion (e.g., every 24 hours) and / or upon the occurrence of a trigger event (e.g., detection of a memory error such as a non-simultaneous memory error, detection that a memory operation requires a higher level of ECC correction, or the like). The scrub process for a VRU may be scheduled such that the set of memory modules / cards 102 are scrubbed within a threshold time (e.g., 24 hours, 48 hours, or the like).

[0071] At 704, the system in one or more embodiments determines whether the identified VRU is in a service state for performing read / write operations. If the identified VRU is in a service state for performing read / write operations (704: Yes), the system may remove the VRU from use (706). The removal of the VRU (e.g., from the service of 1 million consecutive VBAs) may, without limitation, include the removal of the VRU from the free list (identifying the free area of memory of the specific size required) of all 1 million VBAs of the memory system of the identified VRU. Specifically, the system may not allow any VBA of the VRU to be placed on the free list after being released by a write operation. Additionally and / or alternatively, the system may remove any existing data from all VBAs of the VRU currently in use for storing data by moving the data to another location, removing the VBA from the logical to virtual table (LVT) that tracks the current values of the VBAs, and / or removing the VBA from the drift buffer.

[0072] When removing the identified VRU from the service and / or if it is determined that the identified VRU is not in a service state for read / write operations (704: No), the system may initialize the counters in 708 for all ERTs corresponding to the identified VRU (e.g., assign a zero value). In 710, the system may issue write operations using pattern A for each VBA of the identified VRU. Examples of pattern A may include a string of all 1s (to detect all stuck-at-0 bits), a string of all 0s (to detect all stuck-at-1 bits), all 5s (to detect two adjacent bits that are stuck together), and / or the like. When the write operations of pattern A are executed for all VBAs in the VRU, the system may issue read operations for each VBA of the identified VRU in 712 to determine the number of stuck-at-fault bits in the VRU. For example, if pattern A includes a string of all 1s, the read operation may be used to determine the stuck-at-0 bits in the VRU, and if pattern A includes a string of all 0s, the read operation may be used to determine the stuck-at-1 bits in the VRU. Other patterns may be used to identify the number of other stuck faults in the VRU (e.g., stuck-at-X, two or more bits that are stuck together, etc.).

[0073] At 714, the system may issue a write operation using pattern B for each VBA of the identified VRU, where pattern B is different from pattern A. Examples of pattern B may include a string of all 1s (for detecting all stuck-at-0 bits), a string of all 0s (for detecting all stuck-at-1 bits), all 5s (for detecting two adjacent bits in a stuck state with each other), and / or the like. When the write operation of pattern B is performed for all VBAs in the VRU, the system may issue a read operation for each VBA of the identified VRU at 716 to determine the number of stuck-at-fault bits in the VRU. For example, if pattern A includes a string of all 1s and is used to identify stuck-at-0 bits in the VRU, pattern B may include a string of all 0s to determine stuck-at-1 bits in the VRU.

[0074] It should be noted that during the scrub process of the VRU, all 128-bit arrays (not just the 124-bit arrays) of each MRU of the VRU are written / read by ignoring BART during the translation process so that stuck-at-faults may be detected on all bit arrays. Specifically, the translation to the physical page address of each of the 1 million VBAs of the VRU is performed using only CST and MRT. Additionally, the translation during the scrub process also includes the faulty IRU and / or spare IRU. Each scrub process typically leads to the creation of 20 - 24 ERTs (per package).

[0075] At 718, the system may update the ERT counter from step 708 based on the determined stuck - at - 0 and stuck - at - 1 bit values for each bit array of each MRU of the VRU to be scrubbed (i.e., update the counter to indicate the total number of stuck - at faults). At 720, the system may perform media repair procedures (discussed below with respect to FIG. 8).

[0076] At 722, the system may return the VRU to service, for example, by inserting 1 million VBAs of the VRU into the free list. As discussed below, if the VRU is converted to a spare IRU during the media repair procedure, it cannot be returned to service.

[0077] Referring next to FIG. 8, a flowchart showing an example of a media repair procedure performed during each scrub process of the memory system is described. As discussed above, during the scrub process, for each bit - array index / MRU of the VRU being scrubbed, an ERT per package indicating the number of stuck - at faults (e.g., stuck - at - 1, stuck - at - 0, etc.) is constructed. The system may analyze the counters in the ERT to perform the media repair procedure(s) as described in this FIG. 8. The media repair procedure may include bit - array repair (almost always performed), MRU replacement, and / or IRU replacement.

[0078] At 802, the system may determine whether bit - array repair, MRU replacement, and / or IRU replacement needs to be performed on the scrubbed VRU. In certain embodiments, the system may determine whether MRU replacement and / or IRU replacement needs to be performed on the VRU by identifying and analyzing the number of bad bits in the ERT.

[0079] Beats may be identified as bad beats by analyzing the stuck - at bit count for each of the bit arrays within the beat. Specifically, for each of the bit arrays of a beat, the system determines whether the number of stuck - at bits counts is greater than a first threshold (T H ), greater than a second threshold (T L ) but less than the first threshold T H , or is an acceptable number of stuck - at bits. The thresholds T L and T H are determined by the strength of the ECC. An example of T H may be from about 2000 to about 7000 bits per million bits, an example of T L may be from about 200 to about 700 bits per million bits, and an example of an acceptable number of stuck - at bits may be any value less than T L (e.g., less than 100 bits per million bits).

[0080] If there are no more than 4 beats having some stuck bit counts greater than T L (e.g., one bit array of a beat having a stuck - at bit count greater than T H and three bit arrays having stuck - at bit counts greater than T L ), the system may perform only bit array repair as considered below. FIG. 9A shows an example of an ERT created after VRU scrubbing indicating that bit array repair is required.

[0081] In addition and / or alternatively, if some (but not all, and / or less than a threshold number) of the beats have more than 4 bit arrays having stuck - at bit counts greater than T H and / or T LIf the system includes a bit array of more than 11 bits having a larger number of stuck - at bit counts, the system may determine that the beat is a defective MRU. In addition to bit array repair, the system may perform MRU replacement for such a defective MRU. FIG. 9B shows an example of an ERT created after VRU scrubbing indicating that MRU replacement is required.

[0082] If all or a certain threshold number of beats in one ERT for a package are defective (i.e., a bit array of more than 4 bits has a H larger number of stuck - at bit counts and / or a bit array of more than 11 bits has a L larger number of stuck - at bit counts), the system may determine that the IRU containing the defective beat has failed. The system in one or more embodiments performs IRU replacement for such a defective IRU in addition to bit array repair. IRU replacement may also be performed when there is no spare MRU in the package containing the defective MRU that requires replacement. FIG. 9C shows an example of an ERT created after VRU scrubbing indicating that IRU replacement is required.

[0083] The number of bit arrays is provided by way of example only, and other numbers are within the scope of the present disclosure for determining whether bit array repair, MRU replacement, and / or IRU replacement needs to be performed for the VRU.

[0084] At 804, if bit - array repair is required, a system according to one aspect performs bit - array repair procedures by excluding the worst bit - arrays from those used during read / write operations. The system may first identify the worst four bit - arrays and their corresponding indices by analyzing the total number of stuck - at faults observed for each bit - array for each MRU (i.e., column) during ERT. The worst four bit - arrays are the bit - arrays with the largest number of stuck - at faults observed during the scrub process. The system then updates and saves the BART corresponding to the ERT to include the worst four bit - array indices as unused bit - array indices for that MRU. If more than four bit - arrays are defective, the system may call ECC to perform media repair.

[0085] At 806, if MRU / IRU replacement is required for a failed MRU, a system according to one aspect determines whether the failed MRU is a spare MRU and / or is included within a spare IRU. If the failed MRU is a spare MRU and / or is included within a spare IRU (806: Yes), the system may mark the failed MRU (not a spare) as such without replacing the MRU (808). However, if the failed MRU is not a spare MRU and is not included within a spare IRU (806: No), a system according to one embodiment performs MRU replacement (810) and updates the MRU accordingly (812). It should be noted that since performance degradation can occur from improper ordering of D / MG / M values within an IRU's beat, a spare MRU may be considered a candidate for replacement of a failed MRU only if the performance of the MRU will not degrade.

[0086] The system in one aspect performs MRU replacement (810) by first searching for a spare MRU in the same package as the failed MRU that needs to be replaced. If no spare MRU is found, the system in one embodiment also searches for a spare IRU in the same package and divides the spare IRU into multiple spare MRUs. Next, the spare MRUs and / or IRUs in the same package can be used to replace the failed MRU and / or the IRU containing the failed MRU. The system in one embodiment converts a good MRU of the IRU to be replaced (i.e., an MRU that does not need replacement) into a spare MRU. During replacement, in one or more embodiments, the system updates the MRT for that package to swap the physical addresses of the pages in the failed MRU or IRU (i.e., D / MG / G) with those of the spare MRU or IRU.

[0087] However, if the package containing the failed MRU does not contain a spare MRU or IRU, in one or more embodiments, the system replaces the entire IRU with a spare IRU in another package (814). The system according to one aspect converts a good MRU of the IRU to be replaced (i.e., an MRU that does not need replacement) into a spare MRU. During replacement, in one embodiment, the system updates the CST to swap the in-use IRU with the spare IRU and the spare / failed IRU indication in the CST. It should be noted that in one embodiment, a spare IRU from the same package as the failed MRU / IRU is preferred over a spare IRU from another package for replacement.

[0088] If the system cannot find a spare MRU or IRU in any package, in one or more embodiments, the system creates a spare IRU from a VRU to be scrubbed (816) (without returning it to service) and marks the failed MRU as failed. The system updates the CST and MRT in one aspect to indicate that the VRU is being used to create a spare IRU and cannot be used for read / write operations.

[0089] In one or more embodiments, the media may need to be refreshed, for example, once a day. During the refresh operation, the data in the VBA is copied to the new VBA, and then, before the old VBA is returned to the service state, the old VBA is first written all to "1" once, and then, all to "0" the second time.

[0090] The above exemplary embodiments are preferably implemented in hardware, for example, in apparatus and circuitry of a processor, but various aspects of the exemplary embodiments and / or techniques may be implemented in software as well. For example, it will be understood that each block of the flowchart illustrations in FIGS. 4-8, and combinations of blocks in the flowchart illustrations, can be implemented by computer program instructions. These computer program instructions may be provided to a processor or other programmable data processing apparatus to produce a machine, such that the instructions executed on the processor or other programmable data processing apparatus create means for implementing the functions specified in one or more blocks of the flowchart. These computer program instructions may be stored in a computer readable memory or storage media, and the instructions stored in the computer readable memory or storage media may direct a processor or other programmable data processing apparatus to function in a particular manner to produce an article of manufacture that includes instruction means for implementing the functions specified in one or more blocks of the flowchart.

[0091] Accordingly, the blocks of the flowchart illustration support a combination of means for performing the specified function, a combination of steps for performing the specified function, and program instruction means for performing the specified function. It will also be understood that each block of the flowchart illustration, and combinations of blocks in the flowchart illustration, can be implemented by a dedicated hardware-based computer system that performs the specified function or step, or by combinations of dedicated hardware and computer instructions.

[0092] One or more embodiments of the present disclosure may be a system, a method, and / or a computer program product. The computer program product may include computer-readable storage media (s) having thereon computer-readable program instructions for causing a processor to implement aspects of the present disclosure.

[0093] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction-executing device. A computer-readable storage medium can be, for example, but not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, flexible disk, mechanically encoded devices such as punch cards or raised structures in grooves having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium should not be construed herein as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission media (e.g., an optical pulse passing through an optical fiber cable), or an electrical signal transmitted through a wire.

[0094] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices or to an external computer or external storage device via a network, such as, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them for storage on a computer-readable storage medium within respective computing / processing devices.

[0095] Computer-readable program instructions for performing the operations of this disclosure may be written in any combination of one or more programming languages, including assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or object-oriented programming languages such as Smalltalk, C++, or the like, as well as conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, for example, an electronic circuit including a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit for performing aspects of this disclosure.

[0096] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0097] These computer-readable program instructions may be provided to the processor of a computer, other programmable data processing apparatus, or other device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may be stored in a computer-readable storage medium, which may direct a computer, programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium containing the instructions comprises a manufacture including instructions which implement the aspects of the function / act specified in one or more blocks of the flowchart and / or block diagram.

[0098] The computer-readable program instructions may be loaded onto a computer, other programmable apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0099] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by dedicated hardware-based systems that perform the specified functions or acts, or combinations of dedicated hardware and computer instructions.

[0100] Moreover, systems according to various embodiments may include a processor and logic integrated with and / or executable by the processor, the logic being configured to perform one or more of the process steps recited herein. By integrated is meant that the processor has logic embedded therein as hardware logic, such as in an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), and the like. By executable by the processor is meant that the logic is hardware logic, software logic such as firmware, part of an operating system, part of an application program, or some combination of hardware logic and software logic that is accessible by the processor and configured to cause the processor to perform some function upon execution by the processor. The software logic may be stored on local and / or remote memory of any memory type as is known in the art. Any processor known in the art may be used, such as a software processor module and / or a hardware processor, such as an ASIC, an FPGA, a central processing unit (CPU), an integrated circuit (IC), a graphics processing unit (GPU), and the like.

[0101] In addition to the functional elements in the appended claims, all corresponding structures, materials, acts, and equivalents of all means or steps are intended to include any structure, material, or act for performing that function in combination with other claim elements as specifically claimed. The description of the embodiments of the present disclosure has been presented for purposes of illustration and description, but is not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the present disclosure. Embodiments and examples were chosen and described in order to best explain the principles and practical applications of the present disclosure and to enable others of ordinary skill in the art to understand the present disclosure for various embodiments with various modifications as are suited to the particular use contemplated.

[0102] The programs described herein are identified based on the applications in which the programs are implemented in the specific embodiments of the present disclosure. However, it should be recognized that any specific program naming herein is for convenience only, and thus the present disclosure should not be limited to use only in any specific application identified and / or suggested by such naming.

[0103] It will be apparent that the various features of the foregoing systems and / or methodologies may be combined in any manner to yield a plurality of combinations from the descriptions presented above.

[0104] It will further be recognized that embodiments of the present disclosure may be provided in the form of services deployed to benefit users for on-demand provision of services.

[0105] The descriptions of various embodiments of the present disclosure are presented for illustrative purposes and are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles of the embodiments, practical applications, or technical improvements over technologies found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A memory system for storing data, the memory system comprising: one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells; a drift buffer having a plurality of entries for temporarily storing data; a memory controller comprising a translation module, the translation module further comprising a logical-to-virtual translation table (LVT) having a plurality of entries, each entry in the LVT configured to map a logical address to a virtual block address (VBA) and indicate whether corresponding data is located in the drift buffer, the VBA corresponding to a group of the memory cells on the one or more memory cards; wherein each entry in the LVT includes a write wear-level count for tracking the number of write operations to the VBA mapped to each entry and a read wear-level count for tracking the number of read operations to the VBA mapped to each entry, and when the corresponding data is located in the drift buffer, the corresponding entry in the LVT is configured to map to the corresponding entry of the drift buffer in which the corresponding data is stored instead of the corresponding VBA, and the corresponding VBA is retained in association with the corresponding entry of the drift buffer.

2. The write wear-level count in the LVT is programmable to have a write level threshold corresponding to a maximum number of write operations to the VBA, and in response to a write operation in an entry of the LVT that exceeds the write level threshold, the data in the memory card corresponding to the entry of the LVT that exceeds the write level threshold is moved to a new location on the memory card having a different VBA. The memory system of claim 1.

3. The memory system according to claim 2, further comprising a VBA free list that identifies the VBA available for writing data, and in response to a write operation, new VBA is obtained from the VBA free list.

4. The memory system according to claim 3, wherein the system is configured to obtain new VBA from the VBA free list based on the write wear level count.

5. The read wear level count in the LVT is programmable to have a read level threshold corresponding to the maximum number of read operations of the VBA, and in response to a read operation that exceeds the read level threshold in an entry of the LVT, the data in the memory card corresponding to the entry of the LVT that exceeds the read level threshold is written to a new location on the memory card having a different VBA. The memory system according to any one of claims 1 to 4.

6. The memory system according to claim 5, further comprising a VBA free list that identifies the VBA available for receiving write data, and in response to a read operation that exceeds the read level threshold for an entry of the LVT, new VBA is obtained from the VBA free list, and the data in the memory card corresponding to the entry of the LVT that exceeds the read level threshold is written to the new location on the memory card having the different VBA.

7. When the corresponding data is located in the drift buffer, the corresponding VBA is configured to be copied to the corresponding entry in the LVT in response to the corresponding data being deleted from the drift buffer. The memory system according to any one of claims 1 to 6.

8. The memory system includes a drift table having a plurality of entries further comprising, each entry of the drift table is configured to index to one of the plurality of entries in the drift buffer, each entry of the drift table maps a drift buffer index to VBA, the system, in response to writing the data to the memory card, also writes the data to an entry in the drift buffer, and further writes the corresponding VBA and the corresponding logical address to an entry in the drift table indexed to the corresponding entry in the drift buffer, the memory system according to claim 7.

9. A memory system for storing data, the memory system comprising one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells, the one or more memory cards; a memory controller comprising a translation module, the translation module further comprising a logical-to-virtual translation table (LVT) having a plurality of entries, each entry in the LVT being configured to map a logical address to a virtual block address (VBA), the VBA corresponding to a group of the memory cells on the one or more memory cards, a memory controller comprising each entry in the LVT further includes a write wear-level count for tracking the number of write operations to the VBA mapped to each entry, and a read wear-level count for tracking the number of read operations to the VBA mapped to each entry; the memory system the read wear-level count in the LVT is programmable to have a read level threshold corresponding to the maximum number of read operations of the VBA, and in response to a read operation exceeding the read level threshold in an entry of the LVT, the data in the memory card corresponding to the entry of the LVT exceeding the read level threshold is written to a new location on the memory card having a different VBA; Further comprising a VBA free list for identifying the VBA available for receiving write data, in response to a read operation that exceeds the read level threshold for the entry of the LVT, a new VBA is obtained from the VBA free list, and the data in the memory card corresponding to the entry of the LVT that exceeds the read level threshold is written to the new position on the memory card having the different VBA, The system is a memory system configured to obtain a new VBA from the VBA free list based on the write wear level count in response to a read operation that exceeds the read level threshold.

10. A memory system for storing data, the memory system comprising One or more memory cards, each memory card having a plurality of storage chips, and each storage chip having a plurality of dies each having a plurality of memory cells, the one or more memory cards, A memory controller comprising a translation module, the translation module Further comprising a logical to virtual translation table (LVT) having a plurality of entries, each entry in the LVT being configured to map a logical address to a virtual block address (VBA), the VBA corresponding to a group of memory cells on the one or more memory cards, a memory controller Comprising Each entry in the LVT further includes a write wear level count for tracking the number of write operations to the VBA mapped to each entry, and a read wear level count for tracking the number of read operations to the VBA mapped to each entry, The memory system A drift buffer having a plurality of entries for temporarily storing data, and A drift table having a plurality of entries, each entry of the drift table A memory system configured to index to one of the plurality of entries in the drift buffer, each entry of the drift table maps a drift buffer index to VBA, and in response to writing data to the memory card, the system also writes the data to an entry in the drift buffer, and further writes the VBA and the corresponding logical address to an entry in the drift table indexed to the corresponding entry in the drift buffer. **Claim 11** Each entry of the LVT includes a field for indicating whether the drift buffer contains data corresponding to each respective entry of the LVT, and the system is configured in response to a hit on an entry of the LVT indicated by the field that the data is in the drift buffer, and the entry of the LVT indicates an entry in the drift table. The memory system according to claim 8 or 10. **Claim 12** Each entry of the LVT further includes a drift buffer index valid field for indicating whether the drift buffer contains data corresponding to each respective entry of the LVT, and the system is configured to read data from the memory card in response to a request, the system is configured to look up the logical address in the LVT, and in response to finding an entry of the LVT corresponding to the logical address, the system checks the drift buffer index valid field, and in response to the drift buffer index valid field indicating that the requested data is not in the drift buffer, the system uses the VBA from the found entry of the LVT, and in response to the drift buffer index valid field indicating that the requested data is in the drift buffer, the system reads the requested data from the drift buffer. The memory system according to claim 8 or 10. **Claim 13** The system is configured such that, in response to the drift buffer index valid field indicating that the requested data is in the drift buffer, the LVT points to an entry in the drift table, and the system utilizes the information in the drift table to obtain the requested data from the corresponding entry in the drift buffer, the memory system of claim 12.

14. The system is configured such that, in response to data being removed from an entry of the drift buffer, the entry of the LVT corresponding to the entry of the drift buffer from which the data is removed is updated to include the VBA of the entry of the drift buffer being removed from the drift buffer, the memory system of claim 8 or 10.

15. A memory system for storing data, the memory system comprising one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells, the one or more memory cards a memory controller comprising a translation module, the translation module further comprising a logical to virtual translation table (LVT) having a plurality of entries, each entry in the LVT being configured to map a logical address to a virtual block address (VBA), the VBA corresponding to a group of the memory cells on the one or more memory cards, a memory controller and each entry in the LVT further includes a write wear level count for tracking the number of write operations to the VBA mapped to each entry, and a read wear level count for tracking the number of read operations to the VBA mapped to each entry, the memory system comprising A memory system further comprising a chip select table (CST) configured to identify one or more valid storage chips during translation for performing a memory access operation, and a media repair table (MRT) corresponding to each of the storage chips, wherein each MRT is configured to identify one or more storage dies during translation for performing a memory access operation.

16. A method for reading data from one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells, the method comprising: Issuing a request for data located on the one or more memory cards; Looking up a logical address for the requested data in a logical-to-virtual translation table (LVT) having a plurality of entries, each entry in the LVT mapping a logical address to a virtual block address (VBA), the VBA corresponding to a group of memory cells in the one or more memory cards; Checking an entry in the LVT to determine whether the data is located in a drift buffer in response to identifying a location of the logical address of the requested data in an entry in the LVT; Reading the requested data from the drift buffer in response to determining that the data is located within the drift buffer; Obtaining the VBA from the entry in the LVT corresponding to the logical address of the requested data and reading the requested data in the memory card corresponding to the VBA in response to determining that the data is not located in the drift buffer. A method comprising the steps of:

17. The method of claim 16, further comprising updating a field of a read level count in the LVT in response to reading the requested data from the memory card.

18. Comparing the read level count in the entry of the LVT with the field of the read level threshold in the entry of the LVT, and in response to the read level count being equal to or exceeding the read level threshold, writing the data to be read to a new position on the one or more memory cards having different VBAs, the method according to claim 17, further comprising.

19. Updating the LVT with the different VBA in response to writing the data to be read to a new position on the one or more memory cards having different VBAs, the method according to claim 18, further comprising.

20. Further comprising updating the VBA in the entry of the corresponding LVT in response to the data being removed from the drift buffer, the method according to claim 19.

21. Further comprising moving an entry to the head of the drift buffer in response to reading data from the drift buffer, the method according to claim 16.

22. A memory system for reading data from one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells, the memory system comprising: Issuing a request for data located on the one or more memory cards; Looking up a logical address for the requested data in a logical-to-virtual translation table (LVT) having a plurality of entries, each entry in the LVT mapping a logical address to a virtual block address (VBA), the VBA corresponding to a group of memory cells in the one or more memory cards, the looking up; Checking an entry of the LVT to determine whether the data is located in a drift buffer in response to identifying the position of the logical address of the requested data in the entry of the LVT; Reading the requested data from the drift buffer in response to determining that the data is located within the drift buffer; In response to determining that the data is not located in the drift buffer, obtaining the VBA from an entry of the LVT corresponding to the logical address of the requested data, and reading the requested data in the memory card corresponding to the VBA A memory system configured to perform. **Claim 23** A method of writing data to one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells, the method comprising: Issuing a request to write the data onto the one or more memory cards; Obtaining an available virtual block address (VBA) from a VBA free list; Writing the data to a location on the memory card corresponding to the available VBA obtained from the VBA free list; Writing the data to an entry in a drift buffer; Writing the VBA of the available VBA and its corresponding logical address to the available VBA to an entry in a drift table corresponding to the entry in the drift buffer; The method further comprising. **Claim 24** Writing a drift table index that identifies an entry in the drift table corresponding to an entry in the drift buffer to which the data is written, to an entry in a logical-to-virtual translation table (LVT) corresponding to the VBA corresponding to the location on the memory card to which the data is written; Setting a 1-bit to identify that the data is in the drift buffer; The method according to claim 23, further comprising. **Claim 25** A memory system for writing data to one or more memory cards, each memory card having a plurality of storage chips, each storage chip having a plurality of dies having a plurality of memory cells, the memory system comprising: Issuing a request to write the data onto the one or more memory cards; Obtaining an available virtual block address (VBA) from a VBA free list; writing the data to a position on the memory card corresponding to the available VBA obtained from the VBA free list; writing the data to an entry in the drift buffer; writing the VBA of the available VBA and its corresponding logical address to the available VBA to an entry in the drift table corresponding to the entry in the drift buffer; A memory system configured to perform.

Citation Information

Patent Citations

  • Storage device with flash memory and computer

    JP2003216506A

  • Computer system

    JP2007034944A

  • Semiconductor memory device and driving method thereof

    JP2008276832A

  • Memory system

    JP2009211219A

  • Semiconductor storage device

    JP2011128998A