Direct cache hits and transfers in an in-order programmed memory subsystem

By using read cache tables and buffer index tables to pre-read sequential data in the memory subsystem, the problem of low sequential read request performance in the prior art is solved, and efficient direct cache hits and data transfer is achieved.

CN113849424BActive Publication Date: 2025-05-16MICRON TECHNOLOGY INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110718597.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-06-26
Filing Date
2021-06-28
Publication Date
2025-05-16
Estimated Expiration
2041-06-28

AI Technical Summary

Technical Problem

When existing memory subsystems process sequential read requests, there are problems with low cache hit rate and performance bottlenecks, especially when multiple read requests are issued sequentially.

Method used

Direct cache hits and data transfer are achieved by introducing read cache tables and buffer index tables into the memory subsystem, pre-reading sequential data and storing them in the buffers of volatile memory.

Benefits of technology

Improves read performance, reduces latency, and reduces the cost of interrupting write operations, allowing the memory subsystem to handle a large number of sequential read requests more efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113849424B_ABST
    Figure CN113849424B_ABST
Patent Text Reader

Abstract

The present application relates to direct cache hits and transfers in a sequentially programmed memory subsystem. The system includes a buffer and a processing device, the processing device receiving a read request with a logical block address (LBA) value for a memory device, creating a logical transfer unit (LTU) value including the LBA value and mapped to a first physical address of the memory device, and generating a command tag instructing the processing device to retrieve data from the memory device and store the data in a buffer. The command tag includes a first command tag associated with the first physical address and a second command tag associated with a second physical address sequentially after the first physical address. The processor further creates an entry in a read cache table for the buffer. The entry may include a start LBA value set to the first LBA value and a read offset value corresponding to an amount of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate generally to memory subsystems, and more particularly, to direct cache hits and transfers in an in-sequentially programmed memory subsystem. Background Art

[0002] The memory subsystem may include one or more memory devices that store data. The memory devices may be, for example, non-volatile memory devices and volatile memory devices. In general, the host system may utilize the memory subsystem to store data at the memory devices and retrieve data from the memory devices. Summary of the invention

[0003] In one aspect, the present application provides a system comprising: a memory device; a volatile memory comprising a buffer and a read cache table; and a processing device coupled to the memory device and the volatile memory, wherein the processing device will: access a read command having a first command tag, the first command tag comprising a first logical transfer unit (LTU) value and a first buffer address of a first buffer in the buffers, the first LTU value being mapped from a region of a plurality of sequential logical block address (LBA) values ​​to a first physical address in a plurality of sequential physical addresses of the memory device, wherein first data is stored at the first physical address, and wherein the first LTU value contains a first LBA value in the plurality of sequential LBA values; generate a set of command tags instructing a command execution processor of the processing device to retrieve second data from the memory device and store the second data in a set of the buffers, wherein the set of command tags comprises at least a second command tag associated with a second physical address sequentially following the first physical address; and create an entry in the read cache table for the set of the buffers, wherein the entry comprises a start LBA value set to the first LBA value and a read offset value corresponding to an amount of the first data and the second data.

[0004] On the other hand, the present application provides a method, which includes: receiving, by a processing device of a memory subsystem controller, a read request for a first logical block address (LBA) value of a memory device including a first LBA value of an LBA space, wherein the LBA value belongs to a region of multiple sequential LBA values ​​mapped to multiple sequential physical addresses; creating, by the processing device, a first logical transfer unit (LTU) value including the first LBA value, wherein the first LTU value is mapped to a first physical address of the memory device; allocating, by the processing device, a set of buffers in a volatile memory, wherein the capacity of the set of buffers matches the amount of data stored at the first physical address and subsequent physical addresses, wherein the subsequent physical addresses are sequentially numbered within a read offset value starting from the first physical address; generating, by the processing device, a set of command tags instructing a command execution processor of the processing device to retrieve the data from the memory device and store the data in the set of buffers, wherein the set of command tags includes a first command tag associated with the first physical address and a second command tag associated with a second physical address sequentially following the first physical address; and creating, by the processing device, an entry in a read cache table for the set of buffers, wherein the entry includes a starting LBA value set to the first LBA value and the read offset value corresponding to the amount of data.

[0005] On the other hand, the present application provides a non-transitory computer-readable medium storing instructions, which, when executed by a processing device of a memory subsystem controller, causes the processing device to perform multiple operations, including: receiving a read request for a first logical block address (LBA) value of a memory device including an LBA space, wherein the LBA value belongs to a region of multiple sequential LBA values ​​mapped to multiple sequential physical addresses; creating a first logical transfer unit (LTU) value including the first LBA, wherein the first LTU value is mapped to a first physical address of the memory device; allocating a set of buffers in a volatile memory, wherein the capacity of the set of buffers matches the amount of data stored at the first physical address and subsequent physical addresses, wherein the subsequent physical addresses are sequentially numbered within a read offset value starting from the first physical address; generating a set of command tags instructing a command execution processor of the processing device to retrieve the data from the memory device and store the data in the set of buffers, wherein the set of command tags includes a first command tag associated with the first physical address and a second command tag associated with a second physical address sequentially after the first physical address; and creating an entry in a read cache table for the set of buffers, wherein the entry includes a starting LBA value set to the first LBA value and the read offset value corresponding to the amount of data. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] The present disclosure will be more fully understood from the detailed description provided below and the accompanying drawings of various embodiments of the present disclosure.

[0007] Figure 1A An example computing system including a memory subsystem according to an embodiment is shown.

[0008] Figure 1B According to the embodiment Figure 1A Additional details of the memory subsystem.

[0009] Figure 2 is a block diagram illustrating an example of a data structure configured to support region-based mapping, according to various embodiments.

[0010] Figure 3 According to the embodiment Figure 1A-1B A block diagram of the relationship between the command generation processor, translation processor and command execution processor of the memory subsystem controller of the present invention.

[0011] Figure 4 is a flow chart of a method for supporting direct cache hits based on read commands according to an embodiment.

[0012] Figure 5 is a flow chart of a method of supporting direct cache hits according to an embodiment.

[0013] Figures 6A-6C is a flow chart of a method of supporting direct cache hits and transfers according to a related embodiment.

[0014] Figure 7 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION

[0015] Various aspects of the present disclosure relate to direct cache hits and transfers in a sequentially programmed memory subsystem. The memory subsystem may be a storage device, a memory module, or a mixture of a storage device and a memory module. Figure 1A Examples of storage devices and memory modules are described. In general, a host system can utilize a memory subsystem that includes one or more memory components (e.g., memory devices) that store data. The host system can provide data to be stored at the memory subsystem and can request data to be retrieved from the memory subsystem.

[0016] The memory device may be a non-volatile memory device. An example of a non-volatile memory device is a NAND memory device. Figure 1AOther examples of non-volatile memory devices are described. A non-volatile memory device is a package of one or more dies. The dies in the package can be assigned to one or more channels for communication with a memory subsystem controller. Hereinafter, the memory subsystem controller is also referred to as a "controller". Each die can be composed of one or more planes. The planes can be grouped into logical units (LUNs). For some types of non-volatile memory devices (e.g., NAND devices), each plane is composed of a set of physical blocks. Each block is composed of a set of pages. Each page is composed of a set of memory cells ("cells"). A cell is an electronic circuit that stores information. Hereinafter, a block refers to a unit of a memory device for storing data, and may include a group of memory cells, a group of word lines, a word line, or an individual memory cell.

[0017] Data operations may be performed by the memory subsystem. Data operations may be operations initiated by the host. For example, the host system may initiate data operations (e.g., write, read, erase, etc.) to the memory subsystem. The host system may send access requests (e.g., write commands, read commands) to the memory subsystem to store data in a memory device at the memory subsystem and read data from a memory device of the memory subsystem.

[0018] The data to be read or written specified by the host request is referred to as "host data" below. The host request may include logical address information (e.g., logical block address (LBA), namespace) of the host data, which is the location of the host system associated with the host data. The logical address information (e.g., LBA, namespace) may be part of the metadata of the host data. The metadata may also include error handling data (e.g., ECC codewords, parity check codes), data versions (e.g., for distinguishing the age of written data), valid bitmaps (whose LBA values ​​or logical transfer units contain valid data), and the like. For simplicity, if "data" is mentioned below, such data may be understood to refer to at least host data, but may also refer to other data such as media management data and / or system data.

[0019] When performing data operations (e.g., write, read, erase), the memory subsystem may use a striping scheme to process groups of data as units. A LUN stripe is a series of planes that are processed as one unit when writing, reading, or erasing data. Each plane in a LUN stripe may implement the same operation in parallel with all other planes in the LUN stripe. A block stripe is a series of blocks that are processed as units in each plane in a LUN stripe. Blocks in a block stripe have the same block identifier (e.g., block number) in their corresponding planes. A block stripe (hereinafter also referred to as a block set) may be a group of blocks arranged across planes of different dies so that blocks are grouped together for the purpose of data storage. Writing to a block stripe enables more host data to be written and read simultaneously and in parallel across multiple dies. Multiple blocks in one or more block sets may be identified as a data group.

[0020] The host file system can group the host data by location and write the host data sequentially to the memory device of the memory subsystem. Then, the file system can write the host data at different locations to the memory device as parallel sequential streams, each stream has its own location, for example, different host applications can write to the location of their own streams respectively. "Location" can refer to a time location or a spatial location. The memory subsystem controller (e.g., a processing device) usually writes to the media randomly with a portion of the host data (e.g., 4KB), and then uses metadata to map the LBA space to the physical address space of the memory device. However, when a larger host data group (e.g., 100 megabytes (MB) or larger) is written and grouped by data location, the "data group" can be written sequentially as a larger block to one or more block sets across multiple dies. In order to simplify the mapping of such data groups, the LBA values ​​of the region (e.g., the logical address space associated with the data group) can be sequentially sorted in the LBA space and can be mapped to physical addresses that are sequentially sorted in the physical address space. In other words, the data group can be associated with an LBA space region with multiple sequential LBA values, which are sequentially mapped to multiple sequential physical addresses. A memory subsystem that can write (e.g., program) data regions at a time and map these regions accordingly operates in a zone name space (ZNS), e.g., where logical address regions are named / identified as groups. Advantageously, using ZNS for logical-to-physical (LTP) address mapping greatly reduces the amount of metadata that tracks LTP mappings.

[0021] In a memory subsystem, a read request (or read operation) is typically issued one command tag at a time to a translation processor (e.g., translating a logical address into a physical address) of a memory subsystem controller, thereby performing a random read at the granularity specified by the command tag. A command tag, also referred to as a system tag (e.g., systag), contains a logical transfer unit (LTU) value and a buffer address that identifies a buffer (e.g., a slot or entry in a volatile memory) in which the content corresponding to the transfer unit is stored like in a cache. In one embodiment, the LTU value corresponding to a 4 kilobyte (KB) portion of data is a subset of a plurality of sequential LBA values ​​that can be mapped to a physical address via a set of mapping data structures. Therefore, to generate an LTU value, a command generation processor of the controller can combine the LBA value of the read request with an additional LBA value (which can also be received in the read request) that is consecutive to the LBA value. Depending on the LTU type, each LTU value can be translated into a logical block or a logical page. For example, an LTU can correspond to 8KB, 16KB, 32KB, or more data that is increased in increments of 4KB or 8KB of data.

[0022] Because read requests are typically executed one command tag at a time, each read request generates a command message (e.g., a mailbox message in one instance) into a command generation processor of the controller (regardless of whether the read is sequential or not), and generates multiple (e.g., four or more) data structure lookups for mapping the LBA of the read request to a physical location in the memory device, as will be explained in detail. The command message may be used after obtaining a non-volatile memory command, where the command generation processor notifies the translation processor of the command receipt. These data structures (e.g., tables) may be stored in volatile memory. This manner of processing read requests increases overhead and reduces overall performance by increasing latency, particularly under conditions where some read requests are issued sequentially to a physical address space (e.g., of a ZNS) that is written sequentially. Therefore, the sequential data layout for sequential writes is not used to limit the number of lookups that the controller (e.g., the controller's translation processor) must perform to determine the physical address from which to retrieve data to fulfill the read request.

[0023] Aspects of the present disclosure address the above and other deficiencies by: a command generation processor (e.g., a processing device) of a memory subsystem controller generates a set of command tags that instruct a command execution processor to read an amount of data (e.g., sequential data) written sequentially to a set of buffers. The amount of sequential data read from the set of buffers may be significantly greater than the amount requested by a particular read request. By performing this sequential data read lookahead, the command generation processor of the controller may access the sequential data as if it were a cache to fulfill subsequent read requests without having to perform address translation of sequentially numbered LTU values. To do so, the command generation processor may further create and update a read cache table and a read index table to manage the use of the set of buffers as a cache to fulfill these subsequent read requests. Thus, when a read request is received that is known to be within a read offset value of an LBA value of an initial read request, determining the offset within the read offset value of the data enables the location of the requested data within the set of buffers to be determined. The controller may then pass the requested data to the host system in response to a cache "hit" at the set of buffers.

[0024] In various embodiments, the read cache table optionally stores in each entry a region identifier for the region of LBAs in the read request, a starting LBA value set to a first LBA value (e.g., received in an initial read request or read command), and a read offset value. The region identifier is optional because the starting LBA value also identifies the region. Each entry in the read cache table may also optionally store an ending LBA value that identifies the end of the read offset value starting from the starting LBA within the physical address space. The read offset value may be the amount of data that will be read into the set of buffers in a lookahead manner, which may include data corresponding to the initial read request or command. Therefore, the read offset value may be significantly larger than the amount of data mapped to the LTU, such as between 128KB and 2MB. In one example, if the LTU value maps to 16KB of data (e.g., it is a buffer allocation unit offset), then 1MB of data within the read offset value will include 64 data chunks corresponding to 64 LTU values. Therefore, the read lookahead of 1MB of data can save an additional 63 sets of lookups by the translation processor for determining the physical addresses of the additional 63 read requests. This latency reduction is substantial, especially when extrapolated over thousands of read requests.

[0025] In various embodiments, the command generation processor may create and manage a buffer index table to track the LTU value associated with each command tag in the set of command tags. For example, the buffer index table may map the LTU value to the buffer address associated with the LTU value in each command tag. Thus, when a subsequent read request is received, after determining a new LTU value (e.g., by an offset value relative to an initial or first LTU value), the command generation processor may use the new LTU value to index into the buffer index table to determine the corresponding buffer address. The command generation processor may then retrieve the requested data from the identifier buffer and pass the requested data to the host system to implement the subsequent read request. Such indexing within the relatively small read cache table and buffer index table consumes much less processing power and latency than if the translation processor translated each LTU value into a separate physical address and implemented each request or command individually at the granularity of a logical transfer unit.

[0026] Advantages of the present disclosure include, but are not limited to, improved read performance and the elimination of the expensive cost of interrupting write operations to service extremely many read requests (which occur more frequently than write operations), for example, by allowing many read requests to hit in the buffer using read lookahead operations. In addition, the present disclosure illustrates a way to perform direct cache hits and data transfers to reduce latency for sequential read requests from a host system (even if those read requests are interspersed with write operations and / or read requests to other regions). These advantages are synergistic with sequential writes performed by a ZNS-enabled memory device. Other advantages will be apparent to those skilled in the art of memory allocation and error optimization within a memory subsystem discussed below.

[0027] Figure 1A An example computing system 100 is shown that includes a memory subsystem 110 according to some embodiments of the present disclosure. The memory subsystem 110 may include media such as volatile memory (e.g., memory device 140), one or more non-volatile memory devices (e.g., memory device 130), or a combination of such devices. Each memory device 130 or 140 may be one or more memory components.

[0028] The memory subsystem 110 may be a storage device, a memory module, or a mixture of storage devices and memory modules. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).

[0029] The computing system 100 may be a computing device such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (e.g., an airplane, drone, train, car, or other transportation vehicle), a device with Internet of Things (IoT) capabilities, an embedded computer (e.g., an embedded computer included in a vehicle, industrial equipment, or a networked commercial device), or such a computing device that includes a memory and a processing device.

[0030] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to memory subsystems 110 of different types. Figure 1A An example of a host system 120 coupled to one memory subsystem 110 is shown. As used herein, "coupled to" or "coupled with" generally refers to a connection between components or devices, which may be an indirect communication connection or a direct communication connection (e.g., without intervening components or devices), whether wired or wireless, including electrical, optical, magnetic, etc.

[0031] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110, for example, to write data to the memory subsystem 110 and read data from the memory subsystem 110.

[0032] The host system 120 may be coupled to the memory subsystem 110 via a physical host interface that may communicate through a system bus. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, Fibre Channel, a Serial Attached SCSI (SAS), a Double Data Rate (DDR) memory bus, a Small Computer System Interface (SCSI), a Dual In-line Memory Module (DIMM) interface (e.g., a DIMM socket interface supporting Double Data Rate (DDR)), an Open NAND Flash Interface (ONFI), Double Data Rate (DDR), Low Power Double Data Rate (LPDDR), or any other interface. The physical host interface may be used to transfer data between the host system 120 and the memory subsystem 110. The host system 120 may further utilize an NVM Express (NVMe) interface to access components (e.g., memory device 130) when the memory subsystem 110 is coupled to the host system 120 through a PCIe interface. The physical host interface may provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120. As an example, Figure 1A Memory subsystem 110 is shown. In general, host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.

[0033] The memory devices 130, 140 may include any combination of different types of non-volatile memory devices and / or volatile memory devices. Volatile memory devices (e.g., memory device 140) may be, but are not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).

[0034] Some examples of non-volatile memory devices (e.g., memory device 130) include "NAND" (NAND) type flash memory and write-in-place memory, such as three-dimensional cross-point ("3D cross-point") memory. Non-volatile memory cross-point arrays can perform bit storage based on changes in body resistance in conjunction with stackable cross-grid data access arrays. In addition, in contrast to many flash-based memories, cross-point non-volatile memories can perform write-in-place operations, where non-volatile memory cells can be programmed if they have been previously erased. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).

[0035] Each of the memory devices 130 may include one or more arrays of memory cells. One type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), triple-level cells (TLC), and quad-level cells (QLC), may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more arrays of, for example, SLC, MLC, TLC, QLC, or any combination thereof. In some embodiments, a particular memory device may include an SLC portion, and an MLC portion, a TLC portion, or a QLC portion of memory cells. The memory cells of the memory device 130 may be grouped into pages, which may refer to a logical unit of a memory device for storing data. In some types of memory (e.g., NAND), pages may be grouped to form blocks.

[0036] Although non-volatile memory components such as NAND-type flash memory (e.g., 2D NAND, 3D NAND) and 3D cross-point non-volatile memory cell arrays are described, the memory device 130 can be based on any other type of non-volatile memory, such as read-only memory (ROM), phase change memory (PCM), self-selected memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridge RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), or non-(NOR) flash memory, electrically erasable programmable read-only memory (EEPROM).

[0037] The memory subsystem controller 115 (or, for simplicity, controller 115) may communicate with the memory device 130 to perform operations, such as reading data, writing data, or erasing data and other such operations at the memory device 130. The memory subsystem controller 115 may include hardware, such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry with dedicated (i.e., hard-coded) logic to perform the operations described herein. The memory subsystem controller 115 may be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or other suitable processor.

[0038] The memory subsystem controller 115 may include a processor 117 configured to execute instructions stored in a local memory 119. In the example shown, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.

[0039] In some embodiments, local memory 119 may include memory registers that store memory pointers, fetched data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1A The example memory subsystem 110 in FIG. 1 is shown as including a memory subsystem controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115, but may rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).

[0040] In general, the memory subsystem controller 115 may receive commands or operations from the host system 120, and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations, such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address conversion between logical block addresses (e.g., logical block addresses (LBA), namespaces) and physical addresses (e.g., physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include a host interface circuit system for communicating with the host system 120 via a physical host interface. The host interface circuit system may convert commands received from the host system into command instructions to access the memory device 130, and convert responses associated with the memory device 130 into information for the host system 120.

[0041] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that can receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.

[0042] In some embodiments, the memory device 130 includes a local media controller 135 that is used in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller (e.g., the memory subsystem controller 115) can externally manage the memory device 130 (e.g., perform media management operations on the memory device 130). In some embodiments, the memory device 130 is a managed memory device, which is a raw memory device combined with a local controller (e.g., the local media controller 135) for memory management within the same memory device package or memory die. An example of a managed memory device is a managed NAND (MNAND) device.

[0043] In some embodiments, controller 115 includes an error correction code (ECC) encoder / decoder 111. ECC encoder / decoder 111 may respectively perform ECC encoding of data written to memory device 130 and ECC decoding of data read from memory device 130. ECC decoding may be performed to decode ECC codewords, thereby correcting errors in the original read data, and in many cases, also reporting the number of bit errors in the original read data.

[0044] Figure 1B According to the embodiment Figure 1A In one embodiment, a memory subsystem controller 115 (e.g., a processing device, referred to as controller 115 for simplicity) includes one or more registers 112, a command generation processor 122 including a buffer manager 113, a translation processor 123, a command execution processor 124, and a volatile memory 125. For example, processor 117 ( Figure 1A ) may include a command generation processor 123, a translation processor 123 and a command execution processor 124.

[0045] In various embodiments, the volatile memory 125 stores the region mapping data structure 101, the read cache table 127, and the buffer index table 129, as well as Figure 2 In one embodiment, the region map data structure 101 includes a plurality of entries such that each entry has a block set entry identifier that is associated with an entry in the block set map data structure 107, which in turn may be associated with an entry in the page map data structure, which in turn may locate a page in memory, as will be described with reference to FIG. Figure 2Detailed explanation. In some embodiments, the volatile memory 125 includes one or both of a tightly coupled memory (TCM) and a static random access memory (SRAM) device. Storing the read cache table 127 and the buffer index table 129 in the TCM can make the buffer management discussed below as efficient as possible, but they can also be stored in an SRAM device or a combination thereof.

[0046] The memory subsystem 110 may further include a memory device 140A, which may be a dynamic random access memory (DRAM) device or other such volatile memory device, and is typically used to store larger data structures. Such a data structure may be a block set mapping data structure 107, which may map block set identifiers to individual data blocks in a physical address space. The memory device 140A may also be referred to as a shared volatile memory because it is shared by multiple processors for executing instructions and storing data.

[0047] In various embodiments, the memory device 140A may further store a block set mapping data structure 107 and buffers 142, which are shown by way of example to include a first buffer 142A, a second buffer 142B, a third buffer 142C, and so on to an Nth buffer 142N. Any group of these buffers 142 may be considered a set of buffers. The controller 115 may allocate buffers 142 (e.g., by buffer address) to store (e.g., cache) data in order to fulfill a read request. For example, each buffer may be an addressing slot or entry of the volatile memory device 140A. Each buffer may store a data page size or amount of data to which the LTU is mapped.

[0048] Furthermore, as previously discussed with reference to one or more memory devices 130, 140, the physical address spaces of multiple dies (e.g., die A and die B) can be hierarchically organized into planes, blocks, and pages. Thus, for example, each of die A and die B may include plane A and plane B, and each of plane A and plane B may include block A and block B. Block sets (or block stripes) may be defined as groups of blocks arranged across the planes of multiple dies of a memory device. As shown, block set 144 is arranged to include block A of plane A of die A, block A of plane B of die B, and so on, e.g., also including block A of plane C of die C, and other dies (if present and online).

[0049] In various embodiments, the translation processor 123 (and / or a dynamic data placer of the controller 115 coupled to the translation processor 123) dynamically determines a layout for placing data associated with a logical address in cells or memory components (also referred to as "IC dies") of the memory devices 130, 140. The layout specifies a mapping between logical addresses for commands received in the memory subsystem 110 (e.g., from the host system 120) and physical memory locations in the IC dies of the memory subsystem 110.

[0050] For example, based on the availability of IC dies to write, program, store, commit data when input / output scheduling is performed in the memory subsystem 110, the translation processor 123 can determine a layout of a portion of the logical addresses of the LBA space for placing data at a logical address associated with a physical address of the medium of the memory device 130 or 140. When the IC die (including the physical cells within the IC die) is available to commit programming data, a write command is scheduled for execution in the memory subsystem 110; the translation processor 123 generates a portion of the layout for the write command and maps the logical address for the write command for mapping to a memory location within the IC die. The execution of the write command causes the memory subsystem 110 to commit / program the data associated with the write command into the IC die.

[0051] Depending on the availability of the media and / or application of the IC die cross-sequential mapping scheme, the controller 115 may write the data groups sequentially (e.g., provided in a sequential stream of data locations) to fill one IC die at a time, or may write sequentially to multiple IC dies in parallel at a time, for example, to fill IC dies in parallel (if a memory device). The mapping of writes within a region of LBA space may also be done sequentially to simplify the calculations used to perform the mapping, as will be discussed in more detail with reference to the ZNS mapping data structure. When there are multiple IC dies available, the logical addresses (e.g., LBA or LTU values) for commands from multiple write streams may be mapped to the multiple IC dies separately by the dynamically generated portion of the layout so that there are no access conflicts when executing commands from multiple write streams.

[0052] In various embodiments, the translation processor 123 accesses certain ZNS mapping data structures (e.g., the zone mapping data structure 101 and the block set mapping data structure 107) in order to translate LBA values ​​into physical block addresses (PBAs) of a physical address space. The translation processor 123 may be referred to as a flash translation layer (FTL) in the context of flash media (e.g., NOR or NAND flash memory). The mapping employed by the translation processor 123 may generally be assumed to involve some type of block mapping, such as a block-level mapping or a log-level mapping. Where data locations are detected or in a ZNS memory device, the translation processor 123 may map groups of blocks forming zones, such as groups of blocks within a ZNS data structure, which may involve mapping sequentially numbered LBA values ​​to sequentially numbered PAs. Reference Figure 2 The use of ZNS data structures and their interrelationships to map a zone's LBA space to the physical address space of the media is discussed in greater detail.

[0053] In various embodiments, the command generation processor 122 performs command processing, including processing a read or write command received from the host system 120 or generating a read command and a write command based on a read and write request received from the host system 120 or another request agent, respectively. As part of performing a read request, the buffer manager 113 of the command generation processor 122 may allocate a specific number (e.g., a "set") of buffers having a capacity to match the amount of data stored at a first physical address (mapped to the LTU value created to include the LBA value of the read request) and a subsequent physical address, the subsequent physical address being numbered sequentially after the first physical address, for example, within a read offset value that defines a read window size for the memory device. For purposes of explanation, the data stored at the first physical address may be referred to as first data and the data stored at the subsequent physical address may be referred to as second data. In one embodiment, the read offset value is 128KB, thereby enabling a read lookahead of an additional 124KB of second data beyond the first read request, but the range of read offset values ​​may be up to 2MB or more. Various other offset values ​​or read window sizes are contemplated. For example, each subsequent physical address may be incremented by a page number across a read window of a sequential physical address space defined by a read offset value to determine the subsequent physical address.Allocating and tracking buffers by the buffer manager 113 facilitates read lookahead operations.

[0054] In these embodiments, the command generation processor 122 may further generate a set of command tags for instructing the command execution processor 124 of the processing device to retrieve data from the first address and subsequent addresses of the memory device 130 or 140 and store the data in the set of buffers. The command generation processor 122 may further transmit a command group including the set of command tags to the command execution processor 124. In this way, each command tag in the set of command tags includes an LTU value mapped to one of the corresponding sequential physical addresses. Each command tag also includes a buffer address corresponding to a buffer in the buffer 142 stored in the memory device 140A.

[0055] In a related embodiment, the buffer manager 113 may track the use of the buffers 142 and is responsible for locking or releasing the buffers when host commands hit these buffers using any number of buffer management algorithms, thereby tracking the data cached in various buffers by a number of potential applications. The buffer management algorithms may include, for example, a two-three tree algorithm (also known as a 2-3 tree algorithm) in which the buffers 142 are sorted by LTU value (or LBA value), a linked list algorithm, or an N-way cache using a hash algorithm, as well as other cache management algorithms.

[0056] For the purpose of explanation, it is assumed that read commands are intermixed with write commands to more than one region, but three of the read commands include a read command to region zero ("0") having an LBA_0 value, a read command to region 44 having an LBA_44 value, and a read command to region 23 having an LBA_23 value. When performing a read lookahead, the buffer manager 113 may generate a set of command tags to read the lookahead 1MB of data (as a read offset value), which includes first data and second data. The command tags may include a buffer address as a way of allocating one buffer 142 to each LTU value in the set of LTU values ​​identified as being mapped to the read offset value having the first data and the second data.

[0057] Item Number Region Identifier Starting LBA value End LBA value Read offset value 1 0 0 1024 1,024KB 2 44 4400 5424 1,024KB 3 23 2300 2454 1,024KB

[0058] Table 1

[0059] To track and manage the allocation of buffers 142, the buffer manager 113 may further create and update a read cache table 127. For example, the buffer manager 113 may create an entry in the read cache table 127 for each read lookahead operation to track the allocation of buffers 142 to each set of corresponding LTU values ​​tagged by the read lookahead command. Table 1 shows an example of the presentation of the read cache table 127 based on three read commands to the three different regions discussed previously. Each entry may include a region identifier, a starting LBA value (e.g., of the initial or first LBA retrieved according to the read request or command), an ending LBA value, and a read offset value (e.g., 1MB in this example). However, the region identifier is optional because the starting LBA value also identifies the region. The read offset value may be a fixed predetermined system amount of read lookahead data, and the ending LBA value may be an LBA value corresponding to the end of the read offset value starting from the starting LBA value within the physical address space. The starting LBA value and the ending LBA value may define a range of LBA values ​​corresponding to the read offset value in the physical address space, such as a read window size. As mentioned, the read-lookahead data includes first data corresponding to a first LTU value in different instances, and second data corresponding to an additional 63 LTU values ​​of 63 additional data chunks mapped to a corresponding one of the sequential physical addresses of the memory devices 130 or 140. The specific numbers set forth in this example are for purposes of explanation only and may be different in different implementations or scenarios.

[0060] In various embodiments, the buffer manager 113 may further create and update a buffer index table 129 in which sequential read data cached in the buffers is indexed relative to non-contiguous buffer numbers, as shown in Table 2. For example, the data index may refer to an LTU value corresponding to the set of LTU values ​​that is included in a read lookahead command tag generated for a read command to be sent to the command execution processor. For example, each LTU value may be a "data index" indexed relative to a buffer address corresponding to the LTU value. Thus, in the example buffer index table 129 of Table 2, the buffer address may be a "buffer index."

[0061] Data index (for example, for 64 bytes of data) Buffer Index 0 0 1 3 2 4 3 1 … … 63 230

[0062] Table 2

[0063] Metadata that can be used by such buffer management algorithms (e.g., for tracking buffer allocation and usage) includes the LTU / LBA value (based on which the data is ordered), the buffer address (or other buffer identifier for indexing) indicating at which buffer slot the data resides, and a buffer usage count that enables multiple users (e.g., host applications) in a separate read or write path to be tracked jointly. In this way, the buffer manager 113 can manage multiple applications writing and reading multiple areas where any set of commands can write to or read from sequentially stored data, but the allocated buffers may not be sequentially numbered, as shown in Table 2. If the data in the buffer is tracked, hardware acceleration can be used to facilitate data tracking and management in the buffer.

[0064] As another example, after the buffer manager 113 has created or updated the read cache table 127 according to Table 1 and the buffer index table 129 according to Table 2, assume that the host system 120 subsequently issues a read request or command to subsequent sequential LBA values, as shown in Table 3. For example, the sequential LBA value for region zero ("0") may be LBA_0 plus 16KB, followed by LBA_0 plus 32KB, followed by LBA_0 plus 48KB, and so on. The value 16KB may be referred to as a buffer allocation unit offset for the amount of data in the memory device corresponding to the LTU value. The buffer manager 113 may then perform an offset calculation in different cases and determine the LTU value associated with the LBA value retrieved from the subsequent read request. As shown, the read request / command is for sequentially numbered LTU values. The determined LTU values ​​may be, for example, LTU_1, LTU_2, and LTU_3, respectively, as shown in Table 3. Once the LTU value is known, the buffer manager 113 may index within the buffer index table 129 to determine the buffer address corresponding to the LTU value associated with the LBA value in a subsequent read command or request.

[0065]

[0066]

[0067] Table 3

[0068] As further explanation, assume that command generation processor 122 retrieves a second LBA value from a second request received from host system 120. Then, buffer manager 113 may determine, via accessing an entry in read cache table 127, that the second LBA value (LBA_16) has a single buffer allocation unit offset (16KB) relative to the starting LBA value (LBA_0), and therefore corresponds to a second LTU value (LTU_1) in the set of LTU values. Buffer manager 113 may further determine that the second LBA value is within the range of LBA values ​​corresponding to the read offset value (1,024KB=1MB), and therefore is not out of range. Buffer manager 113 may further use the second LTU value to index within buffer index table 127 to retrieve a second buffer address (buffer index 3 in Table 2). Buffer manager 113 may then return a subset of the second data retrieved from a second buffer in the set of buffers corresponding to the second buffer address to host system 120.

[0069] In various embodiments, the buffer manager 113 further uses a flag (e.g., a bit flag) or a counter to track whether any given buffer is being used for a read or write path. This can allow the buffer to satisfy in-flight commands (e.g., already processed) with fast seek times and find specific LBAs with short seek times, so the buffer can be used for cache hits and direct passes to the host system 120 without having to go back to the translation processor 123 for mapping. As long as the region mapping data structure 101 is checked first, consistency due to retrieving data from a buffer that performs similar to a cache should not be an issue, and the command generation processor 122 will continue to do so in the disclosed sequential read optimization. In some embodiments, the controller 115 includes at least a portion of the buffer manager 113. In other embodiments or combinations, the controller and / or the processing device of the host system 120 includes at least a portion of the buffer manager 113. For example, the controller 115 or the processing device of the host system 120 can be configured to execute instructions stored in a memory to perform the operations of the buffer manager 113 described herein. In some embodiments, the buffer manager 113 is implemented in an integrated circuit chip disposed in the memory subsystem 110. In other embodiments, buffer manager 113 is part of an operating system, device driver, or application of host system 120 .

[0070] In these embodiments, the command execution processor 124 orders write and read commands within the channels of the data bus to the memory devices 130, 140. In addition, the command execution processor 124 may retrieve data from a first physical address and from subsequent physical addresses pointed to by the set of command tags of the memory devices 130, 140 in response to receipt of a read command. It should be remembered that each command tag includes an LTU value mapped to a physical address and identifies a buffer address cached within a buffer in the volatile memory device 140A. The command execution processor 124 may further store (e.g., cache) data implementing the read command in an allocated buffer consistent with the corresponding buffer addresses of the set of command tags generated by the command generation processor 122 and included in the command group sent to the command execution processor 124. The command execution processor 124 may further perform, for example, error handling in the physical layer corresponding to the physical address space.

[0071] The translation processor 123 translates the LTU value into a physical address of a physical address space to facilitate the command generation processor 122 to generate a command to the command execution processor 124. Therefore, the translation processor 123 can act as an intermediary between the command generation processor 124 (which receives a memory request with an LBA value and creates an LTU value containing the LBA value) and the command execution processor 124 that needs to know the physical address of the physical layer to implement the command. In the present disclosure, in the read lookahead operation of the sequential read optimization, the conventional use of the translation processor 123 indexing into various ZNS mapping data structures can be eliminated.

[0072] Figure 2 is a block diagram illustrating an example of a data structure configured to support region-based mapping according to various embodiments. The controller 115 may Figure 2 The data structures shown are stored in local memory 119 (e.g., SRAM) or in a memory component (e.g., DRAM) of memory device 140. Controller 115 may also use Figure 2 The data structure configures or implements the media layout (e.g., the layout in which groups of data for a region will be located within a physical address space). Figure 2 In the example, the zone map data structure 201 is configured to provide media layout information for zones in a namespace (eg, an LBA space for ZNS operations). The zone map data structure 201 may be associated with Figure 1B The region map data structure 201 may be the same or similar to the region map data structure 101. The region map data structure 201 may have multiple entries. Each region map entry in the region map data structure 201 identifies information about the region, such as the starting LBA address 211 of the region, the block set identifier 213 of the region, the region cursor value 215 of the region, the status 217 of the region, and the like.

[0073] The host system 120 writes data in the region starting at the LBA of the region start LBA identifier 211. The host system 120 writes data sequentially in the region in the LBA space. After a certain amount of data has been written to the region, the current starting LBA address for writing subsequent data is identified by the region cursor value 215. Each write command to the region moves the region cursor value 215 to the new starting LBA address for the next write command to the region. The status 217 may have values ​​indicating that the region is empty, full, implicitly open, explicitly open, closed, etc. to track the progress of writing to the region.

[0074] exist Figure 2 In the embodiment of the present invention, the logical to physical block mapping data structure 203 is configured to facilitate the translation of LBA addresses to physical addresses in the IC die. The logical to physical block mapping 203 may have multiple entries. The LBA value may be used as or converted into an index (e.g., LTU value) of an entry in the logical to physical block mapping 203. The index may be used to find an entry for an LBA value. Each entry in the logical to physical block mapping 203 identifies the physical address of a memory block in the IC die for an LBA value. For example, the physical address of a memory block in the IC die may include a die identifier 233, a block identifier 235, a page mapping entry identifier 237, and the like. The die identifier 233 identifies a specific IC die (e.g., die A or die B) of the memory device 130, 140 of the memory subsystem 110. The block identifier 235 identifies a specific memory block (e.g., NAND flash memory) within the IC die identified using the die identifier 233. The page mapping entry identifier 237 identifies an entry in the page mapping data structure 205.

[0075] The page mapping data structure 205 may have multiple entries. Each entry in the page mapping 205 may include a page identifier 251 that identifies a page of memory cells within a block of memory cells (e.g., NAND memory cells). For example, the page identifier 251 may include a word line number of the page in the block of NAND memory cells and a sub-block number of the page. In addition, the entry for the page may include a programming mode 253 for the page. For example, the page may be programmed in SLC mode, MLC mode, TLC mode, or QLC mode. When configured in SLC mode, each memory cell in the page will store one data bit. When configured in MLC mode, each memory cell in the page will store two data bits. When configured in TLC mode, each memory cell in the page will store three data bits. When configured in QLC mode, each memory cell in the page will store four data bits. Different pages in an integrated circuit die may have different data programming modes.

[0076] exist Figure 2In the example, the block set data structure 207 stores data of various aspects of the dynamic layout of the control area. The block set data structure 207 can be used with Figure 1B The block set mapping data structure 207 may be the same or similar to the block set mapping data structure 107 of the region. The block set data structure 207 may be a table in one embodiment, which may have multiple entries. Each block set entry in the block set data structure 207 identifies the number / count 271 of integrated circuit dies (e.g., die A and die B) in which the data of the region is stored. For each of the integrated circuit dies for the region, the block set entry of the block set data structure 207 has a die identifier 273, a block identifier 275, a page mapping entry identifier 277, a page mapping offset value, etc.

[0077] Die identifier 273 identifies a specific IC die (e.g., Die A or Die B) among the IC dies of memory subsystem 110 on which subsequent data for the region may be stored. Block identifier 275 identifies a specific memory block (e.g., NAND flash memory or other media) within the IC die identified using die identifier 273 in which subsequent data for the region may be stored. Page map entry identifier 237 identifies a page map entry in page map data structure 205 that identifies a page that may be used to store subsequent data for the region.

[0078] For example, the memory subsystem 110 receives write commands for multiple streams. In an embodiment, each corresponding stream in the multiple streams is configured to write data sequentially in a logical address space in one embodiment; and in another embodiment, the streams in the multiple streams are configured to write data pseudo-sequentially or randomly in a logical address space in one embodiment. Each write stream includes a set of commands marked to write, trim, overwrite a group of data together as a group. In the group, data can be written to the logical space sequentially, randomly, or pseudo-sequentially. Preferably, the data in the group is written to an erase block set, wherein the memory cells in the erase block set store the data of the stream but do not store data from other streams. The erase block set can be erased to remove the data of the stream without erasing the data of other streams.

[0079] For example, each write stream is permitted to write sequentially at the LBAs of the regions in the namespace in the IC die of the memory devices 130, 140 allocated to the memory subsystem 110, but is prohibited from writing data out of order in the LBA (or logical address) space. The translation processor 123 of the memory subsystem 110 identifies multiple physical or erase (or erase) units in the memory subsystem 110 that can be used to write data in parallel.

[0080] The translation processor 123 may select the first command from the plurality of streams for parallel execution in a plurality of physical units available for writing data. Dynamically in response to the first command being selected for parallel execution in the plurality of physical units, the translation processor 123 may generate and store a portion of a layout that maps a logical address identified by the first command in the logical address space to a physical address of a memory unit in the plurality of memory units.

[0081] The command execution processor 124 may execute the first command in parallel by storing data in the memory cells according to the physical addresses. For example, while the first command is being scheduled for execution, execution of the second command may be ongoing in a subset of memory cells of the IC die of the memory subsystem 110. Therefore, the subset of memory cells used to execute the second command is not available for the first command. After scheduling the first command and determining the portion of the layout of the logical addresses for the first command, the first command may be executed in parallel in the plurality of physical cells and / or in parallel with execution of the second command in the remaining physical cells of the memory subsystem 110.

[0082] For example, after identifying a plurality of memory units (e.g., IC dies) that may be used to execute a subsequent command, the translation processor 123 may identify physical addresses that may be used to store data for the subsequent command from the block set data structure 207. The physical addresses may be used to update corresponding entries in the logical-to-physical block mapping data structure 203 for the LBA addresses for the subsequent command.

[0083] For example, when the IC die can freely write data, the translation processor 123 can determine the commands that can be written / programmed to the region in the memory cells in the IC die. Based on the block set data structure 207, the translation processor 123 locates the entry for the region, locates the block identifier 275 and the page mapping entry identifier 277 associated with the identifier 273 of the integrated circuit die, and updates the corresponding fields of the entry in the logical-to-physical block mapping data structure 203 for the LBA of the command for the region using the die identifier 273, the block identifier 275, and the page mapping entry identifier 277.

[0084] Figure 3 According to the embodiment Figure 1A-1B FIG. 1 is a block diagram of the relationship between the command generation processor 122, the translation processor 123, and the command execution processor 124 of the memory subsystem controller 115 of FIG. 1. In various embodiments, the controller 115 includes a shared volatile memory 140B and a command buffer 140C in the shared volatile memory 140B. In one embodiment, the shared volatile memory 140B is a reference Figure 1BIn various embodiments, the command generation processor 122 may receive a first read request from the host system 120 (or other request agent). The first read request may include a first LBA value that corresponds to a first physical address of the memory device 130 or 140 for which the read operation is directed. In servicing the first read request, the command generation processor 122 may create a first logical transfer unit (LTU) value that includes the first LBA value, which will be mapped to the first physical address of the memory device 130, 140.

[0085] In some embodiments, the translation processor 123 may be configured to automatically store (or buffer) the LTU to physical address (PA) mapping 301 in the shared volatile memory 140B when its data is programmed into the memory devices 130, 140. For example, the LTU to PA mapping 301 may be a portion of the logical to physical block mapping data structure 203 and the page mapping data structure 205 that is written when the corresponding physical address is programmed into the memory device 130 or 140. This may provide a quickly accessible data structure that only provides the LTU to PA mapping at the command tag level. In some embodiments, the LTU to PA mapping 301 in the shared volatile memory 140B may be considered a cache to keep the size of this data structure limited.

[0086] Continue to refer Figure 3 , the translation processor 123 may further selectively set a flag 303 (e.g., a bit flag, etc.) in the shared volatile memory 140B. Thus, each entry in the LTU to PA map 301 may include a physical address mapped to an LTU value and a flag. In an alternative embodiment, the address stored in register 112 ( Figure 1B ) may be set. One bit value in the bitmap may correspond to a particular LTU value and, therefore, serve as a flag 303 for the shared volatile memory 140B. The bitmap may be associated with a physical address space that is known to be written sequentially (e.g., for each ZNS operation). The flag 303 (or a bit value within the bitmap) may indicate whether the LTU-to-PA mapping entry is associated with a region of the LBA address space, where the region maps to data that is written sequentially in the memory device 130 or 140. The translation processor 123 may further selectively set a die available flag 305 to indicate that the die at the physical address is available to service the command. In some embodiments, if there is more than one die for a region, there may be more than one flag, one for each region.

[0087] In various embodiments, if the flag 303 is set and the die available flag 305 is set, the command generation processor 122 performs the read lookahead optimization disclosed herein. The read optimization may include, for example, automatically incrementing the first physical address retrieved from the LTU to PA map 301 for the first read request to determine the subsequent physical address within the offset value (e.g., the read window size) of the first physical address. In one embodiment, the automatic increment is performed so that the first physical address is incremented by the page number until the end of the read window size is reached from the first physical address.

[0088] Next, the command generation processor 122 may generate (or update) in the command buffer 140C for instructing the command execution processor 124 to retrieve data from the memory device 130 or 140 and store the data in a set of buffers (described above). Figure 1B The command generation processor 122 may then further transmit a group of commands each including one of the set of command tags to the command execution processor 124 of the processing device. The set of commands may be buffered in the command buffer 140C as Cmd[0], Cmd[1], Cmd[2], and so on until Cmd[n]. In one embodiment, the command generation processor 122 may connect the set of command tags into a command chain (e.g., a series of corresponding commands) and transmit the command chain to the command execution processor 124 in a single command message.

[0089] After the command execution processor 124 stores the data in the corresponding buffer allocated for the read lookahead of the first (or initial) read command, the command generation processor 122 may return the data stored at the first physical address to the host system 120 or other request agent. However, the command generation processor 122 may also further service subsequent read requests or commands directly for subsequent physical addresses outside of the buffer, as described herein. For example, in response to a second read request, the command generation processor 122 may determine that the second LBA value of the second read request corresponds to a second physical address in the subsequent physical addresses. The command generation processor 122 may then retrieve a second subset of the data from a second buffer in the set of buffers having a buffer address associated with a second command tag in the set of command tags, and transmit the second subset of the data to the host system 120 in response to a first of the subsequent read commands.

[0090] Figure 44 is a flow chart of a method 400 for supporting direct cache hits based on read commands according to an embodiment. The method 400 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 400 is performed by Figure 1A-1B The controller 115 (e.g., command generation processor 122) is executed. Although shown in a specific order or sequence, unless otherwise specified, the order of the process can be modified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in different orders, and some processes can be performed in parallel. In addition, in various embodiments, one or more processes can be omitted. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0091] At operation 410, processing logic accesses a read command with a first command tag, wherein the first command tag includes a first logical transfer unit (LTU) value and a first buffer address of a first buffer in buffer 142. For example, the read command may be received from host system 120 or generated by processing logic based on the contents of a read request received from host system 120. The first LTU value is mapped from a region of a plurality of sequential logical block address (LBA) values ​​to a first physical address of a plurality of sequential physical addresses of memory device 130 or 140. In this embodiment, first data is stored at the first physical address, and the first LTU value contains a first LBA value of the plurality of sequential LBA values.

[0092] At operation 420, processing logic generates a set of command tags that instructs the command execution processor of the processing device to retrieve the second data from the memory device and store the second data in a set of buffers. In this embodiment, for example, the set of command tags includes a second command tag associated with a second physical address that is sequentially after the first physical address, a third command tag associated with a third physical address that is sequentially after the second physical address, and so on, until the number of command tags is sufficient to fill the command to read the second data as well as the first data. This read lookahead can be performed without other translation work (performed by the translation processor 123) or read command execution work performed by the command execution processor 124 at the memory device 130 or 140.

[0093] At operation 430, processing logic creates an entry in the read cache table 127 for the set of buffers. For example, the entry may include a region identifier for the region, a starting LBA value set to the first LBA value, and a read offset value corresponding to the amount of the first data and the second data. The entry may further include an ending LBA value corresponding to the end of the read offset value starting from the starting LBA value within the physical address space. An example of the read cache table 127 is shown in Table 1.

[0094] Figure 5 is a flow chart of a method 500 for sequential read optimization according to an embodiment. The method 500 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 500 is performed by Figure 1A-1B The controller 115 (e.g., command generation processor 122) is executed. Although shown in a specific order or sequence, unless otherwise specified, the order of the process can be modified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in different orders, and some processes can be performed in parallel. In addition, in various embodiments, one or more processes can be omitted. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0095] refer to Figure 5 At operation 510, processing logic receives a read request for a first logical block address (LBA) value of a memory device including a first LBA value of an address space of the memory device. At operation 515, processing logic creates a first logical transfer unit (LTU) value including the first LBA value, the first LTU value being mapped to a first physical address of the memory device. The first LTU value need not be sent to the translation processor 123, as long as the translation processor 123 generates a first LTU value in the shared volatile memory 140B ( Figure 3 ) has already created an LTU to PA entry in the LTU to PA mapping 301 of the shared memory 140B. Figure 3 ) is the first physical address indexed by the first LTU value within.

[0096] Continue to refer Figure 5At operation 520, processing logic determines whether a quick find flag is set. This quick find flag may be a flag 303 set in shared volatile memory 140B and associated with the first physical address; or the quick find flag may be a bit value in a register 112 storing a bitmap associated with a ZNS-related physical address space of memory device 130 or 140. In either case, the quick find flag may indicate that the first LTU value is associated with a region of multiple sequential LBA values ​​that are sequentially mapped to multiple sequential physical addresses. At operation 540, processing logic determines whether a die available flag 305 is set, which has been referenced. Figure 3 Be discussed.

[0097] At operation 530, if the fast lookup flag or the die available flag is not set, then processing logic submits the read request via the normal read path, which includes sending the first LTU value to the translation processor 123 to cause the translation processor 123 to perform a lookup within the ZNS data structure that maps the first LTU to the first physical address.

[0098] At operation 550, assuming that the fast find flag and the die available flag have been set relative to the first LTU value, processing logic retrieves the first LTU value from a volatile memory (eg, Figure 3 The shared volatile memory 140B in the shared volatile memory 140B) retrieves the first physical address that has been stored (or buffered) there by the translation processor 123. The first physical address may be indexed within an entry of the LTU-to-PA map 301 in the shared volatile memory 140B.

[0099] At operation 560, processing logic allocates a set of buffers in volatile memory, wherein the capacity of the set of buffers matches the amount of data stored at a first physical address and subsequent physical addresses, the subsequent physical addresses being sequentially numbered within a read window size defined, for example, by a read offset value (Table 1) starting at the first physical address (e.g., the first LTU value maps to the first physical address). The volatile memory storing the buffers may be volatile memory 125, volatile memory device 140A, and / or shared volatile memory 140B. In various embodiments, processing logic determines each subsequent physical address by increasing the first physical address by a page number until the end of the read window size (e.g., the offset value) is reached.

[0100] At operation 570, processing logic generates a set of command tags instructing the command execution processor 124 of the processing device to retrieve data from the memory device and store the data in the set of buffers. The set of command tags may include a first command tag associated with a first physical address and additional command tags associated with subsequent physical addresses.

[0101] At operation 580, processing logic creates an entry in a read cache table for the set of buffers, wherein the entry includes a region identifier for the region, a starting LBA value set to a first LBA value, and a read offset value corresponding to the amount of data. The entry may further include an ending LBA value corresponding to the end of the read offset value within the physical address space starting from the starting LBA value. Processing logic may then use the read cache table to identify that subsequent requests or commands are for LBA values ​​corresponding to physical addresses within the read offset value and are therefore stored within the set of buffers. Thus, processing logic may retrieve a subset of the data stored in the buffers as a cache hit and return the subset of data to the host system 120 or other requesting agent.

[0102] Figures 6A-6C 6 is a flow chart of a method 600 for supporting direct cache hits and transfers according to a related embodiment. The method 600 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 600 is performed by Figure 1A-1B The controller 115 (e.g., command generation processor 122) is executed. Although shown in a specific order or sequence, unless otherwise specified, the order of the process can be modified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes can be performed in different orders, and some processes can be performed in parallel. In addition, in various embodiments, one or more processes can be omitted. Therefore, not all processes are required in every embodiment. Other process flows are possible.

[0103] refer to Fig. 6A At operation 605, processing logic receives a read request for a first logical block address (LBA) value of an address space of a memory device, wherein the first LBA value belongs to a region of a plurality of sequential LBA values ​​mapped to a plurality of sequential physical addresses. At operation 610, processing logic creates a first logical transfer unit (LTU) value including the first LBA value, wherein the first LTU value is mapped to a first physical address of the memory device 130 or 140.

[0104] At operation 615, processing logic allocates a set of buffers in volatile memory (e.g., in buffer 142), where the capacity of the set of buffers matches the amount of data stored at the first physical address and subsequent physical addresses, the subsequent physical addresses being sequentially numbered within the read offset value starting at the first physical address. The volatile memory storing the buffers may be volatile memory 125, volatile memory device 140A, or other shared volatile memory 140B. In various embodiments, processing logic determines each subsequent physical address by increasing the first physical address by the page number until the end of the read offset value is reached.

[0105] At operation 620, the processing logic generates a set of command tags instructing a command execution processor of the processing device to retrieve data from the memory device and store the data in the set of buffers. In this embodiment, the set of command tags includes at least a first command tag associated with a first physical address and a second command tag associated with a second physical address that is sequentially after the first physical address in the subsequent physical addresses. The set of command tags may include additional command tags, for example, a total of up to 64 command tags with sequentially ordered LTU values, where each LTU value corresponds to 16KB of data, and the read offset value of the read-ahead data is 1MB.

[0106] At operation 625, processing logic creates an entry in a read cache table for the set of buffers, wherein the entry includes a region identifier for the region, a starting LBA value set to the first LBA value, and a read offset value corresponding to the amount of data. The entry may further include an ending LBA value corresponding to the end of the read offset value starting from the starting LBA value within the physical address space. An example of the read cache table 127 is shown in Table 1.

[0107] refer to Figure 6B , method 600 is performed in the context of the set of command tags containing a set of LTU values ​​corresponding to a subset of a plurality of sequential physical addresses within a read offset value. With reference to the generation of the set of command tags and the generation of the buffer index table, Figure 6B Portions of method 600 in FIG. 6A may be viewed as additional details.

[0108] At operation 630, processing logic assigns an LTU value from the set of LTU values ​​to each command tag in the set of command tags. The LTU values ​​may be assigned sequentially and may correspond to increments of buffer allocation units starting at a starting LBA value (e.g., 16KB in the above example). At operation 635, processing logic assigns a buffer address of a buffer within the set of buffers to each command tag in the set of command tags. This assignment may be in Fig. 6AThe allocated portion of the set of buffers in the volatile memory discussed at operation 615 of . In one embodiment, the buffer addresses allocated to corresponding sequential LTU values ​​are not necessarily sequential or continuous, so the allocation within the buffer index table can be managed by a buffer management algorithm, as previously discussed.

[0109] At operation 640, processing logic generates a buffer index table in volatile memory to track the LTU value associated with each command tag in the set of command tags that is mapped to the buffer address associated with the LTU value. An example of the buffer index table 129 is shown in Table 2. Processing logic may also track entries of the buffer index table according to a linked list or one of a two-three tree algorithm in which the set of buffers is sorted by LTU value. In this way, processing logic may access the buffer index table 129 after determining a subsequent LTU value associated with the current read request or command and locate a corresponding buffer in the set of buffers that contains the requested data.

[0110] More precisely, reference Figure 6C And using Tables 1 and 2, method 600 is performed to implement subsequent read requests or read commands in cache-type access, thereby pulling data from the set of buffers instead of engaging the translation processor 123 or the command execution processor 124 to pull data from the memory device 130 or 140. At operation 645, processing logic retrieves the second LBA value from the subsequent read request or command received from the host system 120. For example, for the purpose of explanation, the second LBA value is (LBA_16).

[0111] At operation 650, processing logic determines, via accessing an entry in the read cache table 127, that the second LBA value has a buffer allocation unit offset (e.g., 16KB) relative to the starting LBA value (LBA_0), and therefore corresponds to a second LTU value (e.g., LTU_1) in the set of LTU values, and is within the range of LBA values ​​corresponding to the read offset value. At operation 660, processing logic indexes within the buffer index table 129 using the second LTU value to retrieve a second buffer address, such as buffer index 3 in Table 2. Other indexing or addressing values ​​associated with locations within volatile memory are contemplated and may be managed by the buffer manager discussed. At operation 665, in response to a cache hit at the set of buffers, processing logic returns to the host system 120 a subset of data retrieved from a second buffer in the set of buffers corresponding to the second buffer address. In this way, the processing logic can also improve read performance and avoid the expensive cost of interrupting write operations to service many read requests (which occur more frequently than write operations), such as by allowing many read requests to hit in the buffer using read lookahead operations.

[0112] Figure 6C The portion of method 600 depicted in the embodiment of the present invention may be extended to other or subsequent read requests or read commands. For example, the processing logic may retrieve the third LBA value from a subsequent read request received from the host system. The processing logic may further determine, via accessing an entry in the read cache table, that the third LBA value has a buffer allocation unit offset of two times relative to the starting LBA value, and therefore corresponds to a third LTU value in the set of LTU values ​​and is within the range of LBA values ​​corresponding to the read offset value. The processing logic may further use the third LTU value to index within the buffer index table to retrieve a third buffer address. The processing logic may further return to the host system a subset of data retrieved from a third buffer in the set of buffers corresponding to the third buffer address.

[0113] Figure 7 An example machine of a computer system 700 is shown, within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system 700 may correspond to a host system (e.g., Figure 1A ) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1A In some embodiments, the machine may be connected (e.g., using a network) to other machines. The machine may operate in the capacity of a server or a client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment.

[0114] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing (sequentially or otherwise) a set of instructions that specify actions to be taken by the machine. In addition, while a single machine is shown, the term "machine" shall also be taken to include any collection of machines that individually or collectively execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.

[0115] The example computer system 700 includes a processing device 702, a main memory 704 (e.g., read-only memory (ROM), flash memory, dynamic random access memory (DRAM), such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM)), etc.), a static memory 706 (e.g., flash memory, static random access memory (SRAM), etc.), and a data storage system 718, which communicate with each other via a bus 730.

[0116] The processing device 702 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device 702 can also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 702 is configured to execute instructions 726 for performing the operations and steps discussed herein. The computer system 700 may further include a network interface device 708 that communicates via a network 720.

[0117] The data storage system 718 may include a machine-readable storage medium 724 (also referred to as a non-transitory computer-readable medium) on which is stored one or more sets of instructions 726 or software embodying any one or more of the methodologies or functions described herein. The instructions 726 may also reside completely or at least partially within the main memory 704 and / or within the processing device 702 during execution by the computer system 700, the main memory 704, and the processing device 702, which also constitute the machine-readable storage medium. The machine-readable storage medium 724, the data storage system 718, and / or the main memory 704 may correspond to Figure 1A-1B Memory subsystem 110.

[0118] In one embodiment, instruction 726 includes implementing a corresponding Figure 1B The machine-readable storage medium 724 is shown as a single medium in the example embodiment, but the term "non-transitory machine-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. The term "machine-readable storage medium" should therefore be considered to include, but not be limited to, solid-state memory, optical media, and magnetic media.

[0119] Some portions of the previous detailed description have been presented with respect to algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the most effective means for those skilled in the art of data processing to communicate the content of their work to others skilled in the art. Here and in general, an algorithm is conceived as a self-consistent sequence of operations that produce a desired result. The operations are those that require physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals that can be stored, combined, compared, and otherwise manipulated. It has been demonstrated that it is sometimes convenient to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, etc., primarily for common reasons.

[0120] It should be borne in mind, however, that all of these and similar terms are associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the computer system's registers and memories into other data similarly represented as physical quantities within the computer system's memories or registers or other such information storage systems.

[0121] The present disclosure also relates to an apparatus for performing the operations described herein. This apparatus may be specially constructed for the desired purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memories (ROMs), random access memories (RAMs), EPROMs, EEPROMs, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, which are respectively coupled to a computer system bus.

[0122] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used in conjunction with the programs according to the teachings herein, or it may prove convenient to construct a more specialized device to perform the method. The structures for various of these systems will be presented from the description below. In addition, the present disclosure is not described with reference to any particular programming language. It should be appreciated that the teachings of the present disclosure as described herein may be implemented using a variety of programming languages.

[0123] The present disclosure may be provided as a computer program product or software, which may include a machine-readable medium having stored thereon instructions that can be used to program a computer system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, etc.

[0124] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments thereof. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. Therefore, the description and drawings should be viewed in an illustrative sense rather than a restrictive sense.

Claims

1. A memory subsystem comprising: Memory device; volatile memory including a buffer and a read cache table; as well as a processing device coupled to the memory device and the volatile memory, wherein the processing device performs the following operations: accessing a read command having a first command tag, the first command tag including a first logical transfer unit value and a first buffer address of a first buffer of the buffers, the first logical transfer unit value being mapped from a region of a plurality of sequential logical block address values ​​to a first physical address of a plurality of sequential physical addresses of the memory device, wherein first data is stored at the first physical address, and wherein the first logical transfer unit value contains the first logical block address value of the plurality of sequential logical block address values; generating a set of command tags instructing a command execution processor of the processing device to retrieve second data from the memory device and store the second data in a set of the buffers, wherein the set of command tags includes at least a second command tag associated with a second physical address that sequentially follows the first physical address; as well as An entry is created in the read cache table for the set of buffers, wherein the entry includes a starting logical block address value set to the first logical block address value and a read offset value corresponding to an amount of the first data and the second data.

2. The memory subsystem of claim 1 , wherein the processing device further performs the following operations: receiving a read request including the first logical block address value; generating the first logical transfer unit value including the first logical block address value; reading a flag from the volatile memory to determine that the first logical transfer unit value is associated with the region; as well as The read command is filled with the first logical transfer unit value and the first buffer address.

3. The memory subsystem of claim 1, wherein the entry further comprises an ending logical block address value corresponding to an end of the read offset value within a physical address space starting from the starting logical block address value.

4. The memory subsystem of claim 3, wherein each of the buffers will store 8 kilobytes to 32 kilobytes of data, and wherein the read offset value is 128 kilobytes to 2 megabytes of data.

5. The memory subsystem of claim 1 , wherein the processing device further transmits to the command execution processor a command group each including one of the set of command tags, and in response to receiving the set of command tags, the command execution processor retrieves the second data from the memory device and stores the second data in the set of buffers according to corresponding buffer addresses of the set of command tags.

6. The memory subsystem of claim 1 , wherein the set of command tags comprises a set of logical transfer unit values ​​corresponding to a subset of the plurality of sequential physical addresses within the read offset value, and wherein each command tag in the set of command tags comprises: a logical transfer unit value in the set of logical transfer unit values ​​that maps to a physical address in the subset of the plurality of sequential physical addresses; as well as a buffer address of a buffer within the set of buffers; and The processing device further generates a buffer index table in the volatile memory to track the logical transfer unit value associated with each command tag in the set of command tags, the logical transfer unit value being indexed relative to the buffer address associated with the logical transfer unit value.

7. The memory subsystem of claim 6, wherein entries of the buffer index table are tracked according to one of: Linked list; A two-three tree algorithm wherein the buffers are ordered according to logical transfer unit values; or N-way cache using hash algorithm.

8. The memory subsystem of claim 6, wherein the processing device further performs the following operations: retrieving a second logical block address value from a second read request received from the host system; Determining the second logical block address value based on the entry in the read cache table: has a single buffer allocation unit offset relative to the starting logical block address value and therefore corresponds to a second logical transfer unit value in the set of logical transfer unit values; and within a range of logical block address values ​​corresponding to the read offset value; indexing within the buffer index table using the second logical transfer unit value to retrieve a second buffer address; as well as A subset of the second data retrieved from a second buffer in the set of the buffers corresponding to the second buffer address is returned to the host system.

9. A method for supporting direct cache hits and transfers, comprising: receiving, by a processing device of a memory subsystem controller, a read request for a first logical block address value of a logical block address space of a memory device, wherein the logical block address value belongs to a region of a plurality of sequential logical block address values ​​mapped to a plurality of sequential physical addresses; creating, by the processing device, a first logical transfer unit value including the first logical block address value, the first logical transfer unit value mapping to a first physical address of the memory device; allocating, by the processing device, a set of buffers in a volatile memory, wherein the capacity of the set of buffers matches the amount of data stored at the first physical address and subsequent physical addresses, the subsequent physical addresses being sequentially numbered within a read offset value starting from the first physical address; generating, by the processing device, a set of command tags instructing a command execution processor of the processing device to retrieve data from the memory device and store the data in the set of buffers, wherein the set of command tags includes a first command tag associated with the first physical address and a second command tag associated with a second physical address that sequentially follows the first physical address; and An entry is created, by the processing device, in a read cache table for the group of buffers, wherein the entry includes a starting logical block address value set to the first logical block address value and the read offset value corresponding to the amount of data.

10. The method of claim 9, further comprising reading bit values ​​of a bitmap from the volatile memory to determine that the first logical transfer unit value is associated with the region of the logical block address space.

11. The method of claim 9, wherein creating the entry further comprises storing an ending logical block address value within the entry, the ending logical block address value corresponding to an end of the read offset value within a physical address space starting from the starting logical block address value.

12. The method of claim 9 , further comprising transmitting to the command execution processor of the processing device a command group each including one of the set of command tags, and in response to receiving the set of command tags, the command execution processor retrieving the data from the memory device and storing the data in the set of buffers according to corresponding buffer addresses of the set of command tags.

13. The method of claim 9, wherein the set of command tags comprises a set of logical transfer unit values ​​corresponding to a subset of the plurality of sequential physical addresses within the read offset value, and wherein generating the set of command tags further comprises: assigning a logical transfer unit value from the set of logical transfer unit values ​​to each command tag from the set of command tags; as well as assigning a buffer address of a buffer within the set of buffers to each command tag in the set of command tags; as well as Wherein the method further includes generating a buffer index table in the volatile memory to track the logical transfer unit value associated with each command tag in the set of command tags that maps to the buffer address associated with the logical transfer unit value.

14. The method of claim 13, further comprising tracking entries of the buffer index table according to one of: Linked list; A two-three tree algorithm wherein the set of buffers are ordered by logical transfer unit value; or N-way cache using hash algorithm.

15. The method according to claim 13, further comprising: retrieving a third logical block address value from a subsequent read request received from the host system; The third logical block address value is determined via accessing the entry in the read cache table: has a buffer allocation unit offset of two relative to the starting logical block address value and therefore corresponds to a third logical transfer unit value in the set of logical transfer unit values; and within a range of logical block address values ​​corresponding to the read offset value; indexing within the buffer index table using the third logical transfer unit value to retrieve a third buffer address; as well as A subset of the data retrieved from a third buffer of the set of buffers corresponding to the third buffer address is returned to the host system.

16. A non-transitory computer-readable medium storing instructions that, when executed by a processing device of a memory subsystem controller, cause the processing device to perform a plurality of operations, comprising: receiving a read request for a first logical block address value of a logical block address space of a memory device, wherein the logical block address value belongs to a region of a plurality of sequential logical block address values ​​mapped to a plurality of sequential physical addresses; creating a first logical transfer unit logical transfer unit value including the first logical block address, the first logical transfer unit value mapping to a first physical address of the memory device; allocating a set of buffers in a volatile memory, wherein the set of buffers has a capacity matching an amount of data stored at the first physical address and subsequent physical addresses, the subsequent physical addresses being sequentially numbered within a read offset value starting at the first physical address; generating a set of command tags instructing a command execution processor of the processing device to retrieve data from the memory device and store the data in the set of buffers, wherein the set of command tags includes a first command tag associated with the first physical address and a second command tag associated with a second physical address that sequentially follows the first physical address; and An entry is created in a read cache table for the set of buffers, wherein the entry includes a starting logical block address value set to the first logical block address value and the read offset value corresponding to the amount of data.

17. The non-transitory computer-readable medium of claim 16, wherein the plurality of operations further comprises: A command group including one of the set of command tags is transmitted to the command execution processor of the processing device, and in response to receiving the set of command tags, the command execution processor retrieves the data from the memory device and stores the data in the set of buffers according to the corresponding buffer addresses of the set of command tags.

18. The non-transitory computer-readable medium of claim 16, wherein the set of command tags comprises a set of logical transfer unit values ​​corresponding to a subset of the plurality of sequential physical addresses within the read offset value, and wherein generating the set of command tags further comprises: assigning a logical transfer unit value from the set of logical transfer unit values ​​to each command tag from the set of command tags; as well as assigning a buffer address of a buffer within the set of buffers to each command tag in the set of command tags; as well as Wherein the plurality of operations further comprises generating a buffer index table in the volatile memory to track the logical transfer unit value associated with each command tag in the set of command tags that maps to the buffer address associated with the logical transfer unit value.

19. The non-transitory computer-readable medium of claim 18, wherein the plurality of operations further comprises tracking entries of the buffer index table according to one of: Linked list; A two-three tree algorithm wherein the set of buffers are ordered by logical transfer unit value; or N-way cache using hash algorithm.

20. The non-transitory computer-readable medium of claim 18, wherein the plurality of operations further comprises: retrieving a second logical block address value from a subsequent read request received from the host system; The second logical block address value is determined via accessing the entry in the read cache table: has a single buffer allocation unit offset relative to the starting logical block address value and therefore corresponds to a second logical transfer unit value in the set of logical transfer unit values; and within a range of logical block address values ​​corresponding to the read offset value; indexing within the buffer index table using the second logical transfer unit value to retrieve a second buffer address; as well as A subset of the data retrieved from a second buffer in the set of buffers corresponding to the second buffer address is returned to the host system.

Citation Information

Patent Citations

  • Transparent replacement of a system processor

    CN101542433A

  • Data caching in non-volatile memory

    CN102576333A