Managing data dependencies in the transport pipeline of hybrid DIMMs
By building a segmented data cache in the DRAM of hybrid DIMMs and managing data transmission using CAM data structures, the problem of inefficient data dependency management in the memory subsystem is solved, achieving higher performance and efficiency.
Patent Information
- Application Number
- CN202080072207.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-08-26
- Filing Date
- 2020-09-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-09-18
AI Technical Summary
Existing memory subsystems have problems with inefficiency when managing data dependencies in the transfer pipeline in hybrid dual inline memory modules, especially when data transfer delays are long, which can lead to increased dependency chains and delays.
Manage data with different data sizes to improve cache hit rate by building a data cache in DRAM of hybrid DIMMs and splitting it into page cache and sector cache. At the same time, the content addressable memory (CAM) data structure is used to track unprocessed data transmission between the cross-point array memory and DRAM, ensuring the order and dependence of the data transmission.
It improves the performance of hybrid DIMM systems, reduces data transmission delay, enhances the management of data dependency, and ensures the order and efficiency of data transmission.
Smart Images

Figure CN114586018B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to memory subsystems, and more particularly, to managing data dependencies in transfer pipelines in a hybrid dual in-line memory module. Background Art
[0002] The memory subsystem may include one or more memory devices that store data. For example, the memory devices may be non-volatile memory devices and volatile memory devices. In general, the host system may utilize the memory subsystem to store data at the memory devices and retrieve data from the memory devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure will be more fully understood from the detailed description given below and from the accompanying drawings of various embodiments of the present disclosure.
[0004] Figure 1 An example computing system including a memory subsystem according to some embodiments of the present disclosure is described.
[0005] Figure 2 is a flow chart of an example method for managing data dependencies in a delivery pipeline according to some embodiments of the present disclosure.
[0006] Figure 3 is a flow chart of another example method for managing data dependencies in a delivery pipeline according to some embodiments of the present disclosure.
[0007] Figure 4 is a flow chart of an example method for performing operations related to sector eviction transfers according to some embodiments of the present disclosure.
[0008] Figure 5 is a block diagram of an example computer system in which embodiments of the present disclosure may operate. DETAILED DESCRIPTION
[0009] Aspects of the present disclosure relate to managing data dependencies in a transmission pipeline of a hybrid dual in-line memory module (DIMM). The memory subsystem may be a storage device, a memory module, or a hybrid of a storage device and a memory module. Figure 1 Examples of storage devices and memory modules are described. In general, a host system may utilize a memory subsystem that includes one or more components, such as a memory device, that stores data. The host system may provide data to be stored at the memory subsystem and may request data to be retrieved from the memory subsystem.
[0010] The memory subsystem may include nonvolatile and volatile memory devices. One example of a nonvolatile memory device is a "NAND" (NAND) memory device. Another example is a three-dimensional cross-point ("3D cross-point") memory device, which is a cross-point array of nonvolatile memory cells. Figure 1 Other examples of non-volatile memory devices are described. A non-volatile memory device is a package of one or more dies. The dies in the package can be assigned to one or more channels for communicating with a memory subsystem controller. Each die can include a set of memory cells ("cells"). A cell is an electronic circuit that stores information. Depending on the cell type, a cell can store one or more binary information bits and have various logical states related to the number of bits stored. The logical state can be represented by a binary value, such as "0" and "1", or a combination of such values. A non-volatile memory device can include a three-dimensional cross-point ("3D cross-point") memory device, which is a cross-point array of non-volatile memory cells and can be combined with a stackable cross-grid data access array to perform bit storage based on changes in body resistance. In addition, in contrast to many flash-based memories, cross-point non-volatile memories can perform write-in-place operations, in which non-volatile memory cells can be programmed without having to erase the non-volatile memory cells in advance. This nonvolatile memory device may group pages across dies and channels to form management units (MUs).
[0011] The memory subsystem may be a hybrid DIMM that includes a first type of memory device (e.g., 3D crosspoint media) and a second type of memory device (e.g., dynamic random access memory (DRAM)) in a single DIMM package. The first type of memory device (e.g., first memory type) may have a large storage capacity but a high access latency, while the second type of memory device (e.g., second memory type) has a smaller amount of volatile memory but a lower access latency. A cache manager may manage the retrieval, storage, and delivery of data to and from the first type of memory device and the second type of memory device. Data transfer between the first type of memory device (e.g., 3D crosspoint) and the second type of memory device (e.g., DRAM) requires more time to process than the processing speed at which the cache manager processes data access commands (e.g., read access commands and write access commands) from a host system.
[0012] The cache manager allows the second type of memory to act as a cache for the first memory type. Therefore, if the cache hit rate is high, the high latency of the first memory type can be masked by the low latency of the second memory type. For example, a DRAM memory device or other volatile memory can be used as a cache memory for a 3D crosspoint memory device or other non-volatile memory device (such as a storage class memory (SCM)). The host system can utilize a hybrid DIMM to retrieve and store data at a 3D crosspoint memory. The hybrid DIMM can be coupled to the host system via a bus interface (e.g., a DIMM connector). The DIMM connector can be a synchronous or asynchronous interface between the hybrid DIMM and the host system. When the host system provides a data access command (e.g., a read access command), the corresponding data can be returned to the host system from the 3D crosspoint memory or from another memory device of the hybrid DIMM that serves as a cache memory for the 3D crosspoint memory.
[0013] The latency of DRAM may also be longer than that of the cache manager. For example, a cache lookup may require several cycles (e.g., 4 cycles) to determine how data should be moved from one device to another. When multiple data transfers are implemented as data pipes, the throughput can be even higher (e.g., per clock cycle, if not limited by the throughput of the components). Therefore, during the time a data transfer (e.g., a data access operation, such as a read operation, a write operation, a delete operation, etc.) is performed, there may be dozens of lookup results available, resulting in the need to perform more data transfers.
[0014] In conventional memory systems, all data transfers may be arranged (e.g., queued) in an order determined by a cache lookup result (e.g., first-in, first-out, hereinafter referred to as "FIFO") to prevent problems associated with data dependencies. Data dependency is where a data transfer or data access request involves data operated by a previous data transfer or data access request. For example, a cache manager may receive a write access command for a physical address followed by a read access command for the same physical address. If the read access command is executed before the write access command, the read access command will read incorrect data because the write access command has not yet been processed. However, arranging data transfers in order may be undesirable because not all data transfers have data dependencies, and most data transfers may be issued and completed out of order. Completing data access commands out of order may reduce the delays experienced by frequent switching between read and write operations, as well as the delays experienced by switching to a different block or die when an unprocessed data access command to the block or die is still queued.
[0015] Aspects of the present disclosure address the above and other deficiencies by implementing a set of schemes to manage data dependencies. In some embodiments, the DRAM of a hybrid DIMM may be constructed as a data cache that stores data that has been recently accessed and / or highly accessed from a non-volatile memory, so that such data can be quickly accessed by a host system. In one embodiment, the DRAM data cache may be partitioned into two different data caches managed with different data sizes. One of the partitions may include a page cache that utilizes a larger granularity (larger size), and the second partition may include a sector cache that utilizes a smaller granularity (smaller size). Because the page cache utilizes a larger data size, less metadata is used to manage the data (e.g., only a single valid bit for the entire page). The smaller data size of the sector cache uses a larger amount of metadata (e.g., a larger number of valid bits and dirty bits and tags, etc.), but may allow for more granular tracking of host access data, thereby increasing the overall cache hit rate in the DRAM data cache. The hit rate may represent a fraction or percentage of memory access requests associated with data that can be found in the DRAM data cache rather than memory access requests associated with data that is only available in the non-volatile memory. Increasing the hit rate in the DRAM data cache may provide comparable performance to DIMMs having only DRAM memory components, but the presence of non-volatile memory on the DIMM may additionally provide greater capacity memory, lower cost, and support for persistent memory.
[0016] In an illustrative example, to provide coherent memory in a hybrid DIMM, memory access operations and interdependent transfers from one memory component to another memory component can be managed by a cache controller of the hybrid DIMM. In particular, long latency transfers can result in dependency chains, and thus dependent operations can be tracked and executed after the operations (e.g., transfers) on which they depend. A set of schemes can be used to manage these data dependencies and provide improved DIMM performance. A set of content addressable memory (CAM) data structures can be used to track all unprocessed data transfers between a crosspoint array memory and a DRAM. The set of CAMs can include at least a sector CAM and a page CAM. Before performing a sector transfer (i.e., a transfer between a sector cache of a DRAM and a crosspoint array memory), the controller can perform a lookup in the page CAM and the sector CAM to determine whether any unprocessed transfers are to be performed on the same physical host address. If there is a hit, the transfer will not start before the unprocessed transfer at the same physical address is completed. If there is no hit for an unprocessed transfer to or from the same physical address, the transfer can be performed immediately because it does not depend on another transfer.
[0017] Advantages of the present disclosure include, but are not limited to, improved performance of a host system utilizing a hybrid DIMM. For example, cache operations between a first memory component and a second memory component may be internal to the hybrid DIMM. Thus, when data is transferred from a cross-point array memory component to be stored at a DRAM data cache, the transfer of data will not utilize an external bus or interface that the host system also uses when receiving and transmitting write operations and read operations. Additionally, the present disclosure may provide coherent memory in a hybrid DIMM despite long latency data transfers between memory components of the hybrid DIMM.
[0018] Figure 1 An example computing system 100 is illustrated that includes a memory subsystem 110 in accordance with some embodiments of the present disclosure. Memory subsystem 110 may include media such as one or more volatile memory devices (e.g., memory device 140), one or more non-volatile memory devices (e.g., memory device 130), or a combination of such memory devices.
[0019] The memory subsystem 110 may be a storage device, a memory module, or a hybrid of a storage device and a memory module. Examples of storage devices include solid-state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controller (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).
[0020] Computing system 100 may be a computing device such as a desktop computer, a laptop computer, a network server, a mobile device, a vehicle (e.g., an airplane, drone, train, automobile, or other transportation vehicle), an Internet of Things (IoT) enabled device, an embedded computer (e.g., a computer included in a vehicle, industrial equipment, or a networked business device), or such a computing device that includes a memory and a processing device.
[0021] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. In some embodiments, the host system 120 is coupled to memory subsystems 110 of different types. Figure 1 An example of a host system 120 coupled to one memory subsystem 110 is illustrated. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without intervening components), whether wired or wireless, including, for example, electrical connections, optical connections, magnetic connections, etc.
[0022] The host system 120 may include a processor chipset and a software stack executed by the processor chipset. The processor chipset may include one or more cores, one or more caches, a memory controller (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110, for example, to write data to the memory subsystem 110 and read data from the memory subsystem 110.
[0023] The host system 120 may be coupled to the memory subsystem 110 via a physical host interface. Examples of the physical host interface include, but are not limited to, a Serial Advanced Technology Attachment (SATA) interface, a Peripheral Component Interconnect Express (PCIe) interface, a Universal Serial Bus (USB) interface, a Fibre Channel, a Serial Attached SCSI (SAS), a Double Data Rate (DDR) memory bus, a Small Computer System Interface (SCSI), a Dual In-line Memory Module (DIMM) interface (e.g., a DIMM slot interface supporting Double Data Rate (DDR)), etc. The physical host interface may be used to transfer data between the host system 120 and the memory subsystem 110. The host system 120 may further utilize an NVM Express (NVMe) interface to access components (e.g., memory device 130) when the memory subsystem 110 is coupled to the host system 120 through a physical host interface (e.g., a PCIe bus). The physical host interface may provide an interface for passing control, address, data, and other signals between the memory subsystem 110 and the host system 120. Figure 1 Memory subsystem 110 is illustrated as an example. In general, host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.
[0024] Memory devices 130, 140 may include any combination of different types of non-volatile memory devices and / or volatile memory devices. Volatile memory devices (e.g., memory device 140) may be, but are not limited to, random access memory (RAM), such as dynamic random access memory (DRAM) and synchronous dynamic random access memory (SDRAM).
[0025] Some examples of non-volatile memory devices (e.g., memory device 130) include "NAND" (NAND) type flash memory and write-in-place memory, such as a three-dimensional cross-point ("3D cross-point") memory device, which is a cross-point array of non-volatile memory cells. The cross-point array of non-volatile memory can perform bit storage based on changes in body resistance in conjunction with a stackable cross-grid data access array. In addition, in contrast to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, where non-volatile memory cells can be programmed without having to erase the non-volatile memory cells in advance. NAND-type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).
[0026] Each of the memory devices 130 may include one or more arrays of memory cells. One type of memory cell, such as a single level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), three-level cells (TLC), four-level cells (QLC), and five-level cells (PLC), may store multiple bits per cell. In some embodiments, each of the memory devices 130 may include one or more arrays of memory cells, such as SLC, MLC, TLC, QLC, PLC, or any combination of such memory cells. In some embodiments, a particular memory device may include an SLC portion and an MLC portion, a TLC portion, a QLC portion, or a PLC portion of memory cells. The memory cells of the memory device 130 may be grouped into pages that may refer to a logical unit of a memory device for storing data. For some types of memory, such as NAND, pages may be grouped to form blocks.
[0027] Although nonvolatile memory components such as a 3D cross-point array of nonvolatile memory cells and NAND-type flash memory (e.g., 2D NAND, 3D NAND) are described, the memory device 130 may be based on any other type of nonvolatile memory, such as read-only memory (ROM), phase-change memory (PCM), self-select memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridging RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), "NOR" (NOR) flash memory, and electrically erasable programmable read-only memory (EEPROM).
[0028] The memory subsystem controller 115 (or controller 115 for simplicity) can communicate with the memory device 130 to perform operations such as reading data, writing data, or erasing data at the memory device 130, as well as other such operations. The memory subsystem controller 115 may include hardware such as one or more integrated circuits and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry with dedicated (i.e., hard-coded) logic to perform the operations described herein. The memory subsystem controller 115 may be a microcontroller, dedicated logic circuitry (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or other suitable processor.
[0029] The memory subsystem controller 115 may be a processing device that includes one or more processors (e.g., processor 117) configured to execute instructions stored in a local memory 119. In the illustrated example, the local memory 119 of the memory subsystem controller 115 includes an embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.
[0030] In the illustrated example, the local memory 119 of the memory subsystem controller 115 includes embedded memory configured to store instructions for executing various processes, operations, logic flows, and routines that control the operation of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.
[0031] In some embodiments, local memory 119 may include memory registers that store memory pointers, fetched data, etc. Local memory 119 may also include read-only memory (ROM) for storing microcode. Figure 1 The example memory subsystem 110 in FIG. 1 has been described as including a memory subsystem controller 115, but in another embodiment of the present disclosure, the memory subsystem 110 does not include a memory subsystem controller 115, but may rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).
[0032] In general, the memory subsystem controller 115 may receive commands or operations from the host system 120 and may convert the commands or operations into instructions or appropriate commands to achieve the desired access to the memory device 130. The memory subsystem controller 115 may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block addresses (LBA), namespaces) and physical addresses (e.g., physical MU addresses, physical block addresses) associated with the memory device 130. The memory subsystem controller 115 may further include a host interface circuit system to communicate with the host system 120 via a physical host interface. The host interface circuit system may convert commands received from the host system into command instructions to access the memory device 130 and convert responses associated with the memory device 130 into information for the host system 120.
[0033] The memory subsystem 110 may also include additional circuitry or components not illustrated. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the memory subsystem controller 115 and decode the addresses to access the memory device 130.
[0034] In some embodiments, the memory device 130 includes a local media controller 135 that operates in conjunction with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 130. An external controller (e.g., the memory subsystem controller 115) can manage the memory device 130 externally (e.g., perform media management operations on the memory device 130). In some embodiments, the memory subsystem 110 is a managed memory device that includes a raw memory device 130 with control logic (e.g., the local controller 135) on the die and a controller (e.g., the memory subsystem controller 115) for media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.
[0035] In one embodiment, the memory subsystem 110 includes a cache manager 113, which can be used to track and manage data in the memory devices 130 and 140. In some embodiments, the memory subsystem controller 115 includes at least a portion of the cache manager 113. In some embodiments, the cache manager 113 is part of the host system 120, the application, or the operating system. In other embodiments, the local media controller 135 includes at least a portion of the cache manager 113 and is configured to perform the functionality described herein. The cache manager 113 can communicate directly with the memory devices 130 and 140 via a synchronous interface. In addition, data transfers between the memory devices 130 and 140 can be completed within the memory subsystem 110 without accessing the host system 120.
[0036] Memory device 140 may include a data cache that stores data from memory device 130 so that future requests for data can be serviced faster. A cache line is the basic unit of cache storage and may contain multiple bytes and / or data words. Smaller cache line sizes have higher hit rates but require more tag memory than large cache size lines. A tag is a unique identifier for a group of data that can be used to distinguish different areas of mapped memory.
[0037] In some embodiments, all data stored by the memory subsystem 110 may be stored at the memory device 130. Certain data stored at the memory device 130 may also be stored at the data cache of the memory device 140. For example, data that is determined to be more frequently or more recently accessed by the host system 120 may be stored at the data cache to enable faster host access. When the host system 120 provides a read request for data stored at the data cache (i.e., a cache hit), the data may be retrieved from the data cache instead of retrieving the data from the memory device 130. The bandwidth or ability to retrieve data at the data cache may be faster than the bandwidth or ability to retrieve data at the memory device 130.
[0038] The data cache of the memory device 140 may be segmented and include a sector cache 142 for storing small cache lines (hereinafter referred to as "sectors") and a page cache 144 for storing large cache lines (hereinafter referred to as "pages"). The sector cache 142 and the page cache 144 may be managed at different data sizes. The sector cache 142 may utilize a smaller granularity (smaller size), and the page cache 144 may utilize a larger granularity (larger size). In an example, the size of a page may be 2 kilobytes, and the size of a sector may be 64 bytes. A page may include one or more sectors. The page cache 144 may utilize a larger data size requiring less metadata to manage the data (e.g., having only a single valid bit for the entire page). The smaller data size of the sector cache 142 may require a larger amount of metadata (e.g., a larger number of valid bits and / or dirty bits, tags, etc.). The pages in the page cache 144 may be organized into one or more groups. In an example, a page group includes 24 pages. Similarly, sectors in sector cache 142 may be organized into one or more groups. In an example, a sector group includes 16 sectors.
[0039] Memory device 130 may store and manage data at a granularity similar to a sector cache. For example, data may be stored at memory device 130 at a data payload size, which may include one or more sectors in sector cache 142. Thus, data may be transferred between memory device 130 and sector cache 142 at a data payload size (e.g., one or more sectors at a time).
[0040] The cache manager 113 may manage a set of content addressable memory (CAM) data structures that track all outstanding data transfers between the memory device 130 and the memory device 140. The cache manager 113 may include a sector CAM 152 and a page CAM 154. The cache manager 113 may use the sector CAM 152 and the page CAM 154 to track all outstanding data transfers (e.g., read commands, write commands, additional transfers, etc.). For example, before the cache manager 113 performs a data transfer, the cache manager 113 may look up the sector CAM 152 to check if there are any outstanding data transfers to and / or from the same physical address. If the lookup results in a hit, the cache manager 113 does not perform the data transfer until the hit data transfer is completed. If the lookup is a miss, the cache manager 113 may perform the data transfer.
[0041] Each data transfer may be associated with an identifier. For example, each host read command may include an associated read identifier (RID), each host write command may include an associated write identifier (WID), and each additional transfer required to maintain a cache (e.g., sector cache 142, page cache 144, etc.) may have a transfer identifier (XID). The cache manager 113 may include a set of transfer types that are tracked for data dependencies. Examples of transfer types may include read buffer transfers (memory device 140 to read buffer), write buffer transfers (write buffer to memory device 140), page eviction transfers (memory device 140 to memory device 130 pages), sector eviction transfers (memory device 140 to memory device 140 sectors), page fill transfers (memory device 130 to memory device 140), and read miss transfers (memory device 130 to memory device 140 and read buffer). Each of the listed transfer types is a sector transfer, except for page eviction transfers, and sector transfers may be tracked by the sector CAM 152. Page eviction transfers, which are page transfers, may be tracked by page CAM 154. At the start of a transfer, the address indicated by the RID, WID, or XID may be stored in its corresponding CAM (e.g., sector CAM 152 and / or page CAM 154) until the transfer is complete. Thus, any unprocessed transfers are included in the CAM via their associated IDs. Additionally, if more than one sector of a page is included in sector CAM 152 at one time (i.e., unprocessed), the page may also be added to page CAM 154.
[0042] When a new transfer begins, cache manager 113 may search both sector CAM 152 and page CAM 154 to determine if there is already a transfer corresponding to the same sector address or page address in sector CAM 152 or page CAM 154. If the new transfer is a sector transfer, and there is a hit in sector CAM 152, the new transfer may be placed in sector CAM 152 to be executed after the earlier transfer is completed. In this way, multiple sector transfers may be linked together in sector CAM 152 to be executed in sequence. If there is a hit in page CAM 154, the new transfer of the sector may be placed in sector CAM 152, but only executed after the earlier page transfer is completed. Similarly, if the page address of a group of sectors is contained in page CAM 154, the new transfer may wait until the transfer of the group of sectors is completed before execution.
[0043] Sector CAM 152 may have a limited size, which may depend on the maximum number of unprocessed data transfers supported. By way of example, the maximum number of read access commands supported by sector CAM 152 may be 256, and the maximum number of write access commands supported by sector CAM 152 may be 64. If all transfers can be directly bound to a read or write access command, the maximum number of unprocessed transfers may be higher (e.g., 320). However, any number of read access commands and write access commands may be supported by sector CAM 152. Due to moving data from sector cache 142 to page cache 144, and due to multiple dirty sectors (e.g., sectors with dirty bits) of a sector group being evicted, some data transfers may be individually identified and may not be directly bound to a read access command or a write access command. The dirty bit may be used to indicate whether a sector has data that is inconsistent with a non-volatile memory (e.g., memory device 130). For additional data transfers, another set of data transfer IDs (XIDs) may be used. In an example, 192 XIDs may be used. Thus, using the above example, additional sets of data transfer IDs may bring the total number of data transfer IDs to 512. Thus, the depth of sector CAM 152 may be 512.
[0044] To track unprocessed sector transfers, the cache manager 113 may use the sector CAM 152 to record the addresses of unprocessed sector transfers. When a sector transfer is issued, the cache manager 113 may perform a sector CAM 152 lookup. In response to a miss in the sector CAM 152 lookup, the sector address of the issued sector transfer may be recorded in the sector CAM 152 at a location indexed by the ID (e.g., RID, WID, and XID) of the sector transfer. In response to a hit in the sector CAM 152 lookup, the sector address of the issued sector transfer may be recorded in the sector CAM 152 at a location indexed by the ID of the sector transfer, and the hit entry may be invalidated. When the sector transfer is completed, the cache manager 113 may remove the sector transfer from the sector CAM 152 by invalidating the corresponding entry of the sector transfer. The sector CAM 152 may be used twice for each sector transfer (e.g., once for the lookup and once for the write). If a valid bit is part of the sector CAM entry, each sector transfer may also be invalidated. The valid bit may be used to indicate whether a sector is valid (e.g., whether the cache line associated with the sector is allocated for transfers (invalid) or receiving transfers (valid)). The sector CAM 152 may utilize an external valid bit, and the cache manager 113 may perform sector CAM invalidation in parallel with sector CAM lookups and writes. In some implementations, to improve sector CAM 152 lookup performance, two or more sector CAMs may be used. For example, a first sector CAM may be used for sector transfers with even sector addresses, and a second sector CAM may be used for sector transfers with odd sector addresses.
[0045] In addition to sector transfers, cache manager 113 may perform page transfers to move pages from memory device 140 (eg, DRAM) to memory device 130 (eg, storage memory). To track outstanding page transfers, cache manager 113 may use page CAM 154 to record addresses of outstanding page transfers.
[0046] In response to the cache manager 113 receiving a data access command, the cache manager 113 may perform a cache lookup in the sector cache 142 and / or the page cache 144. If the cache lookup results in a hit (e.g., a cache lookup hit), there may be two types of transfers depending on whether the data access command is a read access command or a write access command. For a read access command, the cache manager 113 performs a read buffer transfer, where data is moved from a memory device 140 cache (e.g., sector cache 142, page cache 144, etc.) to a read buffer, which may then be read by the host system 120. For a write access command, the cache manager 113 performs a write buffer transfer, where data is moved from a write buffer to a memory device 140 cache (e.g., sector cache 142 or page cache 144), which has been written with write data by the host system 120.
[0047] The cache manager 113 may use a heuristic algorithm to determine that the group of sectors should be moved to the page cache 144. In this case, an additional transfer may occur, which is defined as a "page fill". For a page fill, the cache manager 113 may select a page based on a sector row and / or page row relationship. The cache manager 113 may then evict the selected page. The miss sectors of the original group of sectors are read from the memory device 130 to fill the newly formed page. The cache manager 113 may then perform the additional transfer. In one example, the cache manager 113 may perform a page evict transfer or a flush transfer, where if the page cache line is dirty (e.g., has a dirty bit), then the entire page cache line may be moved to the memory device 130. In another example, the cache manager 113 may perform a page fill transfer, where the miss sectors are read from the memory device 130 to the page cache line.
[0048] When there is a cache lookup miss, the cache manager 113 may evict the sector-based cache line for the incoming data access command. For example, the eviction may be based on LRU data. The LRU data may include an LRU value that may be used to indicate whether the sector cache line was least recently accessed by the host system 120. For example, when a sector is accessed, the LRU value for the sector may be set to a predetermined value (e.g., 24). The LRU value for every other sector in the sector cache 142 may be decreased by a certain amount (e.g., 1). If the evicted mode is sector-based, all dirty sectors of the evicted sector cache line may be moved to the memory device 130, or if the evicted mode is page-based, may be moved to the page cache 144. The page eviction mode includes selecting a page cache line to evict in order to make room for the evicted sector, and reading the sector that is currently missed in the evicted sector group from the memory device 130. Moving data from sector cache 142 to page cache 144 may be performed by swapping the physical address of the new sector with the physical address of the evicted sector so no actual data transfer occurs.
[0049] The sector eviction mode may include the following data transfers: sector eviction transfer, write buffer transfer, and read miss transfer. For a sector eviction transfer, the cache manager 113 may move a dirty sector to the memory device 130. For a write buffer transfer (in response to a write miss), the data transfer does not begin until an eviction transfer for the same sector is complete. For a read miss transfer, the missed sector is read from the memory device 130 and transferred to the memory device 140 and the read buffer. This data transfer does not begin until an eviction transfer for the same sector is complete. The page eviction mode may include the following transfers: a page fill transfer, a page eviction transfer or clear transfer, a write buffer transfer, and a read miss transfer. For a write buffer transfer and a read miss transfer, the data transfer does not begin until a page eviction transfer is complete.
[0050] The cache manager 113 may use a set of schemes to manage data dependencies and produce high performance. Data dependencies may include read / write transfer dependencies, sector eviction transfer dependencies, page eviction transfer dependencies, and page fill transfer dependencies.
[0051] During a read and write transfer dependency, a read transfer may occur to read data from a sector being written by an unprocessed write transfer command. This may be indicated by a sector CAM lookup hit for the read transfer. The write transfer and the read transfer may go to different data paths, which results in the write transfer being completed after the read transfer, even if the write transfer occurs before the read transfer. The cache manager 113 may hold the read transfer until the unprocessed write transfer is complete by placing the read transfer in the sector block transfer data structure at a location indexed by the sector CAM lookup hit ID. The cache manager 113 may use the same or similar mechanism for write transfers to sectors being read by unprocessed read transfers.
[0052] When a sector transfer is complete, its corresponding sector block transfer entry may be checked in the sector block transfer data structure. If the ID stored in the sector block transfer data structure is different from the ID of the current transfer, the cache manager 113 may unblock the stored transfer. When a sector block transfer has a valid transfer, it may be indicated by a flag that may be checked to see if another transfer should be unblocked.
[0053] Regarding sector eviction transfer dependencies, a sector group may consist of 32 sectors that belong to the same page and share the same tag in the sector-based tag memory. When the cache manager 113 evicts a sector, the cache manager 113 may evict the entire sector group because the tag may be replaced. The cache manager 113 may send all dirty sectors to the memory device 130 and each sector may be moved as a distinct and separate sector eviction transfer. Before a sector group is to be evicted, at least one write transfer may be performed to cause at least one sector to be dirty. When a sector group eviction request occurs before multiple write accesses have completed, it may not be executed by the cache manager 113 and may be held, which also means that all individual sector eviction transfers may be held. To hold multiple sector eviction transfers, a page block transfer data structure may be used.
[0054] In some embodiments, multiple unprocessed sector transfers share the same page address. The cache manager 113 may create a page in the page CAM 154 to represent the multiple unprocessed sector transfers. The page may be referred to as a virtual page or "Vpage" (which does not represent an actual page transfer). The cache manager 113 may use a sector count data structure to track the total number of unprocessed sector transfers represented by a Vpage. When the first sector transfer of a Vpage is issued, the page CAM lookup may be a miss, so the page address of the sector transfer may be written to the page CAM 154 at a location indexed by the transfer ID of the sector, and the sector count indexed by the same ID may be initialized to 1. When the second sector transfer of the Vpage is issued, the page CAM lookup may be a hit, and the sector count indexed by the hit ID may be incremented by 1. When an unprocessed sector transfer belonging to a Vpage is completed, the corresponding sector count may be decremented. The cache manager 113 may know that the sector count is decremented because each sector transfer may record its Vpage in the Vpage data structure. Vpage together with sector count can protect multiple sector transfers so that the associated pages will not be evicted before they are all completed.
[0055] When the first sector eviction transfer is issued, the cache manager 113 may perform a page CAM 154 lookup. If the page CAM 154 lookup is a miss, the cache manager 113 may create a Vpage and initialize a sector count to include the total number of sector eviction transfers. In an example, the sector eviction transfer may use the first sector eviction transfer ID as its Vpage ID. Subsequent sector eviction transfers do not require a Vpage lookup. If the lookup is a hit, the cache manager 113 may perform steps similar to those performed when the lookup is a miss, where the cache manager 113 invalidates the addition of the hit Vpage and may block all sector eviction transfers. The cache manager 113 may perform the blocking by placing the first sector eviction transfer ID in a next transfer data structure of page blocks at a location indexed by the hit ID. The cache manager 113 may also place the first sector eviction transfer ID in another data structure (e.g., a list tail data structure of page blocks) at a location indexed by the newly created Vpage. For the second sector eviction transfer, the cache manager 113 may use the first sector eviction transfer ID as its Vpage ID, and may also use the page blocked list tail data structure to find the location of the page blocked list to place its own transfer ID, thereby forming a linked list with the first sector eviction transfer. The page blocked list tail may also be updated with the second sector eviction transfer ID. The cache manager 113 may process subsequent sector eviction transfers like the second transfer, so that all sector eviction transfers form a linked list and may be unblocked one by one once the first sector eviction transfer is unblocked, which may occur when the sector count of its blocked Vpage is counted down to zero.
[0056] The cache manager 113 may block non-eviction sector transfers until all sector eviction transfers are complete when a sector group is evicted and another non-eviction sector transfer arrives and hits the same sector group. The cache manager 113 may create a new Vpage, invalidate the hit entry, and place the incoming transfer ID in the page blocked next transfer data structure at the location indexed by the hit ID. The cache manager 113 may set the Vpage clear flag associated with the Vpage to indicate that the Vpage represents an eviction. Once the new Vpage is created, the second non-eviction sector address may still be blocked and may form a block list with the first non-eviction sector transfer that hit the clear Vpage (i.e., Vpage clear=1) unless it hits the same sector as the first non-eviction sector transfer. To perform this, a data structure may be used to indicate the newly created Vpage as blocked (e.g., as blocked by a page flag data structure). If the second non-evicted sector hits the first non-evicted sector transfer determined by the sector CAM lookup (i.e., hits), then the second non-evicted sector may be placed in the sector blocked next transfer data structure at the location indexed by the hit ID. If the sector CAM lookup of the second non-evicted sector is a hit, but the Vpage of the hit transfer is not the ID of the first non-evicted sector transfer, then the second non-evicted sector may still be blocked, rather than being placed in the sector blocked next transfer data structure at the location indexed by the hit ID.
[0057] Regarding page eviction transfers, to find the page eviction address for the page eviction transfer, the cache manager 113 may perform a second page CAM lookup, which may be performed based on the sector eviction address found by the sector cache lookup of the first lookup. The second page CAM lookup may be avoided if the page CAM is divided into two segments that can be searched simultaneously. The invalid sectors of the evicted sector group may be filled from the storage memory at the same time as the evicted pages are written to the storage memory (e.g., memory device 130). They are independent, but the incoming sector transfer for the page CAM lookup miss may rely on the page eviction transfer or the clear transfer.
[0058] The cache manager 113 can start the clear transfer with a page CAM search. If the search is a hit, the cache manager 113 can insert an entry into the page CAM 154 at the position indexed by the page eviction transfer ID, and the hit entry can be invalidated. The page eviction transfer ID can also be placed in the next transfer data structure blocked by the page at the position indexed by the hit entry. The sector count indexed by the page eviction transfer can be set to 1. The corresponding Vpage will also be set to clear, to indicate that the inserted entry is an entry for page eviction, so that any transfer that hits this entry can be blocked. If the search is a miss, the operation can be similar to that when it is a hit, except that there may not be a hit entry to invalidate it. The entry used for clearing the transfer in the page CAM 154 can represent an actual page transfer instead of a virtual page. If the corresponding sector CAM search is a miss, the first sector transfer that hits the page to be cleared can be placed in the next transfer data structure blocked by the page. The transfer may also be placed in the tail of the list data structure of the page block at the location indexed by the hit ID so that if the corresponding sector CAM lookup is a miss, then the second sector transfer that hits the cleared page can be linked together. The sector count indexed by the hit ID may also be incremented. If the corresponding sector CAM lookup is a hit, then the incoming transfer may be linked to the hit transfer by placing the incoming transfer ID in the next transfer data structure of the sector block at the location indexed by the hit transfer.
[0059] Regarding page fill dependency propagation, cache manager 113 may handle page fills similar to page evictions with minor differences. In an example, a page eviction may be triggered by a cache lookup miss, while a page fill is triggered by a cache lookup hit. In another example, a page fill may select a page by using the sector row and the least significant bit of the sector-based tag rather than performing another page cache lookup (LRU (least recently used) may still be checked). In another example, a sector transfer may be performed before a page eviction transfer. In yet another example, only the first of multiple page fill transfers performs a page CAM lookup, and the sector count may be updated with the total number of transfers. Subsequent sector transfers do not require a page CAM lookup like a sector eviction transfer. If the page CAM lookup is a miss, the cache manager 113 may create a Vpage. If the page CAM lookup hits a purge Vpage, the Vpage may also be created, but may be blocked by the purge Vpage. If the page CAM lookup hits a non-purge Vpage, the Vpage will not be created, but the associated sector transfer may be performed, linked, or blocked, depending on whether the hitting Vpage is blocked and whether there is a sector CAM lookup hit. The cleaning operation may be similar to the sector eviction and page eviction transfers.
[0060] For a host command from the cache manager 113, the sector CAM 152 may be looked up in the first cycle, written to in the second cycle, and invalidated (hit entry) in the fourth cycle. Since the CAM valid bit may be external to the CAM itself, the sector CAM 152 may be used twice for each host command. Since the peak rate of the sector CAM lookup may be 2 clock cycles, host commands may also be accepted every 2 clock cycles. Page fill and sector evict operations may issue commands every cycle, but the commands may be used for both even and odd sector addresses. Although the first command may have completed the sector CAM update when the second command arrives, the sector CAM pipeline may still be in pipeline hazard because the hit entry has not yet been invalidated.
[0061] For a host command from cache manager 113, Page CAM 154 may be looked up in the first cycle, written to and invalidated (hit entry if necessary) in the fourth cycle. When a second command with the same page address as the first command arrives in the third cycle, Page CAM 154 has not been written or updated by the first command, so it may see potentially stale CAM contents, and its lookup should not be used.
[0062] In an embodiment, the cache manager 113 may use a pipeline feed-forward mechanism. In an example, the first command may obtain a lookup result on whether to update the page CAM 154 in the fourth cycle. The action to be taken may also occur in the fourth cycle. The second command may use the lookup result of the first command, and may take the correct action knowing that it is a command after the first command. Subsequent commands with the same page address as the first command but with a different address from the second command may not require feed-forward because the page CAM content has been updated when the command arrives. However, if the third command has the same address as the second command, feed-forward is still required. To detect when feed-forward is required, the address of the first command may be delayed by 2 clock cycles to be compared with the second command. If it is a match, feed-forward may be used.
[0063] In some embodiments, the memory devices 130, 140 may be organized as an error correction code (ECC) protected codeword (CW) consisting of multiple sectors (e.g., two sectors in a 160-byte CW). Multiple codewords (e.g., a CW for a page) may be protected by parity in a redundant array of independent disks (RAID) manner. To write a sector, the corresponding codeword and parity of the sector may be read by the controller 115. The controller 115 may decode the old CW to extract the old sector. The combined new sector and another old sector (the sector that will not be overwritten) may be encoded into a new CW, which may be used together with the old CW and the old parity to generate new parity. To write two sectors of a CW, the operations performed by the controller 115 may be similar to writing a single sector. This is because the old CW may be read to generate new parity, and the old sector does not need to be extracted. Therefore, it is advantageous to group all writes to the same page together, because the parity only needs to be read and written once, rather than multiple times, once for each sector. There may be two types of transfers involving writing to storage memory (sector evict transfers and page evict transfers), both involving writing sectors belonging to the same page. The controller 115 may be informed that the individual sectors of two evict transfers belong to the same page and may therefore be merged. In an example, a signal may be used to indicate that a sector write transfer is the last sector write belonging to the same page and should trigger an actual storage memory write for all sectors up to the last sector. Write merging may better improve performance when the controller 115 has sufficient buffering to accept pending sector writes up to the last sector while simultaneously performing the actual write of the previously received sector writes.
[0064] Figure 2 is a flow chart of an example method 200 for managing data dependencies in a delivery pipeline according to some embodiments of the present disclosure. The method 200 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 200 is performed by Figure 1 13. The cache manager 113 of the embodiment of the present invention may be executed. Although shown in a particular sequence or order, the order of the processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. In addition, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.
[0065] At operation 210, processing logic receives a data access operation from the host system 120. For example, the memory subsystem controller 115 may receive the data access operation. The memory subsystem controller 115 may be operably coupled to a first memory device and a second memory device. The first memory device may be a memory device 130 (e.g., a cross point array), and the second memory device may be a memory device 140 (e.g., a DRAM). The second memory device may have a lower access latency than the first memory device and may act as a cache for the first memory device. In an example, the second memory device may include a first cache component and a second cache component. The first cache component (e.g., a page cache 144) may utilize a larger granularity than the second cache component (e.g., a sector cache 142). The data access operation may be a read operation or a write operation.
[0066] At operation 220, processing logic determines whether there are any unprocessed data transfers corresponding to the data access operation in the second memory component. For example, processing logic may determine whether the first cache component and / or the second cache component includes any uncompleted or pending transfers associated with the address of the data access operation. The unprocessed data transfers and their associated physical addresses may be stored in at least one of the sector CAM 152, the page CAM 154, or a combination thereof.
[0067] At operation 230, in response to determining that there is no outstanding data transfer associated with the address of the data access operation (e.g., a cache miss), the processing logic may perform the data access operation. At operation 240, in response to determining that there is an outstanding data transfer associated with the address of the data access operation (e.g., a cache hit), the processing logic determines whether an outstanding data transfer from the first memory component to the second memory component is being performed or is scheduled to be performed.
[0068] At operation 250, while an unprocessed data transfer is being performed, processing logic determines a plan to delay the execution of a data access operation corresponding to a cache hit. In an example, processing logic may delay the data access operation until the unprocessed data transfer is completed by placing the data transfer into a sector block transfer data structure at a location indexed by a CAM (e.g., sector CAM 152, page CAM 154) lookup hit ID. When a sector transfer is complete, its corresponding sector block transfer entry may be checked in the sector block transfer data structure. If the ID stored in the sector block transfer data structure is different from the ID of the current transfer, the cache manager 113 may unblock the stored transfer. When a sector block transfer has a valid transfer, it may be indicated by a flag that may be checked to see if another transfer should be unblocked.
[0069] At operation 260, in response to the execution of the unprocessed data transfer, the processing logic may plan the execution of the data access operation. In some embodiments, before performing the operation of copying the data from the first memory component to the second memory component, the processing logic may evict the old segment with the first granularity from the second memory device. For example, the eviction may be based on LRU data.
[0070] Figure 3 300 is a flow chart of an example method 300 for managing data dependencies in a delivery pipeline according to some embodiments of the present disclosure. The method 300 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, dedicated logic, programmable logic, microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 300 is performed by Figure 1 13. The cache manager 113 of the embodiment of the present invention may be executed. Although shown in a particular sequence or order, the order of the processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. In addition, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.
[0071] At operation 310, processing logic maintains a set of host data at memory device 130. At operation 320, processing logic maintains a subset of the host data at memory device 140. Memory device 140 may have lower access latency than memory device 130 and may be used as a cache for memory device 130. Memory device 140 maintains metadata for a first segment of the subset of the host data, the first segment having a first size. In an example, the first segment includes a sector.
[0072] At operation 330, processing logic receives a data operation. The data operation may be a read access operation or a write access operation. At operation 340, processing logic determines that a data structure includes an indication of an unprocessed data transfer associated with a physical address of the data access operation. In an example, the data structure may be at least one of the sector CAM 152, the page CAM 154, or a combination thereof. In an example, processing logic may perform a lookup of at least one of the sector CAM 152 or the page CAM 154 to determine whether the physical address accessed by the data operation has an unprocessed data transfer. For example, the data structure may have a number of entries, each corresponding to an unprocessed data transfer, and each having an associated physical address. In one embodiment, the cache manager 113 may compare the physical address of the data operation with the physical address associated with each entry in the data structure. When the physical address of the data operation matches the physical address associated with at least one of the entries in the data structure, the cache manager 113 may determine that the physical address of the data operation has an unprocessed data transfer.
[0073] At operation 350, in response to determining that an operation to copy at least one first segment associated with a physical address from a first memory component to a second memory component is scheduled to be performed, processing logic may delay the scheduling of the execution of the data access operation until the operation to copy at least one first segment is performed. In an example, processing logic may delay the data access operation until the outstanding data transfer is completed by placing the data transfer into a sector block transfer data structure at a location indexed by a CAM (e.g., sector CAM 152, page CAM 154) lookup hit ID. When the sector transfer is complete, its corresponding sector block transfer entry may be checked in the sector block transfer data structure. If the ID stored in the sector block transfer data structure is different from the ID of the current transfer, the cache manager 113 may unblock the stored transfer. When a sector block transfer has a valid transfer, it may be indicated by a flag, which may be checked to see if another transfer should be unblocked.
[0074] At operation 360, in response to the execution of the unprocessed data transfer, the processing logic may plan the execution of the data access operation. In some embodiments, before performing the operation of copying the data from the first memory component to the second memory component, the processing logic may evict the old segment having the first granularity from the second memory device. For example, the eviction may be based on LRU data.
[0075] Figure 44 is a flow chart of an example method 400 of performing operations related to sector eviction transfers according to some embodiments of the present disclosure. The method 400 may be performed by processing logic, which may include hardware (e.g., a processing device, a circuit system, a dedicated logic, a programmable logic, a microcode, hardware of a device, an integrated circuit, etc.), software (e.g., instructions running or executed on a processing device), or a combination thereof. In some embodiments, the method 400 is performed by Figure 1 13. The cache manager 113 of the embodiment of the present invention may be executed. Although shown in a particular sequence or order, the order of the processes may be modified unless otherwise specified. Therefore, the illustrated embodiments should be understood as examples only, and the illustrated processes may be performed in a different order, and some processes may be performed in parallel. In addition, one or more processes may be omitted in various embodiments. Therefore, not all processes are required in every embodiment. Other process flows are possible.
[0076] At operation 410, processing logic may issue a sector eviction transfer associated with a sector. At operation 420, processing logic may perform a lookup for the sector in the page CAM 154. In response to a page CAM 154 lookup miss, at operation 430, processing logic may generate a Vpage and initialize a sector count to determine the total number of sector eviction transfers. A Vpage may represent multiple unprocessed sector transfers, and a sector count data structure may be used to track the total number of unprocessed sector transfers represented by a Vpage. In an example, a sector eviction transfer may use a first sector eviction transfer ID as its Vpage ID.
[0077] In response to the page CAM 154 finding a hit, at operation 440, processing logic may invalidate the hit Vpage, block all sector eviction transfers, generate a new Vpage, and initialize a sector count to determine the total number of sector eviction transfers. In an example, processing logic may place the first sector eviction transfer ID in a page blocked next transfer data structure at a location indexed by the hit ID, and in another data structure (e.g., a page blocked list tail data structure) at a location indexed by the new Vpage. For the second sector eviction transfer, processing logic may use the first sector eviction transfer ID as its Vpage ID, and may also use the page blocked list tail data structure to find the location of the page blocked list to place its own transfer ID, thereby forming a linked list with the first sector eviction transfer. The page block list tail may also be updated with the second sector eviction transfer ID. Processing logic may handle subsequent sector eviction transfers like the second transfer such that all sector eviction transfers form a linked list and may be unblocked one by one once the first sector eviction transfer is unblocked, which may occur when the sector count of its blocking Vpage is counted down to zero.
[0078] Figure 5An example machine illustrating a computer system 500 within which a set of instructions for causing the machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system 500 may correspond to a host system (e.g., Figure 1 ) that includes or utilizes a memory subsystem (e.g., Figure 1 110), or may be used to execute operations of the controller (eg, execute an operating system to execute operations corresponding to Figure 1 In some embodiments, the machine may be connected (e.g., using a network) to other machines in a LAN, an intranet, an extranet, and / or the Internet. The computer may operate in the capacity of a server or a client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment.
[0079] The machine may be a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a cellular phone, a network appliance, a server, a network router, a switch or a bridge, or any machine capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken by the machine. Further, while a single machine is described, the term "machine" shall also be taken to include any collection of machines that individually or collectively execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0080] The example computer system 500 includes a processing device 502, a main memory 504 (e.g., a read-only memory (ROM), a flash memory, a dynamic random access memory (DRAM), such as a synchronous DRAM (SDRAM) or a Rambus DRAM (RDRAM), etc.), a static memory 506 (e.g., a flash memory, a static random access memory (SRAM), etc.), and a data storage device 518, which communicate with each other via a bus 530. The processing device 502 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. The processing device 502 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device 502 is configured to execute instructions 526 for performing the operations and steps discussed herein. The computer system 500 may further include a network interface device 508 for communicating over a network 520.
[0081] The data storage system 518 may include a machine-readable storage medium 524 (also referred to as a computer-readable medium) on which is stored one or more sets of instructions 526 or software embodying any one or more of the methodologies or functions described herein. During execution of the instructions 526 by the computer system 500, the instructions 526 may also reside, in whole or in part, within the main memory 504 and / or within the processing device 502, which also constitute machine-readable storage media. The machine-readable storage medium 524, the data storage system 518, and / or the main memory 504 may correspond to Figure 1 Memory subsystem 110.
[0082] In one embodiment, instructions 526 include instructions to implement the Figure 1 13. Although the machine-readable storage medium 524 is shown as a single medium in the example embodiment, the term "machine-readable storage medium" should be taken to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be taken to include any medium capable of storing or encoding a set of instructions for execution by a machine and causing the machine to perform any one or more of the methods of the present disclosure. Thus, the term "machine-readable storage medium" should be taken to include, but not limited to, solid-state memory, optical media, and magnetic media.
[0083] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is generally considered here to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. Typically, but not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for common sense reasons, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0084] It should be borne in mind, however, that all of these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may refer to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) numbers within the computer system's registers and memories into physical quantities similarly represented as within the computer system's memories or registers or other such information storage systems.
[0085] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. This computer program may be stored in a computer-readable storage medium, such as, but not limited to, any type of disk, including floppy disks, optical disks, CD-ROMs and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0086] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used with programs according to the teachings herein, or it may prove convenient to construct more specialized equipment to perform the methods. The structures of various these systems will appear as set forth in the description below. In addition, the present disclosure is not described with reference to any particular programming language. It will be appreciated that the teachings of the present disclosure as described herein may be implemented using various programming languages.
[0087] The present invention may be provided as a computer program product or software, which may include a machine-readable medium having instructions stored thereon, which can be used to program a computer system (or other electronic device) to perform a process according to the present invention. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). For example, a machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory device, etc.
[0088] In the foregoing specification, embodiments of the present invention have been described with reference to specific example embodiments of the invention. It will be apparent that various modifications may be made thereto without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. Accordingly, the description and drawings are to be regarded in an illustrative rather than a restrictive sense.
Claims
1. A system comprising: a first memory device, wherein the first memory device comprises a non-volatile memory device; a second memory device comprising a volatile memory device coupled to the first memory device, wherein the second memory device has a lower access latency than the first memory device and is a cache for the first memory device; and a processing device operatively coupled to the first memory device and the second memory device to perform operations comprising: receiving a current data access operation referencing a physical address associated with the first memory device; performing a lookup in a first content addressable memory (CAM) and a second CAM to determine whether the first CAM or the second CAM includes an indication of an outstanding data access operation associated with the physical address referenced by the current data access operation, wherein the first CAM tracks outstanding data access operations stored by a first cache in the second memory device at a first granularity, wherein the second CAM tracks outstanding data access operations stored by a second cache in the second memory device at a second granularity, wherein the second granularity is greater than the first granularity; determining that at least one of the first CAM or the second CAM includes the indication of an unprocessed data access operation associated with the physical address referenced by the current data access operation; determining whether the unprocessed data access operation includes an operation for copying data from the physical address of the first memory device to the second memory device; and In response to determining that the outstanding data access operation includes an operation for copying data from the physical address of the first memory device to the second memory device, determining a plan to delay execution of the current data access operation until the outstanding data access operation is executed.
2. The system of claim 1, wherein: The data access operation includes at least one of a read access operation or a write access operation; and The first memory device is a cross point array memory device.
3. The system of claim 1 , wherein the processing device further performs operations comprising: In response to determining that the first CAM and the second CAM do not contain an indication of an outstanding data transfer for data associated with a physical address of the data access operation, execution of the data access operation is scheduled.
4. The system of claim 1, wherein the plan to delay the execution of the data access operation comprises storing an indication of the data access operation in a transfer data structure.
5. The system of claim 4, wherein the processing device further performs operations comprising: responsive to performance of the operation to copy the data from the first memory device to the second memory device, removing the indication of the unprocessed data transfer from the first CAM or the second CAM by invalidating a corresponding entry of the unprocessed data transfer; retrieving the data access operation from the transfer data structure; and Execution of the data access operation is planned.
6. The system of claim 1, wherein the second memory device stores segments of data at a first granularity and stores segments of data at a second granularity, wherein the second granularity is greater than the first granularity.
7. The system of claim 6, wherein the processing device further performs operations comprising: Prior to performing the operation of copying data from the first memory device to the second memory device, old segments having a first granularity are evicted from the second memory device.
8. A method comprising: maintaining a set of host data at a first memory device of a memory subsystem, wherein the first memory device comprises a nonvolatile memory device; maintaining a subset of host data at a second memory device of the memory subsystem, wherein the second memory device comprises a volatile memory device having lower access latency than the first memory device and functions as a cache for the first memory device, and wherein the second memory device maintains metadata for a first segment of the subset of the host data, the first segment having a first size; receiving a current data access operation referencing a physical address associated with the first memory device; performing a lookup in a first content addressable memory (CAM) and a second CAM to determine whether the first CAM or the second CAM includes an indication of an outstanding data access operation associated with the physical address referenced by the current data access operation, wherein the first CAM tracks outstanding data access operations stored by a first cache in the second memory device at a first granularity, wherein the second CAM tracks outstanding data access operations stored by a second cache in the second memory device at a second granularity, wherein the second granularity is greater than the first granularity; determining that at least one of the first CAM or the second CAM includes the indication of an unprocessed data access operation associated with the physical address referenced by the current data access operation; determining whether the unprocessed data access operation includes an operation for copying at least one first segment from the physical address of the first memory device to the second memory device; as well as In response to determining that the unprocessed data access operation includes an operation for copying at least one first sector from the physical address of the first memory device to the second memory device, scheduling of execution of the data access operation is delayed until the unprocessed data access operation is executed.
9. The method according to claim 8, wherein: The data access operation includes at least one of a read access operation or a write access operation; and The first memory device is a cross point array memory device.
10. The method according to claim 8, further comprising: In response to determining that the first CAM and the second CAM do not contain an indication of an outstanding data transfer for data associated with a physical address of the data access operation, execution of the data access operation is scheduled.
11. The method of claim 8, wherein delaying the plan for the execution of the data access operation comprises storing an indication of the data access operation in a transfer data structure.
12. The method according to claim 11, further comprising: responsive to performance of the operation to copy at least one first segment associated with the physical address from the first memory device to the second memory device, removing the indication of the unprocessed data transfer from the first CAM or the second CAM by invalidating a corresponding entry of the unprocessed data transfer; retrieving the data access operation from the transfer data structure; and Execution of the data access operation is planned.
13. The method of claim 8, wherein the second memory device stores segments of data at a first granularity and stores segments of data at a second granularity, wherein the second granularity is greater than the first granularity, and the method further comprises evicting old segments having the first granularity from the second memory device before performing the operation of copying the at least one first segment from the first memory device to the second memory device.
14. A non-transitory computer-readable storage medium comprising instructions that, when executed by a processing device operatively coupled to a first memory device and a second memory device, perform operations comprising: receiving a current data access operation referencing a physical address associated with the first memory device; performing a lookup in a first content addressable memory (CAM) and a second CAM to determine whether the first CAM or the second CAM includes an indication of an outstanding data access operation associated with the physical address referenced by the current data access operation, wherein the first CAM tracks outstanding data access operations stored by a first cache in the second memory device at a first granularity, wherein the second CAM tracks outstanding data access operations stored by a second cache in the second memory device at a second granularity, wherein the second granularity is greater than the first granularity; determining that at least one of the first CAM or the second CAM includes the indication of an unprocessed data access operation associated with the physical address referenced by the current data access operation; determining whether the unprocessed data access operation includes an operation for copying data from the physical address of the first memory device to the second memory device; and In response to determining that the outstanding data access operation includes an operation for copying data from the physical address of the first memory device to the second memory device, determining a plan to delay execution of the current data access operation until the outstanding data access operation is executed.
15. The non-transitory computer-readable storage medium of claim 14, wherein: The data access operation includes at least one of a read access operation or a write access operation; and The first memory device is a cross point array memory device.
16. The non-transitory computer-readable storage medium of claim 14, wherein the processing device further performs operations comprising: In response to determining that the first CAM and the second CAM do not contain an indication of an outstanding data transfer for data associated with a physical address of the data access operation, execution of the data access operation is scheduled.
17. The non-transitory computer-readable storage medium of claim 14, wherein the plan to delay the execution of the data access operation comprises storing an indication of the data access operation in a transfer data structure.
18. The non-transitory computer-readable storage medium of claim 17, wherein the processing device further performs operations comprising: responsive to performance of the operation to copy the data from the first memory device to the second memory device, removing the indication of the unprocessed data transfer from the first CAM or the second CAM by invalidating a corresponding entry of the unprocessed data transfer; retrieving the data access operation from the transfer data structure; and Execution of the data access operation is planned.
19. The non-transitory computer-readable storage medium of claim 14, wherein the second memory device stores segments of data at a first granularity and stores segments of data at a second granularity, wherein the second granularity is greater than the first granularity.
20. The non-transitory computer-readable storage medium of claim 19, wherein the processing device further performs operations comprising: Prior to performing the operation of copying data from the first memory device to the second memory device, old segments having a first granularity are evicted from the second memory device.
Citation Information
Patent Citations
Batching modified blocks to the same dram page
US20160170887A1
Read operation delay
US20170115891A1
Apparatuses and methods for an operating system cache in a solid state device
US20180107595A1