Storage-level memory access
By using different interfaces between the host and storage-level memory (SCM) for data access, the problem of large SCM write latency is solved, enabling faster data access and lower system power consumption.
Patent Information
- Application Number
- CN202080007137.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-13
- Filing Date
- 2020-06-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-06-15
AI Technical Summary
Storage-level memory (SCM) has a long delay when writing data, making it difficult to replace traditional volatile local memory such as DRAM or SRAM.
By exposing the base address register (BAR) between the host and the SCM, it allows the reading of byte addressable data using the memory device interface, and the operation of writing data using the block device interface. This arrangement allows data to be aggregated or modified in the host's buffer and then written to the SCM at a block size, reducing read and write latency in the SCM.
This method reduces read and write delays in SCM, allowing SCM to be more efficiently used to store byte addressable data, thereby replacing part of DRAM memory and reducing system power consumption and cost.
Smart Images

Figure CN113243007B_ABST
Abstract
Description
Background Art
[0001] Storage - class memory (SCM) has recently been developed as a non - volatile storage option that can provide fine - grained data access (i.e., byte - addressable or cache - line sized). In addition, compared to traditional non - volatile storage devices, such as solid - state drives (SSDs) that use flash memory or hard - disk drives (HDDs) that use rotating disks, SCM typically offers shorter data access latency. SCM can include, for example, memories such as magnetoresistive random - access memory (MRAM), phase - change memory (PCM), and resistive RAM (RRAM).
[0002] Although SCM can allow byte - addressable access to data (i.e., in units smaller than the page size or block size), the time to write data to SCM can be much longer than the time to read data from SCM. This slows down SCM from being used as a more economical and high - performance memory alternative to the memory that is conventionally used in host memories, such as dynamic random - access memory (DRAM) or static random - access memory (SRAM). BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The features and advantages of the embodiments of the present disclosure will become more apparent from the following detailed description and in conjunction with the accompanying drawings. The accompanying drawings and the associated description are provided to illustrate the embodiments of the present disclosure, not to limit the scope of the claimed subject matter.
[0004] Figure 1 is a block diagram of an exemplary environment including a host and a device including storage - class memory (SCM) according to one or more embodiments.
[0005] Figure 2 is a block diagram showing the processing of read requests and write requests according to one or more embodiments.
[0006] Figure 3 is an example of a page table according to one or more embodiments.
[0007] Figure 4 is a state diagram of a page - table entry according to one or more embodiments.
[0008] Figure 5 is a flowchart of a page - table creation process according to one or more embodiments.
[0009] Figure 6 is a flowchart of a write - request process according to one or more embodiments.
[0010] Figure 7 is a flowchart of a read - request process for byte - addressable data according to one or more embodiments.
[0011] Figure 8 is a flowchart of a refresh process from a host memory to SCM according to one or more embodiments.
[0012] Figure 9 is a flowchart of a multi-interface process for a device including SCM according to one or more embodiments.
[0013] Figure 10 is a flowchart of a block write process for a device including SCM according to one or more embodiments. Detailed Description
[0014] Numerous specific details are set forth in the following detailed description in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that various embodiments may be practiced without some of these specific details. In other instances, well-known structures and techniques have not been shown in detail to avoid unnecessarily obscuring the various embodiments.
[0015] Exemplary System Environment
[0016] Figure 1 is a block diagram of an exemplary environment including a host 101 and a device 111 according to one or more embodiments. The host 101 communicates with the device 111 to retrieve data from the device 111 and store data in the device 111. As further described below, the device 111 can be used as a memory device and / or as a storage device for the host 101 via a corresponding device interface. The host 101 and the device 111 can be separate devices, or can be housed together as part of a single electronic device, such as a server, computing device, embedded device, desktop computer, laptop computer or notebook computer, or another type of electronic device, such as a tablet computer, smart phone, network media player, portable media player, television, digital camera or digital video recorder (DVR). In other specific implementations, the host 101 can be a client computer or a storage controller, and the device 111 can be a memory / storage server or a memory / storage node in a network (such as a cloud storage network or a data center). As used herein, a host can refer to a device capable of issuing commands to a device to store data or retrieve data. In this regard, the host 101 can include another storage device capable of executing an application program and communicating with other memory / storage devices, such as a smart data storage device.
[0017] As Figure 1As shown, device 111 includes a storage class memory (SCM) 120 that provides non-volatile storage of data that can be accessed at the byte level (i.e., cache line size) that is less than the page size or block size. SCM 120 can include, for example, chalcogenide RAM (C-RAM), phase change memory (PCM), programmable metallization cell RAM (PMC-RAM or PMCm), ovonic unified memory (OUM), resistive RAM (RRAM), ferroelectric memory (FeRAM), magnetoresistive RAM (MRAM), fast NAND, and / or 3D-XPoint memory. Such SCMs provide faster data read and write than conventional non-volatile storage devices such as flash memory or spinning disks. In some embodiments, in addition to SCM 120, device 111 can also include other types of non-volatile storage devices such as spinning disks or flash memory.
[0018] Although SCMs can provide faster data read and write than conventional forms of non-volatile storage devices, the time to write data to an SCM is typically longer than the time to read data. This is particularly evident in the case of using address indirection in an SCM (such as for wear leveling). As described above, the longer write latency of an SCM can prevent the SCM from being used as a replacement for volatile local memory such as more expensive and higher power dynamic random access memory (DRAM) or static random access memory (SRAM). According to one aspect of the present disclosure, device 111 exposes a base address register (BAR) to host 101 such that read commands for byte-addressable data (e.g., for cache lines or less than page size or block size) can be sent using the memory device interface, while write commands are sent from host 101 using the block device interface to write data to larger data blocks. As discussed in more detail below, data to be written to SCM 120 can be aggregated or modified in buffer 107 of memory 106 of host 101 before being flushed to SCM 120. Host 101 can then send a write command to write the aggregated or modified data block in SCM 120. This arrangement reduces the latency in reading and writing data in SCM 120 such that SCM 120 can be used to store byte-addressable data that could otherwise be stored in memory 106.
[0019] In Figure 1In the example, host 101 includes processor circuitry 102 for executing computer-executable instructions such as the operating system (OS) of host 101. Processor circuitry 102 may include circuitry such as one or more processors for executing instructions and may include, for example, a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), hardwired logic, analog circuitry, and / or combinations thereof. In some embodiments, processor circuitry 102 may include a system on a chip (SoC). In this regard, Figure 1 the example shows that the memory management unit (MMU) 104 is part of processor circuitry 102 or is included in the same housing as processor circuitry 102. Those of ordinary skill in the art will understand that in some embodiments, processor circuitry 102, memory 106, and / or device interface 108 may be combined into a single component or housing. Other embodiments may not include an MMU. As used herein, an MMU may be considered processor circuitry.
[0020] In Figure 1 the example, processor circuitry 102 may access memory 106 via MMU 104 to execute instructions such as instructions for executing a block device interface or a memory device interface for reading data from or writing data to device 111. In this regard, and as discussed in more detail below, buffer 107 of memory 106 may store byte-addressable data for write requests, which is aggregated or buffered to reach the block size of SCM 120.
[0021] Memory 106 serves as the main memory of host 101 and may include, for example, volatile RAM (such as DRAM or SRAM), non-volatile RAM, or other solid state memory. Although the description herein generally refers to solid state memory, it should be understood that solid state memory may include one or more of various types of memory devices such as flash integrated circuits, C-RAM, PC-RAM or PRAM, programmable metallization cell RAM (PMC-RAM or PMCm), OUM, RRAM, NAND memory (e.g., single-level cell (SLC) memory, multi-level cell (MLC) memory (e.g., two or more levels), or any combination thereof), NOR memory, EEPROM, FeRAM, MRAM, other discrete non-volatile memory (NVM) chips, or any combination thereof. In some embodiments, memory 106 may be located external to host 101 but serve as the main memory of host 101.
[0022] The processor circuit 102 also uses the MMU 104 to access the SCM 120 of the device 111 via the device interface 108. In some embodiments, the MMU 104 can access a page table that translates the virtual addresses used by the processor circuit 102 into physical addresses (e.g., byte addresses), where the physical address indicates the location in the memory 106 or SCM 120 where the data for that virtual address is to be stored or retrieved from the memory 106 or SCM 120. In this regard, the MMU 104 can keep track of the location of byte-addressable data. Additionally, the MMU 104 can execute a memory device interface for accessing byte-addressable data (e.g., Figure 2 the memory device interface 10 in
[0023] The device interface 108 allows the host 101 to communicate with the device 111 via a bus or interconnect 110. In some embodiments, the device interface 108 can communicate with the host interface 118 of the device 111 via a standard such as Peripheral Component Interconnect Express (PCIe), Ethernet, or Fibre Channel via the bus or interconnect 110. As discussed in more detail below, the bus or interconnect 110 can include a bus or interconnect that can simultaneously allow commands for byte-addressable data to be executed via the memory device interface and commands for block-addressable data to be executed via the block device interface. In other embodiments, the host 101 and the device 111 can communicate via two or more buses or interconnects, each bus or interconnect providing a memory device interface, a block device interface, or both.
[0024] In this regard, the processor circuit 102 uses multiple logical interfaces to read data from and write data to the SCM 120 of the device 111. To write data and read block-addressable data, the host 101 interfaces with the device 111 using a block device interface or a storage device interface (such as, for example, the Non-Volatile Memory Express (NVMe) interface specification), which can be implemented, for example, by an OS driver executed by the processor circuit 102. To read byte-addressable data, the host 101 interfaces with the device 111 using a memory device interface such as a PCIe Base Address Register (BAR) interface, Gen-Z, Open Coherent Accelerator Processor Interface (OpenCAPI), or Accelerator Cache Coherence Interconnect (CCIX) that can be executed by the processor circuit 102. In some embodiments, the memory device interface can be implemented by the MMU 104 or by other circuitry of the processor circuit 102 such as a hardware accelerator.
[0025] As Figure 1As shown, device 111 includes a host interface 118, a control circuit 112, a memory 116, and an SCM 120. The host interface 118 allows device 111 to communicate with the device interface 108 of host 101 via a bus or interconnect 110. In some specific embodiments, the host interface 118 may communicate with the device interface 108 of host 101 using a standard such as PCIe, Ethernet, or Fibre Channel.
[0026] In addition, the control circuit 112 uses multiple logical interfaces to receive and execute read and write commands from host 101 to access data in SCM 120. To read and write block-addressable data, the control circuit 112 interfaces with host 101 using a block device interface, which may include, for example, an NVMe interface. To read byte-addressable data, the control circuit 112 interfaces with host 101 using a memory device interface. The memory device interface may include, for example, a PCIe BAR interface, Gen-Z, OpenCAPI, or CCIX.
[0027] The control circuit 112 may include circuitry such as one or more processors for executing instructions, and may include, for example, a CPU, GPU, microcontroller, DSP, ASIC, FPGA, hardwired logic, analog circuitry, and / or combinations thereof. In some specific embodiments, the control circuit 112 may include a SoC such that one or both of the host interface 118 and the memory 116 may be combined with the control circuit 112 in a single chip. Similar to the processor circuitry 102 of host 101 discussed above, in some specific embodiments, the control circuit 112 of device 111 may include separate components, such as separate hardware accelerators for implementing the memory device interface and the block device interface.
[0028] The memory 116 of device 111 may include, for example, volatile RAM such as DRAM, non-volatile RAM, or other solid-state memory. The control circuit 112 may access the memory 116 to execute instructions, such as the firmware of device 111, which may include instructions for implementing the memory device interface and the block device interface. In addition, the control circuit 112 may access the memory 116 to obtain data used when executing the firmware of device 111, data to be written into SCM 120, and / or data that has been read from SCM 120.
[0029] One of ordinary skill in the art will know that other specific embodiments may include more than Figure 1more or fewer elements than those shown, and the processes discussed herein may be implemented in other environments. For example, other environments may not include the MMU in host 101, may include a separate MMU or hardware accelerator for implementing the memory device interface, or may include a different number of SCMs or different types of non-volatile storage devices other than SCM 120 in device 111.
[0030] Figure 2 is an exemplary block diagram showing the processing of read requests and write requests by host 101 and device 111 according to one or more embodiments. As Figure 2 shown, memory device interface 10 receives read request A and write request B, while block device interface 12 receives read request C and write request D. The read and write requests may come from an application program executed by processor circuit 102. In Figure 2 the example, memory device interface 10 is implemented by MMU 104 of processor circuit 102, and block device interface 12 is implemented by the OS driver of host 101 executed by processor circuit 102.
[0031] Write request B is initially received by memory device interface 10 but redirected by memory device interface 10 to block device interface 12 because memory device interface 10 is only used to process read requests for byte-addressable data, not write requests. In some embodiments, since the memory mapped to SCM 120 is marked as read-only, MMU 104 hands over control of the write request to the OS of host 101. As described above, SCM 120 generally executes read commands faster than write commands. In the present disclosure, SCM 120 can be used as the local memory of host 101 or as a partial DRAM replacement for read requests, while write requests are executed in the memory 106 of host 101. This generally allows a smaller-sized local memory to be used at host 101, which can reduce the power consumption and cost of the entire system including host 101 and device 111.
[0032] As used herein, read and write requests refer to data accesses at the byte level (i.e., byte-addressable data), such as cache line requests made by an application program executed by processor circuit 102 of host 101. On the other hand, read and write commands refer to commands sent from host 101 to device 111 to access data at the byte level in the presence of a read command from memory device interface 10, or to access data at the block level (i.e., page-addressable data or block-addressable data) from block device interface 12. The page size or block size may correspond to the data unit in the virtual memory managed by the OS of host 101. By Figure 2The block device interfaces 12 and 22 in [device] access data in units of the page size or block size. Some examples of the block size or page size may include 512 bytes, 4 KB, 8 KB, 2 MB, or 4 MB. In contrast, the byte-addressable data accessed by Figure 2 the memory device interfaces 10 and 20 in [device] allows data to be read in units of bytes, including reading a single byte of data from the SCM 120 of device 111.
[0033] As Figure 2 shown, a read request A for byte-addressable data A is repackaged by the memory device interface 10 into a read command A, which is sent to the memory device interface 20 of device 111 to retrieve the byte-addressable data A from the SCM 120. The memory device interface 20 at device 111 receives the read command A and uses an optional logical-to-physical mapping module 24 to identify the physical address in the SCM 120 where the byte-addressable data A is stored. In other embodiments, the logical-to-physical mapping module 24 may be omitted, such as in embodiments where the SCM 120 does not use address indirection of memory technology, such as, for example, wear leveling to more evenly distribute writes throughout the SCM 120. In such embodiments, the memory device interface 20, which may be performed by the control circuit 112 in Figure 1 [device], may perform a read operation A in the SCM 120. When performing the read operation A, the byte-addressable data A returns to the memory device interface 20 and may be temporarily stored in a buffer, such as Figure 1 a buffer in the memory 116 of device 111 in [device], before returning to the memory device interface 10 of the host 101 to complete the read command A.
[0034] The memory device interface 20, which is performed by the control circuit 112 of device 111, is configured to only receive and execute read commands for byte-addressable data. The execution of write commands received by the memory device interface 20 may be blocked or trigger an error at device 111. Such an error may or may not be reported back to the host 101.
[0035] In Figure 2In the example, the write request B is received by the memory device interface 10 of the host 101. In some cases, the write request B may include a request to store byte-addressable data, such as data from a cache flush command, to flush one or more cache lines from one or more caches (e.g., from the L1 / L2 / L3 cache) of the processor circuit 102 of the host 101. In other cases, the write request B may include a request to store block-addressable data or page-addressable data, such as data from an application executed by the processor circuit 102. The memory device interface 10 identifies the write request B as a write request and, in response, redirects the write request B to the block device interface 12. In some embodiments, the write request received by the memory device interface 10 may trigger a fault handler that allows the OS of the host 101 to process the write request via the block device interface 12.
[0036] In the case where the write request B will store byte-addressable data, the block device interface 12 uses the buffer 107 to aggregate or modify one or more portions of a data block including the byte-addressable data to be written to the SCM 120. The block device interface 12 sends a read command for the data block that includes the byte-addressable data to be written to the device 111. The block device interface 22 of the device 111 receives the read command for the block and performs a read operation on the SCM 120, and returns the read block including the byte-addressable data to the block device interface 12 of the host 101. The block device interface 12 buffers the read data block in the buffer 107 and modifies one or more byte-addressable portions of the buffered block for the write request B. In some cases, when the block is stored in the buffer 107, additional write requests for the byte-addressable data included in the buffered block may also be performed.
[0037] Then, the block device 12 sends a write command for the modified block including the byte-addressable data to flush the data for the write request from the buffer 107 to the SCM 120. In some cases, the write command may include additional blocks that have been modified or written, such as data for the write request D. The block device interface 22 of the device 111 receives the write command and uses the optional logical-to-physical mapping module 24 to identify one or more physical addresses in the SCM 120 that store one or more blocks including the data B and D for the write requests B and D. As described above, in other embodiments, the logical-to-physical mapping module 24 may be omitted, such as in embodiments where the SCM 120 does not use address indirect addressing. In such embodiments, it may be by Figure 1The block device interface 22 executed by the control circuit 112 in [device] can identify one or more addresses in the SCM 120 for performing write operations B and D. When performing write operations B and D, the block device interface 22 stores one or more data blocks including data B and data D in the SCM 120. In addition, the block device interface 22 can use the locations where data has been written in the SCM 120 to update the byte addresses of the data used by the memory device interface 20 for inclusion in one or more data blocks.
[0038] After the write operation is completed, one or more write completion indications are returned to the block device interface 22. The block device interface 22 can forward or send the write completion indication to the block device interface 12 of the host 101 to indicate that one or more write commands have been completed, and in addition to providing a block-addressable location for the data, can also provide a new byte-addressable physical address for the data stored in the write operation. In other specific embodiments, the memory device interface 20 can alternatively provide the updated byte-addressable physical address to the memory device interface 10 of the host 101.
[0039] The read request C is also received by the block device interface 12 of the host 101. The data to be retrieved for the read request C is addressed in the form of pages or blocks, which is contrary to the request for byte-addressable data, such as the read request A discussed above. The block device interface 12 repackages the request into a read command C and sends the read command C to the block device interface 22 of the device 111. For its part, the block device interface 22 executes the read command C by using the optional logical-to-physical mapping module 24, which provides the physical address for reading the block-addressable data C from the SCM 120. The block-addressable data C is read from the SCM 120 and returned to the block device interface 22, which passes the data to the block device interface 12 of the host 101 to complete the command. In some cases, the data C can be buffered in the memory of the device 111, such as the memory 116, before being sent to the host 101.
[0040] As will be understood by those of ordinary skill in the art, other specific embodiments may include components or modules different from those shown in the Figure 2 example. For example, other specific embodiments may not include the logical-to-physical mapping module 24, such that the memory device interface 20 and the block device interface 22 access the SCM 120 without using the logical-to-physical mapping module.
[0041] Page Representation Example
[0042] Figure 3 is an example of a page table 16 according to one or more embodiments. The page table 16 can be executed, for example, by Figure 2created by the processor circuit 102 of the host 101 of the memory device interface 10 in []. In some specific embodiments, the processor circuit 102 may maintain multiple page tables, such as those used to map virtual addresses for byte-addressable data and also those used to map virtual addresses for block-addressable data. Additionally, those of ordinary skill in the art will understand that the page table 16 may include other information besides what is shown in the example of Figure 3 , such as, for example, statistical information. The page table 16 may be stored, for example, in the memory 106 of the host 101 or in the memory of the MMU 104 or in other memory of the processor circuit 102.
[0043] As Figure 3 shown, virtual addresses are assigned to different data accessed by the processor circuit 102, and the page table 16 indicates the access type of the data and the physical address where the data is stored in the memory 106 of the device 111 or in the SCM 120. For example, the byte-addressable data for virtual addresses A and C in the page table 16 are stored at physical addresses SCM1 and SCM2 in the SCM 120, respectively. The data for virtual addresses A and C are also indicated in the page table 16 as having read-only access. In this regard, in some specific embodiments, the memory device interface 20 of the device 111 may expose a read-only base address register (BAR) to the memory device interface 10 of the host 101. The access to the byte-addressable data represented in the page table 16 is initially read-only and remains read-only until the data is written or otherwise modified.
[0044] In some specific embodiments, the memory device interface 20 of the device 111 may expose a portion of the BAR as a readable and writable address range mapped to the memory of the device 111 (such as Figure 1 the memory 116 in []). In such specific embodiments, the portion of the BAR mapped to the SCM 120 may remain read-only, while the portion of the BAR exposed to the host 101 and mapped to the memory 116 may allow byte-addressable writes and reads via the memory device interface 10 of the host 101.
[0045] The read / write portion of the BAR mapped to the memory 116 may include, for example, battery-backed or power-fail-safe volatile DRAM or a portion thereof to effectively provide a non-volatile storage device or non-volatile memory. For example, a capacitor or battery may be used to protect the read / write portion of the BAR mapped to the memory 116 from data loss due to a power interruption, and the capacitor or battery may allow the control circuit 112 to transfer data from the BAR-mapped portion of the memory 116 to the SCM 120 after a power loss at the device 111.
[0046] In Figure 3In the example, after a write request for data represented in page table 16 (such as for virtual addresses B and D) has been received, the data for the write request is stored in buffer 107 of memory 106. Figure 3 This is indicated in the example where the allowed access for virtual addresses B and D is read / write access, and the physical addresses storing the data are indicated as Mem.1 and Mem.2 in memory 106 respectively. After the byte-addressable data included in the page or block corresponding to virtual addresses B and D has been rewritten or modified in buffer 107 of memory 106, Figure 2 the memory device interface 10 in can update page table 16. If needed, the byte-addressable data included in the buffered page or block can be modified or read from buffer 107 by the same application that issued the original write request for the data or by a different application. As discussed below with reference to Figure 4 more details, the status of such data remains read / write until the data is flushed from buffer 107 to SCM120.
[0047] Figure 4 is according to one or more embodiments Figure 3 is a state diagram of a page table entry in page table 16, such as an entry for one of virtual addresses A, B, C, or D. As Figure 4 shown, the entry starts in an initial state with read-only access to the physical address in SCM120. This physical address can correspond to the address of the BAR of the memory device interface 10 provided by device 111 to host 101. This allows the memory device interface 10 to directly access the address in SCM120 to read data without using the OS of host 101.
[0048] After a write request for the data represented by the entry is received, the entry moves to a second state. As discussed above, the write request can be processed as a software event by the memory device interface 10 and / or the block device interface 12. This generally allows for a more flexible design and implementation of host-side buffering compared to a hardware solution that may rely entirely on the MMU104.
[0049] In the second state, the block or page including the data for the write request has been retrieved by the block device interface 12 of host 101 and stored in buffer 107 of memory 106 in host 101. The previous or obsolete version of the block may remain in SCM120, but the modified block or page in buffer 107 is the current or valid version of the data for the virtual address. The memory device interface 10 or the block device interface 12 also updates page table 16 to change the access to read / write and maps the virtual address for the entry to the physical address where the data has been written in buffer 107 of memory 106.
[0050] In some embodiments, the memory device interface 10 or the block device interface 12 may recognize that a block or page has not been written to previously, or that the write request is the first write to the block or page. In such embodiments, the data to be written to the block or page may be stored in the buffer 107 without first retrieving the block or page from the device 111. The write request is then executed on the buffered block or page.
[0051] When the entry is in the second state, the block or page of the entry stored in the memory 106 may be modified or rewritten by the same application that issued the write request or by a different application. When the entry is in the second state, data corresponding to the entry may also be read from the physical address in the memory 106, such as byte-addressable data within the buffered block or page. The memory device interface 10 may refer to the entry in the page table 16 in response to read and write requests to modify or read byte-addressable data corresponding to the virtual address stored in the memory 106. Temporarily storing data in the memory 106 generally allows for write operations that are faster than write operations that may be achieved by writing data to the SCM 120. In this regard, buffering the modified byte-addressable data in the memory 106 may be advantageous when the buffered data is reused soon, as it may also be read quickly from the memory 106. The data buffered in the memory 106 may also be read faster than data read from the SCM 120. This is particularly beneficial for cache lines, as cache lines are typically read or modified soon after an initial write.
[0052] In addition, aggregating or modifying data in the memory 106 in a single write operation and using a separate block device interface to flush the aggregated or modified data block is more efficient than performing many smaller write operations in the SCM 120, which has a greater write latency than its read latency. The above use of both the block device interface and the memory device interface with the page table 16, as well as buffering the write data in the buffer 107, may also provide a more efficient arrangement compared to switching the access of the BAR of the SCM 120 from read-only to read / write or switching or temporarily modifying a single interface of the SCM 120 to accommodate byte-addressable data and block-addressable data. Delaying writes to the SCM 120 may improve the performance of the system including the host 101 and the device 111 by allowing for faster writes in the memory 106 of the host 101 and writing the aggregated or modified block to the SCM 120 at a later time when the write latency of the SCM 120 is less critical for a process or thread that may need to wait until the data has been written before continuing execution.
[0053] After the data for an entry has been modified or aggregated into one or more blocks by the block device interface 12, the data for the entry is flushed or demoted from the buffer 107 to the SCM 120 by the block device interface 12 via the block device interface 22 of the device 111. The block device interface 12 of the host 101 updates the entry such that access to the virtual address is unavailable or blocked while the data is being flushed to the SCM 120. In some embodiments, indicating that the virtual address is unavailable in the page table may include removing or deleting the entry for the virtual address or marking the entry as unavailable or obsolete. This ensures data consistency such that different applications do not modify the data in the memory 106 before access to the flushed data in the SCM 120 returns to read-only, which could result in reading an old or obsolete version of the data. Using the memory 106 to temporarily buffer write requests provides asynchronous storage of the data, where the write to the SCM 120 is delayed to improve system performance in terms of input / output operations per second (IOPS), while the aforementioned use of access permissions in the page table 16 allows the data to remain consistent.
[0054] As Figure 4 shown, after the flush or demote operation has been completed, the entry returns to the first state of read-only access in the SCM 120. In some embodiments, the block device interface 22 of the device 111 returns a command completion indication to the block device interface 12 of the host 101 to indicate the completion of the flush command. As discussed above with reference to Figure 2 the block device interface 22 of the device 111 may update the memory device interface 20 of the device 111 with the byte-addressable address of the data written to the SCM 120 for a write operation. Then, the memory device interface 20 of the device 111 may update the memory device interface 10 of the host 101 with the new byte-addressable address of the flushed data such that the memory device interface 10 may update the entry in the page table with the new address of the flushed data in the SCM 120. In other embodiments, the block device interface 12 of the host 101 may receive the updated byte-addressable address from the block device interface 22 of the device 111, which the block device interface 12 may use to update the page table.
[0055] Exemplary Process
[0056] Figure 5 is a flowchart of a page table creation process according to one or more embodiments. Figure 5 The process of
[0057] In block 502, the memory device interface 10 accesses the BAR of the SCM 120. In some embodiments, the control circuit 112 of the device 111 that executes the memory device interface 20 may expose the read-only BAR of the SCM 120 to the memory device interface 10 of the host 101. This allows the memory device interface 10 to have the size and data type information of the SCM 120 for mapping the virtual address of the host 101 to the physical address of the SCM 120, and allows the host 101 to perform direct memory access to the SCM 120 for read operations. Additionally, in some embodiments, the device 111 may also expose the read / write portion of the BAR mapped to the memory 116.
[0058] In block 504, the memory device interface 10 creates a page table that includes a plurality of entries corresponding to memory locations in the SCM 120. More specifically, the entries in the page table correspond to the exposed BAR of the device 111. The page table may include entries for different virtual addresses and mapped physical addresses in the SCM 120. In this regard, the created page table may include an entry for a virtual address or page that allows the memory device 10 to determine the physical location in the SCM 120 of the device 111 for byte-addressable data that is smaller than the page size or block size. The created page table may also include an indication of permitted access to the physical address, as in the case of the page table 16 discussed above. Figure 3 The case of the page table discussed above.
[0059] In block 506, the memory device interface 10 sets the plurality of entries in the page table to read-only. As discussed above, reading data from the SCM 120 is much faster than writing the same amount of data to the SCM 120. Thus, byte-addressable access to the SCM 120 is restricted to read-only access. As discussed in more detail below, writes to the SCM 120 are handled by the block device interface 12 of the host 101 such that data is written to the SCM 120 more efficiently in units of block size (e.g., 512 bytes, 4KB, 8KB, 2MB, or 4MB) rather than writing more times in units smaller than the block size, such as for a single byte of data. Additionally, delaying writes to the SCM 120 can improve the performance of the system including the host 101 and the device 111 by allowing writes to be performed faster in the memory 106 of the host 101 and writing the aggregated or modified data to the SCM 120 at a later time when the impact of the write latency to the SCM 120 does not delay the execution of a process or thread. Figure 6 More specifically discussed, writes to the SCM 120 are handled by the block device interface 12 of the host 101 such that data is written to the SCM 120 more efficiently in units of block size (e.g., 512 bytes, 4KB, 8KB, 2MB, or 4MB) rather than writing more times in units smaller than the block size, such as for a single byte of data. Additionally, delaying writes to the SCM 120 can improve the performance of the system including the host 101 and the device 111 by allowing writes to be performed faster in the memory 106 of the host 101 and writing the aggregated or modified data to the SCM 120 at a later time when the impact of the write latency to the SCM 120 does not delay the execution of a process or thread.
[0060] Figure 6 Is a flowchart of a write request process according to one or more embodiments. Figure 6The process may be performed by, for example, the processor circuitry 102 of the host 101 that implements the memory device interface 10 and the block device interface 12.
[0061] In block 602, the memory device interface 10 or the block device interface 12 receives a write request to write data corresponding to an entry in a page table (e.g., the page table 16 in Figure 3 . The write request may originate from, for example, an application executed by the processor circuitry 102. The write request is for data that is smaller than the page size or the block size and is thus byte-addressable. In such cases, the write request may originate from a flush or demotion of the cache of the processor circuitry 102. For example, one or more cache lines of the L1, L2, or L3 cache of the processor circuitry 102 may have been flushed or demoted.
[0062] In block 604, the data for the write request is written into the buffer 107 of the memory 106 of the device 111 using the block device interface 12. As discussed above, after being redirected from the memory device interface 10 or from another module (such as from a part of the OS of the host 101), a write request for storing byte-addressable data may be received by the block device interface 12 of the host 101. For example, in the case where the memory device interface 10 initially receives the write request, the write request may trigger a fault handler that redirects the write request to the block device interface 12. As discussed above with reference to Figure 2 , the block device interface 12 may also receive write requests for block-addressable data.
[0063] The byte-addressable data in the write buffer 107 for the write request received in block 602 may be aggregated into page-size or block-size units, or the current version of a block or page that includes the byte-addressable data may be read from the device 111 and stored in the buffer 107 to perform the write request. As mentioned above, for a given amount of data, performing a write operation in the SCM 120 takes longer than performing a read operation. Performing the write request in the buffer 107 may result in fewer overall writes to the SCM 120 and result in faster completion of smaller intermediate writes in the memory 106 to improve the efficiency and performance of the host 101 and the device 111. In some embodiments, write requests for data that is already in units of block size or page size may also be buffered in the memory 106 to improve the performance of the write operation by reducing the latency of performing the write operation. As mentioned above, faster completion of the write request may allow processes and threads to continue execution instead of waiting for the data to be written to the SCM 120. In other embodiments, the block-addressable data may alternatively be written to the SCM 120 without using the memory 106 to delay the writing of such data to the SCM 120. Such an arrangement may be preferred in cases where the size of the memory 106 is limited.
[0064] In block 606, the block device interface 12 changes the entry in the page table for the virtual address from read-only access to both read and write access. The block device interface 12 also changes the entry for the first virtual address to point to the location or physical address in the memory 106 where data for the first virtual address is written. As discussed above with reference to Figure 4 the page table entry state diagram, the data buffered in the memory 106 can be written to and read from the memory 106 until the data is flushed or demoted from the memory 106 to the SCM 120 in units of page size or block size. The process of Figure 8 modifying and flushing data from the memory 106 is discussed in more detail below with reference to
[0065] Figure 7 is a flowchart of a read request process for byte-addressable data according to one or more embodiments. Figure 7 The process of
[0066] In block 702, a read request is received by the memory device interface 10 at the host 101 to read byte-addressable data corresponding to an entry in a page table (e.g., Figure 3 the page table 16 in
[0067] ). The read request may come from, for example, an application executed by the processor circuit 102. The read request may be for data smaller than the page size or block size and is thus byte-addressable. In some cases, the read request may originate from a process or thread executed by the processor circuit 102 to load data into the cache of the processor circuit 102.
[0068] As discussed above, by allowing read-only access to the BAR of the SCM 120, the relatively fast read access of the SCM 120 to byte-addressable data can generally be utilized without incurring a greater performance penalty for writing byte-addressable data to the SCM 120. This can allow the host 101 to use a smaller main memory (e.g., memory 106) or save storage space for the host's main memory, which may be located inside or outside the host 101. As described above, in some specific embodiments, the memory 106 may include DRAM or SRAM, which may provide faster read and write access than the SCM 120 but have higher cost and power consumption for a given amount of data storage.
[0069] Figure 8 is a flowchart of a refresh process from a local memory to an SCM according to one or more embodiments. Figure 8 The process of may be performed by, for example, the processor circuit 102 that executes the block device interface 12.
[0070] In block 802, the block device interface 12 receives a write request to write byte-addressable data corresponding to an entry in the page table. The byte-addressable data to be written may include data within a page or block represented by the page table entry. Such write data may come from, for example, a process or thread executed by the processor circuit 102, which may flush or demote a dirty cache line from the cache of the processor circuit 102.
[0071] In block 804, the block device interface 12 reads a data block for the data block or data page represented by the page table entry from the SCM 120. The block or page is stored in the buffer 107 of the memory 106. In addition, the block device interface 12 updates the page table to indicate that the entry or virtual address for the buffered block or page has read / write access and that the data for the entry is at a physical address in the memory 106.
[0072] In block 806, the block device interface 12 modifies the byte-addressable data for the write request by writing the data to the block or page buffered in the buffer 107. As described above, when the block or page is stored in the memory 106, additional write requests and read requests may also be performed on the same byte-addressable data or other byte-addressable portions of the buffered block or page.
[0073] In block 808, the block device interface 12 indicates in the page table that the data for the buffered block or page is not available for reading and writing data to prepare to flush the modified block to the SCM 120. As described above with respect to Figure 4As described in the state diagram, indicating that the data for an entry or virtual address is not available in the page table can help ensure data consistency by not allowing it to be read or written when the data is flushed to the SCM 120.
[0074] In block 810, the block device interface 12 sends a write command to the device 111 to flush or demote the modified data block from the buffer 107 to the SCM 120. In some specific implementations, the block device interface 12 may wait until a threshold number of blocks have been aggregated in the buffer 107, or may wait for a predetermined amount of time without accessing the data in the blocks before flushing one or more modified blocks to the SCM 120 via the block device interface 22 of the device 111. In other cases, the block device interface 12 may flush the aggregated data block in response to reaching a data value for the block in the buffer 107, such as when new write data for a page or block that was not previously stored in the device 111 is collected in the buffer 107. In other specific implementations, flushing one or more data blocks from the buffer 107 may depend on the remaining storage capacity of the buffer 107. For example, the block device interface 12 may flush one or more data blocks from the buffer 107 in response to reaching 80% of the storage capacity of the buffer 107.
[0075] In block 812, the block device interface 12 sets the entry for the modified data block in the page table to read-only in response to the completion of the flush operation. This corresponds to returning from the third state to the first state in the Figure 4 exemplary state diagram. Read-only access may correspond to the read-only BAR of the SCM 120 used by the memory device interface 10. The completion of the flush operation may be determined by the block device interface 12 based on receiving a command completion indication from the block device interface 22 of the device 111.
[0076] In block 814, the block device interface 12 or the memory device interface 10 updates the entry in the page table for the flushed data block to point to the location in the SCM 120 where the block was flushed. In this regard, the new physical address of the data in the SCM 120 may be received by the block device interface 12 as part of the flush command completion indication, or alternatively may be received by the memory device interface 10 via the update process of the memory device interface 20 of the device 111.
[0077] Figure 9 is a flowchart of a multi-interface process of a device including an SCM according to one or more embodiments. Figure 9 The process may be performed by, for example, the control circuit 112 of the device 111.
[0078] In block 902, control circuit 112 receives a write command from host 101 using block device interface 22 to write data in a block to SCM 120. As discussed above, block device interface 22 is also used to read block-addressable data from SCM 120.
[0079] In addition, control circuit 112 receives a read command from host 101 in block 904 using memory device interface 20 to read byte-addressable data from SCM 120. Using two interfaces at device 111 allows SCM 120 to be used by host 101 as a main memory for reading byte-addressable data and as a non-volatile storage device for data blocks.
[0080] In block 906, memory device interface 20 exposes a read-only BAR for SCM 120 to host 101. As described above, the exposed BAR of device 111 may also include a read / write section located in memory 116 of device 111. The BAR may be exposed via, for example, a PCIe bus or an interconnect that allows commands for byte-addressable data to be executed using memory device interface and also allows commands for block-addressable data to be executed using a block device interface (such as, for example, NVMe). The use of the BAR may allow the processor circuit 102 at host 101 to create and update a page table that maps virtual addresses used by an application executed by processor circuit 102 to physical addresses in SCM 120.
[0081] Figure 10 is a flowchart of a device block write process according to one or more embodiments. Figure 10 The process may be performed by, for example, control circuit 112 of device 111.
[0082] In Figure 10 In block 1002 of, control circuit 112 of device 111 receives a write command from host 101 via block device interface 22 to write one or more data blocks in SCM 120. In some embodiments, control circuit 112 may use a logical-to-physical mapping module to determine the locations for writing one or more blocks for the write command in a write operation sent to SCM 120. In other embodiments, control circuit 112 may not use such a logical-to-physical mapping module.
[0083] After receiving an acknowledgment of the completion of the write operation in SCM 120, control circuit 112 updates memory device interface 20 in block 1004 using the physical addresses of the written data. In some embodiments, the updated addresses are shared with host 101 via memory device interface 20 of device 111 such that memory device interface 10 at host 101 may update the page table.
[0084] As discussed above, using the SCM to read byte-addressable data and write in blocks to the SCM allows the SCM to at least replace some of the main memory in the host's main memory while reducing the impact of the SCM's greater write latency. Additionally, using the host's main memory to temporarily buffer the modified byte-addressable data and updating the page table for the buffered data can help ensure that an old or stale version of the data is not read from the SCM.
[0085] Other Embodiments
[0086] Those of ordinary skill in the art will appreciate that the various illustrative logical blocks, modules, and processes described in connection with the examples disclosed herein can be implemented as electronic hardware, computer software, or combinations of both. Additionally, the foregoing processes can be embodied on a computer-readable medium that causes a processor or control circuit to perform or implement certain functions.
[0087] To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, and modules have been generally described above in terms of their functionality. Implementing this functionality as hardware or software depends upon the particular application and design constraints imposed on the overall system. Those of ordinary skill in the art can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.
[0088] The various illustrative logical blocks, units, modules, processor circuits, and control circuits described in connection with the examples disclosed herein can be implemented or performed with a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor can be a microprocessor, but in the alternative, the circuitry can be any conventional processor, controller, microcontroller, or state machine. The processor or control circuit can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, an SoC, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0089] The acts of a method or process described in connection with the examples disclosed herein can be embodied directly in hardware, in a software module executed by a processor or control circuit, or in a combination of both. The steps of the method or algorithm may also be executed in an order alternative to the order provided in the examples. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable medium, an optical medium, or any other form of storage medium known in the art. The exemplary storage medium is coupled to the processor or control circuit such that the circuit can read information from, and write information to, the storage medium. In an alternative, the storage medium may be integral to the processor or control circuit. The circuit and the storage medium may reside in an ASIC or an SoC.
[0090] The foregoing description of the exemplary embodiments of the present disclosure has been provided to enable any person of ordinary skill in the art to make or use the embodiments of the present disclosure. Various modifications to these examples will be readily apparent to those of ordinary skill in the art, and the principles disclosed herein may be applied to other examples without departing from the spirit or scope of the present disclosure. The embodiments are to be considered in all respects only as exemplary and not restrictive. Further, the language used in the following claims in the form of "at least one of A and B" should be understood to mean "only A, only B, or both A and B".
Claims
1. A method for interfacing with a device including a storage - class memory, i.e., SCM, the method comprises: creating a page table that includes a plurality of entries corresponding to memory locations in the SCM of the device; setting the plurality of entries in the page table to read - only; receiving a write request to write byte - addressable data corresponding to a first one of the plurality of entries; writing the byte - addressable data for the write request into a buffer of a memory of a host; receiving a read request to read byte - addressable data corresponding to a second one of the plurality of entries; sending a read command to the device via a memory device interface to read the byte - addressable data for the read request; and redirecting the write request from the memory device interface to a block device interface to execute the write request.
2. The method according to claim 1, further comprising changing the first entry to indicate read and write access, and pointing to a location in the memory of the host where the byte - addressable data for the write request is written.
3. The method according to claim 1, further comprises: reading a data block for the first entry from the SCM; storing the read data block in the buffer of the memory of the host; modifying the data block in the buffer to execute the write request; and sending a write command to the device via a block device interface to flush the modified data block from the buffer of the memory of the host to the SCM.
4. The method according to claim 3, further comprises: indicating that the data for the first entry in the page table is not available for writing data to prepare to flush the modified data block to the SCM; in response to completing flushing the modified data block to the SCM, setting the first entry in the page table to read - only; and updating the first entry in the page table to point to a location in the SCM where the modified data block is flushed.
5. The method according to claim 1, wherein the memory device interface accesses a base - address register, i.e., BAR, of the SCM to read data from the SCM.
6. The method according to claim 1, wherein the write request is executed in the buffer of the memory of the host using an operating system, i.e., OS, and the read command is sent to the device using the memory device interface without using the OS.
7. The method according to claim 1, further comprising communicating with the device via a peripheral component interconnect express, i.e., PCIe, via the memory device interface and the block device interface.
8. A device, the device comprises: a storage - class memory, i.e., SCM, for storing data; and control circuitry configured to: receive a write command from a host via a block device interface and write data in a block to the SCM, and receive a read command from the host and read data in a block from the SCM; and Receive a read command from the host using a memory device interface and read byte-addressable data from the SCM; wherein the block device interface and the memory device interface access the same SCM; wherein the control circuit is further configured to block execution of a write command received via the memory device interface.
9. The apparatus according to claim 8, wherein the control circuit is further configured to expose a base address register of the apparatus, i.e., a BAR, to the host via the memory device interface.
10. The apparatus according to claim 8, wherein the control circuit is further configured to: receive a write command from the host via the block device interface to write a data block in the SCM; and update an address used by the memory device interface for byte-addressable data included in the data block written in the SCM.
11. The apparatus according to claim 8, wherein the control circuit is further configured to communicate with the host via the memory device interface and the block device interface using a peripheral component interconnect express, i.e., PCIe.
12. A host, the host comprising: a memory for storing data; and a processor circuit configured to: receive a write request to write data corresponding to a first entry among a plurality of entries in a page table; write the data for the write request into a buffer of the memory using an operating system of the host, i.e., an OS; receive a read request to read data corresponding to a second entry among the plurality of entries in the page table; and without using the OS, send a read command to a device using a memory device interface to read the data for the read request from a storage class memory of the device, i.e., an SCM.
13. The host according to claim 12, wherein the processor circuit is further configured to change the first entry in the page table to indicate read and write access and point to a location in the memory of the host where the data for the write request is written.
14. The host according to claim 12, wherein the processor circuit is further configured to: read a data block for the first entry from the SCM; store the read data block in the buffer of the memory of the host; modify the data block in the buffer to execute the write request; and send a write command to the device using a block device interface to flush the modified data block from the buffer of the memory of the host to the SCM.
15. The host according to claim 12, wherein the processor circuit is further configured to: indicate that the first entry in the page table is not available for writing data in response to being ready to flush the modified data block to the SCM; set the first entry in the page table to read-only in response to completing flushing the modified data block to the SCM; and update the first entry in the page table to point to a location in the SCM where the modified data block is flushed.
16. The host according to claim 12, wherein the processor circuit is further configured to redirect the write request from the memory device interface to the block device interface to execute the write request.
17. The host according to claim 12, wherein the memory device interface accesses the base address register (BAR) of the SCM to read data from the SCM.
Citation Information
Patent Citations
Method and apparatus to provide both storage mode and memory mode access to non-volatile memory within a solid state drive
US20180004438A1
Persistent memory updating
US20190310796A1