Host system failover via data storage device configured to provide memory services
Through computing fast link connection and nonvolatile memory technology, SSD provides memory services in the memory system, solving the problems of insufficient volatile memory volume and data loss, achieving efficient database change management and failover, and improving system reliability and performance.
Patent Information
- Application Number
- CN202380082375.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2023-11-29
- Publication Date
- 2025-07-08
AI Technical Summary
In prior art In memory systems, especially solid state hard disks (SSDs), the limitation of backup power supplies leads to insufficient amount of volatile memory, affecting non-volatile characteristics, and the risk of data loss during failover is high, so it is impossible to efficiently manage database change data.
Through computing fast link (CXL) connections, SSD provides memory services and storage services, uses volatile memory supported by non-volatile memory or backup power supplies to manage database change data, and uses write-pre-logs and simple sorting tables to achieve continuous data storage and failover.
It improves the data storage capability of SSD in the event of power outage, reduces the delay and data loss during failover, and improves the efficiency and reliability of database operations.
Smart Images

Figure CN120283226A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims priority to U.S. Patent Application No. 18 / 519,565, filed on Nov. 27, 2023, which claims priority to U.S. Provisional Patent Application No. 63 / 385,951, filed on Dec. 2, 2022, the entire disclosure of which is hereby incorporated by reference herein. Technical Field
[0003] At least some of the embodiments disclosed herein generally relate to memory systems, and more particularly (but not limited to), to memory systems configured for use in memory services and storage services. Background Art
[0004] A memory subsystem may include one or more memory devices that store data. The memory devices can be, for example, non-volatile memory devices and volatile memory devices. Generally, a host system can utilize the memory subsystem to store data at and retrieve data from the memory devices. Brief Description of the Drawings
[0005] Embodiments are illustrated in the drawings by way of example and not limitation, in which like reference numerals indicate similar elements.
[0006] Figure 1 Illustrate an example computing system having a memory subsystem in accordance with some embodiments of the present disclosure.
[0007] Figure 2 Show a memory subsystem configured to provide both memory services and storage services to a host system via a physical connection.
[0008] Figure 3 Illustrate a write-ahead log entry that uses the memory services provided by a memory subsystem to track database records stored in the memory subsystem.
[0009] Figure 4 Show a log entry processing in accordance with one embodiment.
[0010] Figure 5 Illustrate a storage string table that uses the memory services provided by a memory subsystem to store changes to database records stored in the memory subsystem.
[0011] Figure 6 Show a method for tracking changes to database records in accordance with one embodiment.
[0012] Figure 7 Shows an example of host system failover based on memory services provided by a memory subsystem according to one embodiment.
[0013] Figure 8 and Figure 9 Describes an active host system that uses in-memory changed data stored in a memory subsystem to take over database operations of another host system according to one embodiment.
[0014] Figure 10 Shows a method for transferring database operations between host systems according to one embodiment. Detailed Description
[0015] At least some aspects of the present disclosure relate to tracking changes to data stored in a storage subsystem using memory services provided by the storage subsystem via a physical connection. The memory subsystem also uses the physical connection to provide storage services for the storage of data in the memory subsystem. Storing in-memory changed data of a database using the memory services provided by the memory subsystem can facilitate the transfer of database operations from one host system to another host system.
[0016] For example, a host system and a memory subsystem (e.g., a solid state drive (SSD)) can be connected via a physical connection according to the Compute Express Link (CXL) computer component interconnect standard. The Compute Express Link (CXL) includes protocols for storage access (e.g., cxl.io) and protocols for cache coherent memory access (e.g., cxl.mem and cxl.cache). Thus, the memory subsystem can be configured to provide both storage services and memory services to the host system using the Compute Express Link (CXL) via a physical connection.
[0017] A typical solid state drive (SSD) is a non-volatile storage device configured or designed to preserve the entire set of data received from a host system in the event of an unexpected power loss. The solid state drive may use volatile memory (such as SRAM or DRAM) as a buffer when processing storage access messages (such as read commands, write commands) received from the host system. To prevent data loss in the event of a power loss, the solid state drive is typically configured with an internal backup power source such that in the event of a power loss, the solid state drive can continue to operate for a limited period of time to save the data buffered in the volatile memory (such as SRAM or DRAM) to non-volatile memory (such as NAND). When the limited period of time is sufficient to ensure that the data is saved in the volatile memory (such as SRAM or DRAM) during a power loss event, the volatile memory supported by the backup power source can be considered non-volatile from the perspective of the host system. Typical implementations of the backup power source (such as capacitors, battery packs) limit the amount of volatile memory (such as SRAM or DRAM) configured in the solid state drive to maintain the non-volatile characteristics of the solid state drive as a data storage device. When implementing the function of this volatile memory via fast non-volatile memory, the backup power source can be eliminated from the solid state drive.
[0018] When the solid state drive is configured with a host interface that supports the Compute Express Link protocol, a portion of the fast volatile memory of the solid state drive can optionally be configured to provide cache coherent memory services to the host system. Such memory services can be accessed via load / store instructions executed over the Compute Express Link connection in the host system at the byte level (such as 64B or 128B). Another portion of the volatile memory of the solid state drive can be reserved for internal use by the solid state drive as buffer memory to facilitate storage services to the host system. Such storage services can be accessed via read / write commands provided by the host system over the Compute Express Link connection at the logical block level (such as 4KB).
[0019] When this solid state drive (SSD) is connected to a host system via a Compute Express Link connection, the solid state drive can be attached and used as both a memory device and a storage device of the host system. The storage device provides storage capacity for the data records of the database that can be addressed by the host system at the block level via read commands and write commands; and the memory device can provide physical memory for the changes to the data records of the database that can be addressed by the host system at the byte level via load instructions and storage instructions.
[0020] Changes to the database can be tracked via a write-ahead log (WAL), a simple sorted table (SST), etc. Before applying the changes to the database, the changes can be written to a non-volatile storage device. The changes recorded in the non-volatile storage device can be used to facilitate the reconstruction of in-memory changes in the event of a crash.
[0021] A write command can be used to save a data block at a storage location identified by a logical block address (LBA). This data block is typically configured to have a predetermined block size of 4KB. However, data having a size smaller than the predetermined block size of the data at the logical block address (such as a write-ahead log entry, a key-value pair in a table added to a simple sorted table) can generally be used to track changes to a database. For example, a change log entry can have a size ranging from a few bytes to a few hundred bytes. It is inefficient to partially modify the data block at the logical block address to store a small amount of changed data or to use the entire block to store a small amount of changed data.
[0022] It is advantageous for a host system to use the memory service provided by a solid-state drive to buffer changed data (such as write-ahead log entries, simple sorted tables). When the accumulated changed data has a size larger than the predetermined block size used for logical block addressing, the data blocks can be packed and written to the logical block address to store the changed data in the non-volatile storage device provided by the solid-state drive.
[0023] From the perspective of the host system, the memory space provided by the solid-state drive through a Compute Express Link connection can be considered non-volatile. The memory allocated by the solid-state drive to provide the memory service through the Compute Express Link connection can be implemented via non-volatile memory or via volatile memory supported by a backup power supply. The backup power supply is configured to be sufficient to ensure that: in the event of an external power interruption of the solid-state drive, the solid-state drive can continue to operate to save data from the volatile memory to the non-volatile storage capacity of the solid-state drive. Thus, in the event of an unexpected power interruption, the data (such as accumulated changed data) in the memory space provided by the solid-state drive is saved without loss.
[0024] After the changes to the database have been continuously committed to the storage device, the data configured to identify the changes may no longer be needed. Thus, the changed data can be discarded. If the changed data is stored in the memory space provided by the solid-state drive, the changed data can be erased from the memory space to provide space to accumulate additional data for identifying additional changes, without writing the changed data to a file in the storage device (such as attached by the solid-state drive). For example, a circular log can be implemented in the memory space provided by the solid-state drive. After the oldest log entry configured to identify a change is written to a file in the storage device, the oldest log entry can be overwritten by the latest log entry.
[0025] In some embodiments, a solid state drive may write changed data from a memory portion of its memory resources allocated to provide memory services to a host system to a storage portion of its memory resources allocated to provide storage services to the host system. The write may be performed without separately retrieving the changed data from the host system because the changed data is already in the faster memory of the solid state drive. This arrangement avoids the need for the changed data to be repeatedly communicated from the host system to the solid state drive for storage in the memory portion and writing to the storage portion.
[0026] It is advantageous to configure a computing system to have multiple host systems operable to independently perform database operations.
[0027] For example, when one of the host systems fails, another host system may take over the operations previously assigned to the failed host system.
[0028] For example, a computing system may be configured with a standby host system. When an active host system performs database operations, in-memory changes to the database may be recorded in a memory device attached to a solid state drive, which also attaches a storage device to the active host system to continuously store the records of the database. Since the active host system may use the memory device to store in-memory data of the database that has not yet been committed to the storage device, both the in-memory data and the continuous record of the database are preserved in the solid state drive. When the active host system fails, the standby host system may continue the database operations from the point where the active host system failed. The standby host system may use the in-memory data and the continuous record of the database in the solid state drive to start the database operations from the same or closest state as when the active host system failed. Thus, the latency and losses caused by the failure of the active host system may be reduced or minimized.
[0029] In some embodiments, cross-mirroring techniques may be used to replicate in-memory contents across multiple host systems. For example, when a host system writes data to its dynamic random access memory (DRAM), the data may be automatically replicated to the dynamic random access memory of another host system via a remote direct memory access (RDMA) network. Thus, when a host system fails, the in-memory contents of the failed host system may be used in the dynamic random access memory of another host system, which may continue the operations previously assigned to the failed host system.
[0030] Alternatively, instead of copying the in-memory content from the dynamic random access memory of a host system to the dynamic random access memory of another host system, the in-memory content can be copied from the dynamic random access memory of the host system to a memory device attached to a solid state drive that is connected to the host system via a Compute Express Link. Optionally, multiple host systems can share the services of the solid state drive; and when one host system fails, another host system can use the in-memory content of the failed host system to continue operation.
[0031] Optionally, the host system can be configured to perform database operations using a memory device attached to a solid state drive that is connected to the host system via a Compute Express Link instead of using its local dynamic random access memory. Thus, the state of the host system when performing database operations can be automatically saved in the solid state drive. Another system can take over the database operations of the host system by attaching the memory device provided by the solid state drive for use in database operations and starting from the state saved in the solid state drive. Host system failover or hot plugging can be performed with reduced or minimal latency. There is no need to copy the in-memory content from the dynamic random access memory of the host system (e.g., to another host system or solid state drive).
[0032] For example, the host system can be connected to a solid state drive via a Compute Express Link. The host system can be configured to use the storage services of the solid state drive to store a persistent copy of the database records managed by the host system. Additionally, the host system can be configured to use the memory services of the solid state drive to store data identifying in-memory changes to the persistent copy (e.g., in the form of write-ahead log entries, simple sorted tables) and other in-memory data related to database operations (e.g., in-memory caches of new or modified database records). When the solid state drive is reconnected (logically or physically) to a replacement host system, the persistent copy of the database records, the data identifying in-memory changes to the persistent copy, and other in-memory data related to database operations become available for use by the replacement host system in the same manner as they were available for use by the previous host system that was replaced. Thus, the replacement host system can replace the previous host system with minimal interruption and latency. In some embodiments, the solid state drive can have two ports that are pre-connected to two host systems respectively. The active host system can use the solid state drive for normal operation through one of the ports, while the replacement host system cannot effectively use the solid state drive through the other port. When the active host system fails, the replacement host system can become active to continue the operation of the failed host system using the existing connection of one of the ports without changing the physical connection between the host system and the solid state drive.
[0033] It is advantageous for the host system to query the solid-state drive using a communication protocol for the memory attachment capabilities of the solid-state drive, such as whether the solid-state drive can provide cache-coherent memory services, how much memory the solid-state drive can attach to the host system's memory when providing memory services, how much of the memory that can be attached to provide memory services can be considered non-volatile (e.g., implemented via non-volatile memory or supported by a backup power supply), what is the access time of the memory that can be allocated by the solid-state drive to memory services, etc.
[0034] The query results can be used to configure the memory allocation in the solid-state drive to provide cache-coherent memory services. For example, a portion of the solid-state drive's fast memory can be provided to the host system for cache-coherent memory access; and the remaining portion of the fast memory can be reserved internally by the solid-state drive. The partitioning of the solid-state drive's fast memory for different services can be configured to balance the benefits of the memory services provided by the solid-state drive to the host system with the performance of the storage services implemented by the solid-state drive for the host system. Optionally, the host system can explicitly request the solid-state drive to set aside a requested portion of its fast volatile memory as memory that can be accessed by the host system via a compute express link, using a cache-coherent memory access protocol, through the connection.
[0035] For example, when the solid-state drive is connected to the host system via a compute express link connection to provide storage services, the host system can send a command to the solid-state drive to query the memory attachment capabilities of the solid-state drive.
[0036] For example, the command to query the memory attachment capabilities can be configured with a different command identifier than a read command; and in response, the solid-state drive is configured to provide a response indicating whether the solid-state drive can operate as a memory device to provide memory services that can be accessed via load instructions and store instructions. Additionally, the response can be configured to identify the amount of available memory that can be allocated and attached as a memory device that can be accessed via a compute express link connection. Optionally, the response can be further configured to include an identification of the amount of available memory that can be considered non-volatile by the host system and can be used by the host system as a memory device. The non-volatile portion of the memory device attached by the solid-state drive can be implemented via non-volatile memory or volatile memory supported by a backup power supply and the non-volatile storage capacity of the solid-state drive.
[0037] Optionally, a solid state drive may be configured with more volatile memory than the amount supported by its backup power supply. After a power interruption to the solid state drive, the backup power supply is sufficient to store data from a portion of the volatile memory of the solid state drive to its storage capacity, but not sufficient to save all of the data in the volatile memory to its storage capacity. Accordingly, a response to a memory attachment capability query may include an indication of the ratio of the volatile and non-volatile portions of the memory that may be allocated by the solid state drive to memory services. Optionally, the response may further include an identification of the access time of the memory that may be allocated by the solid state drive to cache coherent memory services. For example, when a host system requests data from the solid state drive via a cache coherent protocol over a Compute Express Link, the solid state drive may provide the data within a period of time no longer than the access time.
[0038] Optionally, a preconfigured response to this query may be stored at a predetermined location in a storage device attached by the solid state drive to the host system. For example, the predetermined location may be at a predetermined logical block address in a predetermined namespace. For example, the preconfigured response may be configured as part of the firmware of the solid state drive. The host system may use a read command to retrieve the response from the predetermined location.
[0039] Optionally, when the solid state drive has the ability to function as a memory device, the solid state drive may automatically allocate a predetermined amount of its fast volatile memory as a memory device attached to the host system via a Compute Express Link connection. The predetermined amount may be a minimum or default amount configured at the manufacturing facility of the solid state drive or an amount specified by configuration data stored in the solid state drive. Subsequently, the memory attachment capability query may optionally be implemented in the command set of a cache coherent memory access protocol (rather than the command set of a storage access protocol); and the host system may use the query to retrieve parameters specifying the memory attachment capabilities of the solid state drive. For example, the solid state drive may place the parameters into a memory device at a predetermined memory address; and the host may retrieve the parameters by executing a load command with the corresponding memory address.
[0040] It is advantageous for a host system to customize aspects of the memory services of a memory subsystem (such as a solid state drive) for the memory and storage usage patterns of the host system.
[0041] For example, the host system may specify the size of the memory device attached by the solid state drive to the host system such that a set of physical memory addresses configured according to the size may be addressed via the execution of load / store instructions in a processing device of the host system.
[0042] Optionally, the host system may specify time requirements for accessing a memory device via a Compute Express Link (CXL) connection. For example, when a cache request accesses a memory location via the connection, the solid state drive is required to provide a response within an access time specified by the host system when configuring the memory services of the solid state drive.
[0043] Optionally, the host system may specify how much of the memory device attached to the solid state drive is non-volatile such that when the external power supply of the solid state drive fails, the data in the non-volatile portion of the memory device attached to the host system by the solid state drive is not lost. The non-volatile portion may be implemented by the solid state drive via non-volatile memory or volatile memory with a backup power supply to continue the operation of copying data from volatile memory to non-volatile memory during an external power interruption of the solid state drive.
[0044] Optionally, the host system may specify whether the solid state drive will attach a memory device to the host system via a Compute Express Link (CXL) connection.
[0045] For example, the solid state drive may have an area configured to store configuration parameters for a memory device attached to the host system via a Compute Express Link (CXL) connection. When the solid state drive restarts, boots, or powers on, the solid state drive may allocate a portion of its memory resources for the memory device attached to the host system according to the configuration parameters stored in the area. After the solid state drive configures the memory services according to the configuration parameters stored in the area, the host system may access via the cache by executing load instructions and store instructions that identify corresponding physical memory addresses. The solid state drive may configure its remaining memory resources to provide storage services via a Compute Express Link (CXL) connection. For example, a portion of its volatile random access memory may be allocated as a buffer memory for the processing device of the solid state drive; and the host system cannot access and address the buffer memory via load / store instructions.
[0046] When the solid state drive is connected to the host system via a Compute Express Link connection, the host system may send a command to adjust the configuration parameters stored in the area of the attachable memory device. Subsequently, the host system may request the solid state drive to restart to attach a memory device with memory services configured according to the configuration parameters to the host system via the Compute Express Link.
[0047] For example, the host system may be configured to issue a write command (or store command) to save the configuration parameters at a predetermined logical block address (or predetermined memory address) in the area to customize the settings of the memory device configured to provide memory services via a Compute Express Link connection.
[0048] Alternatively, a command having a command identifier different from a write command (or a store instruction) may be configured in a read-write protocol (or a load-store protocol) to instruct the solid state drive to adjust configuration parameters stored in a region.
[0049] Figure 1 Illustrate an example computing system 100 that includes a memory subsystem 110 in accordance with some embodiments of the present disclosure. The memory subsystem 110 may include computer-readable storage media, such as one or more volatile memory devices (e.g., memory device 107), one or more non-volatile memory devices (e.g., memory device 109), or a combination thereof.
[0050] In Figure 1 the memory subsystem 110 is configured as an article (e.g., a solid state drive) that can be used as a component installed in a computing device.
[0051] The memory subsystem 110 further includes a host interface 113 for a physical connection 103 to the host system 120.
[0052] The host system 120 may have an interconnect 121 that connects a cache 123, a memory 129, a memory controller 125, a processing device 127, and a change manager 101. The change manager 101 is configured to accumulate changes for storage in the storage capacity of the memory subsystem 110 using the memory services of the memory subsystem 110.
[0053] The change manager 101 in the host system 120 may be implemented at least in part by instructions executed by the processing device 127 or by logic circuitry or both. Before changes to a database are written to a file in a storage device attached to the host system 120 by the memory subsystem 110, the change manager 101 in the host system 120 may use the memory devices attached to the host system 120 by the memory subsystem 110 to store the changes. Optionally, the change manager 101 in the host system 120 is implemented as part of an operating system 135 of the host system 120, a database manager in the host system 120, or a device driver configured to operate the memory subsystem 110 or a combination of such software components.
[0054] The connection 103 may be according to the Compute Express Link (CXL) standard or other communication protocols that support cache coherent memory access and storage access. Optionally, multiple physical connections 103 are configured to support cache coherent memory access communication and support storage access communication.
[0055] The processing device 127 can be a microprocessor that is configured as a central processing unit (CPU) of a computing device. Instructions (such as load instructions, store instructions) executed in the processing device 127 can access the memory 129 via the memory controller (125) and the cache 123. In addition, when the memory subsystem 110 attaches a memory device to the host system through the connection 103, instructions (such as load instructions, store instructions) executed in the processing device 127 can access the memory device via the memory controller (125) and the cache 123 in a similar manner as accessing the memory 129.
[0056] For example, in response to executing a load instruction in the processing device 127, the memory controller 125 can convert the logical memory address specified by the instruction into a physical memory address to request the cache 123 to perform a memory access to retrieve data. For example, the physical memory address can be in the memory 129 of the host system 120 or in the memory device attached to the host system 120 by the memory subsystem 110 through the connection 103. If the data at the physical memory address is not yet in the cache 123, then the cache 123 can load the data from the corresponding physical address as cache content 131. The cache 123 can provide the cache content 131 to service the memory access request at the physical memory address.
[0057] For example, in response to executing a store instruction in the processing device 127, the memory controller 125 can convert the logical memory address specified by the instruction into a physical memory address to request the cache 123 to perform a memory access to store data. The cache 123 can hold the data of the store instruction as cache content 131 and indicate that the corresponding data at the physical memory address has expired. When the cache 123 needs to free up a cache block (e.g., for loading new data from a different memory address or for holding the data of a store instruction for a different memory address), the cache 123 can flush the cache content 131 from the cache block to the corresponding physical memory address (e.g., in the memory 129 of the host system or in the memory device attached to the host system 120 by the memory subsystem 110 through the connection 103).
[0058] The connection 103 between the host system 120 and the memory subsystem 110 can support a cache-coherent memory access protocol. Cache coherence ensures that changes to copies of data corresponding to a memory address propagate to other copies of the data corresponding to the memory address; and the processing device (e.g., 127) sees load / store accesses to the same memory address in the same order.
[0059] The operating system 135 can include routines of instructions programmed to handle storage access requests from applications.
[0060] In some embodiments, the host system 120 configures a portion of its memory (e.g., 129) to serve as a queue 133 for storing access messages. Such storage access messages may include read commands, write commands, erase commands, and the like. A storage access command (e.g., read or write) may specify a logical block address of a data block in a storage device (e.g., attached to the host system 120 by the memory subsystem 110 via the connection 103). The storage device may retrieve the message from the queue 133, execute the command, and provide the result in the queue 133 for further processing by the host system 120 (e.g., using routines in the operating system 135).
[0061] Typically, a data block addressed by a storage access command (e.g., read or write) has a much larger size than a data unit that can be accessed via a memory access instruction (e.g., load or store). Thus, the storage access command may facilitate batch processing of a large amount of data (e.g., data in a file managed by a file system) simultaneously and in the same manner with the help of routines in the operating system 135. Memory access instructions can be efficiently used for random access of small segments of data without the overhead of routines in the operating system 135.
[0062] The memory subsystem 110 has an interconnect 111 that couples a host interface 113, a controller 115, and memory resources (e.g., memory devices 107, …, 109).
[0063] The controller 115 of the memory subsystem 110 may control the operation of the memory subsystem 110. For example, the operation of the memory subsystem 110 may respond to a storage access message in the queue 133 or to a memory access request from the cache 123.
[0064] In some embodiments, each of the memory devices (e.g., 107, …, 109) includes one or more integrated circuit devices each enclosed in a separate integrated circuit package. In other embodiments, each of the memory devices (e.g., 107, …, 109) is configured on an integrated circuit die; and the memory devices (e.g., 107, …, 109) may be configured in the same integrated circuit device enclosed in the same integrated circuit package. In additional embodiments, the memory subsystem 110 is implemented as an integrated circuit device in an integrated circuit package that encloses the memory devices 107, …, 109, the controller 115, and the host interface 113.
[0065] For example, the memory device 107 of the memory subsystem 110 may have a volatile random access memory 138 that is faster than the non-volatile memory 139 of the memory device 109 of the memory subsystem 110. Thus, the non-volatile memory 139 can be used to provide the storage capacity of the memory subsystem 110 to retain data. At least a portion of the storage capacity can be used to provide storage services to the host system 120. Optionally, a portion of the volatile random access memory 138 can be used to provide cache coherent memory services to the host system 120. The remaining portion of the volatile random access memory 138 can be used to provide buffering services to the controller 115 when processing storage access messages in the processing queue 133 and when performing other operations (such as wear leveling, garbage collection, error detection and correction, encryption).
[0066] When the volatile random access memory 138 is used to buffer data received from the host system 120 before saving it into the non-volatile memory 139, when the power supply of the memory device 107 is interrupted, the data in the volatile random access memory 138 will be lost. To prevent data loss, the memory subsystem 110 may have a backup power supply 105 that may be sufficient to operate the memory subsystem 110 for a period of time to allow the controller 115 to commit the buffered data from the volatile random access memory 138 to the non-volatile memory 139 in the event of an external power interruption of the memory subsystem 110.
[0067] Optionally, the fast memory 138 can be implemented via a non-volatile memory (such as a cross-point memory); and the backup power supply 105 can be eliminated. Alternatively, a combination of fast non-volatile memory and fast volatile memory can be configured in the memory subsystem 110 for memory services and buffering services.
[0068] The host system 120 can send a memory attachment capability query to the memory subsystem 110 via the connection 103. In response, the memory subsystem 110 can provide a response that identifies: whether the memory subsystem 110 can provide cache coherent memory services via the connection 103, how much memory can be attached to provide memory services via the connection 103, how much of the memory available for memory services of the host system 120 is considered non-volatile (e.g., implemented via non-volatile memory or supported by the backup power supply 105), what is the access time of the memory that can be allocated for memory services of the host system 120, etc.
[0069] The host system 120 may send a request to the memory subsystem 110 via the connection 103 to configure the memory services provided by the memory subsystem 110 to the host system 120. In the request, the host system 120 may specify: whether the memory subsystem 110 will provide cache coherent memory services via the connection 103, how much memory will be provided for memory services via the connection 103, how much of the memory provided via the connection 103 is considered non-volatile (e.g., implemented via non-volatile memory or supported by the backup power supply 105), what is the access time of the memory provided to the host system 120 as a memory service, etc. In response, the memory subsystem 110 may partition its resources (e.g., memory devices 107, …, 109) and provide the requested memory services via the connection 103.
[0070] When a portion of the memory 138 is configured to provide memory services via the connection 103, the host system 120 may access the cache portion 132 of the memory 138 via load instructions, store instructions, and the cache 123. The non-volatile memory 139 may be accessed via read commands and write commands that are transmitted via the queue 133 configured in the memory 129 of the host system 120.
[0071] Using the memory services of the memory subsystem 110 provided via the connection 103, the host system 120 may accumulate data identifying changes to the database in the memory of the subsystem (e.g., in a portion of the volatile random access memory 138). When the size of the accumulated change data is higher than a threshold, the change manager 101 may pack the change data into one or more data blocks for one or more write commands addressing one or more logical block addresses. The change manager 101 may be implemented in the host system 120 or the memory subsystem 110, or partially implemented in the host system 120 and partially implemented in the memory subsystem 110. The change manager 101 in the memory subsystem 110 may be implemented at least in part via instructions (e.g., firmware) executed by the processing device 117 of the controller 115 of the memory subsystem 110 or via logic circuits or both.
[0072] Figure 2 A memory subsystem configured to provide both memory services and storage services to a host system via a physical connection is shown according to one embodiment. For example, Figure 2 the memory subsystem 110 and the host system 120 may be implemented in a manner as Figure 1 the computing system 100.
[0073] In Figure 2In, the memory resources (e.g., memory devices 107, …, 109) of the memory subsystem 110 are partitioned into a loadable portion 141 and a readable portion 143 (and in some cases, an optional portion of the buffer memory 149, as in Figure 5 ). The physical connection 103 between the host system 120 and the memory subsystem 110 can support a protocol 145 for load instructions and store instructions to access the memory services provided in the loadable portion 141. For example, the load instructions and store instructions can be executed via the cache 123. The connection 103 can further support a protocol 147 for read commands and write commands to access the storage services provided in the readable portion 143. For example, the read commands and write commands can be provided via a queue 133 configured in the memory 129 of the host system 120. For example, a physical connection 103 that supports Compute Express Link can be used to connect the host system 120 and the memory subsystem 110.
[0074] Figure 2 An example of the same physical connection 103 (e.g., a Compute Express Link connection) configured to facilitate both memory access communication according to protocol 145 and storage access communication according to another protocol 147 is described. Generally, separate physical connections can be used to provide memory access according to a memory access protocol 145 and storage access according to another storage access protocol 147 to the host system 120.
[0075] Figure 3 A write-ahead log entry for tracking database records stored in a memory subsystem using the memory services provided by the memory subsystem according to one embodiment is described. For example, Figure 3 The techniques of Figure 1 and Figure 2 can be implemented in the computing system 100.
[0076] In Figure 3 , the host system 120 can have a database manager 151 configured to perform database operations. The database manager 151 can use the readable portion 143 of the memory subsystem 110 to maintain a persistent copy of the database records 157. To improve database performance, the database manager 151 can use its memory 129 to store cache records 158 used during activity.
[0077] Before changing the records (e.g., 157, 158), the database manager 151 can save a persistent copy of the data identifying the changes. For example, the write-ahead log entry 155 can be used to identify the changes made to the records (e.g., 157, 158). Thus, in the event of a crash, the record changes can be used to perform recovery operations. Additionally, the record changes allow the changes to be rolled back upon request or as desired.
[0078] A typical write-ahead log entry 155 does not have a predetermined fixed size; and its size can be less than a predetermined block size of data that can be addressed via a logical block address. Using a write command to write a data block that is significantly larger than the size of the write-ahead log entry 155 into the readable portion 143 (e.g., in the log file 159) is inefficient. Additionally, writing to a storage system is typically implemented through a storage stack that involves a file system, a basic input / output system (BIOS) driver, a low-level driver, all possible intermediate mappers, and drivers. Therefore, writing to a storage system can be extremely resource-consuming and slow.
[0079] The database manager 151 can store the write-ahead log entry 155 in the loadable portion 141 of the memory subsystem 110 to achieve persistence and accumulation.
[0080] For example, the database manager 151 can generate the write-ahead log entry 155 in the memory 129 of the host system 120 and then move the entry 155 from the host memory 129 to the loadable portion 141 in the memory subsystem 110 to achieve persistence, rather than using a write command to write the write-ahead log entry 155 into the log file 159 in the readable portion 143. After the write-ahead log entry 155 is in the loadable portion 141 to achieve persistence, the database manager 151 can change the cache record 158, as identified by the write-ahead log entry 155.
[0081] After the cache record 158 is stored in the readable portion 143 of the memory subsystem 110 to achieve persistence, there is no need to continuously store the write-ahead log entry 155 that identifies the change. Therefore, in some examples, the write-ahead log entry 155 can be deleted without writing to the log file 159. The memory space in the loadable portion 141 freed from deleting the write-ahead log entry 155 can be used to store additional write-ahead log entries 155.
[0082] In some examples, there are more write-ahead log entries 155 to be saved than can be stored in the loadable portion 141. Therefore, at least a portion of the write-ahead log entry 155 can be written from the loadable portion 141 into the log file 159 in the readable portion 143. After the write-ahead log entry 155 is written into the log file 159, the corresponding write-ahead log entry 155 in the loadable portion 141 can be erased. By grouping the write-ahead log entries 155 for writing into the log file 159 for data blocks, the efficiency of the computing system in implementing the persistence of the write-ahead log entry 155 is improved.
[0083] In some embodiments, the memory subsystem 110 may automatically write the write-ahead log entry 155 from the loadable section 141 to the readable section 143 in response to a request from the database manager 151 or when the aggregated size of the write-ahead log entries 155 is higher than a threshold. Accordingly, the host system 120 does not need to resend the data of the write-ahead log entry 155 with a write command to write the write-ahead log entry 155 to the log file 159.
[0084] In some embodiments, the change manager 101 of the host system 120 and the change manager 101 of the memory subsystem 110 (e.g., implemented via the firmware 153) communicate with each other to save the write-ahead log entry 155 from the loadable section 141 to the readable section 143 and retrieve the write-ahead log entry 155 from the log file 159 for use by the database manager 151.
[0085] Figure 4 Illustrates log entry processing according to one embodiment. For example, Figure 4 the log entry processing of Figure 3 can be implemented using the techniques of Figure 1 and Figure 2 in the computing system 100.
[0086] In Figure 4 , the database manager 151 configured in the host system 120 may generate a log entry (e.g., 173) in the host memory 129 to identify a change to a database record (e.g., 157 or 158). Before changing the database record (e.g., 157 or 158), the change manager 101 may continuously store the log entry 173. The change manager 101 may be implemented as part of the database manager 151, part of the operating system 135, part of the device driver of the memory subsystem 110, or a separate component. For example, the host memory 129 may be volatile. Accordingly, the change manager 101 may move 165 the entry (e.g., 173) from the host memory 129 to the buffer area 161 allocated in the loadable section 141 of the memory subsystem 110 to continuously store the log entry 173. Optionally, the database manager 151 may directly generate the log entry 173 in the buffer area 161.
[0087] Since the buffer area 161 is the non-volatile part of the memory device attached by the memory subsystem 110 to the host system 120, the entries 171,..., 173 in the buffer area 161 may be considered to be continuously stored in the memory subsystem 110. For example, the entries 171,..., 173 in the buffer area 161 may be saved even if an unexpected power interruption occurs in the memory subsystem 110.
[0088] After several log entries 171, …, 173 have accumulated in buffer region 161, change manager 101 may pack 167 at least some of the log entries in loadable section 141 into data blocks 163 and write 169 the data blocks 163 to log file 159 in readable section 143. Optionally, change manager 101 may pack 167 data blocks 163 into an appropriate location within buffer region 161 such that the data blocks 163 can be identified via a series of memory addresses.
[0089] In some embodiments, change manager 101 is partially implemented in memory subsystem 110 to write data blocks 163 directly from buffer region 161 to log file 159 without host system 120 generating the data blocks 163 in host memory 129. Since log entries 171, …, 173 are already in memory subsystem 110, host system 120 does not have to retransmit the data of log entries 171, …, 173 over connection 103 to write data blocks 163.
[0090] In some embodiments, change manager 101 implemented in host system 120 is configured to generate in queue 133 a write command that requests memory subsystem 110 to write data block 163 at a location represented by a logical block address in log file 159 in readable section 143, as in a series of memory addresses in buffer region 161. Since log entries 171, …, 173 are already in memory subsystem 110, host system 120 does not have to retransmit the data of log entries 171, …, 173 over connection 103 to write data blocks 163.
[0091] In some embodiments where memory subsystem 110 is not sufficient to support writing data blocks 163 to log file 159 based on log entries 171, …, 173 in buffer region 161 (e.g., packed into an appropriate location in buffer region 161), change manager 101 in host system 120 may be configured to pack 167 log entries 171, …, 173 in host memory 129 and generate a write command to write data blocks 163 to log file 159 via queue 133 configured in system memory 129.
[0092] For example, Figure 4 the log entries 171, …, 173 in
[0093] In some embodiments, database changes are tracked using simple sorted tables (SSTs). The tables are organized by age based on the time they were most recently created. Each table may contain key-value pairs to identify changes to the database. The tables may be stored in the storage device in ascending order of recency. To improve performance, newly created tables may be kept in memory. The change manager 101 may be configured to place newly created tables in the loadable portion 141 of the memory subsystem 110 in a manner similar to the persistent storage of write-ahead log entries 155 to achieve persistence, as discussed in conjunction with Figure 5 further discussed.
[0094] Figure 5 Illustrate a storage string table for storing changes to database records stored in a memory subsystem using memory services provided by the memory subsystem. For example, Figure 5 the techniques of Figure 1 and Figure 2 may be implemented in the computing system 100 of
[0095] In Figure 5 , the host system 120 may have a database manager 151 configured to perform database operations, as in Figure 3 . The database manager 151 may use the readable portion 143 of the memory subsystem 110 to maintain a persistent copy of the database records 157. To improve database performance, the database manager 151 may use its memory 129 to store cache records 158 that are used while active.
[0096] Prior to changing records (e.g., 157, 158), the database manager 151 may save a persistent copy of the data identifying the changes. For example, a simple sorted table 185 may be used to identify changes made to the records (e.g., 157, 158). Thus, in the event of a crash, the recorded changes may be used to perform recovery operations. Additionally, recording changes allows changes to be rolled back upon request or as desired.
[0097] To improve the efficiency of operations related to the simple sorted table 185, the most recent table 185 may be maintained in memory. For example, the loadable portion 141 may be used to store the table 185 before the most recent table 185 is written to the table file 189.
[0098] Optionally, the change manager 101 may move the most recent table 185 between the host memory 129 and the loadable portion 141 in the memory subsystem 110.
[0099] Since the loadable portion 141 is non-volatile (e.g., implemented via fast non-volatile memory or volatile memory supported by a backup power supply 105), the simple sorted table 185 in the loadable portion may be saved when an unexpected power loss occurs.
[0100] In some embodiments, the memory subsystem 110 may automatically write the simple sort table 185 from the loadable portion 141 to the readable portion 143 in response to a request from the host system 120 or when the aggregated size of the simple sort table 185 is higher than a threshold. Accordingly, the host system 120 does not need to reissue the data of the simple sort table 185 with a write command to write the simple sort table 185 into the table file 189 in the readable portion 143.
[0101] In some embodiments, the change manager 101 in the host system 120 communicates with the change manager 101 in the memory subsystem 110 to save the simple sort table 185 from the loadable portion 141 into the readable portion 143 and retrieve the simple sort table 185 from the table file 189 for use by the database manager 151.
[0102] In some embodiments where the memory subsystem 110 is not sufficient to support writing the table file 189 using the simple sort table 185 stored in the loadable portion 141, the change manager 101 in the host system 120 may be configured to pack the data of the simple sort table 185 into data blocks in the host memory 129 and generate a write command to write the data blocks into the table file 189 via a queue 133 configured in the system memory 129.
[0103] Figure 6 Displays a method for tracking changes to database records. For example, Figure 6 The method may be implemented by Figure 3 、 Figure 4 and Figure 5 The techniques of to be implemented in Figure 1 and Figure 2 The computing system 100 of to continuously store data identifying changes to the database using the memory services of the memory subsystem 110.
[0104] For example, the memory subsystem 110 (e.g., a solid-state drive) and the host system may be connected via at least one physical connection 103. The memory subsystem 110 may optically demarcate a portion (e.g., the loadable portion 141) of its fast memory (e.g., 138) as a memory device attached to the host system 120. The memory subsystem 110 may reserve a portion (e.g., the buffer memory 149) of its fast memory (e.g., 138) as internal memory for its processing device (e.g., 117). The memory subsystem 110 may make a portion (e.g., the readable portion 143) of its memory resources (e.g., the non-volatile memory 139) available as a storage device attached to the host system 120.
[0105] Memory subsystem 110 may have a backup power supply 105, which is designed to ensure that when the power supply of memory subsystem 110 is interrupted, the data stored in at least a portion of volatile random access memory 138 is saved in non-volatile memory 139. Thus, this portion of volatile random access memory 138 can be regarded as non-volatile in the memory service of host system 120.
[0106] The database manager 151 running in host system 120 may use a storage protocol (such as 147) to write records of the database into the storage portion (such as 143) of memory subsystem 110 through connection 103 of host interface 113 of memory subsystem 110. The database manager 151 may include a change manager 101 configured to generate data identifying changes to the database, such as write-ahead log entries 155, simple sorted tables 185, etc. The change manager 101 may use a cache coherence memory access protocol (such as 145) to store the data into the memory portion (such as 141) of memory subsystem 110 through connection 103 of host interface 113 of memory subsystem 110 before changing the database. Since the memory portion (such as 141) is implemented via non-volatile memory or volatile memory 138 with a backup power supply 105, the data stored in the memory portion (such as 141) is persistent. After the data is persistently stored in the memory portion (such as 141) of memory subsystem 110, the database manager 151 may change the database.
[0107] At block 201, host system 120 and memory subsystem 110 communicate with each other using a first protocol for cache coherence memory access (such as 145) and a second protocol for storage access (such as 147) through connection 103 configured between memory subsystem 110 and host system 120.
[0108] At block 203, host system 120 generates first data identifying one or more first changes to the database.
[0109] For example, the first data identifying changes to the database may be in the form of write-ahead log entries 155 or simple sorted tables 185.
[0110] At block 205, host system 120 uses a first protocol for cache coherence memory access (such as 145) to store the first data into the first portion (such as 141) of memory subsystem 110 through connection 103 between memory subsystem 110 and host system 120.
[0111] At block 207, host system 120 generates second data identifying one or more second changes to the database.
[0112] For example, the second data identifying changes to the database may be in the form of additional write-ahead log entries 155 or a simple sorted table 185.
[0113] At block 209, the host system 120 stores the second data into a first portion (e.g., 141) of the memory subsystem 110 via a connection between the memory subsystem 110 and the host system 120 using a first protocol for cache coherent memory access.
[0114] The sizes of the first data and the second data may be small; and it may be inefficient to write the first data and the second data into the memory subsystem 110 respectively using a second protocol (e.g., 147) for storage access of a file. After the changed data (e.g., the first data and the second data) accumulates in the first portion (e.g., 141) of the memory subsystem 110, the changed data may be written into a file 189 (e.g., 159 or 189).
[0115] For example, the first data and the second data may be stored into the first portion (e.g., 141) of the memory subsystem 110 via a store instruction that identifies a memory address in the first portion (e.g., 141) of the memory subsystem 110 executed in the host system 120.
[0116] At block 211, the first data and the second data are written into a second portion (e.g., 143) of the memory subsystem 110 that can be accessed via a second protocol (e.g., 147) for storage access.
[0117] For example, the connection 103 between the host system 120 and the memory subsystem 110 may be a Compute Express Link (CXL) connection.
[0118] For example, the first data and the second data may be written into a second portion (e.g., 143) of the memory subsystem 110 via a write command regarding files (e.g., 159, 189) in the second portion (e.g., 143) of the memory subsystem 110. For example, the write command is configured to identify data to be written at a logical block address in the second portion (e.g., 143) of the memory subsystem 110 by referring to a data block 163 in the first portion (e.g., 141) of the memory subsystem 110. For example, the reference may be based on a series of memory addresses in the first portion (e.g., 141). Writing the first data and the second data into the second portion (e.g., 143) of the memory subsystem 110 may be responsive to the aggregated size of changed data stored in the first portion (e.g., 141) of the memory subsystem 110 exceeding a threshold. After the first data and the second data are stored in the first portion (e.g., 141) of the memory subsystem 110, writing the first data and the second data into the second portion (e.g., 143) of the memory subsystem 110 does not involve additional communication of the first data and the second data via a Compute Express Link (CXL) connection from the host system 120 to the memory subsystem 110.
[0119] For example, before making corresponding changes to the database, the change manager 101 and the database manager 151 may perform write-ahead logging to generate changed data (e.g., write-ahead log entry 155) and continuously store the changed data in the loadable portion 141 of the memory subsystem 110.
[0120] For example, the change manager 101 and the database manager 151 may create a simple sort table 185 in a memory portion (e.g., 141) of the memory subsystem 110 and use the simple sort table 185 in the memory portion (e.g., 141) to track changes to the database.
[0121] The change manager 101 may store changed data (e.g., write-ahead log entry 155, simple sort table 185) from a memory portion (e.g., 141) of the memory subsystem 110 to a storage portion (e.g., 143) of the memory subsystem 110.
[0122] In some embodiments, the change manager 101 is at least partially implemented in the memory subsystem 110 (e.g., via the firmware 153 of the memory subsystem 110). The change manager 101 may write changed data from the memory portion (e.g., 141) to the storage portion (e.g., 143) without separately receiving the changed data after the changed data is stored in the memory portion (e.g., 141).
[0123] In some embodiments, after the size of the change data grows to reach or exceed a predetermined threshold, the change manager 101 in the memory subsystem 110 may automatically write at least a portion of the change data in the memory portion (e.g., 141) to a file (e.g., 159 or 189) in the storage portion (e.g., 143). Alternatively, a write command is sent by the change manager 101 in the host system 120 to the memory subsystem 110 using a second protocol (e.g., 147) for storage access; and in response, the change manager 101 in the memory subsystem 110 may write a block 163 of change data from the memory portion (e.g., 141) to a logical block address in the storage portion (e.g., 143).
[0124] In some implementations, the change manager 101 in the memory subsystem 110 and the change manager 101 of the host system 120 may communicate with each other via the connection 103 to move change data between the memory portion (e.g., 141) and the storage portion (e.g., 143). For example, in response to a request from the host system 120, the change manager 101 in the memory subsystem 110 may read the change data from a file (e.g., 159 or 189) to the memory portion (e.g., 141) to be accessed by the host system 120 using a load command. For example, in response to a request from the host system 120, the change manager 101 in the memory subsystem 110 may write the change data to a file (e.g., 159 or 189) in the memory portion (e.g., 141) so that the host system 120 can then use a read command to access the change data in the file (e.g., 159 or 189).
[0125] The changed data in the memory portion (e.g., 141) can be addressed by the host system 120 using the memory address configured in the load instruction and the store instruction; and the changed data in the storage portion (e.g., 143) can be addressed by the host system 120 using the logical block address configured in the read command and the write command.
[0126] Figure 7 An example of a host system failover 231 based on a memory service provided by the memory subsystem 110 according to one embodiment is shown. For example, data identifying changes to a database may be persistently stored using the memory service of the memory subsystem 110. Figure 3 , Figure 4 , Figure 5 and Figure 6 Technology to Figure 1 and Figure 2 The computing system 100 implements failover 231.
[0127] exist Figure 7In [reference document], the memory subsystem 110 is connected to the host system 120 via the connection 103 for normal operation, as described in Figures 1 to 6 In [reference document]. During normal operation, another host system 220 can be disconnected from the memory subsystem 110 (e.g., via a switch). Alternatively, the memory subsystem 110 can have two ports for connecting to the host systems 120 and 220 respectively. During normal operation, the host system 120 uses the readable portion 143 to manage and operate on the database records 157 via the connection 103; and the host system 220 cannot effectively use the memory subsystem 110 through its connection to the memory subsystem 110 (or can effectively use the memory subsystem 110, but for different tasks, such as managing and operating on a different database). Therefore, the failover 231 does not require physically re-wiring the memory subsystem 110, the host systems 120 and 220.
[0128] For example, the host system 120 can be configured to store in the loadable portion 141 of the memory subsystem 110, in addition to the changes identified in the change file 179, in-memory change data 175 indicating changes to the database records 157 stored in the readable portion 143.
[0129] For example, the change data 175 can include Figure 3 write-ahead log entries 155; and the change file 179 can include a log file 159.
[0130] For example, the change data 175 can include Figure 5 a simple sorted table 185; and the change file 179 can include a table file 189.
[0131] The database records 157, the change file 179, and the in-memory change data 175 as a whole identify the valid state of the database operated by the database manager 151 running in the host system 120. When the host system 120 fails, the memory subsystem 110 can be reconnected to the replacement host system 220 during the operation of the failover 231. The database manager 251 running in the replacement host system 220 can be assigned to perform the database operations previously assigned to the replaced host system 120. Since the memory subsystem 110 preserves the latest state of the database via the database records 157, the change file 179, and the in-memory change data 175, the replacement host system 220 can continue the operations previously assigned to the replaced host system 120 with reduced or minimal loss.
[0132] Optionally, during normal operation of the replaced host system 120 during failover 231, a Remote Direct Memory Access (RDMA) network may be connected between host systems 120 and 220. The memory 129 of the replaced host system 120 may have in-memory content, such as cache records 158. During normal operation of the replaced host system 120, the in-memory content may be copied to the replacement host system 220 via the Remote Direct Memory Access (RDMA) network. Thus, when the replaced host system 120 fails, the replacement host system 220 may have a copy of the in-memory content of the replaced host system 120 and thus be ready to continue operating using the memory subsystem 110. However, configuring the operation of the Remote Direct Memory Access (RDMA) network for memory replication can be an expensive option.
[0133] Since the memory subsystem 110 preserves the state of the database operated by the replaced host system 110, the replacement system 220 may reconstruct the cache records 158 from the data in the memory subsystem 110 without copying at least some of the in-memory content of the replaced host system 120, such as the cache records 158.
[0134] Optionally, some of the in-memory content of the host system 120 (such as cache records 158) may be stored by the host system 120 via memory replication into the loadable section 141. Thus, once the memory subsystem 110 is reconnected to the replacement host system 220, the replacement host system 120 may obtain the in-memory content from the loadable section 141 via replication.
[0135] Optionally, the host system 120 may be configured to directly generate the change data 175 (and optionally, other in-memory content) using the memory space provided by the loadable section 141 (and accessed via use of the cache 123 and the cache coherence memory access protocol 145 to improve performance) instead of its memory 129. Thus, the replicated content between the memory 129 of the host system 120 and the loadable section 141 may be reduced or minimized.
[0136] When the memory subsystem 110 is reconnected to the replacement host system 220, there is no need to copy the corresponding content back from the loadable section 141 to the memory of the replacement host system 220, because the replacement host system 220 may use the memory space provided by the loadable section 141 via the connection 103 (such as a Compute Express Link connection), via its cache 123 and the cache coherence memory access protocol (such as 145).
[0137] In some embodiments, both host systems 120 and 220 are active when performing database operations. Some or all of the database operations assigned to host system 120 may be transferred to another host system 220 (e.g., for failover 231, load balancing, maintenance operations, etc.).
[0138] Figure 8 and Figure 9 illustrate an active host system that takes over database operations of another host system using in-memory changed data stored in a memory subsystem. For example, the techniques of using the memory services of memory subsystem 110 to continuously store data identifying changes to a database can be used to implement the techniques described in Figure 3 , Figure 4 , Figure 5 and Figure 6 for computing systems 100 of Figure 1 and Figure 2 . Figure 8 and Figure 9
[0139] In Figure 8 and Figure 9 , host systems 120, …, 220 are connected to memory subsystems 110, …, 210 via interconnect 183. For example, interconnect 183 may include a switch for connecting memory subsystem 110 to host system 120; and in some examples, multiple memory subsystems 210 may be connected to a host system (e.g., 120 or 220).
[0140] For example, host system 120 may run database manager 151 to operate on database records 157 configured to be continuously stored in the readable portion 143 of memory subsystem 110. Changes made to database records (e.g., 157, 158) managed by database manager 151 in host system 120 are recorded in change file 179 stored in the readable portion 143 of memory subsystem 110 and in the loadable portion 141 of memory subsystem 110. The in-memory data 275 stored in the loadable portion 141 of memory subsystem 110 identifies the latest state of database records (e.g., 157, 158) managed by database manager 151.
[0141] Similarly, host system 220 may run database manager 251 to operate on database records hosted on another memory subsystem (such as 210). Database manager 251 running in host system 220 may have its cache records 258. Optionally, a portion of the database operations of database manager 251 may be performed on database records hosted on memory subsystem 110 (such as a different set of database records than records 157 operated on by database manager 151 running in host system 120).
[0142] In Figure 9 , when host system 120 becomes inaccessible (such as when host system 120 fails or disconnects or shuts down), the content stored in its memory 129 (such as cache records 158) becomes inaccessible and may be considered lost. In response to detecting that host system 120 has become inaccessible, the computing system may assign another host system 220 to manage the database records (such as 157) previously managed by the previous host system 120.
[0143] For example, in response to the host system becoming inaccessible, interconnect 183 of the computing system may be configured to reconnect memory subsystem 110 to an active host system 220 instead of the previous host system 120. The active host system 220 may start database manager 151 to manage the database records 157 stored in the readable portion 143 of memory subsystem 110. Memory subsystem 110 attaches the loadable portion 141 as a memory device to the active host system 220 through connection 103 in a similar or identical manner to how memory subsystem 110 attached the loadable portion 141 as a memory device to the previous host system 220 through connection 103. Thus, memory data 275 becomes available to database manager 151 running in host system 220. Based on the in-memory data 275 and change file 179 in the loadable portion 141, database manager 151 may reconstruct cache records 158.
[0144] Optionally, Figure 8 and Figure 9 Typical host systems (such as 120, 220) in a computing system may be configured to run multiple instances of database managers (such as 151 and 251) in a manner similar to how host system 220 runs database managers 151 and 251 to manage database records hosted on multiple memory subsystems (such as 110 and 210). The computing system may perform load balancing by moving the execution of some instances of the database managers among host systems 120,..., 220.
[0145] For example, when Figure 9When the host system 120 reconnects to the interconnect 183 for operation, the computing tasks of running the database manager 151 in the host system 220 can be reassigned to the host system 120 (as in Figure 8 ) to balance the load imposed on the host systems 120, …, 220. For example, the memory devices provided by the loadable portion 141 of the memory subsystem 110 can be disconnected from the host system 220 and connected to the host system 120; similarly, the storage devices provided by the readable portion 143 of the memory subsystem 110 can be disconnected from the host system 220 and connected to the host system 120; and subsequently, the host system 120 can reconstruct the cache record 158 from the in-memory data 275 in the loadable portion 141, the change file 179, and the database record 157 in the readable portion 143.
[0146] Optionally, the in-memory data 275 can include at least a portion of the cache record 158. Thus, the construction of the cache record 158 in the memory 129 of the host system (e.g., 220 or 120) can be performed via a memory copy from the loadable portion 141 to the memory 129 of the host system (e.g., 220 or 120) without sending read commands to access the readable portion 143. In some embodiments, the cache record 158 is configured to be located in the memory device provided by the loadable portion 141; and thus, there is no need to copy the cache record 158 from the loadable portion 141 to the memory 129 of the host system 120.
[0147] Figure 10 Displays a method for transferring database operations between host systems. For example, Figure 10 The method of Figure 1 and Figure 2 The computing system 100 of Figure 3 , Figure 4 , Figure 5 and Figure 6 The technology of Figure 7 in Figure 8 and Figure 9 in
[0148] At block 301, a first portion (e.g., 141) of the memory subsystem 110 is provided to the first host system 120 as a memory device accessible via a first protocol (e.g., 145) and a second portion (e.g., 143) of the memory subsystem 110 is provided to the first host system 120 as a storage device accessible via a second protocol (e.g., 147) through a connection 103 from the host interface 113 of the memory subsystem 110.
[0149] For example, the memory subsystem 110 can be a solid state drive with a host interface 113 having a Compute Express Link connection 103. The memory subsystem 110 can allocate a loadable portion 141 of its fast memory (e.g., 138) as a memory device attached to a host system (e.g., 120 or 220). The memory subsystem 110 can reserve a portion of its fast memory (e.g., 138) as a buffer memory 149 for its processing device (e.g., 117). The memory subsystem 110 can allocate a readable portion 143 of its memory resources (e.g., non-volatile memory 139) as a storage device attached to a host system (e.g., 120 or 220).
[0150] The memory subsystem 110 can have a backup power supply 105, which is designed to ensure that when the power supply of the memory subsystem 110 is interrupted, at least the data stored in the loadable portion 141 implemented using the volatile random access memory 138 is saved in the non-volatile memory 139. Thus, this loadable portion 141 attached to the host system (e.g., 120 or 220) as a memory device can be considered non-volatile in the memory service of the host system (e.g., 120 or 220).
[0151] For example, a first protocol (e.g., 145) can be configured to perform cache coherent memory access via load instructions and store instructions executed in a host system (e.g., 120 or 220); and a second protocol (e.g., 147) can be configured to perform storage access via read commands and write commands executed in the memory subsystem 110.
[0152] For example, a first protocol (e.g., 145) can be configured to identify access locations via memory addresses at a first data granularity (e.g., 32B, 64B, or 128B) according to load instructions and store instructions; and a second protocol (e.g., 147) is configured to identify access locations via logical block addresses at a second data granularity (e.g., 4KB) according to read commands and write commands.
[0153] At block 303, a first host system 110 running a first database manager 151 writes a first database record 157 to a storage device via a second protocol (e.g., 147).
[0154] For example, a database manager 151 running in a first host system 120 may write records 157 of a database to a storage portion (e.g., 143) of the memory subsystem 110 via a connection 103 of a host interface 113 of the memory subsystem 110 using a storage protocol (e.g., 147). To improve performance, the database manager 151 may have cached records 158 in the memory 129 of the first host system 120. Some of the cached records 158 may be new or updated database records that have not yet been stored in the storage portion (e.g., 143) of the memory subsystem 110.
[0155] At block 305, a first host system 110 running a first database manager 151 stores data (e.g., 175 or 275) identifying changes to a second database record 158 to be written to a storage device in a memory device via a first protocol (e.g., 145).
[0156] The database manager 151 may include a change manager 101 configured to generate data 175 identifying changes to the database, such as write-ahead log entries 155, simple sorted tables 185, etc. The change manager 101 may store the change data 175 to a memory portion (e.g., 141) of the memory subsystem 110 via a connection 103 of a host interface 113 of the memory subsystem 110 using a cache coherence memory access protocol (e.g., 145) before changing the database. Thus, the cached records 158 may be reconstructed using the change data 175.
[0157] In some examples, it is desirable to transfer database operations of a first host system 120 to a second host system 120 for failover 231, load balancing, etc. In some examples, the contents of the memory 129 of the first host system 120 may become inaccessible or lost (e.g., when the first host system 120 fails).
[0158] At block 307, a connection 103 of a host interface 113 of the memory subsystem 110 is connected to a second host system 220 separate from the first host system 120 to provide the second host system 220 with access to the memory device via a first protocol (e.g., 145) and access to the storage device via a second protocol (e.g., 147).
[0159] For example, the memory subsystem 110 may attach the loadable portion 141 as a memory device to the second host system 220 in the same manner as attaching a memory device to the first host system 220 before the transfer (e.g., before a failure of the first host system 220). Similarly, the memory subsystem 110 may attach the readable portion 141 as a storage device to the second host system 220 in the same manner as attaching a storage device to the first host system 220 before the transfer (e.g., before a failure of the first host system 220). The second host system 220 may use the memory device and the storage device to start the second database manager (e.g., Figure 7 251 in Figure 9 151 in
[0160] in the same manner as the first database manager 151 uses the memory device and the storage device attached by the memory subsystem 110) before the transfer.
[0161] At block 309, the second host system 220 running the second database manager (e.g., Figure 7 251 in Figure 9 151 in
[0162] loads data (e.g., 175 or 275) identifying changes to the second database records 158 according to a second protocol (e.g., 147). Figure 7 For example, using the data (e.g., 175 or 275) identifying changes to the second database records 158, the second database manager (e.g., Figure 9 251 in
[0163] 151 in
[0163] may reconstruct the second database records 158 that have been or will be generated in the memory of the first host system 120.
[0163] In some embodiments, the first host system 120 stores in the memory device represented by the loadable section 141 not only data identifying changes to the second database record 158, but also cache records 158 that have been generated or updated by the first database manager 151. For example, the first database manager 151 may generate cache records 158 in its memory 129 and perform a memory copy of the cache records 158 from memory 129 to the loadable section 141. Alternatively, the first database manager 151 may be configured to use the loadable section 141 as the memory for the cache records 158 (instead of using its memory 129).
[0164] Thus, the second host system 220 may copy the second database record 158 previously generated or updated by the first host system 120 from the loadable section 141 to its memory, or directly use the second database record 158 stored in the loadable section 141.
[0165] At block 311, the second host system 220 running a second database manager (such as Figure 7 251 in Figure 9 151 in) services database requests based on the first database record 157 and the second database record 158.
[0166] Generally, the memory subsystem 110 may be a storage device, a memory module, or a combination of a storage device and a memory module. Examples of storage devices include solid state drives (SSDs), flash drives, universal serial bus (USB) flash drives, embedded multimedia controllers (eMMC) drives, universal flash storage (UFS) drives, secure digital (SD) cards, and hard disk drives (HDDs). Examples of memory modules include dual in-line memory modules (DIMMs), small outline DIMMs (SO-DIMMs), and various types of non-volatile dual in-line memory modules (NVDIMMs).
[0167] The computing system 100 may be a computing device such as a desktop computer, a laptop computer, a network server, a mobile device, part of a vehicle (such as an airplane, a drone, a train, an automobile, or other transportation vehicle), an Internet of Things (IoT) enabled device, an embedded computer (such as an embedded computer included in a vehicle, industrial equipment, or a networked commercial device), or such a computing device that includes memory and a processing device.
[0168] The computing system 100 may include a host system 120 coupled to one or more memory subsystems 110. Figure 1Describe an example of a host system 120 coupled to a memory subsystem 110. As used herein, "coupled to" or "coupled with" generally refers to a connection between components, which can be an indirect communication connection or a direct communication connection (e.g., without an intermediary component), whether wired or wireless, including connections such as electrical, optical, magnetic, etc.
[0169] For example, the host system 120 can include a processor chipset (e.g., processing device 127) and a software stack executed by the processor chipset. The processor chipset can include one or more cores, one or more caches (e.g., 123), a memory controller (e.g., controller 125) (e.g., an NVDIMM controller), and a storage protocol controller (e.g., a PCIe controller, a SATA controller). The host system 120 uses the memory subsystem 110 to write data to the memory subsystem 110 and read data from the memory subsystem 110, for example.
[0170] The host system 120 can be coupled to the memory subsystem 110 via a physical host interface 113. Examples of physical host interfaces include (but are not limited to) Serial Advanced Technology Attachment (SATA) interfaces, Peripheral Component Interconnect Express (PCIe) interfaces, Universal Serial Bus (USB) interfaces, Fibre Channel, Serial Attached SCSI (SAS) interfaces, Double Data Rate (DDR) memory bus interfaces, Small Computer System Interface (SCSI), Dual In-line Memory Module (DIMM) interfaces (e.g., DIMM socket interfaces that support Double Data Rate (DDR)), Open NAND Flash Interface (ONFI), Double Data Rate (DDR) interfaces, Low Power Double Data Rate (LPDDR) interfaces, Compute Express Link (CXL) interfaces, or any other interface. The physical host interface can be used to transfer data between the host system 120 and the memory subsystem 110. When the memory subsystem 110 is coupled to the host system 120 via a PCIe interface, the host system 120 can further utilize the Non-Volatile Memory Express (NVMe) interface to access components (e.g., memory device 109). The physical host interface can provide an interface for transferring control, address, data, and other signals between the memory subsystem 110 and the host system 120. Figure 1 Describe the memory subsystem 110 as an example. Generally, the host system 120 can access multiple memory subsystems via the same communication connection, multiple separate communication connections, and / or a combination of communication connections.
[0171] The processing device 127 of the host system 120 can be, for example, a microprocessor, a central processing unit (CPU), a processing core of a processor, an execution unit, etc. In some examples, the controller 125 can be referred to as a memory controller, a memory management unit, and / or a starter. In one instance, the controller 125 controls communication via a bus coupled between the host system 120 and the memory subsystem 110. Generally, the controller 125 can send commands or requests to the memory subsystem 110 to perform desired accesses to the memory devices 109, 107. The controller 125 can further include interface circuitry for communicating with the memory subsystem 110. The interface circuitry can convert responses received from the memory subsystem 110 into information for the host system 120.
[0172] The controller 125 of the host system 120 can communicate with the controller 115 of the memory subsystem 110 to perform operations such as reading data, writing data, or erasing data at the memory devices 109, 107 and other such operations. In some examples, the controller 125 is integrated within the same package as the processing device 127. In other examples, the controller 125 is separate from the package of the processing device 127. The controller 125 and / or the processing device 127 can include hardware such as one or more integrated circuits (ICs) and / or discrete components, buffer memories, cache memories, or combinations thereof. The controller 125 and / or the processing device 127 can be a microcontroller, application specific logic circuitry (such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.), or another suitable processor.
[0173] The memory devices 109, 107 can include any combination of different types of non-volatile memory components and / or volatile memory components. A volatile memory device (such as the memory device 107) can be (but is not limited to) a random access memory (RAM), such as a dynamic random access memory (DRAM) and a synchronous dynamic random access memory (SDRAM).
[0174] Some examples of non-volatile memory components include NAND (or NOT AND) type flash memory and write-in-place memory, such as three-dimensional cross-point (“3D cross-point”) memory. A cross-point non-volatile memory array can perform bit storage based on a change in bulk resistance along with a stacked cross-grid format data access array. Additionally, compared to many flash-based memories, cross-point non-volatile memory can perform write-in-place operations, where non-volatile memory cells can be programmed without first erasing the non-volatile memory cells. NAND type flash memory includes, for example, two-dimensional NAND (2D NAND) and three-dimensional NAND (3D NAND).
[0175] Each of the memory devices 109 may include one or more memory cell arrays. One type of memory cell, such as a single-level cell (SLC), may store one bit per cell. Other types of memory cells, such as multi-level cells (MLC), triple-level cells (TLC), quad-level cells (QLC), and penta-level cells (PLC), may store multiple bits per cell. In some embodiments, each of the memory devices 109 may include one or more memory cell arrays, such as SLC, MLC, TLC, QLC, PLC, or any combination thereof. In some embodiments, a particular memory device may include an SLC portion, an MLC portion, a TLC portion, a QLC portion, and / or a PLC portion of memory cells. The memory cells of the memory devices 109 may be grouped into pages, which may refer to a logical unit of the memory device for storing data. With respect to some types of memories, such as NAND, pages may be grouped to form blocks.
[0176] Although non-volatile memory devices such as 3D cross-point type and NAND type memories (such as 2D NAND, 3D NAND) are described, the memory devices 109 may be based on any other type of non-volatile memory, such as read-only memory (ROM), phase change memory (PCM), self-selecting memory, other chalcogenide-based memories, ferroelectric transistor random access memory (FeTRAM), ferroelectric random access memory (FeRAM), magnetic random access memory (MRAM), spin transfer torque (STT)-MRAM, conductive bridge RAM (CBRAM), resistive random access memory (RRAM), oxide-based RRAM (OxRAM), nor flash memory, and electrically erasable programmable read-only memory (EEPROM).
[0177] The memory subsystem controller 115 (or simply referred to as the controller 115) may communicate with the memory devices 109 to perform operations such as reading data, writing data, or erasing data at the memory devices 109 and other such operations (e.g., in response to commands scheduled on the command bus by the controller 125). The controller 115 may include hardware such as one or more integrated circuits (ICs) and / or discrete components, buffer memory, or a combination thereof. The hardware may include digital circuitry having dedicated (i.e., hard-coded) logic to perform the operations described herein. The controller 115 may be a microcontroller, dedicated logic circuitry (such as a field programmable gate array (FPGA), application specific integrated circuit (ASIC), etc.), or another suitable processor.
[0178] The controller 115 may include a processing device 117 (processor) configured to execute instructions stored in a local memory 119. In the illustrated example, the local memory 119 of the controller 115 includes an embedded memory configured to store instructions for various processes, operations, logic flows, and routines for performing operations of the memory subsystem 110, including handling communications between the memory subsystem 110 and the host system 120.
[0179] In some embodiments, the local memory 119 may include memory registers for storing memory pointers, fetching data, etc. The local memory 119 may also include a read-only memory (ROM) for storing microcode. Although Figure 1 the illustrated example memory subsystem 110 has been shown to include a controller 115, in another embodiment of the present disclosure, the memory subsystem 110 does not include a controller 115 and instead may rely on external control (e.g., provided by an external host or by a processor or controller separate from the memory subsystem).
[0180] Generally, the controller 115 may receive commands or operations from the host system 120 and may convert the commands or operations into instructions or appropriate commands to achieve desired access to the memory device 109. The controller 115 may be responsible for other operations such as wear leveling operations, garbage collection operations, error detection and error correction code (ECC) operations, encryption operations, cache operations, and address translation between logical addresses (e.g., logical block addresses (LBAs), namespaces) and physical addresses (e.g., physical block addresses) associated with the memory device 109. The controller 115 may further include host interface circuitry for communicating with the host system 120 via a physical host interface. The host interface circuitry may convert commands received from the host system into command instructions to access the memory device 109 and also convert responses associated with the memory device 109 into information for the host system 120.
[0181] The memory subsystem 110 may also include additional circuitry or components not shown. In some embodiments, the memory subsystem 110 may include a cache or buffer (e.g., DRAM) and address circuitry (e.g., row decoders and column decoders) that may receive addresses from the controller 115 and decode the addresses to access the memory device 109.
[0182] In some embodiments, the memory device 109 includes a local media controller 137 that operates with the memory subsystem controller 115 to perform operations on one or more memory cells of the memory device 109. An external controller, such as the memory subsystem controller 115, may externally manage the memory device 109 (e.g., perform media management operations on the memory device 109). In some embodiments, the memory device 109 is a managed memory device, which is an original memory device combined with a local controller (such as the local media controller 137) for media management within the same memory device package. An example of a managed memory device is a managed NAND (MNAND) device.
[0183] In one embodiment, an example machine of a computer system in which a set of instructions for causing a machine to perform any one or more of the methodologies discussed herein may be executed. In some embodiments, the computer system may correspond to a host system (e.g., Figure 1 the host system 120) that includes, is coupled to, or utilizes a memory subsystem (e.g., Figure 1 the memory subsystem 110) or may be used to perform the operations discussed above (e.g., execute instructions to perform operations corresponding to those described with reference to Figure 1 ). In alternative embodiments, the machine may be connected (e.g., networked) to other machines on a LAN, intranet, extranet, and / or the Internet. The machine may operate as a server or client machine in a client - server network environment, as a peer machine in a peer - to - peer (or distributed) network environment, or as a server or client machine in a cloud computing infrastructure or environment.
[0184] The machine can be a personal computer (PC), tablet PC, set - top box (STB), personal digital assistant (PDA), cellular phone, network device, server, network router, switch or bridge, network - attached storage facility, or any machine capable of executing a set of instructions (sequentially or otherwise) that specify actions to be taken by the machine. Further, although a single machine is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein.
[0185] An example computer system includes a processing device, a main memory (such as read - only memory (ROM), flash memory, dynamic random access memory (DRAM) (such as synchronous DRAM (SDRAM) or Rambus DRAM (RDRAM), static random access memory (SRAM), etc.), and a data storage system that communicate with each other via a bus (which may include multiple buses).
[0186] The processing device represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, or the like. More specifically, the processing device can be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, or a processor implementing other instruction sets or several processors implementing a combination of instruction sets. The processing device can also be one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, or the like. The processing device is configured to execute instructions for performing the operations and steps discussed herein. The computer system can further include a network interface device for communicating over a network.
[0187] The data storage system can include a machine-readable medium (also referred to as a computer-readable medium) on which one or more sets of instructions or software embodying any one or more of the methodologies or functions described herein are stored. The instructions can also reside, completely or at least partially, within the main memory and within the processing device during execution by the computer system, the main memory and the processing device also constituting machine-readable storage media. The machine-readable medium, the data storage system, and / or the main memory can correspond to Figure 1 the memory subsystem 110.
[0188] In one embodiment, the instructions include instructions for implementing the functionality discussed above (e.g., the operations described with reference to Figure 1 ). Although the machine-readable medium is shown as a single medium in the exemplary embodiment, the term "machine-readable storage medium" should be considered to include a single medium or multiple media storing one or more sets of instructions. The term "machine-readable storage medium" should also be considered to include any medium that is capable of storing or encoding a set of instructions executable by a machine and that causes the machine to perform any one or more of the methodologies of the present disclosure. The term "machine-readable storage medium" should accordingly be considered to include, without limitation, solid-state memory, optical media, and magnetic media.
[0189] Some portions of the foregoing detailed description have been presented in terms of algorithms and symbolic representations of operations on data bits within a computer memory. These algorithmic descriptions and representations are the means used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. An algorithm is here, and generally, conceived to be a self-consistent sequence of operations leading to a desired result. The operations are those requiring physical manipulation of physical quantities. Usually, though not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, principally for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.
[0190] However, it should be remembered that all such and similar terms are to be associated with appropriate physical quantities and are merely convenient labels applied to these quantities. The present disclosure may relate to the actions and processes of a computer system or similar electronic computing device that manipulates and transforms data represented as physical (electronic) quantities within the registers and memories of the computer system into other data similarly represented as physical quantities within the memories or registers of the computer system or other such information storage systems.
[0191] The present disclosure also relates to an apparatus for performing the operations herein. This apparatus may be specially constructed for the intended purpose, or it may comprise a general purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer readable storage medium, such as (but not limited to) any type of disk (including floppy disks, optical disks, CD-ROMs, and magneto-optical disks), read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of media suitable for storing electronic instructions, each coupled to the computer system bus.
[0192] The algorithms and displays presented herein are not inherently related to any particular computer or other device. Various general purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct more specialized devices to perform the method. The structure of various such systems will appear as described in the following description. Additionally, the present disclosure is not described with reference to any particular programming language. It will be understood that various programming languages may be used to implement the teachings of the present disclosure described herein.
[0193] The present disclosure may be provided as a computer program product or software that may include a machine readable medium having stored thereon instructions that may be used to program a computer system (or other electronic device) to perform a process in accordance with the present disclosure. The machine readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, the machine readable (e.g., computer readable) medium includes a machine (e.g., computer) readable storage medium such as read only memory (“ROM”), random access memory (“RAM”), magnetic disk storage media, optical storage media, flash memory components, etc.
[0194] In this description, various functions and operations are described as being performed or caused by computer instructions to simplify the description. However, those skilled in the art will recognize that such expressions mean that the functions are produced by a computer executing instructions by one or more controllers or processors, such as a microprocessor. Alternatively or in combination, the functions and operations may be implemented using dedicated circuitry, with or without software instructions, such as using an application specific integrated circuit (ASIC) or a field programmable gate array (FPGA). Embodiments may be implemented using hardwired circuitry without software instructions or in combination with software instructions. Thus, the techniques are not limited to any specific combination of hardware circuitry and software, nor to any specific source of instructions executed by a data processing system.
[0195] In the foregoing description, embodiments of the present disclosure have been described with reference to specific example embodiments of the present disclosure. It should be understood that various modifications may be made to the present disclosure without departing from the broader spirit and scope of the embodiments of the present disclosure as set forth in the appended claims. The specification and drawings should accordingly be regarded as illustrative rather than restrictive.
Claims
1. A method, comprising: providing a first portion of the memory subsystem to a first host system as a memory device accessible via a first protocol and a second portion of the memory subsystem to the first host system as a storage device accessible via a second protocol through a connection of a host interface of the memory subsystem; writing a first database record into the storage device by the first host system running a first database manager via the second protocol; storing data identifying a change to a second database record to be written into the storage device into the memory device by the first host system running the first database manager via the first protocol; connecting the connection of the host interface of the memory subsystem to a second host system separate from the first host system to provide the second host system with access to the memory device via the first protocol and access to the storage device via the second protocol; loading, by the second host system running a second database manager, the data identifying the change to the second database record according to the second protocol; and serving a database request by the second host system running the second database manager based on the first database record and the second database record.
2. The method of claim 1, wherein the connection is a Compute Express Link connection.
3. The method of claim 2, wherein connecting the connection of the host interface of the memory subsystem to the second host system is in response to a failure of the first host system.
4. The method of claim 3, wherein the first protocol is configured to perform cache coherent memory access via load instructions and store instructions; and the second protocol is configured to perform storage access via read commands and write commands.
5. The method of claim 3, wherein the first protocol is configured to identify an access location via a memory address at a first data granularity; and the second protocol is configured to identify an access location via a logical block address at a second data granularity.
6. The method of claim 3, further comprising: reconstructing, by the second host system, the second database record based on the data identifying the change to the second database record.
7. The method of claim 6, wherein the data identifying the change to the second database record includes write-ahead log entries.
8. The method of claim 6, wherein the data identifying the change to the second database record includes a simple sorted table.
9. The method of claim 6, wherein when the first host system fails, the storage device provided by the memory subsystem does not contain data representing the second database record.
10. The method of claim 3, further comprising: Store at least a portion of the second database record into the memory device by the first host system running the first database manager via the first protocol; and Load the portion of the second database record by the second host system running the second database manager according to the second protocol.
11. A host system, comprising: A memory configured to store instructions representing a database manager; A cache; and A processing device configured to: Execute the instructions to run the database manager; Communicate with the host system through a connection of a host interface from a memory subsystem to: Access a memory device attached to the host system by the memory subsystem via the cache via a first protocol for cache coherent memory access; and Access a storage device attached to the host system by the memory subsystem via a second protocol for storage access; Execute a load instruction to access data identifying a change to a second database record according to the second protocol; and Serve database requests based on a first database record and the second database record stored in the storage device.
12. The host system according to claim 11, wherein the connection is a Compute Express Link connection.
13. The host system according to claim 11, wherein the processing device is further configured to: Reconstruct the second database record based on the data identifying the change to the second database record.
14. The host system according to claim 13, wherein the data identifying the change to the second database record includes a write-ahead log entry or a simple sorted table.
15. The host system according to claim 14, wherein when the storage device is attached to the host system by the memory subsystem, the storage device provided by the memory subsystem does not contain data representing the second database record.
16. The host system according to claim 11, wherein the processing device is further configured to: Load at least a portion of the second database record from the memory device attached to the host system by the memory subsystem according to the second protocol.
17. A non-transitory computer storage medium storing instructions that, when executed in a computing system, cause the computing system to perform a method including: Run a database manager in a host system of the computing system; Communicate with the host system through a connection of a host interface from a memory subsystem to: Access a memory device attached to the host system by the memory subsystem via a first protocol for cache coherent memory access; and Access a storage device attached to the host system by the memory subsystem via a second protocol for storage access; Execute a load instruction to access data identifying a change to a second database record according to the second protocol; and Serve database requests based on a first database record and the second database record stored in the storage device.
18. The non-transitory computer storage medium according to claim 17, wherein the connection is a Compute Express Link connection.
19. The non-transitory computer storage medium according to claim 17, wherein the method further comprises: reconstructing the second database record based on the data identifying the change of the second database record.
20. The non-transitory computer storage medium according to claim 17, wherein the method further comprises: executing a load instruction to access at least a portion of the second database record from the memory device attached to the host system by the memory subsystem according to the second protocol.