Intelligent storage integrated memory storage architecture
Through the integrated memory storage architecture of intelligent memory, the synergy between computed memory semantics solid-state drives and host operating systems, combined with the unified abstract interface layer, the problems of storage performance bottlenecks and resource waste in the existing technology are solved, and an efficient and intelligent memory-storage continuum is realized to meet the complex needs of data-intensive applications.
Patent Information
- Application Number
- CN202510655628.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing storage systems, memory management and operating system mechanisms have failed to fully utilize the potential of new hardware technologies, resulting in performance bottlenecks and waste of resources, such as high latency, write amplification, interface mismatch, CPU pause, inefficient cache efficiency and insufficient hardware utilization.
It provides an integrated memory storage architecture of intelligent memory, including a computed memory semantic solid-state drive, a host operating system and a unified abstract interface layer. Through a unified memory address space, a flexible data placement mechanism that perceives semantics, an embedded computing engine and a semantic-aware intelligent scheduling and offload decision mechanism, a unified, adaptive, and computationally enhanced memory-storage continuum is built.
Overcome the multiple impedance mismatch problems of traditional hierarchical storage in performance, cost, energy efficiency and intelligent processing, and meet the needs of next-generation data-intensive applications for ultra-large capacity, extremely low latency, ultra-high throughput and deep embedded intelligent processing.
Smart Images

Figure CN120216455A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technologies, and particularly to an intelligent memory-storage integrated architecture. Background Art
[0002] Currently, data centers and edge computing are facing an increasingly acute contradiction between the explosive growth of data volume and the performance, capacity, cost, and energy efficiency of storage / memory. New interconnect technologies such as Compute Express Link provide a basis for memory expansion and memoryization of storage devices (such as memory-semantic solid-state drives), enabling solid-state drives to provide access semantics close to memory and huge capacity expansion at a lower cost. At the same time, the development trend of compute-storage technology is to sink some computing tasks to be executed inside the storage device to reduce data movement, latency, and power consumption. However, most of the existing storage systems, memory management, and operating system mechanisms are still based on traditional block interfaces and the main-processor-centered computing model, failing to fully exploit the potential of these new hardware technologies, resulting in performance bottlenecks (such as high latency, write amplification, interface mismatch) and resource waste (such as CPU stalls, low cache efficiency, insufficient hardware utilization). Summary of the Invention
[0003] The intelligent memory-storage integrated architecture provided by this application can construct a unified, adaptive, and compute-enhanced Memory-Storage Continuum to overcome the multiple impedance mismatch problems of traditional hierarchical storage in terms of performance (speed, latency), cost, energy efficiency, and intelligent processing, thereby meeting the comprehensive requirements of next-generation data-intensive applications (such as large-scale language model training / inference, real-time data analysis, high-performance key-value storage) for ultra-large capacity, extremely low latency, ultra-high throughput, and deeply embedded intelligent processing.
[0004] In a first aspect, this application provides an intelligent memory-storage integrated architecture, which includes: a compute-type memory-semantic solid-state drive, configured to provide a unified memory address space based on the Compute Express Link protocol, support byte-granularity access by the host through load / store instructions; and is provided with an internal intelligent hierarchical and granularity management mechanism; and is provided with a flexible data placement mechanism that perceives semantics, and is provided with an embedded computing engine; a host operating system, connected to the compute-type memory-semantic solid-state drive, configured to receive long-latency prompts sent by the compute-type memory-semantic solid-state drive and perform context switching; and is provided with a unified abstract interface layer to uniformly manage byte-granularity memory access to the compute-type memory-semantic solid-state drive, the transmission of data placement semantic prompts, and the request to submit offloading tasks to the embedded computing engine, and is provided with a semantics-aware intelligent scheduling and offloading decision-making mechanism.
[0005] Among them, the computational memory semantic solid-state drive includes: a unified interconnection and memory semantic module, which is used to provide a unified memory address space based on the computational expression link protocol and support the host to access at the byte granularity through load / store instructions; an internal intelligent hierarchical and granularity management mechanism module, which constructs a fine-grained write log at the cache line granularity to absorb byte writes and reduce write amplification; and simultaneously maintains a page granularity read / data cache to utilize spatial locality; and is responsible for managing the merging of logs, garbage collection, and the flow of data between the internal cache and the flash memory; a semantic-aware flexible data placement mechanism module, which is used to utilize data semantic hints and combine the status of its internal write log / data cache to place data with different semantics in the optimal physical area inside the flash memory; an embedded computing engine, which is used to perform hardware acceleration for specific tasks; where the specific tasks include at least one of data compression / decompression, data encryption / decryption, data filtering / aggregation, index assistance, and format conversion.
[0006] Among them, the unified interconnection and memory semantic module is also used to provide the host with a logically continuous and extended physical address space through the mapping of CXL.mem.
[0007] Among them, the internal intelligent hierarchical and granularity management mechanism module is also used to directly and quickly append data to the write log area of the internal DRAM when the host issues a store instruction of the cache line size, and find the latest data copy at a specific address in the log through an optimized index structure; and new writes will logically overwrite the old version at the same address by updating the index.
[0008] Among them, when the host read request misses in the write log and data needs to be obtained from the flash memory, the internal intelligent hierarchical and granularity management mechanism module will read the entire flash page into the page granularity buffer area, and then extract the required cache line from it and return it to the host; and when the data in the write log forms a complete data page ready to be written back to the flash memory after background processing, the data page will be temporarily stored in the page granularity buffer area.
[0009] Among them, the data semantic hints include at least one of data life cycle, access pattern, hot / cold attribute, and object size category.
[0010] Among them, the semantic-aware flexible data placement mechanism module uses the recovery unit provided by the flexible data placement mechanism in the NVMe standard as the basic management unit for physical placement.
[0011] Among them, the host operating system includes: a cooperative context switching module, which is used to receive long-latency hints sent by the computational memory semantic solid-state drive and perform context switching; a semantic-aware intelligent scheduling and offloading decision-making mechanism module, which is used to make data placement decisions, computational offloading decisions, and task coordination.
[0012] Among them, the semantic-aware intelligent scheduling and offloading decision-making mechanism module includes: a data placement decision-making unit, which is used to generate corresponding semantic prompts and transmit them to the compute-in-memory semantic solid-state drive after making data placement decisions; a compute offloading decision-making unit, which is used to identify tasks suitable for offloading to the embedded computing engine and submit tasks through a unified interface; and a task coordination unit, which is used to coordinate the task execution between the host and the embedded computing engine.
[0013] Among them, the data placement decision-making unit is used to make data placement decisions based on explicit application instructions, runtime intelligent analysis, and system status feedback; the compute offloading decision-making unit is used to perform task offloading based on task characteristic analysis, system real-time status, and / or performance / energy efficiency modeling; the task coordination unit is used to adopt preprocessing / postprocessing, pipelining, and / or asynchronous execution collaboration modes.
[0014] The beneficial effects of this application are as follows: Different from the prior art, this application provides an intelligent memory integrated storage architecture, which includes: a compute-in-memory semantic solid-state drive, which is used to provide a unified memory address space based on the compute expression link protocol, support the host to access at the byte granularity through load / store instructions; and is provided with an internal intelligent hierarchical and granularity management mechanism; and is provided with a semantic-aware flexible data placement mechanism, and is provided with an embedded computing engine; a host operating system, connected to the compute-in-memory semantic solid-state drive, which is used to receive long-delay prompts sent by the compute-in-memory semantic solid-state drive and perform context switching; and is provided with a unified abstract interface layer, which uniformly manages byte-granularity memory access to the compute-in-memory semantic solid-state drive, the transmission of data placement semantic prompts, and the request to submit offloading tasks to the embedded computing engine, and sets up a semantic-aware intelligent scheduling and offloading decision-making mechanism. That is, the intelligent memory integrated storage architecture provided by this application can build a unified, adaptive, and compute-enhanced Memory-Storage Continuum to overcome the multiple impedance mismatch problems of traditional hierarchical storage in terms of performance (speed, latency), cost, energy efficiency, and intelligent processing, so as to meet the comprehensive requirements of next-generation data-intensive applications (such as large-scale language model training / inference, real-time data analysis, high-performance key-value storage) for ultra-large capacity, extremely low latency, ultra-high throughput, and deep embedded intelligent processing. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them: Figure 1 It is a schematic structural diagram of an embodiment of the intelligent storage integrated memory storage architecture provided by this application; Figure 2 It is a schematic structural diagram of an embodiment of the computing memory semantic solid-state drive provided by this application; Figure 3 It is a schematic structural diagram of an embodiment of the host operating system provided by this application; Figure 4 It is a schematic structural diagram of an embodiment of the semantic perception intelligent scheduling and offloading decision-making mechanism module provided by this application. Detailed implementation manners
[0016] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. It can be understood that the specific embodiments described herein are only used to explain this application, rather than limiting this application. In addition, it should be noted that for the sake of description, only parts related to this application rather than all structures are shown in the accompanying drawings. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0017] Referring to "embodiment" in this context means that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0018] Refer to Figure 1 , Figure 1 It is a schematic structural diagram of an embodiment of the intelligent storage integrated memory storage architecture provided by this application. This intelligent storage integrated memory storage architecture 100 includes: a computing memory semantic solid-state drive 10 and a host operating system 20.
[0019] The computing memory semantic solid-state drive 10 is used to provide a unified memory address space based on the computing expression link protocol, support the host to perform byte-level access through load / store instructions; and is provided with an internal intelligent hierarchical and granularity management mechanism; and is provided with a semantic-aware flexible data placement mechanism, and is provided with an embedded computing engine.
[0020] Refer to Figure 2 , the computing memory semantic solid-state drive 10 includes: a unified interconnection and memory semantics module 11, an internal intelligent hierarchical and granularity management mechanism module 12, a semantic-aware flexible data placement mechanism module 13, and an embedded computing engine 14.
[0021] The Unified Interconnect and Memory Semantics Module 11 is used to provide a unified memory address space based on the Compute Express Link protocol, supporting byte-granularity access by the host through load / store instructions.
[0022] Among them, the Unified Interconnect and Memory Semantics Module 11 is also used to provide a logically continuous and extended physical address space to the host through the mapping of CXL.mem.
[0023] In some embodiments, the Unified Interconnect and Memory Semantics Module 11 mainly innovates memory access by using the Compute Express Link protocol.
[0024] In traditional computer architectures, the main memory (DRAM) and external storage (such as solid-state drive SSD) are like two independent "worlds". The CPU directly accesses the DRAM in bytes through the low-latency memory bus using load / store instructions, which is what this application calls "memory semantics". Accessing the SSD, on the other hand, requires going through the I / O bus (such as PCIe), relying on the intervention of the operating system and drivers, and data transfer is carried out through read / write commands in units of blocks (usually 4KB or larger), following the "storage semantics". These two very different interfaces, access methods, and data granularities result in a significant performance gap and software processing complexity, forming an "impedance mismatch" in the system.
[0025] The Compute Express Link (CXL) protocol, especially its CXL.mem component, is a key technology designed to break this gap. It gives PCIe devices like SSDs a revolutionary ability: to map their internal storage resources (partially or fully) to the host's physical memory address space. This means that from the perspective of the CPU and the operating system, this part of the SSD storage area seems to be part of the system memory, although their physical locations and performance characteristics (such as latency) are different from those of the DRAM.
[0026] The core changes are reflected in the following aspects: unified address space, memory semantics access, and byte-granularity access.
[0027] The unified address space is mainly reflected in that through the mapping of CXL.mem, the host CPU obtains a logically continuous and extended physical address space, seamlessly including the original DRAM and the newly mapped SSD storage area. This greatly simplifies data management for applications that require extremely large amounts of memory, and software no longer needs to deliberately distinguish whether data is stored in the DRAM or the mapped SSD.
[0028] Memory semantic access is mainly reflected in: The most crucial change lies in the access method. Once the SSD storage is mapped, the CPU can directly use its native load / store instructions to access this space, completely abandoning the complex and costly I / O processes (including operating system calls, driver intervention, NVMe command sending and processing, interrupt handling, etc.) when accessing traditional SSDs. Load / store instructions are extremely lightweight operations directly executed by the CPU.
[0029] Byte-granularity access is mainly reflected in: Load / store instructions natively support precise reading and writing of memory addresses at the byte or word level. This is in sharp contrast to the operations based on fixed block sizes in traditional SSD interfaces. For processing fine-grained data structures such as metadata, small objects, and partial updates of database logs, byte-granularity access is crucial and can significantly improve efficiency.
[0030] Contribution to overcoming impedance mismatch: This CXL-based unified interconnection and memory semantics make a fundamental contribution to overcoming the multiple impedance mismatch problems of traditional hierarchical storage: such as performance improvement, cost-effectiveness, energy efficiency optimization, and the basis for intelligent processing.
[0031] Performance improvement is mainly reflected in: Reducing interface latency, improving fine-grained access efficiency, and laying the foundation for latency hiding.
[0032] Reducing interface latency is mainly reflected in: Load / store instructions bypass the lengthy traditional I / O software stack, greatly reducing the "interface" latency and software overhead of accessing SSDs. Although the physical flash read / write latency remains the main bottleneck, the shortening of the interaction path brings significant performance improvement.
[0033] Improving fine-grained access efficiency is mainly reflected in: Byte-granularity access avoids unnecessary "read-modify-write" whole-block operations when updating small data, reduces write amplification, and improves access efficiency.
[0034] Laying the foundation for latency hiding is mainly reflected in: The memory semantic interface enables the CPU and OS to more precisely track which instruction is waiting for data at which memory address, which is a prerequisite for implementing advanced latency hiding techniques such as "cooperative context switching".
[0035] Cost-effectiveness is mainly reflected in: Achieving low-cost memory expansion. Its core contribution is that it makes it possible to use SSDs with a much lower cost per GB than DRAM to expand the system memory capacity. Application programs can place large data sets (such as AI model parameters, large database caches) that originally had to be placed in expensive DRAM in the SSD area with memory semantics, significantly reducing the hardware cost of building a large memory system.
[0036] The energy efficiency optimization is mainly reflected in: reducing the transmission of invalid data and simplifying the software stack overhead.
[0037] Reducing the transmission of invalid data is mainly reflected in: byte-granularity access allows only the data that is truly needed to be transmitted, avoiding a large amount of invalid data movement caused by reading and writing the entire block under the traditional block interface, thereby saving the energy consumption of the bus and device interface.
[0038] Simplifying the software stack overhead is mainly reflected in: a lighter access method reduces the burden on the CPU in executing I / O-related software instructions, indirectly reducing the CPU power consumption.
[0039] The basis of intelligent processing is mainly reflected in: providing memory operation primitives. At this time, the CPU can directly perform memory operations such as pointer arithmetic, traversing complex data structures, and in-place modification in the mapped SSD area. This provides the necessary basic operation primitives for running more complex algorithms (such as graph algorithms and in-place data conversion) that require fine-grained random access directly on the SSD, making it possible to "sink" more intelligent processing logic closer to the data, which is difficult to efficiently support by the traditional block interface.
[0040] Summary: "Unified interconnection and memory semantics" is the cornerstone of the "Intelligent Memory Integrated Memory Storage Architecture 100". It uses the CXL protocol to seamlessly integrate the SSD into the host's memory map, replaces the traditional cumbersome block I / O operations with the CPU's native and efficient load / store instructions, and endows the key byte-granularity access capability. This transformation directly optimizes the performance (reducing interface latency and software overhead), opens the door to low-cost large-capacity memory expansion, improves energy efficiency by reducing the transmission of invalid data, and lays the foundation for more refined and intelligent data processing on storage devices. Although it cannot eliminate the latency of the physical medium itself, it completely changes the game rules of the host-storage interaction, paving the way for subsequent deeper internal cache optimization, intelligent data placement, compute offloading, and collaborative scheduling and other innovative technologies.
[0041] The internal intelligent hierarchical and granularity management mechanism module 12 constructs a cache-line-granularity fine-grained write log to absorb byte writes and reduce write amplification; at the same time, maintains a page-granularity read / data cache to utilize spatial locality; and is responsible for managing the merging of logs, garbage collection, and the flow of data between the internal cache and flash memory.
[0042] In some embodiments, the internal intelligent hierarchical and granularity management mechanism module 12 is also used to directly and quickly append data to the write log area of the internal DRAM when the host issues a cache-line-sized store instruction, and find the latest data copy at a specific address in the log through an optimized index structure; and new writes logically overwrite the old version at the same address by updating the index.
[0043] In some embodiments, when the host read request misses in the write log and data needs to be fetched from the flash memory, the internal intelligent tiering and granularity management mechanism module 12 reads the entire flash memory page into the page-granularity buffer, and then extracts the required cache line from it and returns it to the host; and when the data in the write log forms a complete data page ready to be written back to the flash memory after background processing, the data page is temporarily stored in the page-granularity buffer.
[0044] In some embodiments, the internal intelligent tiering and granularity management (firmware level) is mainly reflected in: redesigning the internal dynamic random access memory (DRAM) cache management of the solid-state drive. Building a fine-grained write log at the cache line granularity for quickly absorbing byte writes and reducing write amplification; while maintaining a page-granularity read / data cache to utilize spatial locality. The firmware is responsible for managing the merging of the log, garbage collection, and the flow of data between the internal cache and the flash memory, effectively solving the problem of the granularity mismatch between the host byte access and the flash memory page access.
[0045] Modern host CPUs can access the Compute Memory Semantics Solid-State Drive 10 (CMS-SSD) at a fine byte granularity (usually a 64-byte cache line) through the Compute Express Link (CXL) protocol. However, the core storage medium of the CMS-SSD, NAND flash memory, can only be read and written at a larger page (usually 4KB to 16KB) at the physical level. This huge difference between the host access requirement (small granularity) and the flash memory physical limitation (large granularity) is the core problem encountered by traditional SSD designs when facing memory semantics access.
[0046] If left unaddressed and simply mapping the host's byte write requests directly to flash memory page operations will lead to serious consequences: Such as extremely high write amplification: Even if the host only wants to modify a small part (such as 64 bytes) within the page, the SSD must first read the entire flash memory page into the internal cache, modify the corresponding bytes, and then write the entire modified page to a new location in the flash memory (because flash memory cannot be updated in-place). This means that a tiny write may trigger thousands of bytes of flash memory read and write operations, greatly increasing the amount of data written, which is the so-called "write amplification".
[0047] Such as performance bottlenecks: This complex process of "read-modify-write" is not only inefficient but also occupies the flash memory channel for a long time, seriously blocking subsequent write requests and becoming a performance bottleneck.
[0048] Solution: Firmware-level intelligent tiering cache.
[0049] To address this challenge, this application redesigned the dynamic random access memory (DRAM) cache management mechanism inside the CMS-SSD controller and introduced the concept of intelligent layering. Its goal is to build an internal cache system that can efficiently handle byte writes, effectively utilize data locality to serve read requests, and intelligently manage data persistence. This application logically divides the SSD internal DRAM into two key parts that work together: cacheline-granular write log and page-granular read / data cache.
[0050] The cacheline-granular write log is introduced as follows: Location: This is an area specifically designed to quickly absorb byte-granularity (cacheline) write requests from the host.
[0051] The operating modes include: lightning write, efficient indexing, and version overwrite.
[0052] Lightning write is mainly reflected in: when the host issues a store instruction of cacheline size, the data is directly and quickly appended to the write log area of the internal DRAM. Its speed is close to that of the DRAM itself and much faster than operating the flash memory.
[0053] Efficient indexing is mainly reflected in: through an optimized index structure (such as a two-level hash table or a skip list SkipList, also located in the internal DRAM), the firmware can quickly find the latest data copy at a specific address in the log.
[0054] Version overwrite is mainly reflected in: new writes logically overwrite the old version at the same address by updating the index.
[0055] The core advantages are extremely low write latency and elimination of read-modify-write.
[0056] Extremely low write latency is mainly reflected in: the host's write operation is almost instantaneously completed, only waiting for the internal DRAM to write.
[0057] Elimination of read-modify-write is mainly reflected in: fundamentally avoiding the high performance overhead caused by directly operating the flash memory for byte writes.
[0058] The page-granular read / data cache is introduced as follows: Location: It is mainly responsible for caching the entire page of data read from the flash memory and serving as a staging area for background write-back operations.
[0059] The operating modes include: read cache acceleration and write-back transit.
[0060] Read cache acceleration is mainly reflected in: when the host read request misses in the write log and data needs to be retrieved from the flash memory, the firmware reads the entire flash page into this cache area. Subsequently, the required cache line is extracted from it and returned to the host. If the host subsequently accesses other data within the same page, it can directly respond quickly from this cache, effectively utilizing spatial locality.
[0061] Write-back staging is mainly reflected in: when the data in the write log is processed (merged) in the background to form complete data pages ready to be written back to the flash memory, these pages are temporarily stored in this area.
[0062] The core advantages are: accelerating read operations and supporting background write-back.
[0063] Accelerating read operations is mainly reflected in: one flash page read serves multiple accesses to the same page, reducing the number of read operations on the flash memory.
[0064] Supporting background write-back is mainly reflected in: providing the necessary space for the firmware's background data organization and persistence.
[0065] The core data management responsibilities of the firmware (Internal Intelligent Tiering and Granularity Management Mechanism Module 12) are as follows: The SSD firmware acts as an intelligent data dispatcher, responsible for managing these two levels of caches and data interaction with the flash memory: such as, log merging, garbage collection, end-to-end data flow management, and ensuring data consistency.
[0066] Log coalescing is mainly reflected in: this is the key background mechanism for the write log to play its role. The firmware will periodically or when the log space is tight, actively scan the write log to find multiple scattered cache line updates for the same flash page. Then, it intelligently merges these updates (combining the unmodified parts read from the flash memory when necessary) to build the complete and up-to-date version of the page in the internal DRAM cache.
[0067] Garbage collection (Log Cleaning / Destaging) is mainly reflected in: the newly generated data pages after merging are then efficiently written to a new physical location in the flash memory. Once the write is successful, the corresponding old entries and indexes in the write log can be cleared, releasing valuable log space. This is similar to the segment cleaning mechanism in a log-structured file system (LFS), but occurs inside the firmware.
[0068] End-to-end data flow management: The firmware orchestrates the entire process: receiving host requests → serving from the write log / read cache preferentially → triggering flash reads when necessary → performing background log merging → writing the merged pages back to the flash memory → updating the internal address mapping table (FTL).
[0069] Ensuring data consistency is mainly reflected in that the firmware guarantees data integrity and persistence through transactional operations (e.g., ensuring the atomicity of log writing and index updating) and confirmation mechanisms (ensuring that the log is cleared only after the flash write is completed). The battery backup ability of the internal DRAM is crucial for preventing the loss of log data due to power failure.
[0070] The contributions to overcoming impedance mismatch are as follows: This intelligent hierarchical caching and granularity management mechanism makes a core contribution to overcoming multiple impedance mismatches at the firmware level: The contributions at the performance level are as follows: Write latency revolution: Writing logs reduces the write latency of bytes from the microsecond / millisecond level of operating flash memory to the nanosecond level of operating the internal DRAM.
[0071] Significant reduction in write amplification and throughput improvement: Log merging aggregates a large number of small writes into a small number of large writes, greatly reducing write amplification, increasing the effective write throughput, and reducing the wear on the flash memory.
[0072] Alleviating granularity mismatch: It seamlessly converts the granularity difference between host byte access and flash page operations inside the firmware.
[0073] Accelerating read operations: Page-granularity read caching reduces the number of flash memory reads by taking advantage of spatial locality.
[0074] The contributions at the cost level are as follows: Extending the SSD lifespan: The significantly reduced write amplification directly extends the service life of the flash memory medium, reducing the replacement frequency and the total cost of ownership.
[0075] Optimizing the utilization of internal DRAM: The hierarchical design makes the utilization of expensive internal DRAM resources more reasonable and efficient.
[0076] The contributions at the energy efficiency level are as follows: Reducing the energy consumption of flash memory operations: Fewer flash memory programming and erasing operations directly reduce the power consumption of the device during operation.
[0077] Reducing internal data transfer: Compared with the original read-modify-write, log merging can usually organize the data flow more effectively.
[0078] The contributions at the intelligent processing level are as follows: Providing "fresh" data for internal computing: Well-managed internal caches (especially the latest data in the write log) provide the embedded computing engine 14 with the convenience of quickly accessing the latest data, allowing it to operate directly on the efficient cache.
[0079] Summary: "Internal Intelligent Hierarchical and Granularity Management" is the core engine at the firmware level of the Compute-in-Memory Semantic Solid State Drive 10 (CMS-SSD). It fundamentally resolves the core contradiction between the host byte access requirements and the flash physical page operation limitations by ingeniously reconstructing the structure and function of the internal DRAM cache in the SSD - especially by introducing write logs at the cache line granularity, combined with page granularity read caches and intelligent background data merging and cleaning processes. This innovation directly brings about a leap in write performance (low latency, high throughput), a significant reduction in write amplification (improving performance, extending lifespan, reducing costs), and a decrease in operating power consumption (improving energy efficiency), and provides a better data foundation for internal intelligent processing. It is the key internal mechanism that enables the CMS-SSD to achieve high performance and high efficiency.
[0080] The Semantic-Aware Flexible Data Placement Mechanism Module 13 is used to place data with different semantics into the optimal physical area inside the flash memory by leveraging data semantic hints and combining the status of its internal write log / data cache.
[0081] In some embodiments, the data semantic hints include at least one of: data lifecycle, access pattern, hot / cold attribute, and object size category.
[0082] In some embodiments, the Semantic-Aware Flexible Data Placement Mechanism Module 13 uses the recovery units provided by the flexible data placement mechanism in the NVMe standard as the basic management unit for physical placement.
[0083] In some embodiments, Semantic-Aware Flexible Data Placement (firmware level, host boot): Integrate and extend the concept of flexible data placement. The host operating system 20 or application provides data semantic hints (e.g., data lifecycle, access pattern, hot / cold attribute, object size category) to the firmware through a unified interface. The firmware uses these hints and combines the status of its internal write log / data cache to intelligently place data with different semantics (such as write log data, read cache data, data with different lifecycles) into the optimal physical area inside the flash memory (such as different recovery units), further reducing garbage collection interference and write amplification.
[0084] Traditional solid state drive (SSD) firmware is like a diligent but uninsightful warehouse keeper. It only cares about the mapping between logical block addresses (LBAs) and physical flash locations, and knows almost nothing about the stored data content (i.e., the "semantics" of the data). All data, regardless of its nature, is treated equally and randomly mixed in physical flash blocks. When garbage collection (GC) is needed to free up space, the firmware can only make decisions based on underlying physical information (such as the ratio of valid / invalid data in a block), and cannot utilize more valuable upper-layer data characteristics.
[0085] The data semantics mainly lie in: unlocking the key information for efficient storage.
[0086] However, different types of data inherently possess different "personalities" and "fates", which are their semantic information: such as, lifetime, access pattern, size and shape, and "invalidation correlation" characteristics.
[0087] Lifetime: Some data is as short-lived as a meteor (such as temporary caches, session logs) and will quickly become invalid; some data is as persistent as a rock (such as user configurations, core records).
[0088] Access Pattern: Some data is a "social butterfly" and is frequently read and written (hot data); some is "reclusive" and rarely accessed (cold data). The writing methods are also different. Some are written sequentially like flowing water (such as logs), and some are updated randomly at precise points (such as database indexes).
[0089] Object Size: The system may simultaneously process data that is "small and delicate" (small objects) and "huge in size" (large objects).
[0090] Invalidation Correlation: Certain data blocks inherently have the tendency to "live and die together". For example, all cache data belonging to the same user session often becomes invalid together.
[0091] The goal of intelligent placement: Birds of a feather flock together for efficient recycling.
[0092] If the SSD firmware can "understand" these data semantics and accordingly intelligently cluster data with similar characteristics (especially similar lifetimes) and store them in physically adjacent or associated flash regions, it will bring great benefits. Imagine putting all data with very short lifetimes in the same physical block. When they all "reach the end of their lives" collectively, recycling this block becomes extremely easy and efficient because there is almost no "live" data that needs to be laboriously relocated to a new home. This is the core goal of intelligent data placement: maximizing the garbage collection efficiency.
[0093] Technical implementation: An upgraded version of the flexible data placement (FDP) guided by the host and executed by the firmware.
[0094] To achieve this goal, this application draws on and expands the flexible data placement (Flexible Data Placement, FDP) mechanism in the NVMe standard: Base Framework (Integrated with FDP): This application uses the Reclaim Unit (RU) provided by FDP as the basic management unit for physical placement. The SSD firmware divides the internal flash resources into multiple RUs. The host can suggest which logical group the data should be placed into through the Reclaim Unit Handle (RUH).
[0095] Core Upgrade (Expansion - Semantic Awareness): Different from traditional FDP which mainly relies on the host to explicitly specify the RUH, the solution of this application goes further: it includes the host providing "semantic hints" and the firmware making intelligent decisions.
[0096] The host providing "semantic hints" is mainly reflected in: when the host operating system 20 or application writes data (whether it is byte writing or block writing) through the "Unified Abstraction Interface Layer" designed by this application, it can attach "semantic hints" that describe the data characteristics. These hints are higher-level information, such as: Lifecycle: Short / Long (LIFETIME_SHORT / LONG).
[0097] Access Temperature: Hot / Cold (ACCESS_HOT / COLD).
[0098] Write Mode: Sequential / Random (STREAM_SEQUENTIAL / RANDOM).
[0099] Object Size: Small / Large (OBJECT_SMALL / LARGE).
[0100] Even application-customized tags, such as cache metadata, user session X data, etc.
[0101] The firmware's intelligent decision-making is mainly reflected in: after receiving these hints, the CMS-SSD firmware does not execute blindly. It will comprehensively consider the host's intention (semantic hints), its own internal state (such as cache / log situation, wear and remaining space of each RU, current GC pressure, etc.), and intelligently decide which physical RU is most suitable for finally placing the data.
[0102] Host "guides", rather than "forces": In this mode, the host plays the role of an information provider and advisor, passing high-level semantics to the firmware. The final decision-making power for physical placement remains in the hands of the firmware. This not only utilizes the host's application layer knowledge but also avoids the risks and complexities brought by the host directly managing the complex underlying flash.
[0103] Operation examples of intelligent placement: When the CMS-SSD firmware needs to write back data in the internal DRAM cache (whether it is a page after write-log merging or a modified read-cache page) to the flash memory, its decision-making engine will: Interpret the hint: Check the semantic hint associated with this data page.
[0104] Evaluate the status: Combine the internal cache, the log, and the health status of each RU.
[0105] Execute the strategy: For example, for a page with the LIFETIME_SHORT hint? Put it into that reserved RU that is easy to quickly recycle.
[0106] For a page with ACCESS_HOT, if it must be written back, give priority to putting it into the flash memory area with the best performance, and consider reading it back into the cache as soon as possible later.
[0107] For pages with the USER_SESSION_X label from the same application, aggregate them into the same or adjacent RUs to increase the likelihood of their future simultaneous failure.
[0108] For pages generated by write-log merging, decide its "destination" according to the most important semantic hint contained in the log entries in this page.
[0109] Cooperation with the internal cache / log: Intelligent placement mainly occurs at the back end of the data flow - that is, the stage when data is persisted from the internal DRAM to the flash memory. It works in coordination with the write-log and read-cache mechanisms at the front end that are responsible for quickly responding to host requests, complementing each other. The front-end mechanism solves the problems of real-time access and granularity mismatch, while the back-end intelligent placement focuses on optimizing the long-term storage efficiency and management cost of data.
[0110] The core contributions to overcoming impedance mismatch are as follows: The flexible data placement mechanism that perceives semantics provides strong back-end support for overcoming multiple impedance mismatch problems: The contributions at the performance level are as follows: Reduce GC interference and tame long-tail latency: Through efficient GC, its interference with normal read and write operations is significantly reduced, and performance jitter and unpredictable long-tail latency are reduced.
[0111] Improve the continuous write ability: Faster space recycling means better support for high-intensity continuous write loads.
[0112] The contributions at the cost level are as follows: Write amplification reduction and extended SSD lifespan: This is the most core value. Intelligent placement is one of the fundamental means to reduce write amplification to nearly 1. Extremely low write amplification means that the wear of the flash memory medium is significantly slowed down, significantly extending the service life of the SSD, thereby reducing the device replacement cost and the total cost of ownership (TCO).
[0113] Reduce the need for physical redundancy: Efficient garbage collection (GC) may enable the SSD to no longer require such a large internal over-provisioning space to maintain performance, theoretically reducing the manufacturing cost of physical flash memory.
[0114] Contributions in terms of energy efficiency are as follows: Reduce GC energy consumption: Fewer internal data transfers and flash memory erase operations directly reduce the energy consumed by the device during background self-maintenance.
[0115] Increase idle time: Reducing GC interference allows the device to enter the low-power idle state faster and more frequently, reducing the overall operating power consumption.
[0116] Contributions in terms of intelligent processing are as follows: Provide physical data insights: The firmware internally holds the correlation information between data semantics and physical layout. This provides valuable environmental information for future more intelligent internal computing engines. For example, computing tasks can be preferentially scheduled to execute on RUs containing relevant semantic data.
[0117] Drive intelligent caching / prefetching: The firmware can use semantic hints to optimize the management strategy of the internal read cache. For example, more intelligently prefetch data pages with the same semantic tags.
[0118] Summary: "Semantic-aware flexible data placement" is the key backend engine for achieving high-performance, high-efficiency, and long-life CMS-SSDs. It transcends the "blind" storage mode of traditional SSDs. By leveraging the upper-layer data semantic information provided by the host and combining the internal state awareness of the firmware itself, it intelligently plans the layout of data on the physical flash memory medium. Its core goal is to maximize the garbage collection efficiency through "things of a kind come together" (especially clustering by lifecycle). This mechanism directly contributes to significantly reducing write amplification (thereby improving performance, extending lifespan, and reducing costs), reducing GC interference (improving latency and stability), and reducing device energy consumption. It closely collaborates with the front-end internal cache / log mechanism to jointly form a powerful solution to overcome the multiple impedance mismatch problems of traditional storage.
[0119] The embedded computing engine 14 is used for hardware acceleration for specific tasks; wherein, the specific tasks include at least one of data compression / decompression, data encryption / decryption, data filtering / aggregation, index assistance, and format conversion.
[0120] The embedded dedicated computing engines are mainly embodied in integrating lightweight and programmable dedicated computing units (such as small field-programmable gate arrays FPGAs or application-specific integrated circuit ASIC cores) in the solid-state drive controller. These engines are not general-purpose processors but are used for hardware acceleration for specific tasks (such as data compression / decompression, data encryption / decryption, data filtering / aggregation, index assistance, format conversion, etc.).
[0121] The host operating system 20 is connected to the compute-in-memory semantic solid-state drive 10, and is used to receive long-latency hints sent by the compute-in-memory semantic solid-state drive 10 for context switching; and is provided with a unified abstraction interface layer to uniformly manage byte-granularity memory access to the compute-in-memory semantic solid-state drive 10, the transmission of data placement semantic hints, and the request to submit offloading tasks to the embedded computing engine 14, and is provided with a semantic-aware intelligent scheduling and offloading decision-making mechanism.
[0122] In some embodiments, the unified abstraction interface layer is mainly embodied in: drawing on xNVMe to provide a high-level software library. This library shields complex hardware details from upper-layer applications and uniformly manages byte-granularity memory access to the compute-in-memory semantic solid-state drive 10, the transmission of data placement semantic hints, and the request to submit offloading tasks to the embedded computing engine 14.
[0123] The "intelligent memory integrated memory storage architecture 100" and its core hardware, the "compute-in-memory semantic solid-state drive 10 (CMS-SSD)", undoubtedly bring unprecedented powerful functions but also introduce significant underlying complexities. These complexities are embodied in: New access mode: Instead of a single block device, it supports byte-granularity access (load / store instructions) of memory semantics through Compute Express Link (CXL).
[0124] New control dimension: It increases the ability to provide data placement semantic hints, and requires upper-layer software (applications or operating systems) to pass information about data life cycles, access patterns, etc. to the firmware.
[0125] New computing paradigm: It integrates the embedded computing engine 14, providing the possibility of offloading computing tasks to be executed inside the storage device, which requires corresponding task submission and management interfaces.
[0126] Potential cooperation details: Although "cooperative context switching" is mainly handled by the operating system and firmware, the upper-layer software library may need to provide configuration options or status query interfaces to optimize its behavior.
[0127] If every application developer is required to directly understand and operate these underlying hardware details - for example, delving deep into the CXL protocol, mastering flash media characteristics, adapting to specific firmware interfaces, and learning the instruction set of the embedded computing engine 14 - it will undoubtedly set extremely high technical barriers. This will greatly impede the popularization and application of this advanced hardware technology, making it difficult to translate its powerful potential into actual productivity.
[0128] Solution: Build a "translator" that shields complexity.
[0129] To solve this problem, this application must build a bridge between the complex hardware and upper-layer applications, which is the "Unified Abstraction Interface Layer". This interface layer usually exists in the form of a software library (for example, it is a.so file in the Linux system and a.dll file in Windows). Its core goals are: Encapsulate underlying details: Encapsulate all complex operations related to specific hardware implementations (such as CXL protocol interaction, firmware command encoding, memory mapping management, task packaging, etc.) inside the library.
[0130] Provide a stable and easy-to-use interface (API): Provide a set of higher-level, semantically clear, stable, and easy-to-use function or method calls to upper-layer applications. Application developers only need to learn and use this set of standard interfaces to utilize the various advanced functions of the CMS-SSD.
[0131] Ensure portability: Ideally, this set of interfaces should have the ability to be cross-platform or across different vendors' CMS-SSD hardware (if the hardware follows a common specification). Applications only need to program against this set of interfaces to run on different compatible hardware without having to rewrite the code for each piece of hardware.
[0132] Learn from and go beyond: Starting from the successful experience of xNVMe.
[0133] The xNVMe project provides valuable experience for this application. It has successfully provided a unified API for the originally fragmented NVMe storage access paths (such as the traditional POSIX interface, asynchronous libaio / io_uring, user-space SPDK, etc.). The cleverness of xNVMe lies in that it does not force the introduction of a new abstraction layer, but rather acts like a flexible "adapter" that intelligently maps the upper-layer unified calls to the best available or user-specified underlying I / O paths in the system while maintaining extremely low performance overhead.
[0134] The "Unified Abstraction Interface Layer" of this application draws on the concept of xNVMe - providing a unified API, shielding underlying differences, and maintaining flexibility and low overhead. However, the challenges in this application are greater because it needs to manage not only different I / O paths but also three entirely new dimensions: the semantics of memory access, the transfer of data placement semantics, and the management of compute offloading tasks.
[0135] The detailed functions of the Unified Abstraction Interface Layer are as follows: This software library needs to have the following core functions: a unified access interface, management of data placement semantic hints, management of compute offloading tasks, and device feature query and configuration.
[0136] The unified access interface is mainly reflected in: Memory Semantic Interface: Provide interfaces similar to standard memory operations (such as memcpy), pointer access, or more advanced data structure interfaces (for example, a CMS-SSD-aware persistent key-value store). The underlying layer of the library will automatically translate these operations into efficient CXL load / store instructions to access the CMS-SSD area mapped to the memory space. It will handle details such as address translation and necessary cache synchronization.
[0137] (Optional) Block Interface Compatibility: To facilitate the migration of legacy applications, the library can also provide traditional, block-based read / write interfaces. The underlying layer still accesses through CXL but emulates the semantics of block devices to ensure compatibility.
[0138] The management of data placement semantic hints is mainly reflected in: Standardized Hints: Define a set of clear semantic tags related to application scenarios (such as "temporary data", "hot data", "archived data", etc.).
[0139] API Integration: Provide a concise API that allows applications to easily attach these semantic tags when performing write operations. For example, write_with_hint(address, data, size, HINT_TEMPORARY | HINT_HOT);.
[0140] Underlying Transfer: The library is responsible for translating these high-level semantic tags into underlying signals that the CMS-SSD firmware can understand (possibly through specific CXL messages, register writes, or other mechanisms).
[0141] The management of compute offloading tasks is mainly reflected in: Task Description and Submission: Provide an API that enables an application to clearly describe the computing tasks it wants to offload (e.g., specify the ID of the firmware built-in function to be executed, the addresses and sizes of input / output data on the CMS-SSD, and other specific parameters), and submit the tasks to the library.
[0142] Interface Encapsulation: The library is responsible for packing the task requests from the application layer into a command format that conforms to the interface specification of the CMS-SSD embedded computing engine 14 and sending them to the device through an appropriate CXL channel (which may be CXL.io or a specific "doorbell" mechanism).
[0143] Result Retrieval and Synchronization: Provide synchronous or asynchronous APIs to wait for the completion of the computation, retrieve the result data, and handle possible errors.
[0144] Device Feature Query and Configuration are mainly reflected in: Capability Discovery: Provide an API that allows an application to query the specific capabilities of the connected CMS-SSD hardware (e.g., which data placement semantic tags are supported, what computing acceleration functions are available, how large the internal cache is, etc.).
[0145] Behavior Configuration: Allow an application or system administrator to configure some default behaviors of the library (e.g., default data placement policies, priorities of computing tasks, etc.).
[0146] Implementation Considerations are as follows: Form: Usually a user-space shared library or static library, which is convenient for application linking and deployment.
[0147] Interaction: May need to communicate with specific device drivers in the operating system kernel (for resource management, interrupt handling, etc.), or directly interact with the hardware through a user-space driver mechanism (such as UIO, VFIO) in specific scenarios (if permissions and design allow).
[0148] Intelligence: The library itself can detect hardware characteristics at runtime and dynamically select the optimal underlying execution path according to the system load and configuration to achieve a certain degree of adaptive optimization.
[0149] The Core Contributions to Overcoming Impedance Mismatch are as follows: The "Unified Abstraction Interface Layer" plays a crucial enabling and catalytic role in overcoming multiple impedance mismatches: mainly in performance impedance mismatch (indirect contribution), cost impedance mismatch (core contribution), energy efficiency impedance mismatch (indirect contribution), and intelligent processing impedance mismatch (core contribution).
[0150] Performance impedance mismatch (indirect contribution) is mainly reflected in that by simplifying the utilization of underlying high-performance features (such as low-latency byte access and compute offloading), application developers can more easily and widely integrate these features into applications, thus fully unleashing the potential of the hardware and indirectly overcoming the performance bottlenecks caused by accessing new hardware using traditional interfaces.
[0151] Cost impedance mismatch (core contribution) is mainly reflected in: Significantly reducing software costs: This is the most direct and important contribution. It liberates developers from the heavy underlying hardware adaptation work, significantly reducing the software engineering costs of developing, testing, debugging, and long-term maintenance and support for this new type of hardware.
[0152] Promoting the ecosystem and reducing hardware costs: Standardized or widely adopted interfaces can promote competition and interoperability among hardware manufacturers, accelerate the maturity of the ecosystem, and ultimately may reduce the cost of the hardware itself through economies of scale and competition.
[0153] Energy efficiency impedance mismatch (indirect contribution) is mainly reflected in that by simplifying the invocation of energy-saving features such as compute offloading, it makes it easier for applications to build and deploy more energy-efficient systems.
[0154] Intelligent processing impedance mismatch (core contribution) is mainly reflected in: Providing a standardized intelligent entry: It provides a stable and easy-to-use "invocation entry" for applications to use the intelligent processing function (compute offloading) embedded in the storage device. This makes it more realistic and efficient to deeply integrate intelligent computing into storage-intensive applications.
[0155] In summary, the "unified abstraction interface layer" plays an indispensable role of "translator" and "adapter" between upper-layer application software and underlying complex innovative hardware. By providing a set of carefully designed high-level, stable, and easy-to-use APIs, it effectively shields the underlying complexity in the processes of accessing memory semantic storage, transmitting data placement hints, and managing compute offloading tasks. Its core value lies in greatly reducing the threshold and cost of software development and adaptation, and enhancing the usability and portability of the system. This enables the revolutionary performance, cost, energy efficiency, and intelligence advantages brought by the "intelligent memory integrated storage architecture 100" to be truly adopted and utilized by a wide range of applications, and is the key software infrastructure to promote the development and maturity of the entire technology ecosystem.
[0156] See Figure 3 , the host operating system 20 includes: a cooperative context switching module 21 and a semantic-aware intelligent scheduling and offloading decision mechanism module 22.
[0157] The cooperative context switching module 21 is used to receive long-latency hints sent by the compute-in-memory semantic solid-state drive 10 and perform context switching.
[0158] In some embodiments, cooperative context switching (OS level): When the firmware of the Compute-in-Memory Semantic Solid State Drive 10 predicts that an access will encounter significant latency (such as flash access or internal computing tasks), it sends a "long latency hint" to the host through an extended CXL response, triggering the operating system to perform opportunistic context switching, effectively hiding the latency.
[0159] In a traditional computing model, when the central processing unit (CPU) executes an instruction to access memory (such as load or store), if the required data is not in the cache, the CPU's execution pipeline usually stalls and can only passively wait for the data to return from the main memory (such as dynamic random access memory, DRAM). Although this waiting time exists, it is relatively controllable.
[0160] However, when the present application introduces a new type of storage device, such as the Compute-in-Memory Semantic Solid State Drive 10 (CMS-SSD), the situation has changed fundamentally. Although such devices provide the semantics of memory access through interfaces such as Compute Express Link (CXL), enabling the CPU to access it as if it were memory, the physical latency of its underlying technology (such as flash access or internal embedded computing) is much higher than that of DRAM. This means that once an access request fails to hit the cache inside the CMS-SSD (such as its own DRAM), the CPU may face an extremely long stall.
[0161] A more difficult problem is that this "memory-like" access through CXL.mem is transparent to traditional operating systems. The operating system cannot pre-perceive this potential long latency as it does with traditional driver- and interrupt-based block device I / O. Therefore, it cannot proactively perform context switching when detecting I / O waiting and schedule CPU resources to other ready threads as usual. As a result, the CPU core seems to be "stuck" on that memory instruction accessing the CMS-SSD, wasting precious computing cycles.
[0162] Based on this, the present application proposes a "tacit" cooperation between the operating system and the solid state drive.
[0163] The "cooperative context switching" mechanism is designed to break this dilemma. It establishes an innovative cooperation channel between the operating system and the firmware of the Compute-in-Memory Semantic Solid State Drive 10. Its core idea is: Let the firmware, which knows its own state best, actively send a "warning signal" to the host when it anticipates an upcoming long latency operation. The operating system then captures this signal and regards it as an opportunity to perform a task switch, thereby efficiently using the time that would otherwise be wasted due to CPU stalls to execute other computing tasks. This effectively hides the long latency of memory semantic storage access.
[0164] Technical implementation process: The specific implementation process of this cross-layer collaboration mechanism is as follows: The process of firmware prediction and sending "long delay prompt" is as follows: Timing: When the host CPU sends a memory access request to the CMS-SSD through CXL.mem.
[0165] Prediction: The CMS-SSD firmware quickly checks its internal state (such as cache hit, flash queue situation, background task status, whether to start internal computing, etc.), and estimates the time required to process this request.
[0166] Decision-making and signal: If the firmware predicts that the delay will significantly exceed the preset threshold (this threshold is usually equivalent to the time overhead required for the host to perform a context switch), it will not return data after waiting for the operation to complete according to the normal process. Instead, it immediately sends a special "no data response" message back to the host through the CXL.mem protocol. The key of this message is that it embeds a clear "long delay prompt" signal. The meaning of this signal is: "Request received, but I need a long time to process, please note, host." The process of the host receiving the signal and triggering a hardware exception is as follows: Receiving: The CXL controller on the host side receives this special response with the "long delay prompt".
[0167] Notifying the core: The CXL controller associates this signal with the CPU core that initially issued the request.
[0168] Precise triggering: When this CPU core is about to "retire" the load / store instruction that causes the long delay access (this is done to avoid misprocessing speculative execution instructions), it recognizes this "long delay prompt" signal and triggers a specific type of hardware exception (which can be called "CXL long delay exception"). This exception mechanism ensures that this application can accurately locate the instruction that causes the problem.
[0169] The process of the operating system taking over is as follows: The process of exception handling and task switching is as follows: Responding to the exception: The processor pre-registered by the operating system for this new type of exception is activated to take over the control of the CPU.
[0170] Status saving and concession: The exception handler quickly saves the complete execution context (register status, instruction pointer, etc.) of the currently interrupted thread, marks it as the "waiting for long delay store response" state, and then removes it from the CPU's run queue.
[0171] Opportunistic Scheduling: The exception handler immediately invokes the operating system scheduler. At this time, the scheduler clearly knows that the current thread has been paused due to an external long delay. Therefore, it can immediately select another thread with the highest priority from the ready queue and put it into operation, maximizing the utilization of CPU resources.
[0172] The task recovery and access retry process is as follows: Background Processing: Meanwhile, the CMS-SSD continues to execute that time-consuming operation (e.g., reading flash memory or performing internal calculations) in the background without interference.
[0173] Data Ready: When the operation is completed and the data is ready, the CMS-SSD sends a standard memory data response to the host through CXL.mem.
[0174] Thread Wake-up and Retry: When the operating system scheduler decides to resume running the previously interrupted thread, it restores the context of the thread and points the execution flow back to the load / store instruction that caused the exception. The thread resumes execution from that instruction. Since the required data is very likely to have been prepared by the CMS-SSD and may have been cached in the CPU cache, host memory, or the internal DRAM of the CMS-SSD at this time, this memory access will be completed very quickly, no longer triggering a long-delay exception, and the thread can continue to execute seamlessly.
[0175] The core contributions to overcoming impedance mismatch are as follows: The "cooperative context switching" mechanism makes a key contribution to overcoming the impedance mismatch problem of traditional hierarchical storage through the above-mentioned sophisticated cross-layer collaboration: Contribution to performance impedance mismatch: Hiding Long-Tail Latency: The most core contribution. It converts the long access latency (whether it is flash read / write or internal calculation) that would originally cause the CPU to pause for a long time into a time window for running other threads, greatly improving the effective utilization rate of the CPU and the overall system throughput. For applications that require quick response, this can significantly improve the user experience.
[0176] Improving Concurrency Efficiency: Allows multiple threads to access the CMS-SSD more comfortably in parallel. The waiting of one thread does not block other threads from using the CPU, and at the same time, other threads can continue to send requests to the CMS-SSD, which helps to saturate the utilization of the internal parallelism of the device and the CXL link bandwidth.
[0177] Contribution to cost impedance mismatch: Enhancing Hardware Value: By reducing the ineffective waiting time of the CPU, the actual utilization efficiency of the expensive CPU resources is improved, resulting in a higher return on hardware investment for the overall system.
[0178] The contribution of energy efficiency impedance mismatch lies in: Reducing CPU Idle Power: It avoids the idling or inefficient waiting of the CPU during long latency periods, reduces this part of the power consumption, and helps improve the performance per watt of the entire system.
[0179] The contribution of intelligent processing of impedance mismatch lies in: Enabling LongerIn-Storage Computation: This mechanism can not only hide the data access latency, but also hide the latency caused by the execution of long-running computational tasks by the embedded computing engine 14 inside the CMS-SSD. This provides the key latency tolerance ability for offloading more complex computational tasks (such as data preprocessing, compression, partial query logic, etc.) to the inside of the storage device, enabling the host CPU to work asynchronously with the internal computing engine of the device and giving full play to the advantages of "intelligent storage integration".
[0180] In summary, "cooperative context switching" is the key technology to achieve a high-performance and high-efficiency "intelligent storage integration memory storage architecture 100". It is not a simple task switch, but an intelligent latency hiding mechanism based on device prediction and system cooperation. It uses the extended interconnection protocol signals and specific hardware exceptions as information bridges to accurately convert the inevitable long latency during memory semantic storage access into the "prime time" for effective scheduling by the operating system, thus systematically alleviating the constraints of storage latency on CPU performance, energy efficiency, and intelligent potential.
[0181] The semantic-aware intelligent scheduling and offloading decision-making mechanism module 22 is used to make data placement decisions, computing offloading decisions, and task cooperation.
[0182] Refer to Figure 4 , the semantic-aware intelligent scheduling and offloading decision-making mechanism module 22 includes: a data placement decision unit 221, a computing offloading decision unit 222, and a task cooperation unit 223.
[0183] The data placement decision unit 221 is used to generate corresponding semantic hints and transmit them to the computational memory semantic solid-state drive 10 after making data placement decisions. Among them, the data placement decision unit 221 is used to make data placement decisions based on explicit application instructions, runtime intelligent analysis, and system status feedback.
[0184] The computing offloading decision-making unit 222 is used to identify tasks suitable for offloading to the embedded computing engine 14 and submit the tasks through a unified interface. The computing offloading decision-making unit 222 is used to perform task offloading based on task characteristic analysis, the real-time state of the system, and / or performance / energy efficiency modeling.
[0185] The task coordination unit 223 is used to coordinate the task execution between the host and the embedded computing engine 14. The task coordination unit 223 is used to adopt preprocessing / postprocessing, pipelining, and / or asynchronous execution cooperation modes.
[0186] In some embodiments, semantic-aware intelligent scheduling and offloading decision-making (OS / runtime level): The operating system or runtime system (which can be combined with application-level hints or online profiling) is responsible for: Data placement decision: Determine which data resides in the host memory and which resides in the CMS-SSD, and generate corresponding semantic hints to be passed to the firmware.
[0187] Computing offloading decision: Identify tasks suitable for offloading to the CMS-SSD embedded computing engine 14 (for example, tasks with strong data locality, computationally intensive and parallelizable, and involving a large amount of storage I / O), and submit the tasks through a unified interface.
[0188] Task coordination: Coordinate the task execution between the host CPU and the CMS-SSD embedded computing engine 14. For example, the host processes the control flow and complex logic, and the CMS-SSD processes data-intensive computations.
[0189] When this application has a high-speed host memory (DRAM) and a computational memory semantic solid-state drive 10 (CMS-SSD) with memory access capabilities, a large capacity, and embedded computing capabilities, the system evolves into a powerful heterogeneous memory / storage system. However, it is not easy to control this system. If all data is simply piled up on the CMS-SSD, or if it completely relies on application developers to manually move data between the DRAM and the CMS-SSD, its potential cannot be fully exploited, but may instead increase development complexity and even lead to performance degradation. Similarly, in the face of the embedded computing capabilities provided by the CMS-SSD, how to wisely decide which computing tasks should be processed by the traditional host CPU and which should be "sunk" to be executed inside the storage device also becomes a crucial optimization problem.
[0190] Core idea: Make the system more "understand" the application and automatically optimize.
[0191] The "Semantic-Aware Intelligent Scheduling and Offloading Decision" mechanism is designed precisely to address these challenges. It serves as the "intelligent brain" and "command center" of the entire "Intelligent Memory Storage Architecture 100". Its core mission is to utilize deeper information (i.e., "semantics") to dynamically and automatically optimize two key things through the operating system (OS) or the runtime system of a specific application: Where data should be placed. (Intelligent Data Placement) Where computations should be executed. (Intelligent Computation Offloading) Its goal is to maximize the performance, energy efficiency, and resource utilization of the entire system. The key here lies in "semantic awareness" - the decision-making process no longer solely relies on underlying physical metrics (such as access speed, bandwidth occupancy rate), but rather delves deeper into understanding and leveraging the data characteristics of upper-layer applications and the intent information of computational tasks.
[0192] The executors of the decision: the collaboration between the operating system and the runtime system.
[0193] This "intelligent brain" can function at two levels: Operating System (OS) level: It has a global view and understands the resource usage of all processes in the system. It can formulate global data placement and task offloading strategies, which are relatively transparent to application programs. However, this usually requires modifying the operating system kernel and may lack a fine-grained understanding of the internal logic of specific applications.
[0194] Runtime system level: Such as Java Virtual Machine (JVM), Database Management System (DBMS), big data processing frameworks (such as Spark), or machine learning frameworks (such as TensorFlow), etc. They have a deeper understanding of the data structures they manage and the computational patterns they execute. Therefore, they can make optimization decisions that are more in line with application requirements and more precise. However, this usually requires development for different runtime environments.
[0195] In practice, a hybrid mode may be the most effective: The operating system provides a basic resource management and scheduling framework, while the runtime system provides more specific application semantic information and optimization suggestions, and the two work together.
[0196] The specific content and execution of intelligent decision-making include: intelligent data placement, intelligent computation offloading, and task coordination.
[0197] The content of intelligent data placement is as follows: Goal: Place the "right data" in the "right position". The core principle is to prioritize placing the "hot" data that is accessed most frequently and has the highest latency requirements in the fastest host DRAM; while the "cold" data with lower access frequencies, large volumes, or less latency sensitivity is placed in the lower-cost and high-capacity CMS-SSD.
[0198] Basis for decision-making: Explicit application hints: Application developers can directly "tell" the system through a unified interface (such as the aforementioned software library API) which data are critical hotspots and which are suitable for long-term storage in CMS-SSD.
[0199] Online Profiling during runtime: The system automatically monitors the data access patterns during runtime, such as access frequency and recency: Tracking which memory pages are frequently accessed and which have been recently accessed (similar to the principle of cache eviction algorithms).
[0200] Dynamic "temperature" identification: Based on historical access records, dynamically "label" the data to distinguish between hot data and cold data.
[0201] System status feedback: Dynamically adjust the strategy according to the current system resources (such as whether DRAM is tight and whether CMS-SSD access is congested).
[0202] Execution methods include: Transparent page migration: Once the decision is made, the system (OS or Runtime) will automatically migrate the selected data pages between DRAM and CMS-SSD and update the corresponding address mappings (such as page tables), remaining as transparent as possible to the upper-layer applications.
[0203] Guiding firmware optimization: When deciding on data placement (especially for newly allocated or written data to CMS-SSD), the system will generate corresponding semantic hints (such as "hot data", "temporary data", etc.) and pass them through the interface to the CMS-SSD firmware, enabling the firmware to also optimize at the physical storage level (such as placing hot data in flash regions with better performance).
[0204] The content of intelligent computing offloading is as follows: Goal: Assign the computing tasks to the most suitable computing unit - whether it is the powerful but general-purpose host CPU or the CMS-SSD embedded engine close to the data and possibly with specific acceleration capabilities. - To achieve the best performance or energy efficiency.
[0205] Basis for decision-making: Task characteristic analysis: Understand the characteristics of the computing task itself: Data Locality: Where is the data that the computing task needs to process mainly stored? If a large amount of data is on the CMS-SSD, offloading it can reduce data movement.
[0206] Computational complexity and type: Is the task computationally intensive or I / O intensive? Can its computational mode (such as simple filtering, aggregation, encryption / decryption, compression) be efficiently accelerated by the dedicated computing engine of the CMS-SSD? Parallel potential: Is the task suitable for parallel processing? Can the internal parallel computing resources of the CMS-SSD be utilized? System real-time status: Is the host CPU currently busy? Is the computing engine of the CMS-SSD idle? Is the CXL link bandwidth sufficient?
[0207] (Optional) Performance / energy efficiency modeling: Estimate the execution time and energy consumption of the task at different locations through a simple model to assist in making the optimal decision.
[0208] Execution methods include: Task identification: Identify candidate computing tasks suitable for offloading through compile-time analysis, explicit marking by developers (using specific APIs), or runtime dynamic detection, etc.
[0209] Submit for execution: Once the decision to offload is made, the system calls the unified interface API to submit the task description and the required data information (usually the address of the data on the CMS-SSD) to the CMS-SSD for execution.
[0210] The content of task collaboration is as follows: Goal: Ensure that the host CPU and the computing engine of the CMS-SSD can cooperate correctly and efficiently to jointly complete a complex application logic.
[0211] Common collaboration modes: Preprocessing / Postprocessing: The host is responsible for preparing data and task control, calling the CMS-SSD to execute the core calculation, and then retrieving the results for subsequent processing.
[0212] Pipeline operation: The data stream passes through the host (for complex logical judgment) and the CMS-SSD (for data-intensive processing) in sequence.
[0213] Asynchronous execution: After the host submits the task, it does not need to wait and can continue to process other transactions, and later obtains the results through callbacks or polling, etc.
[0214] Support mechanisms are as follows: Dependency Management: The system needs to manage the dependencies between host tasks and offloading tasks to ensure the correct execution order (e.g., using synchronization mechanisms such as semaphores, Future / Promise).
[0215] Data Sharing and Consistency: The hardware cache coherence provided by CXL.mem is the foundation. At the system level, more advanced memory allocation and synchronization services may also be required to ensure the correctness of application data.
[0216] Control Flow Coordination: Usually, the host CPU dominates the overall application control flow and decides the next action based on the execution status and results of CMS-SSD tasks.
[0217] Core Contribution to Overcoming Impedance Mismatch: "Semantic-Aware Intelligent Scheduling and Offloading Decision" is the key driving force for overcoming multiple impedance mismatch problems: The contribution to performance impedance mismatch lies in: minimizing access latency through intelligent data placement, accelerating processing using dedicated hardware through computing offloading, and enabling parallel work between the host and devices through task coordination, comprehensively improving system speed and throughput.
[0218] The contribution to cost impedance mismatch lies in: maximizing the value of DRAM and only using it to store the most critical hot data, enabling a more economical DRAM configuration to cooperate with a large-capacity CMS-SSD to achieve the target performance, thereby reducing the overall hardware cost.
[0219] The contribution to energy efficiency impedance mismatch lies in: reducing data movement (a key effect of intelligent placement and computing offloading) is the main way to reduce energy consumption. At the same time, allocating computing tasks to dedicated engines with higher energy efficiency ratios further improves the system energy efficiency.
[0220] The contribution to intelligent processing impedance mismatch lies in: the core value. It transforms the "near-data processing" process that originally required developers to manage laboriously into a process that is automatically and intelligently decided and executed by the system, greatly enhancing the practicality and efficiency of in-memory intelligent processing, and paving the way for upper-layer applications to build more complex innovative applications that utilize in-memory computing capabilities.
[0221] Summary: The "Semantic-Aware Intelligent Scheduling and Offloading Decision" mechanism is the "intelligent brain" for the efficient operation of the "Intelligent Memory-Integrated Memory Storage Architecture 100". It makes dynamic and intelligent decisions about the data storage location and computing execution location at the operating system or runtime level by deeply understanding the information (semantics) at the application level. This not only optimizes the utilization of heterogeneous resources, directly improving performance, reducing costs, and enhancing energy efficiency, but more importantly, it truly transforms the computing capabilities embedded in storage into an easy-to-use and automatically optimized system-level capability, which is the core software enabling technology for achieving the leap from "storage" to "intelligent storage".
[0222] The coordination between the cooperation mechanism and the unexpected technical effects is as follows: Coordination between writing logs and data placement: Fine-grained writing logs naturally aggregate frequent small write operations, which not only reduces write amplification but also facilitates the firmware to perform data placement with awareness of semantics. For example, an entire log segment can be placed as a whole in a specific flash memory area according to its lifecycle or access frequency hint.
[0223] Coordination between context switching and internal caching: The write log and read cache inside the compute-in-memory semantic solid-state drive 10 first attempt to service requests. Only when the internal cache misses and flash memory access or long compute tasks need to be executed will long latency hints and potential context switching be triggered. This mechanism makes context switching a "last resort" rather than a frequent operation, improving the effectiveness of switching, while the internal cache provides the first layer of fast response.
[0224] Coordination between compute offloading and data stream: The firmware can utilize its knowledge of the internal data layout (guided by placement hints) and cache status to optimize the data access path for offloading tasks. For example, data can be compressed or filtered directly in the internal DRAM cache or write log, avoiding unnecessary flash memory reads and writes, which cannot be achieved by traditional host CPUs.
[0225] Balance between the unified interface and system complexity: Although the underlying hardware and firmware integrate multiple complex mechanisms (hierarchical caching, data placement, compute offloading, cooperative scheduling), the unified abstract interface layer provides a relatively simple and consistent view for upper-layer applications, reducing the threshold for development and use.
[0226] How each technical means in the "Intelligent Memory Integrated Storage Architecture 100" specifically contributes to overcoming the impedance mismatch problems in the four dimensions of traditional hierarchical storage in terms of performance (speed, latency), cost, energy efficiency, and intelligent processing: 1. Overcoming performance impedance mismatch.
[0227] Problem manifestation: There is a huge access latency gap between the host memory (DRAM) and the solid-state drive (SSD); there is a bandwidth bottleneck between the host CPU and the storage device; performance jitter and long-tail latency caused by internal operations (such as garbage collection) in the SSD; I / O amplification and inefficiency caused by the mismatch between byte access and page access granularity.
[0228] Contribution of technical means: Unified compute expression link protocol interface and memory semantics: Providing a byte-granularity access interface reduces the protocol overhead and latency caused by traditional block interface conversion, making SSD access closer to memory semantics.
[0229] Internal intelligent stratification and granular management: Write log: Absorbs byte write requests at a speed close to that of internal DRAM, greatly reducing write latency and eliminating the performance bottleneck of host direct writing to Flash. Through background merging, the total number of writes and data volume to Flash are reduced, alleviating the impact of write amplification on throughput.
[0230] Data cache: Caches hot data pages read from flash memory to reduce read latency.
[0231] Granularity management: It effectively bridges the host's byte access requirements and the flash memory's page access physical limitations at the firmware level, reducing inefficient I / O and performance loss caused by granularity mismatch.
[0232] Semantic-aware flexible data placement: By physically aggregating data with similar life cycles, garbage collection efficiency is significantly improved, the interference of garbage collection operations on foreground I / O is reduced, and performance jitter and long-tail latency are reduced.
[0233] Embedded dedicated computing engine: Offloads computing tasks to a location close to the data, reducing the latency of data traveling back and forth between the host and the device, and increases the processing speed of specific computing tasks through dedicated hardware acceleration.
[0234] Cooperative context switching: Specifically targets unavoidable long delays (such as flash memory access and internal calculations), hides long-tail delays by switching CPU execution threads, and improves CPU utilization and system responsiveness.
[0235] 2. Overcome cost-impedance mismatch.
[0236] Problem manifestations: The cost per bit of DRAM is much higher than that of SSD; the large-capacity DRAM cache configured to make up for the insufficient performance of SSD is very expensive; SSD requires a large amount of internal overprovisioning and host-side overprovisioning to ensure performance and lifespan, which increases the actual cost of use; the cost of developing and adapting software for complex storage hierarchies is high.
[0237] Technical contribution: Unified computing expression link protocol interface and memory semantics: It makes it possible to use lower-cost SSDs to expand memory capacity, reducing the high hardware costs caused by relying solely on DRAM expansion.
[0238] Semantic-aware flexible data placement: By significantly reducing write amplification, the service life of the SSD is significantly extended and the frequency of device replacement is reduced. More importantly, it enables high performance to be achieved without a large amount of host-side over-provisioning, greatly improving the effective capacity-cost ratio of the SSD.
[0239] Unified Abstraction Interface Layer: Reduces the development and maintenance costs for applications to adapt to this new type of complex hardware, and improves software portability.
[0240] Internal Intelligent Hierarchical and Granularity Management: Reduces write amplification, indirectly extends the SSD lifespan, and lowers the total cost of ownership (TCO).
[0241] 3. Overcome the energy efficiency impedance mismatch.
[0242] Problem manifestation: Frequent data transfer between different storage levels consumes a large amount of energy; the host CPU executes storage-related calculations (such as compression and verification) that could be offloaded inefficiently and with high power consumption; inefficient garbage collection operations inside the SSD increase the device active time and consume more power; high write amplification leads to more flash write operations and increases energy consumption.
[0243] Contributions of technical means: Internal Intelligent Hierarchical and Granularity Management: The write log merging mechanism reduces the total amount and number of data finally written to the flash memory, and lowers the flash write power consumption.
[0244] Semantic-aware Flexible Data Placement: Efficient garbage collection reduces unnecessary internal data transfer and the number of flash erasure writes, and lowers the device operating power consumption.
[0245] Embedded Dedicated Computing Engine: Offloads computing tasks to a dedicated hardware with more optimized power consumption, reduces the transmission energy consumption of data across the CXL link, and improves computing energy efficiency.
[0246] Cooperative Context Switching: By improving the utilization rate of the CPU during long latency waiting, it avoids CPU idling and indirectly improves the overall system energy efficiency.
[0247] 4. Overcome the intelligent processing impedance mismatch.
[0248] Problem manifestation: Traditional storage devices are "passive", and all data processing (such as filtering, aggregation, format conversion, encryption and decryption, compression and decompression) requires data to be transferred to the host CPU, which limits the ability of near-data processing and increases latency and bandwidth consumption.
[0249] Contributions of technical means: Embedded Dedicated Computing Engine: Directly endows the storage device with endogenous computing capabilities. Allows specific types of computing tasks (especially those that are data-intensive, suitable for stream processing, or hardware-accelerated) to be directly completed inside the storage device, realizing the integration of computing and storage.
[0250] Semantic-aware Flexible Data Placement: Provides an optimization foundation for the internal computing engine. For example, the firmware knows which data blocks contain specific types of or recently accessed data, and can guide the computing engine to access and process this data more efficiently.
[0251] Semantic-aware Intelligent Scheduling and Offloading Decision-making: Provides a mechanism to effectively map the intelligent processing requirements at the application level to the embedded computing capabilities of the CMS-SSD. It determines which computations should be executed closer to the data, realizing an end-to-end intelligent processing flow.
[0252] Unified Abstract Interface Layer: Provides an upper-layer application with a standard way to call these intelligent processing functions embedded in storage.
[0253] Summary: Through the synergistic effect of its various technical means, the intelligent memory storage architecture 100 systematically solves the multiple impedance mismatch problems of traditional hierarchical storage in terms of performance, cost, energy efficiency, and intelligent processing. It is not a single-dimensional optimization, but rather constructs a new infrastructure layer that integrates the characteristics of memory and storage, embeds intelligent computing capabilities, and can adaptively optimize through hardware innovation and software collaboration.
[0254] 4. Technical Effects (Effects / Results) This "intelligent memory storage architecture 100" is expected to bring the following comprehensive and creative and non-obvious effects: Extreme Performance and Capacity Expansion: Achieves TB-level memory semantic capacity expansion through the computational expression link protocol. At the same time, by using internal intelligent caching, write log merging, cooperative context switching, and computational offloading, it significantly reduces access latency (especially write latency and long-tail read latency), improves the effective throughput, and approaches the ideal memory performance.
[0255] Unprecedented Resource Efficiency: Significantly reduces the device-level write amplification, ensures performance and lifespan without large-scale host over-provisioning, and greatly improves the actual available capacity and lifespan of the solid-state drive. Computational offloading reduces the host CPU load and data transfer overhead, and significantly improves the overall system energy efficiency.
[0256] Deeply Intelligent Data Processing: Seamlessly embeds specific computational tasks (such as compression, encryption, filtering, preprocessing, etc.) into the storage access process, realizing the true meaning of "computation close to data", and providing new optimization possibilities for upper-layer applications (such as databases, big data analysis, artificial intelligence), such as directly executing partial query predicates or data conversions at the storage layer.
[0257] Unification of Adaptability and Usability: The system can perform adaptive optimization (data placement, cache management, task scheduling) based on the data semantic hints provided by the application and the internal monitoring status, while simplifying the development complexity of upper-layer applications through a unified abstract interface.
[0258] Enhanced Sustainability: By improving hardware utilization (reducing over-provisioning), extending device lifespan (reducing write amplification), and reducing overall energy consumption (reducing data movement, offloading computations).
[0259] Summary: This solution is not a simple stacking of technologies. Instead, through the core innovation of the computational memory semantic solid-state drive 10, combined with deep software-hardware collaboration (OS, firmware, runtime), it achieves unification and intelligence in interfaces, data management, and computational paradigms, aiming to fundamentally solve the multiple impedance mismatches between memory and storage, forming an efficient, intelligent, and adaptive memory-storage continuum. The close cooperation among its technical means (such as write logging and data placement, internal caching and context switching, computational offloading and internal data flow) brings systematic gains beyond single-technology optimization and unexpected synergistic effects.
[0260] In several implementation manners provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation manners described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there may be other division manners in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0261] If the integrated unit in the above-mentioned other implementation manners is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various implementation manners of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.
[0262] The above are only the embodiments of the present application, and do not thus limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall similarly be included within the patent protection scope of the present application.
Claims
1. An intelligent storage integrated memory storage architecture, characterized in that: The smart storage integrated memory storage architecture includes: A computational memory semantic solid-state drive, which is used to provide a unified memory address space based on a computational expression link protocol, support byte-granularity access by the host through load / store instructions; and is provided with an internal intelligent tiering and granularity management mechanism; and is provided with a semantically aware flexible data placement mechanism, and is provided with an embedded computing engine; A host operating system is connected to the computational memory semantic solid-state drive and is used to receive long-delay prompts sent by the computational memory semantic solid-state drive and perform context switching; and a unified abstract interface layer is provided to uniformly manage byte-granularity memory access to the computational memory semantic solid-state drive, the delivery of data placement semantic prompts, and the submission of requests for offloading tasks to the embedded computing engine, as well as the provision of a semantically-aware intelligent scheduling and offloading decision-making mechanism.
2. The smart storage integrated memory storage architecture according to claim 1 is characterized in that: The computational memory semantic solid state drive comprises: Unified interconnect and memory semantics module, which is used to provide a unified memory address space based on the computational expression link protocol, and supports byte-granularity access by the host through load / store instructions; Internal intelligent tiering and granularity management mechanism module, which builds a cache line-sized fine-grained write log to absorb byte writes and reduce write amplification; while maintaining a page-sized read / data cache to exploit spatial locality; and is responsible for managing log merging, garbage collection, and data flow between the internal cache and flash memory; A semantically aware flexible data placement mechanism module that uses data semantic hints and the status of its internal write log / data cache to place data with different semantics into the optimal physical area inside the flash memory. An embedded computing engine is used for hardware acceleration of specific tasks; wherein the specific tasks include at least one of data compression / decompression, data encryption / decryption, data filtering / aggregation, index assistance, and format conversion.
3. The smart-storage integrated memory storage architecture according to claim 2, characterized in that: The unified interconnect and memory semantics module is also used to provide a logically continuous and extended physical address space to the host through the mapping of CXL.mem.
4. The smart-storage integrated memory storage architecture according to claim 2, characterized in that: The internal intelligent tiering and granularity management mechanism module is also used to directly and quickly append data to the write log area of the internal DRAM when the host issues a storage instruction of the cache line size, and to find the latest data copy of a specific address in the log through an optimized index structure; and the new write will logically overwrite the old version of the same address by updating the index.
5. The smart-storage integrated memory storage architecture according to claim 2, characterized in that: When the host read request does not hit in the write log and needs to obtain data from the flash memory, the internal intelligent tiering and granularity management mechanism module reads the entire flash memory page into the page granularity cache area, and then extracts the required cache line from it and returns it to the host; And when the data in the write log is processed in the background to form a complete data page ready to be written back to the flash memory, the data page will be temporarily stored in the page granularity cache area.
6. The smart-storage integrated memory storage architecture according to claim 2, characterized in that: The data semantic hints include: at least one of data life cycle, access mode, hot and cold attributes, and object size category.
7. The smart-storage integrated memory storage architecture according to claim 2, characterized in that: The semantically-aware flexible data placement mechanism module utilizes the recovery unit provided by the flexible data placement mechanism in the NVMe standard as the basic management unit for physical placement.
8. The smart-storage integrated memory storage architecture according to claim 1, characterized in that: The host operating system includes: A collaborative context switching module, used for receiving a long delay prompt sent by the computational memory semantic solid state drive and performing context switching; The semantically-aware intelligent scheduling and offloading decision mechanism module is used to make data placement decisions, computing offloading decisions, and task collaboration.
9. The smart-storage integrated memory storage architecture according to claim 8, characterized in that: The semantically-aware intelligent scheduling and unloading decision-making mechanism module includes: A data placement decision unit is used to generate corresponding semantic hints and transmit them to the computational memory semantic solid state drive after making a data placement decision; The computing offloading decision unit is used to identify tasks suitable for offloading to the embedded computing engine and submit the tasks through a unified interface; The task coordination unit is used to coordinate the task execution between the host and the embedded computing engine.
10. The smart-storage integrated memory storage architecture according to claim 9, characterized in that: The data placement decision unit is used to make data placement decisions based on explicit instructions from applications, runtime intelligent analysis, and system status feedback; The computation offloading decision unit is used to perform task offloading based on task characteristic analysis, system real-time status and / or performance / energy efficiency modeling; The task coordination unit is used to adopt pre-processing / post-processing, pipeline operation and / or asynchronous execution coordination mode.
Citation Information
Patent Citations
Computer memory expansion device and operation method thereof
CN116134475A
Dynamic hybrid flash memory translation layer address mapping method and system based on popularity perception
CN119149445A
Computer Memory Expansion Device and Method of Operation
US20210374080A1
Transactional memory support for compute express link (CXL) devices
US20230022544A1
Systems, methods, and apparatus for computational device communication using a coherent interface
US20240338315A1
Cited By
AI-driven adaptive storage layering and cache prefetching system
CN120428926A