Systems, methods, and apparatus for caching on storage device

By distinguishing the data access priorities of operating systems and applications, using scoring mechanisms and different cache replacement strategies, the problem of data access latency on storage devices is solved, and efficient data access and storage performance improvements are achieved.

CN120066990APending Publication Date: 2025-05-30SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411709612.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-24
Filing Date
2024-11-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When caching on storage devices, it is difficult for the prior art to effectively distinguish the access priorities of operating systems and applications, resulting in an increase in data access latency.

Method used

By determining the operational relevance of data to the operating system or application, computing the data score, and writing the data to the storage medium based on the score, using different cache replacement strategies to optimize data access.

Benefits of technology

It realizes optimized cache management of operating system and application data, reduces data access latency, and improves the performance of storage devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066990A_ABST
    Figure CN120066990A_ABST
Patent Text Reader

Abstract

A method may include determining that data is related to operation of an operating system; determining a score of the data; and writing the data to the memory medium based on the score. The data may include at least one page table. The at least one page table may include one or more entries; one or more entries correspond to accessed data above a threshold; and the method may further include writing data corresponding to the accessed data above the threshold from the storage medium to the memory medium and / or storing data corresponding to the accessed data above the threshold in the memory medium. The one or more entries may correspond to accessed data below a threshold; and the method may further include modifying data corresponding to the accessed data below the threshold from the memory medium to the storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Citation of Related Applications

[0002] This application claims the benefit and priority of U.S. Provisional Patent Application Serial No. 63 / 603,629, filed on November 28, 2023, which is incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to storage devices, and more particularly to systems, methods, and apparatuses for caching on storage devices. Background Art

[0004] A page table is a data structure used by an operating system (OS) that can be used to store mappings between virtual addresses and physical addresses. System memory can be used to store page tables and other data structures. If a storage device is used as extended memory, the storage device can be used to store some data structures.

[0005] The above information disclosed in this background art section is only for enhancing the understanding of the background of the principles of the present invention, and thus may include information that does not constitute the prior art. Summary of the Invention

[0006] In some aspects, the techniques described herein relate to a method that includes: determining that data is related to an operation of an operating system; determining a score for the data; and writing the data to a storage medium based on the score. In some aspects, the data is first data; the score is a first score; and the method further includes: determining that second data is related to an operation of an application; determining a second score for the second data; and writing the data to the storage medium based on the score. In some aspects, the first data uses a first cache; and the second data uses a second cache. In some aspects, the first cache applies a cache replacement policy different from that of the second cache. In some aspects, the data includes at least one page table. In some aspects, the at least one page table includes one or more entries; the one or more entries correspond to data accessed above a threshold; and the method further includes writing data corresponding to the data accessed above the threshold from the storage medium to a memory medium. In some aspects, the at least one page table includes one or more entries; the one or more entries correspond to data accessed above a threshold; and the method further includes storing data corresponding to the data accessed above the threshold in a memory medium. In some aspects, the at least one page table includes one or more entries; the one or more entries correspond to data accessed below a threshold; and the method further includes modifying data corresponding to the data accessed below the threshold from the memory medium to the storage medium.

[0007] In some aspects, the techniques described herein relate to a system that includes a host device, the host device including one or more circuits configured to associate virtual addresses with physical addresses on a memory device; and the memory device including a storage medium and a memory medium; wherein the memory device is configured to perform one or more operations including: receiving data related to an operation of an operating system; determining a score for the data; and writing the data to the storage medium based on the score. In some aspects, the data is first data; the score is a first score; and the memory device is further configured to perform one or more operations including: receiving second data related to an operation of an application; determining a second score for the second data; and writing the data to the storage medium based on the score. In some aspects, the first data uses a first cache; and the second data uses a second cache. In some aspects, the first cache applies a cache replacement policy different from that of the second cache. In some aspects, the data includes at least one page table for associating virtual addresses with physical addresses. In some aspects, the at least one page table includes one or more entries; the one or more entries correspond to data accessed above a threshold; and the memory device is further configured to perform one or more operations including writing data corresponding to the data accessed above the threshold from the storage medium to the memory medium. In some aspects, the at least one page table includes one or more entries; the one or more entries correspond to data accessed above a threshold; and the memory device is further configured to perform one or more operations including storing data corresponding to the data accessed above the threshold in the memory medium. In some aspects, the at least one page table includes one or more entries; the one or more entries correspond to data accessed below a threshold; and the memory device is further configured to perform one or more operations including modifying data corresponding to the data accessed below the threshold from the memory medium to the storage medium.

[0008] In some aspects, the techniques described herein relate to an apparatus that includes a memory medium; a storage medium; and at least one circuit configured to perform one or more operations, the one or more operations including: receiving a data structure related to an operation of an operating system; determining a score of the data structure; and writing at least a portion of the data structure to the memory medium based on the score. In some aspects, the score is a first score; and the at least one circuit is further configured to perform one or more operations, the one or more operations including: receiving data related to an operation of an application; determining a second score of the data; comparing the first score and the second score; and writing the data to the storage medium based on the second score. In some aspects, the data structure uses a first cache; and the data related to the operation of the application uses a second cache. In some aspects, the first cache applies a cache replacement policy different from that of the second cache. In some aspects, the data structure related to the operation of the operating system and the data related to the operation of the application use a cache including at least one of a type and a priority level. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings are not necessarily to scale, and in all the drawings, for illustrative purposes, elements of similar structure or function are generally represented by the same reference numeral or portions thereof. The drawings are only intended to facilitate the description of the various embodiments described herein. The drawings do not depict every aspect of the teachings disclosed herein and do not limit the scope of the claims. To prevent the drawings from becoming obscure, not all components, connections, etc. may be shown, and not all parts may have reference numerals. However, the pattern of component configurations can be readily apparent from the drawings. The drawings, together with the description, illustrate example embodiments of the present disclosure and, together with the description, are used to explain the principles of the present disclosure.

[0010] Figure 1 An embodiment of a storage device scheme according to an example embodiment of the present disclosure is shown.

[0011] Figure 2 Another embodiment of a storage device scheme according to an example embodiment of the present disclosure is shown.

[0012] Figure 3 Another embodiment of a storage device scheme according to an example embodiment of the present disclosure is shown.

[0013] Figure 4 An example of a page table walk according to an example embodiment of the present disclosure is shown.

[0014] Figure 5 An example memory request according to an example embodiment of the present disclosure is shown.

[0015] Figure 6Shows an example address range according to an example embodiment of the present disclosure.

[0016] Figure 7 Shows an example of a register for caching on a storage device according to an example embodiment of the present disclosure.

[0017] Figure 8 Shows an example of a unified cache according to an example embodiment of the present disclosure.

[0018] Figure 9a Shows an example of an operating system cache according to an example embodiment of the present disclosure.

[0019] Figure 9b Shows an example of an application cache according to an example embodiment of the present disclosure.

[0020] Figure 10 Shows an example flowchart for caching on a storage device according to an example embodiment of the present disclosure. Detailed Description

[0021] In some embodiments, the storage device can be used as a device memory (e.g., a memory expander for a host). When the storage device is regarded as a device memory, the host can write data that would normally be written to the device memory to the storage device. Examples of the types of data that can be written to the storage device can include page tables. For example, an application and / or the OS on the host can use virtual addresses to reference the memory on the storage device. However, the storage device can use physical addresses to access the memory. To facilitate the translation of virtual addresses to physical addresses, a data structure called a page table can be used to store the mapping of virtual addresses to physical addresses. In some embodiments, one or more entries of the page table can be stored on the storage device.

[0022] In some embodiments, when data such as page table entries is stored on the storage device, the host may experience increased latency when accessing the memory on the storage device rather than the device memory on the host (e.g., accessing the device memory is generally faster than accessing the memory on the storage device). In some embodiments, the storage device can mitigate some of this latency by utilizing a memory medium (e.g., a cache medium) to store frequently accessed regions of the memory. According to the embodiments of the present disclosure, mechanisms for improving the cache performance of data on the storage device can be used. For example, in some embodiments, methods for distinguishing OS and application accesses and placing higher-priority data on the cache medium can be used. In some embodiments, methods for minimizing the occurrence of moving higher-priority data to slower memories can be used. Additionally, in some embodiments, methods for allowing the OS to convey important information to the storage device to increase the device cache hit rate can be used.

[0023] Figure 1 An embodiment of a storage device solution according to an example embodiment of the present disclosure is shown. Figure 1 The illustrated embodiment may include one or more host devices 100 and one or more storage devices 150 configured to communicate using one or more communication connections 110.

[0024] In some embodiments, the host device 100 may be implemented with any component or combination of components that can utilize one or more features of the storage device 150. For example, the host may be implemented with one or more of a server, a storage node, a compute node, a central processing unit (CPU), a workstation, a personal computer, a tablet computer, a smart phone, etc. or a combination thereof.

[0025] In some embodiments, the storage device 150 may include a communication interface 130, a memory 180 (some or all of which may be referred to as device memory), one or more compute resources 170 (which may also be referred to as computational resources), a device controller 160, and / or device functional circuitry 190. In some embodiments, the device controller 160 may control the overall operation of the storage device 150, including any operations, features, etc. described herein. For example, in some embodiments, the device controller 160 may parse, process, invoke, etc. commands received from the host device 100.

[0026] In some embodiments, the device functional circuitry 190 may include any hardware for implementing the main functions of the storage device 150. For example, the device functional circuitry 190 may include storage media, such as magnetic media (e.g., if the storage device 150 is implemented as a hard disk drive (HDD) or a tape drive), solid state media (e.g., one or more flash devices), optical media, etc. For example, in some embodiments, the storage device may be at least partially implemented as a NAND flash-based solid state drive (SSD), persistent memory (PMEM) (such as cross-grid non-volatile memory), memory with body resistance change, phase change memory (PCM), or any combination thereof. In some embodiments, the device controller 160 may include a media conversion layer, such as a flash translation layer (FTL) for connecting to one or more flash devices. In some embodiments, the storage device 150 may be implemented as a compute storage drive, a compute storage processor (CSP), and / or a compute storage array (CSA).

[0027] As another example, if the storage device 150 is implemented as an accelerator, the device functional circuitry 190 may include one or more accelerator circuits, memory circuits, and the like.

[0028] The computing resources 170 may be implemented with any component or combination of components that can perform operations on data that may be received, stored, and / or generated at the storage device 150. Examples of computing engines may include combinational logic, sequential logic, timers, counters, registers, state machines, complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), embedded processors, microcontrollers, central processing units (CPUs), such as complex instruction set computer (CISC) processors (e.g., x86 processors), and / or reduced instruction set computer (RISC) processors, such as ARM processors, graphics processing units (GPUs), data processing units (DPUs), neural processing units (NPUs), tensor processing units (TPUs), etc., which may run instructions stored in any type of memory and / or implement any type of execution environment, such as containers, virtual machines, operating systems (such as Linux), extended Berkeley packet filter (eBPF) environments, etc., or combinations thereof.

[0029] In some embodiments, the memory 180 may be used, for example, by one or more of the computing resources 170 to store input data, output data (e.g., computation results), intermediate data, transformed data, etc. The memory 180 may be implemented, for example, with volatile memory (such as dynamic random access memory (DRAM), static random access memory (SRAM), etc.) and any other type of memory (such as non-volatile memory).

[0030] In some embodiments, the memory 180 and / or the computing resources 170 may include software, instructions, programs, code, etc. that can be executed, run, etc. using one or more computing resources (e.g., hardware (HW) resources). Examples may include software implemented in any language such as assembly language, C, C++, binary code, FPGA code, one or more operating systems, kernels, environments such as eBPF, etc. The software, instructions, programs, code, etc. may be stored in a repository in, for example, the memory 180 and / or the computing resources 170. In some embodiments, the software, instructions, programs, code, etc. may be downloaded, uploaded, sideloaded, pre-installed, built-in, etc. into the memory 180 and / or the computing resources 170. In some embodiments, the storage device 150 may receive one or more instructions, commands, etc. to select, enable, activate, run, etc. the software, instructions, programs, code, etc. Examples of computing operations, functions, etc. that may be implemented by the memory 180, computing resources 170, software, instructions, programs, code, etc. may include any type of algorithms, data movement, data management, data selection, filtering, encryption and / or decryption, compression and / or decompression, checksum calculation, hash value calculation, cyclic redundancy check (CRC), weight calculation, activation function calculation, training, inference, classification, regression, etc. for artificial intelligence (AI), machine learning (ML), neural networks, etc.

[0031] In some embodiments, the communication interface 120 at the host device 100, the communication interface 130 at the storage device 150, and / or the communication connection 110 may be implemented using any type of interface, protocol, etc. and / or utilize one or more interconnections, one or more networks, a network of networks (e.g., the Internet), etc. or a combination thereof. For example, one or more of the communication connection 110 and / or interface 120 and / or 130 may be implemented using and / or utilize any type of wired and / or wireless communication medium, interface, network, interconnection, protocol, etc., including Peripheral Component Interconnect Express (PCIe), NVMe, NVMe over Fabric (NVMe-oF), Compute Express Link (CXL), and / or coherence protocols (such as CXL.mem, CXL.cache, CXL.io, etc.). Gen-Z, Open Coherent Accelerator Processor Interface (OpenCAPI), Cache Coherent Interconnect for Accelerators (CCIX), etc., Advanced eXtensible Interface (AXI), Direct Memory Access (DMA), Remote DMA (RDMA), RDMA over Converged Ethernet (ROCE), Advanced Message Queuing Protocol (AMQP), Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Fibre Channel, InfiniBand, Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, any generation of wireless network including 2G, 3G, 4G, 5G, 6G, etc., any generation of Wi-Fi, Bluetooth, Near Field Communication (NFC), etc. or any combination thereof. In some embodiments, the communication connection 110 may include one or more switches, hubs, nodes, routers, etc.

[0032] In some embodiments, storage device 150 may be implemented in any physical form factor. Examples of form factors may include 3.5 inches, 2.5 inches, 1.8 inches, etc., storage device (e.g., storage drive) form factors, M.2 device form factors, enterprise and data center standard form factors (EDSFF) (which may include, for example, E1.5, E1.1, E3.5, E3.1, E3.5 2T, E3.1 2T, etc.), add-in card (AIC) (e.g., PCIe card (e.g., PCIe expansion card) form factors, including half-height (HH), half-length (HL), half-height, half-length (HHHL), etc.), next generation small form factor (NGSFF), NF1 form factor, compact flash (CF) form factor, secure digital (SD) card form factor, personal computer memory card international association (PCMCIA) device form factor, etc., or combinations thereof. Any computing device disclosed herein may be connected to a system using one or more connectors, such as SATA connectors, SCSI connectors, SAS connectors, M.2 connectors, EDSFF connectors (e.g., 1C, 2C, 4C, 4C+, etc.), U.2 connectors (which may also be referred to as SSD form factor (SSF) SFF-8639 connectors), U.3 connectors, PCIe connectors (e.g., card edge connectors), etc.

[0033] Any storage device disclosed herein may be used in conjunction with one or more personal computers, smart phones, tablet computers, servers, server chassis, server racks, data zones, data centers, edge data centers, mobile edge data centers, and / or any combination thereof.

[0034] In some embodiments, storage device 150 may be implemented with any device that may include or may access memory, storage media, etc., to store data that may be processed by one or more computing resources 170. Examples may include memory expansion and / or buffer devices, such as CXL type 2 and / or CXL type 3 devices, and CXL type 1 devices that may include memory, storage media, etc.

[0035] Figure 2 Another embodiment of a storage device scheme according to an example embodiment of the present disclosure is shown. Figure 2 The elements shown in may be related to Figure 1Elements similar to those shown in [description], where similar elements may be indicated by reference numerals that end with the same numbers, letters, etc. and / or contain the same numbers, letters, etc. In some embodiments, the host device 100 may include an application module 210; and the storage device 150 may include an interface 120, a controller 160, a memory medium 260 (e.g., a cache medium), and / or a storage medium 270. In some embodiments, the interface 120 and / or the controller 160 may be implemented on one or more circuits of the storage device 150. In some embodiments, one or more circuits may include one or more FPGAs, ASICs, and / or SOCs.

[0036] In some embodiments, the storage medium 260 may be a relatively fast memory, such as DRAM, and the storage medium 270 may be a slower non-volatile memory, such as NAND flash. In some embodiments, the memory medium 260 may be used as a cache to store accessed data above a threshold in the faster memory. In some embodiments, the application module 210 may run an application that can access data from the storage device 150 (e.g., send a request to the storage device 150). For example, in some embodiments, the application module 210 may request data from the storage device 150 using an I / O block access request 220 to retrieve data from the storage medium 270. In some embodiments, the application program module 210 may use a memory access request received at the controller 160 to retrieve data from the memory medium 260. Specifically, in some embodiments, in response to receiving a memory access request 230, the storage device 150 may send the request to the controller 160 to check the memory medium 260 for data corresponding to the request. In some embodiments, in response to a cache hit (e.g., data is found on the memory medium 260), the data may be returned from the memory medium 260. In some embodiments, in response to a cache miss (e.g., data is not found on the memory medium 260), the controller 160 may copy the data from the storage medium 270 to the memory medium 260 and return the data from the memory medium 260.

[0037] In some embodiments, the storage device 150 may be advertised as system memory (e.g., device memory). In other words, the storage device 150 may appear to the host device 100 as an additional memory node and be managed by the OS non-uniform memory access (NUMA) memory management. In some embodiments, if the storage device 150 appears to the host device 100 as a memory node, the host device 100 may store data (such as one or more of its data structures) on the storage device 150. In some embodiments, at least a portion of the data structure, such as a page table (e.g., one or more entries of the page table), may be stored on the storage device 150.

[0038] In some embodiments, the translation of virtual addresses to physical addresses can be managed by a memory management hardware unit (MMU). In some embodiments, the MMU can use a cache (e.g., a translation lookaside buffer (TLB)) to store recently accessed page table entries. However, in some embodiments, the number of virtual addresses (e.g., when the host device 100 is attached to a storage device with a large memory capacity) may require a large number of entries in the page table, where the number of entries may not fit into the TLB. Thus, in some embodiments, the page table entries can be stored partially in the system memory (such as the storage device 150). In some embodiments, searching for page table entries on the storage device 150 rather than on the TLB may affect the overall system performance (e.g., page table lookups may be slower on the storage device 150 than on the TLB).

[0039] Figure 3 Another embodiment of a storage device scheme according to an example embodiment of the present disclosure is shown. Figure 3 The elements shown in Figure 1 and Figure 2 may be elements similar to the elements shown in Figure 3 including a CPU 310, an MMU 320, a TLB 330, one or more CPU caches 340, and / or a system memory 350. In some embodiments, the CPU 310, the MMU 320, the TLB 330, one or more CPU caches 340, and / or the system memory 350 may be implemented on a host device (e.g., Figure 1 and Figure 2 the host device 100 in Figure 2 ). In some embodiments, the storage device 150 may further include one or more circuits (e.g., design logic 370), SRAM 362, and / or DRAM 364. In some embodiments, the SRAM 362 and / or the DRAM 364 may be part of the memory medium 260 in Figure 2The controller 160 therein. In some embodiments, the cache controller 378 may include a cache placement unit 380.

[0040] In some embodiments, the MMU 320 may be responsible for some memory operations of the CPU 310. For example, the MMU 320 may be responsible for translating the virtual addresses used by the CPU 310 into physical addresses. In some embodiments, the MMU 320 may use the TLB 330 for some virtual-to-physical translations of addresses. For example, in some embodiments, the TLB 330 may store the most recent translations of virtual addresses to physical addresses. In some embodiments, the TLB 330 may be a part of the MMU 320. In some embodiments, the TLB 330 may store translations between the CPU 310 and one or more CPU caches 340, between one or more CPU caches 340 and the system memory 350, and / or between different levels of one or more CPU caches 340. In some embodiments, when the storage device 150 is used as extended memory, the TLB 330 may also store the translations between the host and the storage device 150. In some embodiments, when a request containing a virtual address is received, the MMU 320 may search the TLB 330 for the virtual address. In some embodiments, if the virtual address is found in the TLB 330, a TLB hit occurs, and the TLB 330 may return the corresponding physical address. In some embodiments, if the virtual address is not found in the TLB 330, e.g., a TLB miss occurs, the page table may be searched. If the address is found in the page table, in some embodiments, the address may be written into the TLB 330.

[0041] In some embodiments, data in one or more CPU caches 340 can be accessed to reduce latency on the host. For example, in some embodiments, some or all of the page tables can be included in one or more CPU caches 340 and / or system memory 350. In some embodiments, the entries of the page table can be grouped into one or more page tables. In other words, the page table can be a multilevel page table, where one or more page table entries are stored in multiple page tables. In some embodiments, the multilevel page table can be hierarchical. In some embodiments, a virtual address can be searched in the top level page table. In some embodiments, if the virtual address is found in the top level page table, the next level page table can be searched. In some embodiments, if the virtual address is found in the next level page table, the lower level page tables can be searched until the last level page table is searched. In some embodiments, if the last level page table does not contain the virtual address, an error, such as a page fault, can be returned. The above process can be referred to as page table traversal. In some embodiments, the page table traversal can be performed by hardware. In some embodiments, if, for example, the multilevel page table has four levels, the page table traversal may require four memory accesses to retrieve the page table entry from the last level page table. In some embodiments, a bitmap is stored in the page table entry, which indicates the presence and / or accessibility of the page in memory, as shown in Table 1.

[0042] bit function _PAGE_PRESENT The page resides in memory and is not swapped out _PAGE_PROTNONE The page resides but is not accessible _PAGE_RW Set if the page can be written _PAGE_USER Set if the page is accessible from user space _PAGE_DIRTY Set if the page is written _PAGE_ACCESSED Set if the page is accessed

[0043] Table 1

[0044] In some embodiments, the storage device 150 may include one or more types of cache media, such as SRAM 362 and DRAM 364. In some embodiments, SRAM 362 and DRAM 364 may each have their own controllers, such as SRAM controller 374 and DRAM controller 376, respectively, for handling communications between the cache controller 378 and SRAM 362 and DRAM 364. For example, in some embodiments, a request for data may be passed by the cache controller 378 to the SRAM controller 374 to search for data on the SRAM 362. In some embodiments, SRAM 362 and DRAM 364 may not be exposed to the host. In other words, the storage device 150 may determine where the data is located. In some embodiments, a memory request may be received by the cache controller 378. In some embodiments, the cache controller 378 may send the request to SRAM 362 and DRAM 364. In some embodiments, if the data is found on the SRAM 362 or DRAM 364, the data may be returned from the SRAM 362 or DRAM 364. In some embodiments, if the data is not found on the SRAM 362 or DRAM 364, the request may be sent to the storage medium 270 using the storage medium I / F 384. In some embodiments, the cache controller 378 may also be responsible for looking up, inserting, and evicting data blocks from the cache media and for managing cache metadata. In some embodiments, the cache controller 378 may maintain a cache policy (e.g., a cache placement policy) for managing the device cache. In some embodiments, the MMU 320 may be responsible for including additional attributes in the memory requests sent to the storage device 150.

[0045] In some embodiments, the storage device 150 may include a cache policy engine or cache access predictor 382. In some embodiments, the cache access predictor 382 may assist the cache controller 378 in improving the cache hit rate. In some embodiments, the cache access predictor 382 may be used to predict future accesses and issue prefetch or eviction commands to the cache controller 378. In some embodiments, the cache controller 378 may provide information about incoming memory requests from the host to the cache access predictor 382 and respond to queries regarding the status of data blocks in the cache.

[0046] Figure 4Shows an example of page table traversal according to an example embodiment of the present disclosure. In some embodiments, a virtual address may include a level 1 offset 410, a level 2 offset 420, a level 3 offset 430, a level 4 offset 440, and / or an offset 470. In some embodiments, in order to find a physical address using a multi-level page table, a page table base register (PTBR) 450 may be used as a starting position. In some embodiments, the PTBR 450 using the level 1 offset 410 may be used to access a page table entry (PTE) 452 in the level 1 page table. In some embodiments, the base address from the PTE 452 and the level 2 offset 420 may be used to access a PTE 454 in the level 2 page table. In some embodiments, the base address from the PTE 454 and the level 3 offset 430 may be used to access a PTE 456 in the level 3 page table. In some embodiments, the PTE 456 and the level 4 offset 440 may be used to access a PTE 458 in the level 4 page table. In some embodiments, the PTE 458 and the offset 470 may be used to obtain a physical address. The physical address may include a frame number 460 and an offset 472. In some embodiments, in order to obtain a physical address from a virtual address, in this example, four memory accesses may be required. In the above example, the multi-level page table includes 4 levels. However, within the scope of the present disclosure, the page table may have a different number of levels. In some embodiments, each process may have its own page table.

[0047] In some embodiments, a page table may be divided into one or more page tables. In some embodiments, one or more of the page tables may be stored on a storage device, such as Figure 1 the storage device 150 in. For a physical address to be searched in the page table, the page table stored on the storage device may be searched. In some embodiments, since the storage device may not be as fast as the device memory on the host, the host may experience latency from accessing the storage device. Therefore, in order to minimize latency, in some embodiments, the storage device may ensure that data above a threshold for access (such as a page table) may be stored in a cache medium rather than in a storage medium on the storage device.

[0048] In some embodiments, when a memory access request is sent by, for example, the OS, the request may include a host physical address, an opcode (e.g., read or write), and other attributes. In some embodiments, a source identifier and a score may also be included in the request. For example, the source identifier may identify where the request was received. For example, if the request is received due to a TLB miss, the source identifier may identify the TLB as the source of the miss. In some embodiments, if the request is received due to a data cache miss, the cache may be identified as the source. In some embodiments, other source identifiers may be used to identify the source of the request. For example, if a memory request is initiated from a data cache miss (i.e., a last-level cache miss) or a TLB miss, the MMU may notify the storage device. In some embodiments, the MMU unit may include additional bits to indicate this information in the request sent to the storage device. In some embodiments, a memory access protocol may be used to provide this information to the storage device. In some embodiments, the source identifier and the score may be added to the protocol or integrated into the current protocol using reserved bits. In some embodiments, one of the attributes may be access type information. In some embodiments, the access type may indicate whether the memory request belongs to the OS or a user's application. In some embodiments, in addition to this information, bits may also be used to indicate the priority score of the memory access. For example, using two bits for the priority attribute, up to four categories (e.g., highest importance, high importance, low importance, lowest importance) may be provided.

[0049] In some embodiments, the MMU may include additional information in the request to the storage device to better manage its device cache. For example, a page table level ID may be included in the memory request. In some embodiments, it may be more efficient to cache entries at a higher level of the page table (e.g., Figure 4 level 1 410 or level 2 420 in

[0050] ). In some embodiments, the MMU may include this information in each memory request. In some embodiments, the memory device may use this information to prioritize the entries in the cache accordingly.

[0051] In some embodiments, additional attributes of each incoming memory request may be included. These additional attributes may carry the necessary information to inform the storage device whether the memory access belongs to the OS or an application. In some embodiments, the attributes may carry information about the importance of the data passed by the software or the OS to the device cache controller. In some embodiments, the MMU may be responsible for including the additional attributes in the memory requests sent to the memory device.

[0052] In some embodiments, one or more circuits on the host (e.g., Figure 3 the MMU 320 in ) may be modified to receive the additional information in the request. For example, in addition to the host physical address, the opcode (e.g., read or write), and other attributes, the source identifier and the score may also be included in the request. The MMU may use this additional information to place the data in the cache and the storage device.

[0053] Figure 5 An example memory request according to an example embodiment of the present disclosure is shown. In some embodiments, the storage device may receive a memory request. In some embodiments, the memory request may include attributes such as the host physical address 510, the memory opcode (read / write) 520, and / or other attributes 550. In some embodiments, the memory request may further include a source ID 530 and a priority score 540. In some embodiments, the source ID 530 may indicate where the memory request originated. For example, if the request originated from a TLB miss, the source ID 530 may have a value indicating a TLB miss, e.g., a first value. If the request originated from a cache miss, the source ID 530 may have a value indicating a cache miss, e.g., a second value. In some embodiments, the source ID 530 may be used to set the priority of the data based on where the request originated. In some embodiments, the request may include a priority score 540. In some embodiments, the priority score 540 may have values of highest priority, high priority, low priority, and / or lowest priority. In some embodiments, the priority score 540 may be used to determine the priority of the data. In some embodiments, the source ID 530 and the priority score 540 may be used to determine the priority of the data.

[0054] Figure 6 An example address range according to an example embodiment of the present disclosure is shown.

[0055] In some embodiments, the OS can know what address ranges belong to the OS and what address ranges are allocated for applications. For example, the OS can have a starting range 610 and an ending range 620. The application can have a starting range 630 and an ending range 640. For data operations from the OS, the address of the data can be between the starting range 610 and the ending range 620. For data operations used by the application, the address of the data can be between the starting range 630 and the ending range 640. Thus, the OS and the application can have separate memory ranges allocated for their respective operations.

[0056] Figure 7 An example of a register for caching on a storage device according to an example embodiment of the present disclosure is shown.

[0057] In some embodiments, one or more control status registers (CSRs) 710 can be used to notify the host of the address ranges of the OS and one or more applications. When a physical address is received, one or more CSRs 710 can be checked to determine the priority of the data corresponding to the physical address. For example, if the access is within the address range of the OS, the OS can be given priority access over, for example, application access. In some embodiments, a software-based solution can be used, where the OS provides information to the memory device by sending one or more CSR commands.

[0058] In some embodiments, a method can be used in which a cache predictor logic uses the OS page table stored on the device to predict future accesses.

[0059] In some embodiments, the system device can identify the OS data structures stored on the device and apply a different caching scheme compared to the data belonging to the application. For example, slowing down access to some of the OS-related data structures (such as page tables) that may reside on the storage medium, and thus, access to the data structures can be slow, affecting the overall performance of the system. This can be because the system equally processes all accesses from the application and the OS from the perspective of the device cache, demoting some of the OS data from the device cache to a slower medium (e.g., the storage medium) to favor less critical data belonging to the application.

[0060] In some embodiments, cache techniques that prioritize OS data structures over application data can be used. For example, when an OS data structure is migrated to slower memory (e.g., from cache memory to storage media), the performance degradation can be high. In some embodiments, the memory device can distinguish between OS and application memory accesses. In some embodiments, OS data structures (such as page tables) resident in the device cache can be used to perform data prefetching and eviction. For example, some of the accesses initiated from the OS can reveal information about future memory accesses, which can be used to increase the device cache hit rate. In some embodiments, the hardware can utilize information about the activity level of memory regions (i.e., pages) to evict idle pages from the device cache to make room for more active pages (i.e., hot pages).

[0061] In some embodiments, the storage device can attempt to minimize the occurrence of demoting OS data structures to slower memory. In some embodiments, this can include using a dedicated cache for OS data structures or using methods that do not evict OS data from the cache media to support application data. For example, the storage device can prefer OS data and application data with a high priority score and place / retain that data in the cache to ensure that it has a lower latency than other data (e.g., application data with a low priority score). In some embodiments, the OS can, for example, use a CSR to transfer important information to the storage device to increase the device cache hit rate.

[0062] In some embodiments, mechanisms can be used to improve the performance of an application through cache policies for hierarchical memory devices. For example, the page tables stored in the storage device can be tracked to identify future data accesses. In some embodiments, this information can be used to prefetch data from slower memory to faster memory (i.e., the cache). In some embodiments, methods can be used that use the page table entries stored in the storage device to identify unused (i.e., idle) pages and evict them from the device cache.

[0063] In some embodiments, a software-based method using the OS can be used. In some embodiments, the method may not require hardware support from the host. In some embodiments, information about different memory regions including a memory region belonging to the OS data structure and a memory region belonging to application data can be transmitted by the OS to the storage device. In some embodiments, the OS and the application can occupy different physical address ranges in the system memory. In some embodiments, once the memory region is defined and set by the OS, the device driver can notify the device by sending a corresponding command. In some embodiments, an add command that notifies the device of the start address and the end address of the region belonging to the OS data structure can be used. When receiving the command, the device can update its internal registers to store the information and use them for future memory references. In some embodiments, one such information can be the range of physical addresses (start address and end address) belonging to different memory regions. In some embodiments, by having the memory range information, the device can filter incoming memory requests based on the physical address of the incoming memory request. In some embodiments, if the OS updates the memory address range (e.g., expands one of the address ranges), it can notify the device by sending a new command to update the device-side registers. In some embodiments, if the device does not receive such a command, the device can ignore the filtering step and process all incoming memory requests equally.

[0064] In some embodiments, the CSR register that specifies the memory region can be exposed to the host system software using a set of memory-mapped addresses. In some embodiments, the exposed CSR can be part of one or more memory address ranges advertised by the storage device. In some embodiments, the CSR location can be at a fixed or partially programmable location (e.g., only the base location is fixed).

[0065] In some embodiments, a hardware-based method using the MMU can be used. In some embodiments, the MMU can be configured to add additional attributes to identify the source, and a priority level can be added to the entries of the MMU, such as page table entries. In other embodiments, a software-based method using the CSR as described above can be used.

[0066] In some embodiments, two cache schemes can be introduced for the device-side cache. For example, in some embodiments, a unified cache can be used. As shown below in Figure 8 the OS and the application can be co-located in the same cache. In some embodiments, a separate cache for the OS and a separate cache for application data can be used.

[0067] Figure 8shows a first cache scheme. In Figure 8 , the OS and the application are co-located in the same cache. In some embodiments, to distinguish between the OS and the application, two attributes may be added. In some embodiments, the first attribute may indicate whether the data belongs to the OS or the application, such as type 840. In some embodiments, a second attribute may be added that includes the priority level of the data set by the OS or the MMU, such as priority level 850. In some embodiments, this information may be supplementary to standard cache metadata, such as a valid bit (e.g., valid 810), a tag bit (e.g., tag 820), and replacement policy information (e.g., a least recently used (LRU) counter) (e.g., replacement policy 830). In some embodiments, the device cache controller may utilize this information in different ways. For example, one approach may be to use this information to evict blocks from the cache to make room for new blocks. In some embodiments, the cache replacement policy may attempt to keep OS data in the cache for a longer time than application data. In some embodiments, this policy may override the baseline cache eviction policy, such as LRU or first in first out (FIFO). In some embodiments, the cache controller may use a hybrid approach that takes into account both the baseline policy and the priority information to select candidates for cache eviction. For example, the cache controller may consider the priority level to decide which block to evict. If all blocks have the same priority, the controller may use the baseline policy (e.g., LRU) to break the tie.

[0068] Figure 9a and Figure 9b shows a second cache scheme. In some embodiments, there may be two different caches: one for application data (e.g., Figure 9a ), and one for OS-related data (e.g., Figure 9b)。In some embodiments, the two caches can have different attributes, such as different sizes, associativities, etc. In some embodiments, each cache can use a different cache policy, such as different replacement policies and write policies (e.g., one can use a write-back policy while the other uses a write-through policy). In some embodiments, each cache can use a different memory technology (e.g., DRAM, SRAM, etc.). In some embodiments, each cache can share attributes and have different attributes. For example, the two caches can have valid bits (e.g., valid bits 910 and 950), tag bits (e.g., tags 920 and 960), and / or replacement policy information (e.g., replacement policies 930 and 970). In some embodiments, the application data cache can have a priority 940. In some embodiments, since the OS-related data cache can contain data that can be considered more important than other data, all OS-related data can be stored in the cache. This ensures that OS-related data stays in the faster memory, thus reducing the latency of OS-related operations.

[0069] In some embodiments, the page table access can be an indicator of upcoming memory accesses. In some embodiments, the physical page number found in the last-level page table entry can be the exact physical address that will be accessed by the host later. In some embodiments, to utilize this knowledge, a method of using the page table information resident in the device can be used to issue prefetch and eviction commands to improve the device cache hit rate. In some embodiments, some of the attributes in the page table entry can carry some useful information for the cache prediction logic to prefetch or evict blocks from the cache.

[0070] In some embodiments, for the prefetch mechanism, the cache predictor logic can use the attributes in the page table entry to issue a prefetch command to bring the data into the cache in advance. For example, the cache predictor logic can issue a prefetch command with the physical address (i.e., frame number) extracted from the page table entry. In some embodiments, the eviction mechanism, the eviction command, can be based on the activity level of the page table entry.

[0071] In some embodiments, the device cache controller may notify the cache predictor logic of page table memory accesses. In some embodiments, additional bits that separate OS memory accesses from application data may be integrated in each memory request. In some embodiments, the cache predictor logic may use attributes in the page table entries to issue prefetch commands to bring the data into the cache in advance. For example, the cache predictor logic may issue a prefetch command with a physical address (e.g., frame number) extracted from the page table entry. In some embodiments, addresses adjacent to the physical address specified in the page table entry may be prefetched. In some embodiments, if the prefetched address already exists in the cache, the cache controller may notify the cache predictor logic and discard the prefetch request.

[0072] In some embodiments, the cache predictor logic may evict blocks from the cache to improve cache efficiency. In some embodiments. Similar to prefetching, the cache predictor logic may use some attributes in the page table entries to make eviction decisions.

[0073] In some embodiments, attributes (such as the _PAGE_ACCESSED attribute) may be used to decide whether to retain or evict a block from the cache. In some embodiments, if a page is accessed, this attribute may be set by the storage device. For example, a zero bit may indicate that the page has not been accessed. In some embodiments, the cache predictor logic may use a timer-based eviction strategy where pages that have not been accessed within a certain time window are evicted from the cache. In some embodiments, another attribute (such as the _PAGE_DIRTY attribute) may be used to make eviction decisions. For example, this attribute may be set when a page is written to. In some embodiments, the cache predictor may only evict those pages that are not dirty.

[0074] In some embodiments, the cache controller may use priority information to decide whether to cache certain data. In some embodiments, a flexible policy that allows caching based on the available empty blocks in the cache may be used. For example, if the cache has many empty blocks, it may allow data with all different priorities to be stored in the cache. However, when the cache is half full or nearly full, the policy may change to only allow the highest priority blocks to be stored in the cache.

[0075] In some embodiments, a timer-based replacement policy may be used. In some embodiments, the cache controller may use a timer to evict blocks from the cache after a certain number of cycles. In some embodiments, the cache controller may select a longer timer period for blocks with higher priority to allow those blocks to stay in the cache for a longer time.

[0076] Figure 10 A flowchart of prefetching data according to an example embodiment of the present disclosure is shown. For example, at block 1010, according to an embodiment, a storage device may receive memory address information. For example, the storage device may receive prefetch memory address information from a host device. In some embodiments, the memory address information may be information related to application data used by the host. In some embodiments, the host may send other information that the storage device may use to determine the address of the cache data to be loaded onto the storage device. In some embodiments, the address information may correspond to data for which any logic may be used to determine the next data. In some embodiments, the address information may include one or more addresses. In some embodiments, the address information may be an indication of an address that the storage device may convert to an address on the storage device. In some embodiments, the storage device may use a table to convert the address information on the storage device. In some embodiments, the data for determining the address on the storage device may be sent by the host, an internal process, or the storage device itself.

[0077] At block 1020, according to an embodiment, the storage device may store the address information in a buffer (e.g., a prefetcher queue). In some embodiments, the prefetcher may include a buffer. In some embodiments, the prefetcher may receive the address information from the host and use the address information to fill the buffer. In some embodiments, the buffer may be a circular buffer or some other queue for storing address information. In some embodiments, the storage device may load the address to be retrieved from the storage device. In some embodiments, the buffer may include a message from the host in the storage device. In some embodiments, the buffer may receive an indication of an address that may be used to determine the actual address. Although a first-in, first-out (FIFO) queue is described, in some embodiments, the prefetcher may be an ordered list available for storing address information on the storage device. In some embodiments, the buffer may contain other information for retrieving the address on the storage device.

[0078] At block 1030, according to an embodiment, data may be loaded from a storage medium to a memory medium based on the memory address information. For example, if the buffer contains a memory address, that memory address may be used to load data from the storage medium to the cache medium. In some embodiments, the buffer may contain other information for determining the address information on the storage medium. For example, the buffer may contain an address range.

[0079] Thus, in some embodiments, the access latency of the SSD using cache technology can be minimized. In some embodiments, application performance can be improved by prioritizing critical information and non-critical information for caching. In some embodiments, prefetching and eviction can be used to improve cache performance. In some embodiments, the total cost of ownership can be reduced by using storage devices to provide large memory capacities (e.g., extended memory).

[0080] In some embodiments, the cache medium can be accessed by software using load and / or store instructions, while the storage medium can be accessed by software using read and / or write instructions.

[0081] In some embodiments, a memory interface and / or protocol (such as any generation of Double Data Rate (DDR) (e.g., DDR4, DDR5, etc.), DMA, RDMA, Open Memory Interface (OMI), CXL, Gen-Z, etc.) can be used to access the cache medium, while a storage interface and / or protocol (such as Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), NVMe, NVMe-oF, etc.) can be used to access the storage medium.

[0082] Although some embodiments may be described in the context of a cache medium that can be implemented with a cache medium such as DRAM, in other embodiments, other types of media (e.g., storage media) can be used for the cache medium. For example, in some embodiments, some or all of the memory medium 260 can be implemented with a medium other than a cache medium that can have one or more relative characteristics (e.g., relative to the storage medium 270), which can make one or both of the memory medium 260 more suitable for its corresponding function. For example, in some embodiments, the storage medium 270 can have a relatively high capacity, low cost, etc., while some or all of the memory medium 260 can have a relatively low access latency, which can make it relatively more suitable for use as a cache.

[0083] The storage device 150 and any other device disclosed herein can be used in conjunction with one or more personal computers, smart phones, tablet computers, servers, server chassis, server racks, data areas, data centers, edge data centers, mobile edge data centers, and / or any combination thereof.

[0084] Any functionality described herein, including any user functionality, device functionality, etc. (e.g., any control logic) can be implemented in hardware, software, firmware, or any combination thereof, including, for example, hardware and / or software combinational logic, timing logic, timers, counters, registers, state machines, volatile memory (such as DRAM and / or SRAM), non-volatile memory (including flash memory), persistent memory (such as cross-grid non-volatile memory, memory with bulk resistance change, PCM, etc.) and / or any combination thereof, complex programmable logic devices (CPLDs), FPGAs, ASICs, central processing units (CPUs), including CISC processors (such as x86 processors) and / or RISC processors, such as ARM processors, graphics processing units (GPUs), neural processing units (NPUs), tensor processing units (TPUs), data processing units (DPUs), etc., that execute instructions stored in any type of memory. In some embodiments, one or more components can be implemented as a system-on-chip (SoC).

[0085] Some of the embodiments disclosed above have been described in the context of various implementation details, such as devices implemented as storage devices that can use specific interfaces, protocols, and / or the like, but the principles of the present disclosure are not limited to these or any other specific details. For example, some functionality has been described as being implemented by certain components, but in other embodiments, the functionality can be distributed among different systems and components at different locations and have various user interfaces. Certain embodiments have been described as having specific processes, operations, etc., but these terms also cover embodiments in which the specific processes, operations, etc. can be implemented with multiple processes, operations, etc., or embodiments in which multiple processes, operations, etc. can be integrated into a single process, step, etc. A reference to a component or element can refer only to a part of the component or element. For example, a reference to a block can refer to the entire block or one or more sub-blocks. The use of terms such as "first" and "second" in this disclosure and the claims can be used only for the purpose of differentiating the elements they modify and may not indicate any spatial or temporal order, unless it is obvious from the context. In some embodiments, a reference to an element can refer to at least a part of the element, e.g., "based on" can refer to "at least partially based on", etc. A reference to a first element may not imply the existence of a second element. The principles disclosed herein have independent utility and can be embodied separately, and not every embodiment can utilize every principle. However, the principles can also be embodied in various combinations, some of which can amplify the benefits of the individual principles in a synergistic manner. The various details and embodiments described above can be combined to produce additional embodiments in accordance with the inventive principles of this patent disclosure.

[0086] In some embodiments, a portion of an element may refer to less than the element or all of the elements. A first portion of an element and a second portion of the element may refer to the same portion of the element. The first portion of the element and the second portion of the element may overlap (e.g., a portion of the first portion may be the same as a portion of the second portion).

[0087] In the embodiments described herein, the operations are example operations and may involve various additional operations not explicitly shown. In some embodiments, some of the shown operations may be omitted. In some embodiments, one or more of the operations may be performed by components other than those shown herein. Additionally, in some embodiments, the chronological order of the operations may be changed. Further, the figures are not necessarily drawn to scale.

[0088] The principles disclosed herein may have independent utility and may be embodied separately, and not every embodiment may utilize every principle. However, the principles may also be embodied in various combinations, some of which may amplify the benefits of the individual principles in a synergistic manner.

[0089] In some embodiments, the latency of a storage device may refer to the delay between the storage device and the processor when accessing memory. Additionally, the latency may include delays caused by hardware, such as the read / write speed of accessing the storage device, and / or the structure of an arrayed storage device that generates individual delays when reaching the respective elements of the array. For example, a first storage device in the form of DRAM may have a faster read / write speed than a second storage device in the form of a NAND device. Further, the latency of the storage device may change over time based on conditions such as relative network load and the performance of the storage device over time, as well as environmental factors (such as changing temperature that affects the delay on the signal path).

[0090] Although some example embodiments may be described in the context of specific implementation details, such as a processing system that may implement a NUMA architecture, storage devices and / or pools that may be connected to the processing system using an interconnect interface and / or protocol such as CXL, etc., the principles are not limited to these example details and may be implemented using any other type of system architecture, interface, protocol, etc. For example, in some embodiments, any type of interface and / or protocol may be used to connect one or more storage devices, including Peripheral Component Interconnect Express (PCIe), Non-Volatile Memory Express (NVMe), NVMe-over-fabric (NVMe oF), Advanced eXtensible Interface (AXI), UltraPath Interconnect (UPI), Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Remote Direct Memory Access (RDMA), RDMA over Converged Ethernet (ROCE), Fibre Channel, InfiniBand, Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, etc. or any combination thereof. In some embodiments, the interconnect interface may be implemented using one or more memory semantics and / or memory coherence interfaces and / or protocols, the one or more memory semantics and / or memory coherence interfaces and / or protocols including one or more CXL protocols, such as CXL.mem, CXL.io, and / or CXL.cache, Gen-Z, Cache Coherent Interconnect for Accelerators (CAPI), Cache Coherent Interconnect for Accelerators (CCIX), etc. or any combination thereof. Any storage device may be implemented using one or more of any type of storage device interface, including DDR, DDR2, DDR3, DDR4, DDR5, LPDDRX, Open Memory Interface (OMI), NVLink, High Bandwidth Memory (HBM), HBM2, HBM3, etc.

[0091] In some embodiments, any one of or components in a storage device, memory pool, host, etc. can be implemented in any physical and / or electrical configuration and / or form factor, such as a stand-alone device, an add-on card such as a PCIe adapter or expansion card, an insertion device such as a connector and / or slot that can be inserted into a server chassis (e.g., a connector on the backplane and / or midplane of a server or other device). In some embodiments, any one of or components in a storage device, memory pool, host, etc. can be implemented using any connector configuration for an interconnect interface (such as a SATA connector, SCSI connector, SAS connector, M.2 connector, U.2 connector, U.3 connector, etc.) for a form factor of a storage device (such as 3.5 inches, 2.5 inches, 1.8 inches, M.2, enterprise and data center SSD form factor (EDSFF), NF1, etc.). Any device disclosed herein can be implemented in whole or in part using a server chassis, server rack, data area, data center, edge data center, mobile edge data center, and / or any combination thereof and / or used in conjunction therewith. In some embodiments, any one of or components in a storage device, memory pool, host, etc. can be implemented as a CXL Type 1 device, CXL Type 2 device, CXL Type 3 device, etc.

[0092] In some embodiments, any function described herein (including, for example, any logic for implementing tiering, device selection, etc.) can be implemented in hardware, software, or a combination thereof, including combinational logic, sequential logic, one or more timers, counters, registers, and / or state machines, one or more CPLDs, FPGAs, ASICs, CPUs (such as CISC processors (such as x86 processors) and / or RISC processors (such as ARM processors), GPUs, NPUs, TPUs, etc.), executing instructions stored in any type of memory, or any combination thereof. In some embodiments, one or more components can be implemented as a system on a chip (SOC).

[0093] In this disclosure, numerous specific details are set forth in order to provide a thorough understanding of the disclosure, but the disclosed aspects may be practiced without these specific details. In other instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the subject matter disclosed herein.

[0094] References to "one embodiment" or "an embodiment" in the present specification mean that the particular features, structures, or characteristics described in connection with the embodiment may be included in at least one embodiment disclosed herein. Thus, the phrases "in one embodiment", "in an embodiment", "according to one embodiment" (or other phrases with similar meanings) that appear throughout the present specification may not necessarily all refer to the same embodiment. In addition, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the word "exemplary" means "serving as an example, instance, or illustration". Any embodiment described herein as "exemplary" should not be construed as necessarily being preferred or advantageous over other embodiments. Additionally, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Further, depending on the context discussed herein, singular terms may include the corresponding plural forms, and plural terms may include the corresponding singular forms. Similarly, hyphenated terms (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) may occasionally be used interchangeably with their corresponding non-hyphenated versions (e.g., "two dimensional", "predetermined", "pixel specific", etc.), and capitalized entries (e.g., "Counter Clock", "Row Select", "PIXOUT", etc.) may be used interchangeably with their corresponding non-capitalized versions (e.g., "counter clock", "row select", "pixout", etc.). Such occasional interchangeable use should not be regarded as inconsistent with each other.

[0095] In addition, depending on the context discussed herein, singular terms may include the corresponding plural forms, and plural terms may include the corresponding singular forms. It should also be noted that the various figures (including component diagrams) shown and discussed herein are for illustrative purposes only and are not drawn to scale. For example, for clarity, the dimensions of some elements may be exaggerated relative to other elements. Additionally, if deemed appropriate, reference numerals are repeated in the figures to indicate corresponding and / or similar elements.

[0096] The terms used herein are for the purpose of describing some example embodiments only and are not intended to limit the claimed subject matter. As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. When used in this specification, the terms "comprises" and / or "comprising" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.

[0097] When an element or layer is referred to as being "on," "connected to," or "coupled to" another element or layer, it can be directly on, connected or coupled to the other element or layer, or intervening elements or layers may be present. In contrast, when an element is referred to as being "directly on," "directly connected to," or "directly coupled to" another element or layer, no intervening elements or layers are present. The same reference numerals always refer to the same elements. As used herein, the term "and / or" can include any and all combinations of one or more of the associated listed items.

[0098] As used herein, the terms "first," "second," etc. are used as labels for the nouns that follow and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless explicitly defined as such. Additionally, the same reference numerals may be used across two or more figures to refer to components, assemblies, blocks, circuits, units, or modules having the same or similar functionality. However, such usage is merely for the sake of simplicity of illustration and ease of discussion; it does not mean that the construction or architectural details of such components or units are the same in all embodiments, or that such commonly referenced parts / modules are the only way to implement some of the example embodiments disclosed herein.

[0099] The term "module" can refer to any combination of software, firmware, and / or hardware configured to provide the functionality described herein in connection with the module. For example, software can be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any of the embodiments described herein can include, for example, components, hardwired circuitry, programmable circuitry, state machine circuitry, and / or firmware that stores instructions executed by the programmable circuitry, either individually or in any combination. A module can be embodied, jointly or singly, as circuitry that forms part of a larger system, such as, by way of example and not limitation, an integrated circuit (IC), a system on a chip (SoC), components, and the like. Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, e.g., one or more modules of computer program instructions, encoded on a computer storage medium for execution by, or to control the operation of, a data processing apparatus. Alternatively or additionally, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver apparatus for execution by the data processing apparatus. A computer storage medium can be, or include, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Moreover, while a computer storage medium is not a propagated signal, a computer storage medium can be the source or destination of computer program instructions encoded in an artificially generated propagated signal. The computer storage medium can also be one or more separate physical components or media (such as, for example, multiple CDs, disks, or other storage devices), or be included in one or more of them. Additionally, the operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0100] Although this specification may include many specific implementation details, the implementation details should not be construed as limiting the scope of any claimed subject matter, but rather as descriptions of features specific to particular embodiments. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination within a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Moreover, although the features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be deleted from the combination, and the claimed combination can be directed to a sub-combination or variation of a sub-combination.

[0101] Similarly, although the operations are depicted in the drawings in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order, or that all of the illustrated operations be performed, to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Additionally, the separation of various system components in the foregoing embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems generally can be integrated together in a single software product or packaged into multiple software products.

[0102] Accordingly, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired result. Additionally, the processes depicted in the figures need not be in the particular order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.

[0103] Although certain exemplary embodiments have been described and shown in the drawings, it should be understood that such embodiments are merely illustrative and that the scope of the disclosure is not limited to the embodiments described or shown herein. The invention may be modified in arrangement and detail without departing from the inventive concept, and such changes and modifications are considered to fall within the scope of the appended claims.

Claims

1. A method for caching on a storage device, comprising: Determine that the data is relevant to the operation of the operating system; determining a score for the data; as well as The data is written to a memory medium based on the score, wherein the memory medium is used as an extended memory.

2. The method according to claim 1, in, The data is first data; wherein the score is a first score; and Wherein, the method further comprises: determining that the second data is related to the operation of the application; determining a second score for the second data; and The data is written to a storage medium based on the score.

3. The method according to claim 2, in, The first data uses a first cache; wherein the second data uses a second cache; and The first cache applies a cache replacement strategy different from that of the second cache.

4. The method according to claim 2, in, The first data and the second data use a cache including at least one of a type and a priority level.

5. The method according to claim 1, wherein: The data includes at least one page table.

6. The method according to claim 5, in, The at least one page table includes one or more entries; wherein the one or more entries correspond to data accessed above a threshold; and The method further includes writing data corresponding to the accessed data above a threshold from a storage medium to the memory medium.

7. The method according to claim 5, in, The at least one page table includes one or more entries; wherein the one or more entries correspond to data accessed above a threshold; and The method further includes storing data corresponding to the accessed data above a threshold in the memory medium.

8. The method according to claim 5, in, The at least one page table includes one or more entries; wherein the one or more entries correspond to data accessed below a threshold; and The method further includes modifying data corresponding to the accessed data below a threshold from the memory medium to a storage medium.

9. A system for caching on a storage device, comprising: a host device comprising one or more circuits configured to associate a virtual address with a physical address on a memory device; and The memory device includes a storage medium and a memory medium; The memory device is configured to perform one or more operations, wherein the one or more operations include: receiving data related to the operation of the operating system; determining a score for the data; and The data is written to the memory medium based on the score, wherein the memory medium serves as expanded memory on the host device.

10. The system according to claim 9, in, The data is first data; wherein the score is a first score; and The memory device is further configured to perform one or more operations, wherein the one or more operations include: receiving second data related to the operation of the application; determining a second score for the second data; and The data is written to the storage medium based on the score.

11. The system according to claim 10, in, The first data uses a first cache; wherein the second data uses a second cache; and The first cache applies a cache replacement strategy different from that of the second cache.

12. The system according to claim 10, in, The first data and the second data use a cache including at least one of a type and a priority level.

13. The system according to claim 9, wherein: The data includes at least one page table for associating the virtual address with a physical address.

14. The system according to claim 13, in, The at least one page table includes one or more entries; wherein the one or more entries correspond to data accessed above a threshold; and The memory device is further configured to perform one or more operations, wherein the one or more operations include writing data corresponding to the accessed data above a threshold from the storage medium to the memory medium.

15. The system according to claim 13, in, The at least one page table includes one or more entries; wherein the one or more entries correspond to data accessed above a threshold; and The memory device is further configured to perform one or more operations, wherein the one or more operations include storing data corresponding to the accessed data above a threshold in the memory medium.

16. The system according to claim 13, in, The at least one page table includes one or more entries; wherein the one or more entries correspond to data accessed below a threshold; and The memory device is further configured to perform one or more operations, the one or more operations comprising modifying data corresponding to the accessed data below a threshold from the memory medium to the storage medium.

17. A device for caching on a storage device, comprising: Storage media; Storage media; and At least one circuit is configured to perform one or more operations, the one or more operations comprising: receiving a data structure related to the operation of the operating system; determining a score for the data structure; and At least a portion of the data structure is written to the memory medium based on the score, wherein the memory medium functions as an expansion memory.

18. The device according to claim 17, in, The score is a first score; as well as The at least one circuit is further configured to perform one or more operations, wherein the one or more operations include: receiving data related to the operation of the application; determining a second score for the data; comparing the first score and the second score; and The data is written to the storage medium based on the second score.

19. The device according to claim 17, in, The data structure uses a first cache; wherein data related to the operation of the application uses a second cache; and The first cache applies a cache replacement strategy different from that of the second cache.

20. The apparatus according to claim 17, in, The data structures related to the operation of the operating system and the data structures related to the operation of the application use a cache including at least one of a type and a priority level.