System and method for high performance prefetching
By using per-file data structure and user-level runtime in the operating system, cross-layer low-interference and high-performance prefetching is achieved, which solves the inefficiency of existing storage systems in terms of fast system consistency and data loading, and improves the efficiency of data prefetching and the performance of storage systems.
Patent Information
- Application Number
- CN202411571028.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-27
- Filing Date
- 2024-11-06
- Publication Date
- 2025-05-27
AI Technical Summary
Existing storage systems have inefficiencies in rapid system consistency and data loading, especially when prefetching data, which is prone to inefficiency and bottlenecks.
By introducing per-file data structures, such as bitmaps, applications can transparently track and share file status in the page cache, and achieve cross-layer low-interference and high-performance prefetching through prefetch progress communication between user-level runtime and the operating system.
This method improves the efficiency of data prefetching, reduces cache misses and I/O bottlenecks, and avoids system calls and locking bottlenecks, optimizing the performance of the storage system.
Smart Images

Figure CN120045528A_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application Serial No. 63 / 603,115, filed on November 27, 2023, entitled “CROSS-LAYERED AND LOW-INTERFERENCE HIGH-PERFORMANCE PREFETCHING,” which is incorporated herein by reference. Technical Field
[0003] The present disclosure relates generally to operating systems, and more particularly to systems, methods, and apparatus for prefetching data in an operating system. Background Art
[0004] Storage systems face the problem of continuous need for faster systems. Consistency and loading data for such systems are also issues for such storage systems.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the invention and therefore it may contain information that does not constitute prior art. Summary of the invention
[0006] According to some embodiments, the method and apparatus may include requesting, by the application, a data structure from an operating system (OS) via a runtime level interface for prefetching data for the application. The method may also include receiving, by the application, the data structure from the OS via a runtime level interface. In some embodiments, the method may include prefetching, by the application, data to be used in the application based on information in the data structure.
[0007] In some embodiments, the data structure may be a per-file data structure. According to some embodiments, the per-file data structure may be a bitmap. In some embodiments, the bits in the bitmap may be mapped to blocks in the files of the per-file data structure. The per-file data structure may be updated by the OS during read, write, and pre-fetch operations. In some embodiments, the application may use an application-level copy of the per-file data structure to check cached pages. The method may include updating the per-file data structure by the application based on the application's use of the data, and sending the updated per-file data structure to the OS. In some embodiments, the application may use a counter to track and update the state of the data to be used.
[0008] According to some embodiments, a computing device may include a processor, a memory storing instructions, and when executed by the processor, the instructions configure the device to: request a data structure from an OS via a runtime level interface by an application for prefetching data for the application. The device may also receive the data structure from the OS via the runtime level interface by the application. The device may also prefetch data to be used in the application based on information in the data structure.
[0009] In some embodiments, the data structure may be a per-file data structure. The per-file data structure may be a bitmap. The bits in the bitmap may be mapped to blocks in the file of the per-file data structure. The per-file data structure may be updated by the OS during read, write, and pre-fetch operations. In some embodiments, the application may use an application-level copy of the per-file data structure to check cached pages. In some embodiments, the memory stores instructions that, when executed by the processor, may further configure the device to: update the per-file data structure by the application based on the application's use of the data, and send the updated per-file data structure to the OS.
[0010] The system may include an OS. The system may also include a memory. In some embodiments, the system may include a user-level runtime having a runtime-level interface, the runtime-level interface being configured to request a data structure from the OS via the runtime-level interface by the application for prefetching data for the application. The system also includes receiving the data structure from the OS via the runtime-level interface by the application. In some embodiments, the system may include prefetching data to be used in the application based on information in the data structure by the application.
[0011] In some embodiments, the data structure may be a per-file data structure. In some embodiments, the per-file data structure may be a bitmap. The bits in the bitmap may be mapped to blocks in the files of the per-file data structure. The per-file data structure may be updated by the OS during read, write, and pre-fetch operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To easily identify the discussion of any particular element or act, the most significant digit or digits in a reference number refer to the figure number that first introduces the element.
[0013] Figure 1 An exemplary prefetch OS according to an example embodiment of the present disclosure is shown.
[0014] Figure 2 Exemplary code for an exemplary function call according to an example embodiment of the present disclosure is shown.
[0015] Figure 3 An exemplary prefetch system according to an example embodiment of the present disclosure is shown.
[0016] Figure 4 An example embodiment of a user-level runtime that performs predictive prefetching according to an example embodiment of the present disclosure is shown.
[0017] Figure 5 A flow chart for initialization and prediction-based prefetching according to an example embodiment of the present disclosure is shown.
[0018] Figure 6 A flow chart for managing data structures used in prefetching according to an example embodiment of the present disclosure is shown.
[0019] Figure 7 is an example schematic diagram of a system for prefetching data according to an example embodiment of the present disclosure. DETAILED DESCRIPTION
[0020] In the following detailed description, many specific details are set forth in order to provide a thorough understanding of the present disclosure. However, those skilled in the art will appreciate that the disclosed aspects may be practiced without these specific details. In other instances, well-known methods, processes, components, and circuits are not described in detail to avoid obscuring the subject matter disclosed herein.
[0021] References throughout this specification to "one embodiment" or "an embodiment" mean that the specific features, structures, or characteristics described in conjunction with the embodiment may be included in at least one embodiment disclosed herein. Therefore, the phrases "in one embodiment" or "in an embodiment" or "according to an embodiment" (or other phrases with similar meanings) that appear in various places throughout this specification may not necessarily all refer to the same embodiment. In addition, specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In this regard, as used herein, the word "exemplary" means "used as an example, instance, or illustration". Any embodiment described herein as "exemplary" should not be interpreted as necessarily being preferred or advantageous over other embodiments. In addition, specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. In addition, depending on the context discussed herein, a singular term may include a corresponding plural form, and a plural term may include a corresponding singular form. Similarly, hyphenated terms (e.g., "two-dimensional", "pre-determined", "pixel-specific", etc.) may occasionally be used interchangeably with corresponding non-hyphenated versions (e.g., "two dimensional", "pre determined", "pixel specific", etc.), and capitalized terms (e.g., "Counter Clock", "Row Select", "PIXOUT", etc.) may be used interchangeably with corresponding non-capitalized versions (e.g., "counter clock", "row select", "pixout", etc.). Such occasional interchangeable usage should not be considered inconsistent with each other.
[0022] In addition, depending on the context discussed herein, singular terms may include corresponding plural forms, and plural terms may include corresponding singular forms. It should also be noted that the various drawings (including component diagrams) shown and discussed herein are for illustrative purposes only and are not drawn to scale. Similarly, various waveforms and timing diagrams are shown only for illustrative purposes. For example, for clarity, the size of some elements may be enlarged relative to other elements. In addition, if it is considered appropriate, the reference numerals are repeated in the drawings to indicate corresponding and / or similar elements.
[0023] The terms used herein are used only for the purpose of describing some example embodiments and are not intended to limit the claimed subject matter. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that when used in this specification, the terms "comprises" and / or "comprising" specify the presence of the features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. As used herein, the terms "first", "second", etc. are used as labels for the nouns preceding them, and do not imply any type of ordering (e.g., space, time, logic, etc.) unless explicitly defined as such. In addition, the same reference numerals may be used across two or more figures to refer to parts, components, blocks, circuits, units or modules having the same or similar functions. However, such use is only for simplicity of description and ease of discussion; this does not mean that the construction or architectural details of such components or units are the same in all embodiments, or that such commonly referenced parts / modules are the only way to implement some example embodiments disclosed herein.
[0024] It should be understood that when an element or layer is referred to as being on, "connected to" or "coupled to" another element or layer, it can be directly on, directly connected to or coupled to another element or layer, or there can be intermediate elements or layers. In contrast, when an element is referred to as being "directly on," "directly connected to" or "directly coupled to" another element or layer, there are no intermediate elements or layers. The same reference numerals always refer to the same elements. As used herein, the term "and / or" includes any and all combinations of one or more associated listed items.
[0025] As used herein, the terms "first", "second", etc. are used as labels for the nouns that precede them and do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) unless explicitly defined as such. In addition, the same figure mark may be used across two or more figures to refer to parts, components, blocks, circuits, units, or modules with the same or similar functions. However, such use is only for simplicity of illustration and ease of discussion; it does not mean that the construction or architectural details of such components or units are the same in all embodiments, or that such commonly referenced parts / modules are the only way to implement some example embodiments disclosed herein.
[0026] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the subject matter belongs. It will be further understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art, and will not be interpreted in an idealized or overly formal sense unless explicitly so defined herein.
[0027] As used herein, the term "module" refers to any combination of software, firmware, and / or hardware configured to provide the functionality described herein in conjunction with the module. For example, software may be embodied as a software package, code, and / or instruction set or instructions, and the term "hardware" as used in any embodiment described herein may include, for example, an assembly, hardwired circuits, programmable circuits, state machine circuits, and / or firmware storing instructions executed by programmable circuits, either alone or in any combination. Modules may be collectively or individually embodied as circuits that form part of a larger system, such as, but not limited to, an integrated circuit (IC), a system on a chip (SoC), an assembly, and the like.
[0028] Prefetching can be used to load data into applications by the operating system (OS) for faster access by loading into random access memory (RAM) from local or remote storage drives. Prefetching sometimes fails to retrieve complete files. Some prefetching systems may prefetch data for applications inefficiently and cause bottlenecks or conflicting locks / access to files. Data retrieval systems that retrieve data before use may not often retrieve complete files. Lack of complete files can lead to inefficiencies as more and more data files are retrieved by applications above the OS.
[0029] Embodiments described herein include methods and systems for user-level software to obtain visibility of cache and input / output (I / O) prefetch status by transparently tracking and sharing the status of each file in the page cache, as well as the prefetch progress between the OS and the user-level runtime. Embodiments include parallel prefetching and I / O via, for example, a range bitmap with a scalable range tree. By having a range tree with per-node ranges and per-node locks, multiple application threads using the same or different file descriptors can concurrently access non-conflicting ranges of a file. Alternative data structures such as arrays, linked lists, records, hash tables, graphs, binary trees, Adelson-Velsky and Landis (AVL) trees, etc. may also be used.
[0030] Embodiments include a cross-layer system design that enables visibility of I / O prefetch status between applications and the OS. In some aspects, the cross-layer system may include a system that communicates between the OS or system-level software and a higher-level software application that works with an application or user interface. Some embodiments described herein may track and share prefetch progress between the OS and user-level software. In some embodiments, the user-level cross-layer prefetch runtime may transparently capture the prefetch and / or lost I / O page status for each file in the cache. In addition, some embodiments may capture I / O access patterns and automatically issue prefetch requests for lost data pages to reduce cache misses and I / O. In some embodiments, the disclosed system may avoid and / or reduce prefetch system calls, thereby reducing system call and locking bottlenecks. In addition, according to some embodiments, prefetch bottlenecks on remote and decomposed storage devices may be reduced.
[0031] Figure 1 An exemplary prefetch OS 100 is shown according to an example embodiment of the present disclosure. Figure 1 An exemplary high-level depiction of per-file data structures and cross-layer visibility between the OS and runtime levels implemented by the embodiments described herein is shown. According to some examples, the prefetch OS 100 may include an application 102, a transparent interception library 104, a user-level runtime 106, an OS module 108, and a storage device 110. The user-level runtime 106 can optimize the input / output prefetching of each file of the application. For the files of the application, the user-level runtime 106 can use the transparent interception library 104 to modify the application programming interface (API) call, and change the arguments passed (by the runtime to the OS) as a layer to the portable OS interface (POSIX) I / O, and the prefetch call issued by the application 102.
[0032] When the file is opened, the user-level runtime 106 can initialize one or more files and their file descriptor structures. When a system call or I / O request (e.g., read() 116 or write()) to a file is issued, the predictor for the embodiments described herein can identify the access pattern of the application to the file and recommend the bytes to be pre-fetched. The predictor can include a prediction of the future use of the application to the file, and how much data or which data should be pre-fetched based on the usage pattern. Then, the user-level runtime 106 can issue a readahead_info() 114 system call to pre-fetch blocks, and derive the per-file cache state of the OS module 108. The predictor can use the same application thread, and the pre-fetch call can be issued using a dedicated background thread. In some embodiments, the user-level runtime 106 can also be used to adapt to the memory budget and perform active pre-fetching and cache eviction.
[0033] OS-level components (such as OS modules 108) can maintain cache states and prefetch related information, which can be exported to the user-level runtime 106. OS modules 108 can separate regular I / O paths and prefetch paths to reduce contention between threads issuing regular I / O and threads issuing prefetch operations. In some aspects, contention can include file locks, page cache 122 locks, log locks, and memory manager locks.
[0034] When a file is opened, the OS module 108 may initialize one or more data structures, such as a per-file page cache bitmap 124. For an I / O call (such as read() 116 or write()), the OS module 108 may update one or more data structures (such as a per-file page cache bitmap 124) after fetching or evicting a page from the cache. The OS module 108 may handle the readahead_info() 114 system call by first checking if the requested data block is fully or partially present or absent. The OS module 108 may then adjust the prefetch request based on the data block status. The OS module 108 may then issue the request. Upon return from the I / O call, the OS module 108 may export the per-file cache bitmap 118 of the file to the user-level runtime 106. The OS module 108 may include various optimizations and flexibility in the I / O prefetch path and prefetch parameters, thereby removing some static limitations. Figure 2 Exemplary code for an exemplary function call 200 according to an example embodiment of the present disclosure is shown.
[0035] Figure 3 An exemplary prefetch system 300 according to an exemplary embodiment of the present disclosure is shown. Applications 301 and user-level runtimes 302 can benefit from visibility into cache status and progress / status of prefetch requests because it allows them to evaluate the effectiveness of prefetches and adjust future requests accordingly. The embodiments described herein implement cache visibility and prefetch awareness with minimal performance overhead and no application changes. The embodiments described herein may include a portion of one or more APIs, such as readahead_info() 114. For example, an info parameter may be included, which may be a data structure that stores a bitmap per file and other information about the file. This is used as an example and is not intended to be limiting in any way.
[0036] The readahead_info() 114 system call can extend existing OS prefetch calls, such as Figure 2302 can use readahead_info() 114 system call and information structure to read every file bitmap 326-328 for prediction and future pre-fetch operation. In certain embodiments, every file bitmap 326-328 can be stored in user-level runtime 302, and is updated according to the use and prediction method described herein. In addition, in certain embodiments, every file bitmap 326-328 can be sent to OS module 304.
[0037] Embodiments described herein may include depicting I / O prefetch and regular I / O paths when possible. Some embodiments may use a stateful per-file bitmap in the OS module 304 at the OS level, for example, to track cache status and improve prefetch efficiency. For example, a stateful state may include the state of accessed blocks or pages of a file used by one or more threads. This may be used in conjunction with a per-file cache tree. Each bit in the bitmap may represent or map to a block in the file (default), and the bitmap may be an unsigned long integer array, for example, which may grow and / or shrink with file size. The bitmap may also be imported into the user-level runtime 302. In other embodiments, integers, short integers, or any other type of data type may be used.
[0038] Update the bitmap or data structure in OS module 304
[0039] The per-file (or per-index node) bitmap or data structure may be continuously updated during read, write, and prefetch operations. A slow path and / or a fast path may be used to reduce contention for a single large per-file cache tree lock between regular I / O and prefetch operations.
[0040] During I / O operations such as read() and write(), the OS module 304 may use a slow path (relative to the fast prefetch lookup below) that involves checking the existence of the requested block in the cache by traversing a large array or data structure (e.g., Xarray) per file of pointers. The walk may be completed using a page vector that records the availability of multiple blocks in the cache. When a block is missing from the cache, a read request may be issued for it, and the cache bitmap per file (or per index node) may be updated. During this process, a cache tree lock may be maintained, resulting in contention with concurrent prefetch operations and affecting prefetch validity. A cache tree as used herein may be one or more hierarchical cache structures in which data may be concurrently accessed by different threads. A lock may include locking or controlling access to a node of a cache tree structure to ensure thread feasibility. A lock may be, for example, a global lock, a read / write lock, or even a fine-grained lock.
[0041] Fast prefetch lookup
[0042] User-level runtime 302 can use the user-level copy of every file bitmap 326-328 to check the pages of high-speed cache.When issuing readahead_info 330-332, OS module 304 can use the fast path of the bitmap search for file by obtaining the read-write lock of bitmap, to reduce the lock contention with conventional I / O.When more pages are requested and inserted in the high-speed cache, readahead_info 330-332 can use the slow path by obtaining the write lock.Using simple bitmap operation can be used to help increase the speed of operation.In order to further reduce contention, OS module 304 can update every file bitmap 322-324 once after completing the traversal, instead of updating for each page.
[0043] Some applications, such as molecular simulations or databases, can use per-thread file descriptors to allow concurrent access to shared files across multiple threads or processes. These threads can use their file descriptors to read from or write to specific areas of the file. For example, some applications use per-thread file descriptors for concurrent I / O to share log and database files between client threads and background threads.
[0044] According to some embodiments, application threads Thread 1 303-Thread 2 305 can check the per-file cache bitmap, and a collection of dedicated auxiliary threads can issue actual pre-fetch requests to the OS. Therefore, concurrent updates and accesses to each index node (an index node can be a data structure in some file systems that describes a file system object (such as a file or directory)) bitmap can be performed in serial order using a read-write lock (rw-lock). Some embodiments can utilize per-index node bitmap locks to allow concurrent access to shared files across threads. In order to reduce concurrency bottlenecks across application threads, some embodiments can maintain a separate bitmap for each thread or file descriptor, and import the cache bitmap of the file from the OS module 304.
[0045] Some embodiments may include a concurrent per-file range tree. A range tree may be an ordered tree data structure (such as a binary tree) that maintains a list of data points. The range tree may use a private or shared file descriptor for each thread to track the range of blocks accessed by each thread. Each node of the range tree may represent a continuous range of blocks, wherein each range / node has its own lock. Each node may also embed a bitmap or data structure, wherein each bit represents, for example, a block in a range. The range of a block may be dynamically increased or decreased based on the range of one or more blocks accessed by each thread and the corresponding bitmap. By maintaining a range tree with per-node ranges and per-node locks, multiple application threads using the same or different file descriptors may concurrently access non-conflicting ranges of a file, thereby reducing scalability bottlenecks, and avoiding duplication of bitmaps. In some embodiments, threads accessing overlapping blocks may share bitmaps, and may benefit from the perception of pages already in cache, thereby reducing redundant prefetch requests and associated overhead.
[0046] Low-overhead prediction and prefetching
[0047] Embodiments can provide efficient cache prefetching that can adapt to different access patterns without high overhead. For example, a small data structure such as a per-file bitmap can be considered that uses only a few kilobytes or megabytes per file. To achieve this, some embodiments can first detect the access pattern of the file by intercepting I / O operations (e.g., such as POSIX) and decide the number of blocks to prefetch. According to some embodiments, the pattern detector can recognize a wide range of access patterns, including sequential, random, forward / backward strides, and changes in access patterns.
[0048] The user-level runtime 302 may use one or more n-bit counters for each file to detect access patterns. The counters may indicate the level of sequentiality (e.g., when several consecutive pages are accessed, or conversely when pages are accessed in a random sequence), and may represent files in several different states, such as:
[0049] Highly random (000, access distance exceeds the maximum prefetch distance of 128KB).
[0050] Random (001, random but within 128KB distance).
[0051] Partially random (010, a mix of sequential and random access).
[0052] Possibly - sequential (011, frequent sequence alternating with random access).
[0053] Sequential (100, sequential but with strides). And
[0054] Definitely (110) order.
[0055] Additional states or variations may be considered. For example, in some embodiments, a subset such as 2 or 3 states may be considered.
[0056] When a read or write operation occurs, the user-level runtime can increment or decrement the value of a counter depending on whether the previous access was sequential, and the value of the counter can determine the blocks to be prefetched. According to some embodiments, when a file is opened, the system can start in a "deterministically random" state, meaning that no blocks are prefetched. However, as sequential access increases, the prefetched blocks can also increase by 2n, for example, where n is the value of the access pattern counter. To avoid issuing prefetch calls for blocks that are already in the page cache, the user-level runtime can check the cache bitmap and modify prefetch requests for blocks that are not in the cache.
[0057] The user-level runtime 302 can recognize various I / O access modes, such as sequential, random, forward / backward stride, etc. The predictor can intercept each I / O and quickly evaluate the conversion or oscillation of the file between access modes. In addition, the number of bits used for each file counter can be configured to improve pre-fetch accuracy. According to some embodiments, a 3-bit counter can provide optimal performance for different workloads with varying access patterns without over-prefetching. In order to optimize prediction and reduce the overhead of pattern detection, once a stable state (e.g., deterministically sequential or random) is reached, the user-level runtime 302 can delay predictions for the next n accesses. The user-level runtime 302 can utilize the cross-layer capabilities and runtime prefetching of the OS module 304 to reduce frequent system calls and related overhead.
[0058] File descriptor prefetch
[0059] An access pattern detector may be maintained for each descriptor, and a user space file descriptor structure containing block range information and access pattern counters may be used. In one example, thread 1 311 and thread 2 313 may access file 1 318 on local / remote storage 306 simultaneously. For example, thread 1 311 may submit request 1 310 via user-level runtime 302. Request 1 310 may use fadvise() in a request for 6Mb of file 1 data, starting at offset 0. In response to the request, OS module 304 may provide only 3Mb of the requested 6Mb to prefetched request 1 315. In response, user-level runtime 302 may use readahead 1 330 to request the remaining 3Mb of file 1 starting at offset 3Mb. In response, prefetched readahead 1 317 may be sent by local / remote storage 306.
[0060] Prefetched readahead1 317 may submit the remaining 3Mb of file 1 318 in response to readahead1 330 having a request for 3Mb starting at offset 3Mb. In a parallel manner, thread 2 313 may access file 1 318 because thread 2 313 is accessing different portions of file 1 318. Thus, request 3 314 and request 4 316 are requests for 32Kb at offset 8Mb and 64Kb at offset 8.32Mb of file 1 318, respectively. Readahead2 332 then requests these corresponding regions of file 1 318 via OS module 304. Thus, prefetched readahead2 321 submits the requested 96Kb of total memory for use by thread 2 313.
[0061] Figure 4 An example embodiment 400 of a user-level runtime 406 that performs predictive prefetching in coordination with an OS module 408 according to an example embodiment of the present disclosure is shown. Each file descriptor may have some additional metadata created to track its access history. For example, when thread 1 303 accesses file 1 in a random pattern and thread 2 305 accesses file 1 in a sequential pattern, prefetching may occur only for non-overlapping areas of the file accessed by thread 1 303, as shown. Each "1" in the per-file descriptor access history 1 402 and per-file descriptor access history 2 404 shows an access in, for example, bytes in a page of the file. Per-file descriptor access history 2 404 shows 4 pages accessed in a row. Per-file descriptor access history 1 402 shows a trace of predictor worker 1 that tracks random access patterns. Per-file descriptor access history 2 404 shows a trace of predictor worker 2 that tracks sequential access patterns.
[0062] In the case of overlapping accesses across file descriptors, the user-level runtime 406 can use cache state awareness to avoid redundant prefetches while ensuring cache hits. In some embodiments, the user-level runtime 406 can use a bit in the cache bitmap for the entire range in the range tree and prefetch the entire range when memory is insufficient.
[0063] Memory-aware aggressive prefetching and eviction
[0064] Embodiments described herein include improving I / O performance for varying memory availability and changing access patterns, including high intensity I / O phases and low intensity I / O phases. Embodiments may delegate prefetch control to a user-level runtime, which may use page cache awareness to adjust prefetches and evictions based on a memory budget set by an application, container, virtual machine (VM), or system administrator. Some embodiments may utilize the available memory budget to aggressively prefetch from the start of an application and reduce higher forced cache misses (e.g., blocks loaded into the cache for the first time). This may be considered an improvement to Oses that employ incremental prefetching, which may suffer from higher initial cache misses.
[0065] Aggressive I / O prefetching
[0066] In some embodiments, the user-level runtime can constantly monitor memory usage and adjust prefetching aggressiveness accordingly to make the best use of available memory. The user-level runtime can check system memory availability from the beginning, and set higher and lower thresholds to indicate when to stop active prefetching and when to stop all prefetching, respectively. In some embodiments, system administrators can use configuration files to customize these values. When a file is opened, the user-level runtime can assume that it is sequential, and prefetch some blocks (e.g., 2MB) before issuing enough I / O to detect access patterns. When the prediction is correct and the file is marked as "definitely" sequential, the embodiments described herein can issue larger prefetch requests (when budget allows), thereby accelerating access to the file and reducing cache misses. When the prediction is incorrect, the user-level runtime can default to conventional prefetching, and stop prefetching when the file is detected as random. In some embodiments, the user-level runtime can be extended to support customized prefetching strategies and window sizes based on the priority of the file.
[0067] Active recycling
[0068] To aggressively reclaim cache pages, embodiments described herein may employ a multi-pronged approach. The user-level runtime 302 may maintain a memory budget per process and monitor active and inactive files using a Least Recently Used (LRU) technique. When the memory budget is low, inactive file cache pages may be forcibly evicted. Additionally, for larger files, in addition to page-level evictions performed by the OS LRU, embodiments described herein may also use per-file cache state to evict infrequently accessed pages by using an API that tells the kernel how it expects to use the file handle, e.g., fadvice().
[0069] Supports memory mapped I / O
[0070] Memory mapped I / O can be used by applications for read-intensive workloads to reduce system call and data copy overhead between the OS and user space. However, there may be challenges in predicting and prefetching memory mapped I / O without explicit I / O calls. To address this problem, embodiments described herein may utilize cross-layer cache bitmap states. When an application issues a memory mapped I / O call (e.g., mmap()), the user-level runtime may call a background thread that periodically queries the cache state of the index node. Using the cache state, the access pattern detector may determine the number of pages to be prefetched.
[0071] Optimizing OS prefetching path
[0072] According to some embodiments, the OS may be extended to allow higher-level layers such as a user-level runtime to dynamically increase prefetch limits using one or more readahead_info system call information structures. For example, a user-level runtime may issue prefetch requests up to 1.2GB that match the NVMe bandwidth. Larger requests may not adversely affect blocking I / O because the OS virtual file system layer may limit any I / O request to a maximum value (e.g., a 2MB request), and existing OS prefetch congestion control may postpone prefetch requests that delay the dispatch of blocking I / O (such as reads or writes). The embodiments described herein may reduce the overhead of prefetch operations for iterating pointer arrays, and reduce cache hit costs, use delineation paths for prefetching and blocking I / O operations, and reduce contention for cache tree locks, and speed up prefetch performance.
[0073] Some embodiments may be implemented regardless of the underlying OS file system and storage devices that may use a virtual file system. Some embodiments may be implemented without requiring any modifications to the application code. The application may link a user-level runtime that intercepts POSIX I / O operations to predict access patterns and prefetches accordingly. The user-level runtime may also implement a transparent interception library or API call for prefetching data using the readahead_info system call, accessing cache status using a bitmap, obtaining per-process memory usage from the OS, and for the user-level runtime to communicate with the OS to relax strict prefetch restrictions.
[0074] Figure 5 A flow chart 500 for initialization and prediction-based prefetching according to an example embodiment of the present disclosure is shown. Although the example routine depicts a specific sequence of operations, the sequence may be changed without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not substantially affect the functionality of the routine. In other examples, different components of the example device or system that implements the routine may perform functions substantially simultaneously or in a specific sequence.
[0075] The method may start at box 502, where one or more applications are initialized. Initialization may occur when a user opens an application. Initialization may include initializing files and file descriptor structures and starting the process of the application in OS module 108. Initialization may also occur when the system draws various applications for use based on its needs. In box 504, the method may attach a prefetch library to the application, as described in the embodiments of this article. Attaching the prefetch library enables the user-level runtime to access data in the OS, which allows more efficient and accurate prefetching via an extended API. The prefetch library may include a user-level runtime 106, which includes an OS module 108. The prefetch library may also include a user-level runtime 302 and an OS module 304 element. In box 506, the method may record the read and write of the application when the application starts its process. As described herein, one or more counters may be used. Counters may be used to detect access patterns, for example, indicating when a page is accessed sequentially or randomly.
[0076] In box 508, the user-level runtime can execute one or more prediction algorithms to actively pre-fetch based on the pattern of use of application data or files. If the read access pattern is sequential or strided, the user-level runtime can be separated. For example, the file can be split into pages, where each page is 4 kilobytes. When the file is accessed page by page (1234), the method can identify it as sequential. When the file is accessed in a simple mode (such as every other page), this can be identified as a strided mode. For example, the pattern can be indicated and predicted as sequential, strided, and random.
[0077] In box 510, the method may call a function to pre-fetch predicted data. For example, in some embodiments, readahead_info() may be called. As an exemplary function, the user-level runtime library may use readahead_info() to pre-fetch data from the OS module. The function call may enable complete retrieval of file data. In some embodiments, additional data specified in the function call may be retrieved from one or more files of the application. In box 512, the method may use one or more bitmaps to know what file data is in memory. In addition, a data structure may be used to track what file data is in memory. For example, arrays, linked lists, records, hash tables, graphs, binary trees, AVL trees, etc. may be used to track data retrieved from files. In some embodiments, when the application continues to run and further prediction and tracking is desired, the method may return to box 506.
[0078] Figure 6A flowchart 600 for managing a data structure used in prefetching according to an example embodiment of the present disclosure is shown. In block 602, the method may request a data structure from an OS via a runtime level interface by an application for prefetching data for an application. In some embodiments, the data structure may be a per-file data structure, such as a bitmap. Each bit may indicate a page being accessed in a prefetched file. In block 604, the method may receive a data structure from an OS via a runtime level interface by an application. In some embodiments, a system-level function call may be extended so that prefetching may include retrieving the remaining pages of a complete file or a file of an application. In block 606, the method may prefetch data to be used in an application based on information in a data structure. For example, the information may indicate that a complete file should be retrieved. Tracking application use of file data and passing a data structure or bitmap between the application layer and the OS creates a cross-layer awareness of the OS cache state, which enables the runtime to perform more accurate and efficient prefetching, thereby avoiding insufficient or excessive prefetching, and avoiding redundant system calls. When data is shared across one or more threads in a large-scale application, an embodiment may implement concurrent I / O prefetching across non-conflicting blocks. Utilizing pre-application access patterns, per-file cache status, and available free memory, embodiments may adjust prefetch aggressiveness to reduce cache misses and alleviate I / O bottlenecks.
[0079] In block 608, the method may update the per-file data structure based on application usage of the data. In block 610, the method may send the updated per-file data structure to the OS.
[0080] The data structure can be a per-file data structure. The per-file data structure can also be a bitmap. In some embodiments, each bit in the bitmap can represent or map to one or more blocks in a file associated with the per-file data structure. In some embodiments, the per-file data structure can be updated by the OS during read, write, and pre-fetch operations. In some embodiments, an application can use an application-level copy of the per-file data structure to check cached pages. In some embodiments, an application can use a counter to track and / or update the state of data to be used according to the methods described herein.
[0081] Figure 7 7 is an example schematic diagram of a system 700 for prefetching data according to an example embodiment of the present disclosure. The system 700 includes a processing circuit 702 coupled to a memory 704, a storage device 710, a user interface 706 (or GUI), and a network interface 712. In an embodiment, the components of the system 700 can be communicatively connected via a system bus 708. The system 700 can be designed to manage prefetch or data retrieval operations, and shared data structures can help manage and predict prefetch / data retrieval operations.
[0082] The system bus 708 may include any interface and / or protocol including Peripheral Component Interconnect Express (PCIe), Non-Volatile Memory Express (NVMe), NVMe over Fabric (NVMe-oF), Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Remote Direct Memory Access (RDMA), RDMA over Converged Ethernet (ROCE), Fibre Channel, InfiniBand, Serial ATA (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), iWARP, Hypertext Transfer Protocol (HTTP), CXL, etc. or any combination thereof.
[0083] Processing circuit 702 may be implemented as one or more hardware logic components and circuits. For example, but not limited to, illustrative types of hardware logic components that may be used include field central processing units (CPUs), field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), general purpose microprocessors, microcontrollers, digital signal processors (DSPs), etc., or any other hardware logic component that can perform calculations or other information manipulations. Processing circuit 702 may be used to actively retrieve and pre-fetch data, manage and predict the use of data to be retrieved by user-level applications.
[0084] The memory 704 may be volatile (eg, RAM, etc.), non-volatile (eg, ROM, flash memory, etc.), or a combination thereof. In one configuration, computer readable instructions for implementing one or more embodiments disclosed herein may be stored in the storage device 710 .
[0085] In another embodiment, the memory 704 is configured to store software for managing pre-fetching and retrieving data. Software should be broadly interpreted as meaning any type of instruction, whether referred to as software, firmware, middleware, microcode, hardware description language or other. Instructions may include code (e.g., in source code format, binary code format, executable code format, or any other suitable code format). When executed by the processing circuit 702, the instructions cause the processing circuit 702 to perform the various processes described herein. Specifically, the instructions, when executed, cause the processing circuit 702 to pre-fetch, predict, and store the reading and writing of application-level software.
[0086] The storage device 710 may be a solid state device (SSD), a magnetic storage device, an optical storage device, etc., and may be implemented as, for example, flash memory or other memory technology, CD-ROM, Digital Versatile Disks (DVD), or any other medium that may be used to store the desired information. The storage device 710 may store pre-fetch instructions 714 executed according to the flowchart 500 and pre-fetch analysis instructions 716 executed according to appropriate prediction, read and write tracking techniques discussed. The network interface 712 allows the system 700 to communicate with a cloud server network for purposes such as receiving data, sending data, etc.
[0087] It should be understood that the embodiments described herein are not limited to Figure 7 The specific architecture shown in , and other architectures may be equally used without departing from the scope of the disclosed embodiments.
[0088] The description of various embodiments of the present teachings has been presented for illustrative purposes, but is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or technical improvements to technologies found in the marketplace, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
[0089] Although what is considered to be the best state and / or other examples have been described above, it should be understood that various modifications may be made therein, and the subject matter disclosed herein may be implemented in various forms and examples, and the present teachings may be applied to many applications, only some of which are described herein. The appended claims are intended to claim any and all applications, modifications, and variations that fall within the true scope of the present teachings.
[0090] The components, steps, features, objects, benefits and advantages discussed herein are merely illustrative. None of them and the discussion related to them are intended to limit the scope of protection. Although various advantages have been discussed herein, it should be understood that not all embodiments must include all advantages. Unless otherwise stated, all measurements, values, rated values, positions, dimensions, sizes and other specifications set forth in this specification (including in the appended claims) are approximate, not precise. They are intended to have a reasonable range that is consistent with the functions to which they are related and consistent with the conventions in the field to which they belong.
[0091] Many other embodiments are also contemplated. These include embodiments with fewer, additional and / or different components, steps, features, objects, benefits and advantages. These also include embodiments in which components and / or steps are arranged and / or ordered differently.
[0092] Aspects of the present disclosure are described herein with reference to call flow diagrams and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each step of the flow diagram and / or block diagram and the combination of the blocks in the call flow diagram and / or block diagram can be implemented by computer-readable program instructions.
[0093] These computer-readable program instructions can be provided to a processor of a computer, a special-purpose computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create components for implementing the functions / actions specified in one or more boxes of the call flow process and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can instruct the computer, programmable data processing device, and / or other equipment to function in a specific manner, so that the computer-readable storage medium having instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes of the call flow and / or block diagram.
[0094] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, so that a series of operating steps are performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable apparatus, or other device implement the functions / actions specified in one or more boxes of the call flow process and / or block diagram.
[0095] The flow charts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each frame in the call flow process or block diagram may represent a part of a module, segment or instruction, which includes one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the frame may not occur in the order indicated in the figure. For example, the two frames shown in succession can actually be executed substantially simultaneously, or the frames can sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each frame in the block diagram and / or the call flow diagram and the combination of frames in the block diagram and / or the call flow diagram can be implemented by a dedicated hardware-based system that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.
[0096] Although the foregoing has been described in conjunction with exemplary embodiments, it should be understood that the term "exemplary" is merely meant to be an example, rather than the best or optimal. Unless stated immediately above, nothing stated or illustrated is intended or should be construed as conferring any component, step, feature, object, benefit, advantage, or equivalent to the public, whether or not stated in the claims.
[0097] It should be understood that the terms and expressions used herein have the common meaning consistent with the corresponding investigation and research fields corresponding to these terms and expressions, unless otherwise specified herein. Relational terms such as first and second can be used only to distinguish one entity or action from another entity or action, without necessarily requiring or implying any actual such relationship or order between these entities or actions. The term "comprises", "comprising" or any other variation thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including the list of elements not only includes those elements, but also may include other elements that are not explicitly listed or inherent to such process, method, article or device. In the absence of further constraints, an element preceded by "a" or "an" does not exclude the presence of additional identical elements in the process, method, article or device including the element.
[0098] An abstract of the present disclosure is provided to allow the reader to quickly ascertain the nature of the technical disclosure. The abstract is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. In addition, in the foregoing detailed description, it can be seen that various features are grouped together in various embodiments for the purpose of simplifying the present disclosure. The method of the present disclosure should not be interpreted as reflecting the intention that the claimed embodiments have more features than those expressly recited in each claim. On the contrary, as reflected in the following claims, the subject matter of the invention lies in less than all the features of a single disclosed embodiment. Therefore, the appended claims are hereby incorporated into the detailed description, with each claim itself as a separately claimed subject matter.
Claims
1. A method for prefetching, comprising: requesting, by the application via the application interface, from the operating system OS a data structure for retrieving data for the application; receiving, by the application from the OS via the application interface, the data structure; as well as Data to be used in the application is retrieved by the application based on the information in the data structure.
2. The method according to claim 1, wherein: The data structures include data structures associated with files.
3. The method according to claim 2, wherein: The data structure associated with the file includes a bitmap.
4. The method according to claim 3, wherein: The bits in the bitmap map to blocks in the file of the data structure associated with the file.
5. The method according to claim 4, wherein: The data structures associated with files are modified by the OS based on read, write, and retrieve operations.
6. The method according to claim 2, wherein: The application checks the cached page using the application's copy of the data structure associated with the file.
7. The method according to claim 1, in, The application uses counters to track and update the status of the data.
8. The method according to claim 2, further comprising: modifying, by the application, the data structure associated with the file based on use of the data by the application; as well as The modified data structure associated with the file is sent to the OS.
9. A computing device comprising: processor; as well as A memory storing instructions, which, when executed by the processor, configure the device to: requesting, by the application via the application interface, from the operating system OS a data structure for retrieving data for the application; receiving, by the application from the OS via the application interface, the data structure; as well as Data to be used in the application is retrieved by the application based on the information in the data structure.
10. The computing device according to claim 9, wherein: The data structures include data structures associated with files.
11. The computing device according to claim 10, wherein: The data structure associated with the file includes a bitmap.
12. The computing device according to claim 11, wherein: The bits in the bitmap map to blocks in the file of the data structure associated with the file.
13. The computing device of claim 12, wherein: The data structures associated with files are modified by the OS based on read, write, and retrieve operations.
14. The computing device of claim 9, wherein: The application checks the cached page using the application's copy of the data structure associated with the file.
15. The computing device of claim 9, wherein: The memory stores instructions which, when executed by the processor, further configure the apparatus to: modifying, by the application, the data structure associated with a file based on use of the data by the application; and The modified data structure associated with the file is sent to the OS.
16. A system for prefetching, comprising: Operating system OS; Memory; as well as An application having an application interface, the application being configured to: requesting a data structure from the OS via the application interface for retrieving data for the application; receiving the data structure from the OS via the application interface; as well as Data to be used in the application is retrieved based on the information in the data structure.
17. The system of claim 16, wherein: The data structures include data structures associated with files.
18. The system of claim 17, wherein: The data structure associated with the file includes a bitmap.
19. The system of claim 18, wherein: The bits in the bitmap map to blocks in the file of the data structure associated with the file.
20. The system of claim 19, wherein: The data structures associated with files are modified by the OS based on read, write, and retrieve operations.