Multi-model cache management algorithm based on online learning
By combining multi-model cache management algorithms with LRU, MRU and Xgboost policies, dynamically selecting cache replacement strategies and optimizing dirty page replacement, the problem of insufficient cache hit rate and adaptability in the existing technology is solved, and the performance and life of SSD is improved.
Patent Information
- Application Number
- CN202510583630.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-19
AI Technical Summary
The existing cache replacement strategy based on machine learning has shortcomings in cache hit rate and adaptability. Offline learning is difficult to adapt to dynamic load changes, while the decision-making of online learning is coarse and has large computing overhead, resulting in a degradation of SSD performance.
A multi-model cache management algorithm based on online learning is adopted, combined with LRU, MRU and Xgboost offline learning strategies, and a cache replacement strategy is dynamically selected through the Q table, and a dirty page replacement is optimized using the active write back module to reduce the delay and write amplification effects.
Improves cache hit rate, extends the life of SSD, and adapts to the resource limitations of embedded controllers, improving system performance.
Smart Images

Figure CN120508511A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of solid-state drive cache management, and in particular to a multi-model cache management algorithm based on online learning. Background Art
[0002] NAND flash-based solid-state drives (SSDs) are widely used in personal computers, mobile and embedded devices (such as smartphones, digital cameras, and smart homes), as well as data centers, high-performance computing, and enterprise storage due to their low power consumption, excellent random read / write performance, small size, and shock resistance. High-performance, large-capacity SSDs typically include an internal DRAM cache to reduce response time and wear on the back-end flash media, thereby improving SSD performance and extending its lifespan. Due to the limited space and high cost of DRAM, the cache capacity is typically much smaller than the data size of the workload. When the cache fills up, different caching strategies can significantly impact the cache hit rate and SSD performance.
[0003] An advanced approach to cache replacement research uses machine learning (ML) to assist the system in replacing cached data. Existing ML-based cache replacement strategies utilize models that can be categorized as online reinforcement learning and offline learning. Online reinforcement learning methods continuously adjust strategies based on real-time feedback during runtime, enabling them to adapt to dynamically changing workloads. Their advantage lies in their ability to adaptively optimize in diverse application scenarios, improving cache hit rates. Offline learning methods are typically trained based on historical access data and use pre-trained models for decision-making during deployment. This approach avoids the exploration costs of online learning and is more advantageous in environments with limited computing resources, while also ensuring relatively stable performance.
[0004] However, both online and offline learning have certain flaws when applied to cache systems: the limitation of offline learning is its lack of adaptability to dynamic changes in workloads. Once the model is trained, it may be difficult to maintain a high cache hit rate when faced with different data patterns. Therefore, the model needs to be updated regularly to maintain good performance. In addition, the decision granularity, that is, the basic unit of data corresponding to the label generated by the model, such as page granularity and request granularity, has higher accuracy but also incurs greater overhead. Online learning is limited by its mechanism of frequent communication between the host and the device. Its decision granularity is usually large, and it cannot directly decide whether to retain or evict each request, resulting in a decrease in accuracy. Summary of the Invention
[0005] To solve the existing problems, the present invention provides a multi-model cache management algorithm based on online learning. The specific solution is as follows:
[0006] A multi-model cache management algorithm based on online learning includes the following steps:
[0007] S1, the SSD controller receives the I / O request from the upper-layer application and captures the request characteristics;
[0008] S2, after processing a fixed number of pages, aggregates the above request features to generate window-level indicators;
[0009] S3, the window-level indicators are discretized and mapped into a finite state space;
[0010] S4, Q table calculates the expected return of each strategy under the current state based on the discretized state characteristics, i.e. the predicted value;
[0011] S5, the online learning module dynamically selects a cache replacement strategy from the basic strategy pool based on the predicted value;
[0012] S6, an additional priority eviction list is set in the cache. The active write-back module monitors the busyness of the backend flash memory chip in real time and asynchronously writes the cold dirty pages in the priority eviction list back to the flash memory, reducing cache replacement latency and write amplification effect, allowing seamless switching of various cache replacement strategies.
[0013] Preferably, the request features captured in step S1 include the access address, read / write type, timestamp for calculating the time increment, and the length of the continuous access sequence, and the window-level indicators in step S2 include the hit rate change rate, response time increment, read / write ratio, and ratio of sequential request pages.
[0014] Preferably, the basic strategy pool in step S4 includes three strategies: LRU strategy, MRU strategy and Xgboost-based offline learning strategy;
[0015] The LRU strategy is a universal strategy under the assumption of temporal locality. It maintains a linked list of pages sorted by access time. Each time a page is accessed, it is moved to the head of the linked list. When a page is eliminated, the page at the end of the linked list, that is, the page that has not been accessed for the longest time, is selected.
[0016] The MRU strategy is a targeted optimization strategy for sequential and burst loads, giving priority to eliminating recently accessed pages;
[0017] The Xgboost-based offline learning strategy uses Xgboost (Extreme Gradient Boosting), an offline learning strategy that continuously corrects the prediction error of the previous model during iterative training. At the same time, it introduces regularization terms to suppress overfitting, thereby extracting feature associations with high generalization capabilities from historical data. By analyzing the global pattern of historical load, it predicts future access probabilities.
[0018] Preferably, the Q table update of the online learning module in step S4 follows the Bellman equation, and its immediate reward is calculated by weighted calculation of the hit rate change rate, response time increment and load matching degree, and step S5 is to give priority to the corresponding strategy through the predicted value and threshold mechanism calculated in step S4 to reduce the strategy switching frequency.
[0019] Preferably, the selection principle of the cold pages in the priority eviction list in step S6 is as follows: by collecting the characteristics of each page in the cache, including the recency most recent access time, frequency access frequency and page size, these characteristics are used to predict the future reuse distance of the page, and according to the size of the future reuse distance, the pages are divided into three categories: COLD pages, i.e. cold pages, COOL pages, i.e. warm pages and HOT pages, i.e. hot pages; among them, COLD pages are inserted at the end of the priority eviction list to accelerate elimination because they have no access value for a long time; COOL pages have no repeated access requirements in the short term, but may be accessed again in the future, so they are inserted at the head of the priority eviction list as a buffer; HOT pages are retained in the LRU list and continue to be managed by the temporal locality strategy; pages that have not been sampled use the LRU logic by default to ensure that they will not be deleted by mistake until they are sampled or evicted from the cache according to the LRU method.
[0020] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is run, the algorithm described in any one of the above items is executed.
[0021] The present invention also discloses a computer system, including a processor and a storage medium, wherein a computer program is stored on the storage medium, and the processor reads and runs the computer program from the storage medium to execute any of the algorithms described above.
[0022] The beneficial effects of the present invention are:
[0023] 1. Improve cache hit rate through dynamic strategy switching;
[0024] 2. The active write-back module reduces dirty page replacement delay and extends SSD life;
[0025] 3. The total memory usage of the Q table and Xgboost model is small, adapting to the resource limitations of the embedded controller. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1It is a flowchart of the present invention. DETAILED DESCRIPTION
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0029] Existing machine learning-based caching strategies are categorized into online reinforcement learning and offline learning. Offline learning relies on historical data to train models, making it difficult to adapt to dynamic load changes. Online learning, while capable of dynamically adjusting strategies, suffers from coarse decision granularity and high computational overhead. Furthermore, traditional strategies are prone to I / O conflicts during dirty page replacement, resulting in degraded SSD performance. To address these issues, this paper proposes a multi-model cache management algorithm based on online learning.
[0030] In order to solve the above problems, the present invention discloses a multi-model cache management algorithm based on online learning. The algorithm consists of three parts: an online learning module, a basic strategy pool, and an active write-back module.
[0031] The basic strategy pool integrates multiple basic cache strategies, including the LRU strategy based on LRU design, the MRU strategy based on MRU design, and the offline learning strategy based on Xgboost.
[0032] The online learning module uses the Q-table to dynamically decide which basic strategy the cache system should use based on the feedback.
[0033] Furthermore, the active write-back module can write dirty data, determined to be cold by these basic policies, back to the flash memory, reducing I / O conflicts caused by cache replacement of dirty data and improving SSD response time. Therefore, this algorithm can leverage the advantages of both online and offline learning, taking advantage of both the adaptability of online learning and the high accuracy of offline learning, effectively improving system performance and SSD lifespan.
[0034] Specifically, if Figure 1 ,A multi-model cache management algorithm based on online learning, including the following steps:
[0035] S1, the SSD controller receives the I / O request from the upper-layer application and captures the request characteristics.
[0036] The general process of SSD cache serving upper-layer I / O is as follows:
[0037] The upper-layer application issues an I / O request: The request issued by the upper-layer application is transmitted to the SSD through the host interface logic HIL. The request includes the logical address LBA, data length, and operation type.
[0038] The SSD controller receives a request. If the request matches the cache, the SSD controller reads or updates the data directly from the cache. If the request misses the cache, the SSD controller writes the data to the cache or reads the data from the backend flash array into the cache, depending on the read or write type. The cache plays a crucial role in SSDs, as its performance directly determines the read and write performance of the SSD.
[0039] When data streams enter the SSD, the system first captures the request characteristics through the monitoring module, including the access address, read and write type, timestamp used to calculate the time increment, and the length of the continuous access sequence.
[0040] S2: After processing a fixed number of pages, the above request features are aggregated to generate window-level indicators; the window-level indicators include the hit rate change rate, response time increment, read-write ratio, and ratio of sequentially requested pages.
[0041] Window-level metrics refer to performance parameters collected by the system within a fixed time window or a fixed data volume window. For example, in SSD cache management, after processing every 1,000 page requests, the following metrics are collected:
[0042] Hit rate change rate (the difference between the current window hit rate and the previous period);
[0043] Response time delta (change in average response time for the current window);
[0044] Read-write ratio (ratio of read requests to write requests);
[0045] Sequential access ratio (the ratio of consecutive logical address requests to total requests).
[0046] The role of these indicators is to reflect the load characteristics of the system in a certain period of time (such as burst writes, sequential scans, hotspot accesses, etc.), and is the basis for dynamically adjusting the cache strategy.
[0047] S3, the window-level indicators are discretized and mapped into a finite state space;
[0048] Discretization refers to dividing continuous indicator values into a limited number of discrete levels, such as converting the numerical range into three categories: "low, medium, and high".
[0049] Example:
[0050] Hit rate change rate:
[0051] Decline: rate of change ≤ -5%; stable: -5% < rate of change < +5%; rise: rate of change ≥ +5%.
[0052] Sequential access ratio:
[0053] Low: <30%; Medium: 30%-70%; High: >70%.
[0054] The purpose of discretization is:
[0055] Reduce complexity: simplify infinite possible continuous values into a finite number of categories, reducing computation and storage requirements;
[0056] Adapting to resource limitations: Embedded devices (such as SSD controllers) have limited memory, and the state space is controllable after discretization;
[0057] Enhanced robustness: Avoid interference from noisy data on decision-making (e.g., small fluctuations do not trigger strategy switching).
[0058] Mapping to a finite state space combines multiple discretized metrics into a unique state identifier, with all possible states forming a finite set. This is used to drive Q-table decisions—each state corresponds to a row in the Q-table, recording the expected return (Q value) of each strategy (LRU / MRU / Xgboost); and to enable dynamic adaptation—the system selects the optimal strategy based on the current state (e.g., switching to MRU when the sequential access ratio is high).
[0059] S4, the Q-table calculates the expected return (i.e., predicted value) of each strategy in the basic strategy pool under the current state based on discretized state features. As a lightweight decision-making engine, the Q-table maps the SSD operating state into discrete features and records the predicted value of each strategy under different states. Its core advantages are reflected in three aspects: First, the Q-table avoids the computational overhead of complex models by discretizing the state space, requiring only table lookups and simple arithmetic operations during the decision-making process; second, its extremely low memory footprint allows it to reside entirely in the SSD controller's on-chip cache, ensuring nanosecond-level decision latency; and third, discrete state encoding enhances the interpretability of system behavior, facilitating analysis of strategy selection logic.
[0060] The performance of cache replacement policies is highly dependent on the access characteristics of the workload. To cover diverse application scenarios, this solution's basic policy pool integrates two classic policies—Least Recently Used (LRU) and Most Recently Used (MRU)—as well as a predictive policy based on offline learning. This section theoretically analyzes the design principles and applicable scenarios of LRU and MRU, and illustrates their complementary nature in mixed workload environments.
[0061] 1. LRU strategy: LRU - a universal strategy under the assumption of temporal locality
[0062] The LRU strategy is based on the principle of temporal locality in computer architecture, which states that recently accessed data is more likely to be accessed again in the near future. Its core operation is to maintain a linked list of pages sorted by access time: each time a page is accessed, it is moved to the head of the linked list. When a page is removed, the page at the end of the linked list (i.e., the page that has not been accessed in the longest time) is selected.
[0063] Theoretically, the optimization goal of LRU is to minimize the "furthest future access time" (an approximation of the ideal case of the Belady algorithm), so it performs well in loads with strong temporal locality. For example: (1) Transaction Processing (OLTP) systems: such as banking transactions and online order processing, their access patterns show repeated reading and writing of data in a short period of time. (2) Key-Value Stores: Most requests are concentrated on a small amount of hot data, and LRU can efficiently maintain the hot page set through simple time sorting.
[0064] However, LRU's limitations manifest themselves in cyclic access patterns. If a workload periodically accesses a set of pages that exceeds the cache capacity (such as sequentially scanning a large dataset), the LRU will frequently evict and reload pages, causing the cache hit rate to plummet. This phenomenon is known as "LRU thrashing," and its root cause is the failure of the temporal locality assumption.
[0065] 2. MRU strategy: MRU - targeted optimization of sequential and burst loads
[0066] The MRU strategy uses the opposite elimination logic of LRU—it prioritizes the most recently accessed pages. The theoretical basis of this strategy lies in the characteristics of two special types of loads: (1) Sequential access loads: For example, in scenarios such as full table scans, once data is written, it is no longer accessed in the short term. Eliminating the latest page can prevent the cache from being occupied by short-term invalid data. (2) Burst write loads: In high-concurrency write scenarios (such as peak traffic in message queues), the most recently written data may lose its value due to the rapid overwriting of subsequent requests. MRU reduces cache pollution by promptly eliminating such pages.
[0067] The optimization effect of MRU can be explained by the theory of write amplification suppression. In NAND flash memory, frequent modifications to the same physical page can lead to write amplification, which in turn shortens the device lifespan. MRU quickly eliminates newly written dirty pages, prompting them to be written back to the backend target flash memory as soon as possible, thereby reducing the cost of cache replacement and alleviating SSD congestion. When the online learning module selects the MRU strategy, newly requested pages entering the cache are directly migrated to the priority eviction list.
[0068] However, MRU performs poorly in random read-intensive workloads. If the access pattern lacks spatial or temporal locality, MRU's inverse elimination logic will destroy the residency of potential hotspot data, resulting in a hit rate significantly lower than LRU.
[0069] 3. Offline learning strategy based on Xgboost: Use offline models to predict the hotness of requests
[0070] Xgboost is an efficient and flexible gradient boosting algorithm that is very popular in machine learning and data science, especially for modeling structured data. First, Xgboost uses multithreading and parallel computing techniques, making training very fast, especially for large datasets, which is consistent with our application scenario. Second, Xgboost automatically handles missing values, effectively accounting for missing values through the selection of split points, making it more convenient and accurate in caching systems. Because the workload includes a large number of new and cold requests, caching systems inherently face a certain degree of missing request features. Furthermore, Xgboost does not require feature normalization, which reduces the overhead of feature collection.
[0071] Offline learning strategies predict future access probabilities by analyzing global patterns of historical load. Their theoretical advantages are: (1) capturing long-term patterns, such as daily data backups and weekly statistics tasks, which are cyclical behaviors that LRU / MRU cannot handle. (2) modeling complex associations, such as nonlinear patterns like “friends of friends” jump access in social networks and collaborative filtering in recommendation systems.
[0072] Xgboost (Extreme Gradient Boosting) is an efficient ensemble learning algorithm that can accurately capture nonlinear access patterns in complex loads by constructing multiple decision trees and optimizing the gradient boosting process. Its core principle is to continuously correct the prediction error of the previous model during iterative training, while introducing regularization terms to suppress overfitting, thereby extracting feature associations with high generalization capabilities from historical data. In the SSD cache scenario, the deployment of Xgboost demonstrates multiple advantages: First, after lightweight pruning and fixed-point quantization of the model, the inference latency can be compressed to less than 5 microseconds, fully adapting to the real-time requirements of the embedded controller; second, Xgboost can effectively distinguish data pages with different life cycles, providing a global perspective for decision-making for dynamic elimination strategies.
[0073] The policy decision function of the online learning module based on the Q table is defined as:
[0074]
[0075] Among them, r t+1 And set to:
[0076]
[0077] The online learning module updates the Q value of the strategy through the Bellman equation, where the learning rate α controls the Q value update speed, the discount factor γ measures the importance of future benefits, and the immediate reward r is composed of the hit rate change rate ΔH, the response time increment ΔR and the load matching degree Match(S) weighted, where ΔH and ΔR are calculated by the hit rate H of the previous cycle. t-1 and response time R t-1 Normalization avoids dimensional differences from interfering with decision making, and ε is a minimum value to prevent division by zero errors. a represents the action space, i.e., the three basic strategies, and a′ represents the next state s t+1 The following actions may be performed.
[0078] S5, the online learning module dynamically selects a cache replacement strategy from the basic strategy pool based on the predicted value.
[0079] When the difference between the LRU and other strategies' scores is less than 5%, the system prioritizes LRU to reduce computational overhead. This threshold design reduces the frequency of strategy switching while ensuring performance. The exploration rate ε is initially set high and decays exponentially over the training cycle, gradually transitioning from exploration to utilizing the optimal strategy. Match (S) measures the degree of match between the strategy and the current load characteristics (calculated based on the read-write ratio and the proportion of sequentially requested pages). When the combined score difference between LRU and strategy C is less than 5%, the system prioritizes LRU to reduce the model's computational overhead.
[0080] S6, an additional priority eviction list is set in the cache. The active write-back module monitors the busyness of the backend flash memory chip in real time and asynchronously writes the cold dirty pages in the priority eviction list back to the flash memory, reducing cache replacement latency and write amplification effect, allowing seamless switching of various cache replacement strategies.
[0081] The selection principle for cold pages in the priority eviction list is as follows: by collecting the characteristics of each page in the cache, including the recency of the most recent access time, the frequency of access, and the page size, these characteristics are used to predict the future reuse distance of the page. Based on the size of the future reuse distance, the pages are divided into three categories: COLD pages (cold pages), COOL pages (warm pages), and HOT pages (hot pages). Among them, COLD pages are inserted at the end of the priority eviction list to accelerate their elimination because they have no access value in the long term. COOL pages have no repeated access needs in the short term, but may be accessed again in the future, so they are inserted at the head of the priority eviction list as a buffer. HOT pages remain in the LRU list and continue to be managed using the temporal locality strategy. Unsampled pages use the LRU logic by default to ensure that they are not mistakenly deleted until they are sampled or evicted from the cache according to the LRU method.
[0082] This design achieves fine-grained cache management through a hierarchical tagging mechanism: the rapid elimination of COLD pages reduces cache space occupied by invalid data, while the buffering of COOL pages prevents the excessive eviction of potentially valuable data under sudden loads. For example, periodically generated temporary files are often marked as COLD pages, and their timely cleanup significantly reduces write amplification. In machine learning training tasks, frequently accessed batches of data blocks are identified as HOT pages and remain in the cache for a long time to improve read throughput.
[0083] The active writeback module is a key performance optimization component of this solution. It aims to reduce write latency during dirty page replacement and optimize SSD lifespan. Its core principle is to predict when low-value pages will be eliminated and write dirty data back to the flash memory in advance. During cache replacement, the cost of replacing dirty pages is much higher than replacing clean pages, so active writeback can significantly reduce system latency. This solution adds a priority eviction list to the active writeback module. Pages prioritized for eviction by the basic policy are moved from the main list to the priority eviction list. When the active writeback module detects a dirty page, it adds it to the asynchronous writeback queue. Furthermore, the active writeback module monitors the backend flash chip's busyness in real time (i.e., the sum of the external request wait time and the current remaining internal execution time on that flash chip). If a chip with low busyness is found, the active writeback module writes back the dirty pages found in the priority eviction list. Since these dirty pages have been identified as cold by the basic policy, they are optimal writeback targets. Active writeback avoids flash transaction blocking and improves flash chip utilization efficiency. Furthermore, through active writeback, dirty data can be converted to clean data, significantly reducing the cost of cache replacement and thus improving the overall performance of the SSD. Through the above design, a multi-model intelligent cache management algorithm effectively improves SSD performance.
[0084] The cost analysis of this application is as follows:
[0085] 1. Online Learning Model
[0086] Time overhead: Online learning uses a Q-table as the decision-making unit. Q-table lookups use direct addressing, so their time complexity is O(1). The time complexity of Q-table updates depends on the number of base policies, and is O(K), where K is the number of base policies. Since this solution has three base policies, the time complexity of Q-table updates is actually O(1).
[0087] Space overhead: The state of the Q table consists of discretized state features, so its number of states is is the discretization level of the i-th feature. Since the number of features in this scheme is 4 and the number of discretization levels of each feature is 3, the Q-table of this scheme generates a total of 81 entries, and each entry occupies 4B. Therefore, the space overhead of the Q-table of this scheme is extremely low.
[0088] 2. Offline Learning Model
[0089] Time overhead: The maximum tree depth (max_depth) of the Xgboost model is 5, and there are 6 features in total. Therefore, the inference time of the Xgboost model is within 5 microseconds, which is acceptable for modern large-capacity and high-performance SSDs.
[0090] Space overhead: The number of subtrees in the Xgboost model is 40, and its space overhead is within 350KB.
[0091] The online module of this solution is based on a lightweight Q-learning decision maker, which can dynamically select the basic strategy based on feedback. The basic strategy pool presets three basic strategies: LRU, MRU and a strategy based on an offline learning model. The three can be seamlessly switched through an additional priority eviction list set in the cache. The active write-back module can monitor the dirty pages in the priority eviction list and actively write back these dirty pages according to the load of the back-end flash memory. Since the pages in the priority eviction list are identified as cold pages by the basic strategy, the active write-back module can avoid writing back hot data to the greatest extent. The present invention solves the long-standing problems in the field of SSD cache, such as the difficulty of dynamic load adaptation, high cost of dirty page replacement, and large model overhead, through multi-model dynamic switching, policy-aware active write-back and hierarchical list design.
[0092] This solution uses an online learning module as the control center, integrating a lightweight Q-learning decision maker and a state monitoring unit. By collecting indicators such as cache hit rate, response time, sequential access ratio of access sequences, and read-write ratio in real-time historical cycles, it dynamically switches the basic strategies in the basic strategy pool, thereby ensuring the adaptability and dynamic adjustment capabilities of this solution.
[0093] This solution designs multiple basic strategies for the online module to switch dynamically in real time. These basic strategies can handle different load patterns: Strategy A (LRU) can handle loads with high temporal locality, and Strategy B (MRU) prioritizes the eviction of requests that have recently entered the cache. Therefore, it can handle situations where a large amount of cold data (such as one-time access data) flushes the cache in a short period of time. When the complexity of the load access pattern exceeds the processing capabilities of Strategy A and Strategy B, the online module can switch the cache strategy to Basic Strategy C - an offline learning strategy based on Xgboost. Strategy C can use the trained Xgboost model to infer the tags of pages in the cache, classify the pages in the cache according to the tags, and prioritize the eviction of cold tag pages.
[0094] The active write-back module designed in this solution triggers active write-back when the chip is idle, improving flash memory utilization and reducing cache replacement costs, thereby enhancing SSD write performance and extending SSD lifespan. Furthermore, the active write-back module integrates with the policies in the basic policy pool to ensure that the data actively written back is cold data as determined by these basic policies, thus avoiding the exacerbated write amplification caused by the active write-back of hot data. This significantly improves system performance.
[0095] The present invention also discloses a computer-readable storage medium and a computer system. The computer-readable storage medium stores a computer program, which, when executed, executes any of the algorithms described above. The computer system includes a processor and a storage medium, wherein the storage medium stores the computer program. The processor reads and executes the computer program from the storage medium to execute any of the algorithms described above.
[0096] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.
[0097] In embodiments of the present invention, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of medium. Combinations of the above should also be included within the scope of computer-readable media.
[0098] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0099] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multi-model cache management algorithm based on online learning, characterized in that: The following steps are involved: S1, the SSD controller receives the I / O request from the upper-layer application and captures the request characteristics; S2, after processing a fixed number of pages, aggregates the above request features to generate window-level indicators; S3, the window-level indicators are discretized and mapped into a finite state space; S4, Q table calculates the expected return of each strategy in the basic strategy pool under the current state based on the discretized state characteristics, i.e. the predicted value; S5, the online learning module dynamically selects the cache replacement strategy for the next cycle from the basic strategy pool based on the predicted value; S6: An additional priority eviction list is set up in the cache to collect pages that are identified as needing priority eviction by each cache strategy. The active write-back module monitors the busyness of the backend flash memory chip in real time and asynchronously writes the cold and dirty pages in the priority eviction list back to the flash memory, reducing cache replacement latency and write amplification effect, allowing various cache replacement strategies to be switched seamlessly.
2. The algorithm according to claim 1, characterized in that: The request features captured in step S1 include the access address, read / write type, timestamp used to calculate the time increment, and the length of the continuous access sequence. The window-level indicators in step S2 include the hit rate change rate, response time increment, read / write ratio, and ratio of sequential request pages.
3. The algorithm according to claim 1, characterized in that The basic strategy pool in step S4 contains three strategies: LRU strategy, MRU strategy, and Xgboost-based offline learning strategy; The LRU strategy is a universal strategy under the assumption of temporal locality. It maintains a linked list of pages sorted by access time. Each time a page is accessed, it is moved to the head of the linked list. When a page is eliminated, the page at the end of the linked list, that is, the page that has not been accessed for the longest time, is selected. The MRU strategy is a targeted optimization strategy for sequential and burst loads, giving priority to eliminating recently accessed pages; The Xgboost-based offline learning strategy uses Xgboost extreme gradient boosting, an offline learning strategy that continuously corrects the prediction error of the previous model during iterative training. At the same time, it introduces regularization terms to suppress overfitting, thereby extracting feature associations with high generalization capabilities from historical data. By analyzing the global pattern of historical load, it predicts future access probabilities.
4. The algorithm according to claim 3, characterized in that In step S4, the Q-table update of the online learning module follows the Bellman equation, and its immediate reward is calculated by weighted calculation of the hit rate change rate, response time increment, and load matching degree. In step S5, the corresponding strategy is selected preferentially based on the predicted value and threshold mechanism calculated in step S4 to reduce the frequency of strategy switching.
5. The algorithm according to claim 1, characterized in that The selection principle of cold pages in the priority eviction list in step S6 is as follows: by collecting the characteristics of each page in the cache, including the recency of the most recent access time, the frequency of access and the page size, these characteristics are used to predict the future reuse distance of the page, and the pages are divided into three categories according to the size of the future reuse distance: COLD pages, i.e. cold pages, COOL pages, i.e. warm pages and HOT pages, i.e. hot pages; among them, COLD pages are inserted at the end of the priority eviction list to accelerate their elimination because they have no access value in the long term; COOL pages have no repeated access requirements in the short term, but may be accessed again in the future, so they are inserted at the head of the priority eviction list as a buffer; HOT pages are retained in the LRU list and continue to be managed by the temporal locality strategy; pages that have not been sampled use the LRU logic by default to ensure that they will not be deleted by mistake until they are sampled or evicted from the cache according to the LRU method.
6. A computer-readable storage medium, characterized in that: The medium stores a computer program, and when the computer program is run, the algorithm according to any one of claims 1 to 5 is executed.
7. A computer system, characterized in that: The method comprises a processor and a storage medium, wherein the storage medium stores a computer program, and the processor reads and runs the computer program from the storage medium to execute the algorithm according to any one of claims 1 to 5.
Citation Information
Cited By
Page caching strategy switching method and device, electronic equipment and storage medium
CN122195943A