Cache management method, system and device and computer readable storage medium

Through the hot and cold request classifier and three-level linked list management built by machine learning, the cache data location is adjusted according to the access frequency and request characteristics, which solves the problem of low SSD cache hit rate and improves cache performance and device life.

CN120669924APending Publication Date: 2025-09-19SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510897436.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

How to improve the cache hit rate in SSDs based on NAND flash technology? Considering the high cost of DRAM and limited cache space, the cache hit rate is low.

Method used

A hot and cold request classifier is built through machine learning. The hot and cold classification results of data operation requests are generated based on access frequency, recency and interval value. A three-level linked list is used to manage the data in the cache, and the position of the target logical page address in the cache is adjusted according to the hot and cold attributes.

Benefits of technology

Improves cache hit rate, reduces flash memory read and write operations, extends SSD lifespan, and improves performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669924A_ABST
    Figure CN120669924A_ABST
Patent Text Reader

Abstract

The invention discloses a cache management method, system and device and a computer readable storage medium, and relates to the technical field of caches. Determining an access frequency, a near factor and an interval value of a target logic page address corresponding to the data operation request; obtaining a pre-trained cold and hot request classifier, wherein the cold and hot request classifier is established based on machine learning; processing the near factor, the access frequency and the interval value based on a cold and hot request classifier to generate an initial cold and hot classification result of the data operation request; adjusting target data of the target logic page address in the cache according to the initial cold and hot classification result; the near factor represents the number of requests between two times of continuous access to target logic page addresses; the interval value represents the number of the current requests and the number of the subsequent requests between the same access target logical page addresses. The near factor, the interval value and the access frequency are processed based on machine learning, so that data storage in the cache is consistent with cold and hot attributes, and the cache hit rate is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of cache technology, and more specifically, to a cache management method, system, device, and computer-readable storage medium. Background Art

[0002] Solid-state drives (SSDs) based on NAND flash memory technology are used in storage systems. To improve data transmission efficiency, SSDs incorporate a portion of DRAM (Dynamic Random Access Memory) as a cache, providing a performance bridge between the host I / O interface and the NAND flash memory. This cache temporarily stores frequent read and write requests. When a request matches a logical page address (LPA) in the cache, it responds quickly, significantly shortening response times and reducing read and write operations on the flash memory, further improving SSD performance and extending its lifespan. However, due to the high cost of DRAM and limited cache space, cache hit rates are low.

[0003] In summary, how to improve the cache hit rate is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0004] The purpose of this application is to provide a cache management method that can, to a certain extent, solve the technical problem of how to improve the cache hit rate. This application also provides a cache management system, an electronic device, and a computer-readable storage medium.

[0005] In order to achieve the above objectives, this application provides the following technical solutions:

[0006] A cache management method, comprising:

[0007] Get data operation request;

[0008] Determining an access frequency, a recency factor, and an interval value of a target logical page address corresponding to the data operation request;

[0009] Obtain a pre-trained hot and cold request classifier, where the hot and cold request classifier is built based on machine learning;

[0010] Processing the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request;

[0011] Adjusting target data of the target logical page address in the cache according to the initial hot-cold classification result;

[0012] The recency factor represents the number of requests between two consecutive accesses to the target logical page address; and the interval value represents the number of requests between a current request and a subsequent request for the same access to the target logical page address.

[0013] In an exemplary embodiment, the processing of the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request includes:

[0014] According to a generation condition that a hot and cold classification result is negatively correlated with the recency factor and the interval value and positively correlated with the access frequency, the hot and cold request classifier processes the recency factor, the access frequency, and the interval value to generate an initial hot and cold classification result of the data operation request.

[0015] In an exemplary embodiment, the processing of the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request includes:

[0016] inputting the recency factor, the access frequency and the interval value into each hot and cold decision tree in the hot and cold data classifier;

[0017] receiving an initial hot and cold attribute score output by the hot and cold decision tree;

[0018] Processing is performed according to the initial hot and cold attribute scores to generate an initial hot and cold classification result of the data operation request.

[0019] In an exemplary embodiment, receiving the initial hot and cold attribute scores output by the hot and cold decision tree includes:

[0020] receiving an initial hot and cold attribute score output by the hot and cold decision tree according to a hot and cold attribute score generation formula;

[0021] The formula for generating the cold and hot attribute scores includes:

[0022] Y = (1 / R) * F * (1 / J);

[0023] Wherein, Y represents the initial hot / cold attribute score; R represents the recency factor; F represents the access frequency; and J represents the interval value.

[0024] In an exemplary embodiment, adjusting the target data of the target logical page address in the cache according to the initial hot / cold classification result includes:

[0025] Determining a size value of the data operation request;

[0026] Determining an access cycle and a distance value of the target logical page address; the distance value represents the number of requests for the same logical page address between two consecutive accesses to the target logical page address;

[0027] generating a target hot / cold classification result of the data operation request based on the initial hot / cold classification result, the size value, the access cycle, and the distance value;

[0028] According to the target hot / cold classification result, the target data of the target logical page address in the cache is adjusted.

[0029] In an exemplary embodiment, adjusting the target data of the target logical page address in the cache according to the target hot / cold classification result includes:

[0030] Analyzing the target hot and cold classification results;

[0031] In response to the target cold-hot classification result being a hot classification result, moving the target data corresponding to the target logical page address to the head of a first linked list in the cache;

[0032] In response to the target hot / cold classification result being an intermediate classification result, moving the target data corresponding to the target logical page address to the head of the second linked list in the cache;

[0033] In response to the target hot / cold classification result being a cold classification result, the target data corresponding to the target logical page address is moved to the head of the third linked list in the cache.

[0034] In an exemplary embodiment, the present invention further comprises:

[0035] In response to the first linked list being full, moving the data at the tail of the first linked list to the head of the second linked list;

[0036] In response to the second linked list being full, moving the data at the tail of the second linked list to the head of the third linked list;

[0037] In response to the third linked list being full, the data at the end of the third linked list is stored in the hard disk.

[0038] A cache management system, comprising:

[0039] Request acquisition module, used to obtain data operation requests;

[0040] a parameter determination module, configured to determine an access frequency, a recency factor, and an interval value of a target logical page address corresponding to the data operation request;

[0041] A classifier acquisition module is used to obtain a pre-trained hot and cold request classifier, which is built based on machine learning;

[0042] a classification result generating module, configured to process the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request;

[0043] a data adjustment module, configured to adjust the target data of the target logical page address in the cache according to the initial hot / cold classification result;

[0044] The recency factor represents the number of requests between two consecutive accesses to the target logical page address; and the interval value represents the number of requests between a current request and a subsequent request for the same access to the target logical page address.

[0045] An electronic device, comprising:

[0046] memory for storing computer programs;

[0047] A processor is configured to implement the steps of any of the above cache management methods when executing the computer program.

[0048] A computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above cache management methods.

[0049] The present application provides a cache management method, which includes obtaining a data operation request; determining the access frequency, recency factor, and interval value of a target logical page address corresponding to the data operation request; obtaining a pre-trained hot and cold request classifier, wherein the hot and cold request classifier is built based on machine learning; processing the recency factor, access frequency, and interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request; adjusting the target data of the target logical page address in the cache according to the initial hot and cold classification result; wherein the recency factor represents the number of requests between two consecutive accesses to the target logical page address; and the interval value represents the number of requests between a current request and a subsequent request for the same access to the target logical page address. In this application, the recency factor represents the number of requests between two consecutive accesses to the target logical page address, and the interval value represents the number of requests between the current request and the subsequent request for the same access target logical page address. Since the number of other requests between the target logical page addresses will affect the hit situation of the data in the cache, for example, the greater the number of requests, the lower the hit rate of the data in the cache. The access frequency determines the hit rate of the data itself in the cache, and the hot and cold request classifier built based on machine learning has a strong generalization ability that can reduce overfitting, so it can generate a more accurate initial hot and cold classification result. In this way, the target data of the target logical page address in the cache can be adjusted more accurately, so that the storage of the data in the cache corresponds to the hot and cold attributes of the data, thereby improving the hit rate of the cache. The cache management system, electronic device and computer-readable storage medium provided by this application also solve the corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.

[0051] Figure 1 A flowchart of a cache management method provided in an embodiment of the present application;

[0052] Figure 2 Flowchart for random forest machine learning;

[0053] Figure 3 Training a graph for a random forest machine learning model

[0054] Figure 4 This is a schematic diagram of the cache of the solid-state drive;

[0055] Figure 5 This is a management diagram of a three-level linked list;

[0056] Figure 6This is a schematic diagram of updating a three-level linked list;

[0057] Figure 7 A schematic diagram of the structure of a cache management system provided in an embodiment of the present application;

[0058] Figure 8 A schematic diagram of the cache management system that works with FTL;

[0059] Figure 9 A schematic diagram of the structure of a solid-state drive designed in conjunction with the cache management solution of this application;

[0060] Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application;

[0061] Figure 11 Another structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0062] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0063] See also Figure 1 , Figure 1 A flowchart of a cache management method provided in an embodiment of the present application.

[0064] A cache management method provided in an embodiment of the present application may include the following steps:

[0065] Step S101: Obtain a data operation request.

[0066] In actual applications, you can first obtain a data operation request for operating data. The data operation request can be a request to operate on the server log, or a request to operate on the server's operation fault data, etc. This application specifically defines the type of data operation request, etc.

[0067] Step S102: Determine the access frequency, recency factor and interval value of the target logical page address corresponding to the data operation request; the recency factor represents the number of requests between two consecutive accesses to the target logical page address; the interval value represents the number of requests between the current request and the subsequent request between the same access target logical page address.

[0068] In practical applications, considering that during the processing of a data operation request, the logical page address (LPA) is required to operate on the target data corresponding to the data operation request, the operation information on the target data can be reflected in the operation on the target logical page address corresponding to the data operation request. Therefore, the access information of the target logical page address corresponding to the data operation request can be determined, so as to determine the hot and cold properties of the target data with the help of the access information of the target logical page address.

[0069] In an exemplary embodiment, considering that access frequency represents the number of accesses to the same logical page address within a period of time, it can represent the hot and cold properties of requests over a period of time. For example, a small number of accesses to the LPA indicates a cold request; otherwise, the request is a hot request, meaning that a hot request with a high access frequency may be accessed multiple times. The recency factor represents the number of requests between two consecutive accesses to the target logical page address. When the number of intervening requests is not greater than the cache size, the request will be re-accessed after the first access hit. Due to temporal locality, the request can be re-accessed in the future. Conversely, if there are many requests between two consecutive access requests and the amount of data in these intervening requests is greater than the cache size, the requested data will be evicted the next time it reaches the cache. Therefore, the recency factor can reflect the hot and cold properties of the request. The interval value represents the number of requests between the current request and the subsequent request for the same target logical page address. The fewer the number of intervening requests, the faster the request is re-accessed. Therefore, the interval value can also reflect the hot and cold properties of the request. Therefore, the access frequency, recency factor, and interval value of the target logical page address corresponding to the data operation request can be determined, so that the hot and cold classification result can be subsequently determined based on the access frequency, recency factor, and interval value.

[0070] Step S103: Obtain a pre-trained hot and cold request classifier, which is built based on machine learning.

[0071] Step S104: Process the recency factor, access frequency, and interval value based on the hot / cold request classifier to generate an initial hot / cold classification result of the data operation request.

[0072] In actual applications, the present application pre-builds a hot and cold request classifier based on machine learning for hot and cold evaluation of logical page addresses. Therefore, after the access frequency, recency factor, and interval value of the target logical page address corresponding to the data operation request are obtained, the recency factor, access frequency, and interval value can be processed based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request. It should be noted that the hot and cold attributes of the data operation request are consistent with the hot and cold attributes of the target logical page address and the hot and cold attributes of the target data. This will not be emphasized in the subsequent description. For example, when the initial hot and cold classification result indicates that the hot and cold attributes of the data operation request are hot, the hot and cold attributes of the target logical page address are also hot, and the hot and cold attributes of the target data are also hot. For another example, when the initial hot and cold classification result indicates that the hot and cold attributes of the data operation request are cold, the hot and cold attributes of the target logical page address are also cold, and the hot and cold attributes of the target data are also cold.

[0073] In an exemplary embodiment, considering that the hot and cold attributes are positively correlated with the access frequency and negatively correlated with the recency factor and the interval value, in the process of processing the recency factor, access frequency and interval value based on the hot and cold request classifier to generate the initial hot and cold classification results of the data operation request, the recency factor, access frequency and interval value can be processed based on the hot and cold request classifier according to the generation condition that the hot and cold classification results are negatively correlated with the recency factor and interval value and positively correlated with the access frequency to generate the initial hot and cold classification results of the data operation request.

[0074] In an exemplary embodiment, the type of machine learning can be flexibly determined according to the application scenario. For example, in the process of processing the recency factor, access frequency and interval value based on the hot and cold request classifier to generate the initial hot and cold classification results of the data operation request, the recency factor, access frequency and interval value can be input into each hot and cold decision tree in the hot and cold data classifier; the initial hot and cold attribute scores output by the hot and cold decision tree are received; and the initial hot and cold attribute scores are processed to generate the initial hot and cold classification results of the data operation request. In other words, the initial hot and cold classification results can be generated with the help of the decision tree, and the decision tree can be run offline. It should be noted that the training process of the hot and cold decision tree can be combined with the training principle through the access frequency, recency factor and interval value, such as the training process and the process of producing C function, such as Figure 2 and Figure 3 As shown, the ML model is trained offline and converted into an "if-else" statement in a C function. Initially, the classification conditions of each decision tree in the random forest are converted into "if-else" conditions, and the judgment result is the judgment result of the decision tree. Then, all judgment results and votes are ensured to produce the final label. In this process, the features are used as parameter inputs of the function, and this method is used to abstract the model into a C function. The label value representing the hot and cold attribute score is obtained through the function return, and the label value is used for cache management.

[0075] In an exemplary embodiment, in the process of receiving the initial hot and cold attribute scores output by the hot and cold decision tree, the initial hot and cold attribute scores output by the hot and cold decision tree according to the hot and cold attribute score generation formula may be received; the hot and cold attribute score generation formula includes:

[0076] Y = (1 / R) * F * (1 / J);

[0077] Among them, Y represents the initial hot and cold attribute score; R represents the recency factor; F represents the access frequency; and J represents the interval value.

[0078] Thus, high initial hot / cold attribute scores indicate that high-priority requests can be easily placed in the cache. A low recency factor indicates good workload locality, making 1 / R more cost-effective. Requests with high frequency exhibit higher heat and subsequently generate large initial hot / cold attribute scores. Recency factor and frequency values ​​are collected in the feature set. The interval value is the number of I / O requests between the current request and subsequent requests accessing the same LPA. The fewer the number of interval requests, the faster the request is re-accessed. Therefore, requests within a small range are more cost-effective than 1 / J.

[0079] Step S105 : adjusting the target data of the target logical page address in the cache according to the initial hot / cold classification result.

[0080] In actual applications, after determining the initial hot and cold classification result of the data operation request, the target data of the target logical page address in the cache can be adjusted according to the initial hot and cold classification result. For example, if the initial hot and cold classification result indicates that the hot and cold attributes of the data operation request are hot, the target data of the target logical page address in the cache can be kept unchanged. For another example, if the initial hot and cold classification result indicates that the hot and cold attributes of the data operation request are cold, the target data corresponding to the target logical page address can be written from the cache to the solid-state drive, etc.

[0081] To understand the cache, assume that the cache process of the solid-state drive is as follows Figure 4As shown in the figure, in actual applications, when the host file system sends read and write commands to the SSD, the Flash Translation Layer (FTL) handles operations such as cache management, map conversion, garbage collection (GC), and wear leveling (WL). Furthermore, a portion of the DRAM is used as a cache to temporarily store requests. This cache can alleviate the mismatch in transmission performance between the host and the flash memory. When the host issues a read request, data from the Nand Flash can be pre-read into the cache. Therefore, when the application requests the same data, the SSD can read it from the cache instead of accessing the flash memory, reducing read latency.

[0082] In an exemplary embodiment, the size value of the request, the access cycle of the target logical page address, and the distance value also affect the hot and cold properties of the request. For example, given the request volume, a smaller I / O request will result in a large number of requests, and the data in the cache will be frequently updated. In addition, the data accessed by the request takes up a large amount of cache space. When a new request sequence arrives, the locality will be very poor, and many requests will face cache misses. Smaller I / O writes take less time, thereby reducing write overhead and reducing the impact of write operations on SSD read performance. Therefore, smaller requests have high cache priority; for example, frequency cannot reflect temporal locality well, and the recency factor only provides information on the two most recent accesses. If there are high-frequency requests in the cache, it is not feasible to exclude requests based solely on access frequency when replacing the cache, but the measurement cycle can. To address this issue, the term "distance" is defined as the ratio of the number of requests between the first and most recent access to the same LPA to the LPA's frequency over the duration of the LPA. A smaller period indicates a smaller average interval between access requests, while a larger frequency and a larger period indicate a larger average interval between access requests. The smaller the frequency, the higher the period. The cache should prioritize requests with smaller periods while evicting requests with larger periods to improve the hit rate. For example, distance is the number of requests to a unique LPA between two accesses to the same LPA. For two consecutive accesses to the same LPA, there are other LPAs that can be accessed multiple times, rather than just once within this time interval. Therefore, when adjusting the target data in the cache for the target logical page address based on the initial hot / cold classification results, it is necessary to determine the size of the data operation requests; determine the access period and distance value of the target logical page address; the distance value represents the number of requests to the same logical page address between two consecutive accesses to the target logical page address; generate a target hot / cold classification result for the data operation requests based on the initial hot / cold classification results, the size value, the access period, and the distance value; and adjust the target data in the cache for the target logical page address based on the target hot / cold classification results.

[0083] It should be noted that the distance value is different from the recency factor. The recency factor represents the total number of requests between two identical LPAs, while the distance value represents the number of unique requests between two identical LPAs. For ease of understanding, assume that the logical page addresses accessed within a period of time are numbered 122334455661. With number 1 as the statistical object, the recency factor is 10, and the distance value is 5. That is, the distance value is the result of deduplicating the recency factor.

[0084] In an exemplary embodiment, in the process of adjusting the target logical page address of the target data in the cache according to the target hot and cold classification result, a three-layer linked list can be designed to organize different types of requests, and then manage the data accordingly. For example, the target hot and cold classification result can be parsed; in response to the target hot and cold classification result being a hot classification result, the target data corresponding to the target logical page address is moved to the head of the first linked list in the cache; in response to the target hot and cold classification result being an intermediate classification result, the target data corresponding to the target logical page address is moved to the head of the second linked list in the cache; in response to the target hot and cold classification result being a cold classification result, the target data corresponding to the target logical page address is moved to the head of the third linked list in the cache. To facilitate understanding of this process, it is assumed that the first linked list is represented by L1, the second linked list is represented by L, and the third linked list is represented by L0. The management process of the linked lists is as follows: Figure 5 shown.

[0085] In a specific application scenario, in the process of applying a linked list, considering that as the data moves, the linked list may be full, corresponding processing is required at this time, that is, in response to the first linked list being full, the data at the end of the first linked list is moved to the head of the second linked list; in response to the second linked list being full, the data at the end of the second linked list is moved to the head of the third linked list; in response to the third linked list being full, the data at the end of the third linked list is stored in the hard disk.

[0086] In this way, this application designs a three-level linked list to manage the data in the cache, and distinguishes the data with different hot and cold attributes in the cache through different linked list management; the cold cache data in the linked list is deleted first and saved in the Nand Flash, while the hot data is stored in the cache as much as possible. Compared with the conventional linked list, the three-layer linked list provides enhanced retrieval capabilities by accurately replacing the requests with the lowest access frequency in the cache, thereby improving the cache hit rate and preventing flash memory degradation.

[0087] In order to understand the update process of the three-level linked list, assuming that the data operation request is an IO request, the update process is as follows: Figure 6 As shown, the following process is included:

[0088] Receive IO requests;

[0089] Determine whether the cache hits;

[0090] If it hits, update the valid linked list, update the linked lists L0, L1 and L2; determine whether it is a write request, if so, update the cache data, otherwise end;

[0091] If there is no hit, determine whether the cache is full;

[0092] If the cache is full, determine whether the IO linked list is empty. If not, delete the L0 tail data and then write the cache data. If it is empty, determine whether the L1 linked list is empty. If the L1 linked list is empty, delete the L2 tail data and then write the cache data. If the L1 linked list is not empty, delete the L1 tail data and then write the cache data.

[0093] If the cache is not full, write the cache data;

[0094] After writing the cache data, the linked lists L0, L1 and L2 are updated and then the process ends.

[0095] The present application provides a cache management method, which includes obtaining a data operation request; determining the access frequency, recency factor, and interval value of a target logical page address corresponding to the data operation request; obtaining a pre-trained hot and cold request classifier, wherein the hot and cold request classifier is built based on machine learning; processing the recency factor, access frequency, and interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request; adjusting the target data of the target logical page address in the cache according to the initial hot and cold classification result; wherein the recency factor represents the number of requests between two consecutive accesses to the target logical page address; and the interval value represents the number of requests between a current request and a subsequent request for the same access to the target logical page address. In this application, the recency factor represents the number of requests between two consecutive accesses to the target logical page address, and the interval value represents the number of requests between the current request and the subsequent request between the same access target logical page address. Since the number of other requests between the target logical page addresses will affect the hit situation of the data in the cache, for example, the more requests, the lower the hit rate of the data in the cache, the access frequency determines the hit rate of the data itself in the cache, and the hot and cold request classifier built based on machine learning has strong generalization ability and can reduce overfitting, so it can generate more accurate initial hot and cold classification results, so that the target data of the target logical page address in the cache can be adjusted more accurately, so that the storage of data in the cache corresponds to the hot and cold attributes of the data, thereby improving the hit rate of the cache.

[0096] See also Figure 7 , Figure 7 A schematic diagram of the structure of a cache management system provided in an embodiment of the present application.

[0097] An embodiment of the present application provides a cache management system, which may include:

[0098] Request acquisition module 101, used to obtain data operation requests;

[0099] A parameter determination module 102 is configured to determine an access frequency, a recency factor, and an interval value of a target logical page address corresponding to a data operation request;

[0100] The classifier acquisition module 103 is used to obtain a pre-trained hot and cold request classifier, which is built based on machine learning;

[0101] The classification result generating module 104 is used to process the recency factor, access frequency and interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request;

[0102] A data adjustment module 105 is configured to adjust the target data of the target logical page address in the cache according to the initial hot / cold classification result;

[0103] The recency factor represents the number of requests between two consecutive accesses to the target logical page address; the interval value represents the number of requests between the current request and the subsequent request for the same access target logical page address.

[0104] In a cache management system provided in an embodiment of the present application, a classification result generation module may include:

[0105] The classification result generating unit is used to process the recency factor, access frequency and interval value based on the hot and cold request classifier in accordance with the generation condition that the hot and cold classification result is negatively correlated with the recency factor and the interval value and positively correlated with the access frequency, so as to generate an initial hot and cold classification result of the data operation request.

[0106] In a cache management system provided in an embodiment of the present application, a classification result generation unit may be used to:

[0107] Input the proximity factor, access frequency, and interval value into each hot and cold decision tree in the hot and cold data classifier;

[0108] Receive the initial hot and cold attribute scores output by the hot and cold decision tree;

[0109] Processing is performed based on the initial hot and cold attribute scores to generate the initial hot and cold classification results of the data operation request.

[0110] In a cache management system provided in an embodiment of the present application, a classification result generation unit may be used to:

[0111] receiving the initial hot and cold attribute scores output by the hot and cold decision tree according to the hot and cold attribute score generation formula;

[0112] The formula for generating the hot and cold attribute scores includes:

[0113] Y = (1 / R) * F * (1 / J);

[0114] Among them, Y represents the initial hot and cold attribute score; R represents the recency factor; F represents the access frequency; and J represents the interval value.

[0115] In an embodiment of the present application, a cache management system is provided, wherein a data adjustment module may include:

[0116] a size value determining unit, configured to determine a size value of a data operation request;

[0117] a parameter determination unit, configured to determine an access cycle and a distance value of a target logical page address; the distance value represents the number of requests for the same logical page address between two consecutive accesses to the target logical page address;

[0118] A hot and cold classification unit, configured to generate a target hot and cold classification result of the data operation request based on an initial hot and cold classification result, a size value, an access cycle, and a distance value;

[0119] The data adjustment unit is used to adjust the target data of the target logical page address in the cache according to the target hot and cold classification result.

[0120] In a cache management system provided by an embodiment of the present application, a data adjustment unit can be used to:

[0121] Analyze the target hot and cold classification results;

[0122] In response to the target cold-hot classification result being a hot classification result, moving the target data corresponding to the target logical page address to the head of the first linked list in the cache;

[0123] In response to the target hot / cold classification result being an intermediate classification result, moving the target data corresponding to the target logical page address to the head of the second linked list in the cache;

[0124] In response to the target hot / cold classification result being a cold classification result, the target data corresponding to the target logical page address is moved to the head of the third linked list in the cache.

[0125] In a cache management system provided in an embodiment of the present application, the data adjustment unit may also be used to:

[0126] In response to the first linked list being full, moving the data at the tail of the first linked list to the head of the second linked list;

[0127] In response to the second linked list being full, moving the data at the tail of the second linked list to the head of the third linked list;

[0128] In response to the third linked list being full, the data at the end of the third linked list is stored in the hard disk.

[0129] It should be noted that the cache management system and the structure of the solid state drive in this application can be flexibly configured as needed. For example, the cache management system that cooperates with the FTL can be configured as follows: Figure 8 As shown, it includes IO request module, feature collection module, Agent module, training model module, generation model function module, Cache management module and update Cache module;

[0130] IO request module, used to receive read and write commands from the host and send corresponding read and write request commands to the FTL;

[0131] Feature collection module, used to collect feature information related to read and write commands;

[0132] Agent module, used to linearly train machine learning models through agent modules;

[0133] The training model module is used to obtain corresponding results based on the information of the feature collection module;

[0134] Generate model function module, which is used to train machine learning model through feature information and labels to obtain hot and cold classification results;

[0135] Cache management module, used to manage cache information through three-level linked lists and model output results

[0136] The Cache Update module is used to update, add, delete, and perform other operations on the data in the cache according to the instructions of the Cache Management module.

[0137] The cache management device effectively processes read and write commands from the host. Through a management method that combines offline machine learning with a three-level linked list, it improves the cache hit rate, enhances the read and write speed, and improves the performance of the solid-state drive.

[0138] The internal structure of the solid state drive can be as follows Figure 9 As shown in the figure, the algorithm combining the three-level linked list with offline random forest machine learning improves the cache hit rate of the solid-state drive, and the offline algorithm releases the performance of the solid-state drive.

[0139] This application also provides an electronic device and a computer-readable storage medium, both of which have the corresponding effects of the cache management method provided in the embodiment of this application. Figure 10 , Figure 10 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0140] An electronic device provided in an embodiment of the present application includes a memory 201 and a processor 202. The memory 201 stores a computer program, and when the processor 202 executes the computer program, the steps of the cache management method described in any of the above embodiments are implemented.

[0141] See also Figure 11Another electronic device provided in an embodiment of the present application may further include: an input port 203 connected to the processor 202 for transmitting commands inputted from the outside to the processor 202; a display unit 204 connected to the processor 202 for displaying the processing results of the processor 202 to the outside world; and a communication module 205 connected to the processor 202 for enabling communication between the electronic device and the outside world. The display unit 204 may be a display panel, a laser scanning display, etc. The communication method adopted by the communication module 205 includes but is not limited to Mobile High-Definition Link (MHL), Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), wireless connection: Wireless Fidelity (WiFi), Bluetooth communication technology, Bluetooth low energy communication technology, and communication technology based on IEEE802.11s.

[0142] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the cache management method described in any of the above embodiments are implemented.

[0143] The computer-readable storage medium involved in this application includes random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD-ROMs (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.

[0144] A computer program product provided by an embodiment of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the cache management method described in any of the above embodiments.

[0145] For descriptions of the relevant portions of the cache management system, computer program product, electronic device, and computer-readable storage medium provided in the embodiments of the present application, please refer to the detailed description of the corresponding portions of the cache management method provided in the embodiments of the present application, which will not be repeated here. In addition, portions of the above-mentioned technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.

[0146] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0147] The above description of the disclosed embodiments will enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A cache management method, characterized in that: include: Get data operation request; Determining an access frequency, a recency factor, and an interval value of a target logical page address corresponding to the data operation request; Obtain a pre-trained hot and cold request classifier, where the hot and cold request classifier is built based on machine learning; Processing the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request; Adjusting target data of the target logical page address in the cache according to the initial hot / cold classification result; The recency factor represents the number of requests between two consecutive accesses to the target logical page address; and the interval value represents the number of requests between a current request and a subsequent request for the same access to the target logical page address.

2. The cache management method according to claim 1, wherein: The processing of the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request includes: According to a generation condition that a hot and cold classification result is negatively correlated with the recency factor and the interval value and positively correlated with the access frequency, the hot and cold request classifier processes the recency factor, the access frequency, and the interval value to generate an initial hot and cold classification result of the data operation request.

3. The cache management method according to claim 2, wherein: The processing of the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request includes: inputting the recency factor, the access frequency and the interval value into each hot and cold decision tree in the hot and cold data classifier; receiving an initial hot and cold attribute score output by the hot and cold decision tree; Processing is performed according to the initial hot and cold attribute scores to generate an initial hot and cold classification result of the data operation request.

4. The cache management method according to claim 3, wherein: The receiving the initial hot and cold attribute scores output by the hot and cold decision tree includes: receiving an initial hot and cold attribute score output by the hot and cold decision tree according to a hot and cold attribute score generation formula; The formula for generating the cold and hot attribute scores includes: Y = (1 / R) * F * (1 / J); Wherein, Y represents the initial hot / cold attribute score; R represents the recency factor; F represents the access frequency; and J represents the interval value.

5. The cache management method according to claim 1, wherein: The step of adjusting the target data of the target logical page address in the cache according to the initial hot / cold classification result includes: Determining a size value of the data operation request; Determining an access cycle and a distance value of the target logical page address; the distance value represents the number of requests for the same logical page address between two consecutive accesses to the target logical page address; generating a target hot / cold classification result of the data operation request based on the initial hot / cold classification result, the size value, the access cycle, and the distance value; According to the target hot / cold classification result, the target data of the target logical page address in the cache is adjusted.

6. The cache management method according to claim 5, characterized in that: The step of adjusting the target data of the target logical page address in the cache according to the target hot / cold classification result includes: Analyzing the target hot and cold classification results; In response to the target cold-hot classification result being a hot classification result, moving the target data corresponding to the target logical page address to the head of a first linked list in the cache; In response to the target hot / cold classification result being an intermediate classification result, moving the target data corresponding to the target logical page address to the head of the second linked list in the cache; In response to the target hot / cold classification result being a cold classification result, the target data corresponding to the target logical page address is moved to the head of the third linked list in the cache.

7. The cache management method according to claim 6, wherein: Also includes: In response to the first linked list being full, moving the data at the tail of the first linked list to the head of the second linked list; In response to the second linked list being full, moving the data at the tail of the second linked list to the head of the third linked list; In response to the third linked list being full, the data at the end of the third linked list is stored in the hard disk.

8. A cache management system, characterized in that: include: Request acquisition module, used to obtain data operation requests; a parameter determination module, configured to determine an access frequency, a recency factor, and an interval value of a target logical page address corresponding to the data operation request; A classifier acquisition module is used to obtain a pre-trained hot and cold request classifier, which is built based on machine learning; a classification result generating module, configured to process the recency factor, the access frequency, and the interval value based on the hot and cold request classifier to generate an initial hot and cold classification result of the data operation request; a data adjustment module, configured to adjust the target data of the target logical page address in the cache according to the initial hot / cold classification result; The recency factor represents the number of requests between two consecutive accesses to the target logical page address; and the interval value represents the number of requests between a current request and a subsequent request for the same access to the target logical page address.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the cache management method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the cache management method according to any one of claims 1 to 7 are implemented.