Management Method, Device and Storage Medium of Shared Cache
By dynamically determining partition sizes and eviction algorithms based on the access characteristics of different classes of IO requests, the method optimizes shared cache performance by enhancing hit rates and reducing data flushing in shared cache systems.
Patent Information
- Application Number
- CN202411307486.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-02
- Filing Date
- 2022-07-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The prior art usually relies on artificially specified or simple strategies when partitioning cache partitions in shared caches, resulting in poor cache performance and the use of the same phase-out algorithm fails to effectively optimize the cache requirements of multiple entities.
By determining the access characteristics of each category of IO requests, the partition size and phase-out algorithm are dynamically configured according to the hit rate and cache size relationship to optimize shared cache performance.
It improves the overall cache performance of shared cache, improves the hit rate of reading data and the hit rate of writing data, reduces the number of times of flashing disks to the back-end storage unit, and optimizes the utilization of cache resources.
Smart Images

Figure CN119292962B_ABST
Abstract
Description
[0001] This application is a divisional application of a Chinese application with an application date of July 29, 2022 and an application number of 202210908210.5. Both this application and the Chinese application with an application number of 202210908210.5 claim the priority of a Chinese patent application with an application number of 202210197738.6 and an invention title of "A Method and Device for Managing Shared Cache" filed on March 2, 2022, and the entire content thereof is incorporated herein by reference. Technical Field
[0002] This application relates to the field of computer technologies, and in particular, to a method, a device, and a storage medium for managing a shared cache. Background Art
[0003] A shared cache refers to sharing cache resources among multiple entities (such as multiple applications, multiple clients, or multiple processing cores (cores) of a computing device, etc.) in a cache architecture to meet different cache requirements. For example, for a central processing unit (CPU) including multiple cores in a computing device, the multiple cores share the last level cache (LLC) of the CPU. For another example, in a distributed cache system, the data of a single client is scattered and cached in different cache nodes. In contrast, in a scenario with multiple clients, the multiple clients share the cache resources in a single cache node.
[0004] When multiple entities share a cache, to avoid cache contention, the cache shared by the multiple entities can be partitioned so that different partitions are used to cache the data to be cached by different entities.
[0005] However, when currently partitioning cache partitions in a shared cache, the partition size is usually specified manually or determined by a simple priority policy, and in addition, the same cache eviction algorithm is used for the entire shared cache, resulting in poor performance of the shared cache. Summary of the Invention
[0006] This application provides a method, a device, and a storage medium for managing a shared cache to improve the cache performance in the shared cache.
[0007] To achieve the above object, this application provides the following technical solutions:
[0008] In a first aspect, the present application provides a method for managing a shared cache, which is used to manage the shared cache. Among them, the shared cache is used to cache data requested by multiple I / O requests. The multiple I / O requests correspond to K categories, and the shared cache corresponds to N replacement algorithms. The method includes: determining the access characteristics of the I / O requests of each category in the K categories for accessing the shared cache. According to the access characteristics of the I / O requests of the K categories and the hit rate of the shared cache, determining the partition size and replacement algorithm of the I / O requests of each category in the shared cache. Configuring the cache size of the I / O requests of each category in the shared cache as the determined partition size of the I / O requests of each category in the shared cache, and configuring the replacement algorithm of the I / O requests of each category in the shared cache as the determined replacement algorithm of the I / O requests of each category in the shared cache. Among them, the access characteristics of the I / O requests of each category in the K categories for accessing the shared cache are: the relationship between the hit rate and the cache size of the I / O requests of each category in the K categories under the N replacement algorithms respectively.
[0009] Through the method for managing the shared cache provided by the present application, by obtaining the access characteristics of the input / output (I / O) requests of each category, the partition size and cache replacement algorithm are determined for the I / O requests of each category in the shared cache, so as to improve the cache performance suitable for each type of I / O request, thereby improving the cache performance in the entire shared cache. Specifically, in the scenario of reading data, the hit rate of the shared cache can be improved, thereby improving the efficiency of data reading. In the scenario of writing data, by increasing the write hit rate, the number of times the shared cache flushes to the backend storage unit is saved.
[0010] In a possible design, the above determining the access characteristics of the I / O requests of each category in the K categories for accessing the shared cache includes: for the first replacement algorithm among the N replacement algorithms, simulating in the shared cache the hit rate of the I / O requests of each category in the K categories applying the first replacement algorithm in caches of different sizes, so as to obtain the relationship between the hit rate and the cache size. Among them, the first replacement algorithm is any one of the N replacement algorithms.
[0011] In another possible design, the above determining the access characteristics of the I / O requests of each category in the K categories for accessing the shared cache includes: for the first replacement algorithm among the N replacement algorithms, according to the reuse distance of each I / O request in the first type of I / O requests and different cache sizes, determining the hit rate of the first type of I / O requests applying the first replacement algorithm in caches of different sizes, so as to obtain the relationship between the hit rate and the cache size. Among them, the first replacement algorithm is any one of the N replacement algorithms, and the first type of I / O requests is any one of the K categories of I / O requests.
[0012] The above two possible designs provide two methods for determining the access characteristics of the IO requests of each of the K categories to access the shared cache.
[0013] In another possible design approach, the method for determining the partition size and the replacement algorithm of the IO requests of each category in the shared cache according to the access characteristics of the IO requests of the K categories and the hit rate of the shared cache includes: determining the hit rates of the IO requests of the K categories in the shared cache for each combination according to the X hit rates corresponding to the X cache sizes determined for the IO requests of each category under each replacement algorithm. The cache size corresponding to the IO requests of each category when the hit rate of the IO requests of the K categories in the shared cache is the largest is determined as the partition size of the IO requests of each category in the shared cache, and the replacement algorithm corresponding to the IO requests of each category when the hit rate of the IO requests of the K categories in the shared cache is the largest is determined as the replacement algorithm of the IO requests of each category in the shared cache. Among them, for any one of the K categories of IO requests, the X cache sizes and the N replacement algorithms form X * N combinations, each combination includes a cache size and a replacement algorithm, and the X cache sizes are X cache sizes preset for the cache corresponding to the IO requests of each category.
[0014] Through this possible design, it is possible to determine the partition size and the replacement algorithm corresponding to the IO requests of each category when the hit rate of the IO requests of the K categories in the shared cache is the largest. Furthermore, based on the determined partition size and replacement algorithm, configuring the cache size and the replacement algorithm of the IO requests of each of the K categories in the shared cache can achieve the purpose of optimizing the cache size and the replacement algorithm of the IO requests of each of the K categories in the shared cache by jointly solving the two factors of the replacement algorithm and the cache size that affect the cache performance. Thus, the optimized cache size and the replacement algorithm of the IO requests of each of the K categories in the shared cache can improve the overall cache performance of the shared cache.
[0015] In another possible design approach, before determining the access characteristics of the IO requests of each of the K categories to access the shared cache, the above method further includes: obtaining a plurality of IO requests. Dividing the IO requests into K categories according to the characteristics of the addresses of the data accessed by the plurality of IO requests or according to the category tags carried in the plurality of IO requests.
[0016] When classifying IO requests into K categories according to the characteristics of the addresses of the data accessed by multiple IO requests, the entity initiating the IO requests does not need to label the category tags for classification for the IO requests, so that no additional resource overhead is generated for the entity initiating the IO requests. Moreover, since the method provided in this application does not intrude into the upper layer of the cache (i.e., the entity initiating the IO requests), the method provided in this application can be applied to a general cache system for diverse customers without ecological support.
[0017] In another possible design, if the above shared cache is the LLC of the CPU in the computing device, then the above multiple IO requests are the IO requests initiated by multiple processing cores in the CPU.
[0018] In another possible design, if the above shared cache is the cache in the cache node, then the above multiple IO requests are the IO requests initiated by multiple computing nodes accessing the cache node.
[0019] In another possible design, if the above shared cache is a cache pool composed of caches in multiple nodes, then the above multiple IO requests are the IO requests initiated by multiple computing nodes accessing the cache pool.
[0020] Through the above three possible designs, the method provided in this application can be applied to multiple scenarios.
[0021] In another possible design, the above access characteristics are characterized by the hit rate curve (HRC) or miss rate curve (MRC) of the IO requests.
[0022] In another possible design, the determination of the access characteristics of the IO requests in each of the K categories accessing the shared cache includes: periodically determining the access characteristics of the IO requests in each of the K categories accessing the shared cache. For the access characteristics of the IO requests in each of the K categories accessing the shared cache determined in the first period, the determination of the partition size and replacement algorithm of the IO requests in each category in the shared cache according to the access characteristics of the IO requests in the K categories and the hit rate of the shared cache includes: determining the partition size and replacement algorithm of the IO requests in each category in the shared cache in the first period according to the access characteristics of the IO requests in the K categories determined in the first period and the hit rate of the shared cache. Here, the first period is any period for determining the access characteristics of the IO requests in each of the K categories accessing the shared cache.
[0023] Through the above possible designs, the shared cache can determine the cache size and eviction algorithm for the I / O requests of each of the K categories periodically determined by the computing device, so that it can periodically adjust the cache size and eviction algorithm of the I / O requests of each category in the shared cache, thereby achieving an improvement in the hit rate of the I / O requests in the shared cache in the time domain, and thus improving the overall cache performance of the shared cache in the time domain.
[0024] In a second aspect, the present application provides a management device for a shared cache. The management device for the shared cache is used to execute any one of the methods provided in the first aspect above. The present application can perform a functional module division on the management device for the shared cache according to any one of the methods provided in the first aspect above. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. Exemplarily, the present application can divide the management device for the shared cache into a determination unit, a configuration unit, etc. according to functions. The descriptions of the possible technical solutions and beneficial effects executed by each of the above divided functional modules can all refer to the technical solutions provided in the first aspect or its corresponding possible designs above, and will not be elaborated here.
[0025] In a third aspect, the present application provides a computing device, which is used to manage a shared cache. The computing device includes: a memory, and one or more processors configured to read program instructions stored in the memory to execute any one of the methods provided in the first aspect and any of its possible design manners.
[0026] In a fourth aspect, the present application provides a computer-readable storage medium, which includes program instructions. When the program instructions run on a computer or a processor, the computer or the processor is caused to execute any one of the methods provided in any of the possible implementation manners in the first aspect.
[0027] In a fifth aspect, the present application provides a computer program product, which when running on a computing device causes any one of the methods provided in any of the possible implementation manners in the first aspect to be executed.
[0028] It can be understood that any of the above provided devices, computer storage media, or computer program products, etc. can be applied to the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods, and will not be elaborated here.
[0029] In the present application, the name of the above management device for the shared cache does not constitute a limitation on the device or functional module itself. In actual implementation, these devices or functional modules can appear under other names. As long as the functions of each device or functional module are similar to those of the present application and fall within the scope of the claims of the present application and their equivalent technologies. Brief Description of the Drawings
[0030] Figure 1 It is a schematic diagram of a scenario for a shared cache;
[0031] Figure 2 It is another schematic diagram of a scenario for a shared cache;
[0032] Figure 3 It is yet another schematic diagram of a scenario for a shared cache;
[0033] Figure 4 It is a schematic diagram of the hardware structure of a computing device provided by an embodiment of the present application;
[0034] Figure 5 It is a schematic flowchart of a method for managing a shared cache provided by an embodiment of the present application;
[0035] Figure 6 It is a schematic diagram of an IO request initiated by an entity to access a shared cache provided by an embodiment of the present application;
[0036] Figure 7 It is a schematic flowchart of a method for generating a classifier provided by an embodiment of the present application;
[0037] Figure 8 It is a schematic diagram of the classification of an IO request provided by an embodiment of the present application;
[0038] Figure 9 It is a schematic flowchart of another method for managing a shared cache provided by an embodiment of the present application;
[0039] Figure 10 It is a schematic flowchart of yet another method for managing a shared cache provided by an embodiment of the present application;
[0040] Figure 11 It is a schematic diagram of the structure of a shared cache management device 110 provided by an embodiment of the present application. Detailed Description of the Embodiments
[0041] To more clearly understand the embodiments of the present application, some terms or technologies involved in the embodiments of the present application are described below:
[0042] 1), Cache
[0043] When the computing speed of the computing unit does not match the memory access speed of the computing unit accessing the data stored in the storage unit, the computing unit will generate a large amount of idling while waiting to access data in the storage unit. To solve this problem, cache technology came into being.
[0044] Among them, a cache refers to a memory that can perform high-speed data exchange. Generally, the cache is located between a storage unit (such as an external memory, simply referred to as an external storage) and a computing unit (such as a CPU). Among them, the external storage can be, for example, the hard disk of a device, and the hard disk can be, for example, a solid state disk (SSD) or a hard disk drive (HDD), etc.
[0045] On the one hand, by preloading high-value data from the storage unit into the cache in advance, when the computing unit accesses this high-value data, it can directly read it from the cache that can perform high-speed data exchange. In this way, the efficiency of the computing unit accessing memory data can be improved, thereby reducing the idle time of the computing unit and increasing the execution speed of the application program. Among them, high-value data can be, for example, data that the computing unit has recently accessed, or data that the computing unit accesses relatively frequently, or working set data, etc., which is not limited to this. Here, the working set can be understood as a data set composed of all the data required by an application program.
[0046] On the other hand, when an entity writes data to the storage unit, the data to be written can be first written into the cache, and then the cache regularly writes the cache data to the external storage (referred to as flushing the disk). Among them, the entity can be, for example, an application, a computing node, a client device, or a core of a CPU, etc., which is not limited to this. In this way, for the entity, writing data into the cache can be considered that the input / output (IO) operation for writing data has been completed, which can improve the response speed of the entity.
[0047] Optionally, the cache of the device can generally be implemented by the memory of the node and / or the SSD of the node. Among them, the memory, also known as the main memory, is the storage space directly addressable by the CPU. As an example, the memory can be a dynamic random access memory (DRAM), etc., which is not limited to this.
[0048] 2) Shared cache
[0049] To meet different cache requirements, the cache resources shared among multiple entities in the cache architecture are called shared caches. Among them, the multiple entities can be, for example, multiple applications, multiple computing nodes, multiple client devices, or multiple processing cores in the CPU of a computing device, etc., which is not limited to this. Among them, an application refers to an application program running on a computing device, a computing node is a computing node that needs to access the shared cache, a client device can be, for example, a client device of a storage system, a client device of an application program, etc., and a computing device can be any device including a CPU such as a general computer, a laptop, a tablet, a server, etc., and this is not limited.
[0050] The shared cache can be the third-level cache of a CPU that includes multiple processing cores in a computing device. In this case, the multiple entities that initiate IO requests to the shared cache are the multiple processing cores in the CPU.
[0051] As an example, refer to Figure 1 , Figure 1 which shows a schematic diagram of a scenario of the shared cache. As Figure 1 shown, the CPU 10 of the computing device includes 4 processing cores, namely processing core 1, processing core 2, processing core 3, and processing core 4. Among them, each of the 4 processing cores of the CPU 10 is configured with an independent level 1 cache and a level 2 cache, and the 4 processing cores of the CPU 10 can all access the third-level cache of the CPU 10, that is, the third-level cache of the CPU 10 is the shared cache of these 4 processing cores.
[0052] The shared cache can also be the cache in a cache node in a cache system. In this case, the multiple entities that initiate IO requests to the shared cache are multiple client devices that access the cache node, or multiple computing nodes that access the shared cache. Among them, the specific form of the client device can also be any computing node with computing and processing functions, which is not limited here.
[0053] Among them, the cache system can be a distributed cache system. It can be understood that in a distributed cache system, the data of a computing node can be cached in the caches of multiple cache nodes in the distributed cache system. Conversely, the cache of a cache node in the distributed cache system can cache the data of multiple computing nodes. Therefore, the cache of the cache node in the distributed cache system that is used to cache the data of multiple computing nodes is the shared cache of these multiple computing nodes.
[0054] The above cache system can also be a centralized cache system. Generally, a centralized cache system includes one cache node. In this case, the data of multiple computing nodes are all cached in the cache of this cache node. Therefore, the cache of the cache node in the centralized cache system is the shared cache of these multiple computing nodes.
[0055] It should be understood that the above cache nodes can be independent devices. In this case, the caches of the independent devices are all used as the caches of the cache nodes. Alternatively, the functions implemented by the above cache nodes are implemented by functional modules in the independent device, that is, in addition to implementing the functions of the cache nodes, the independent device can also implement other functional purposes. In this case, part of the cache of the independent device is used as the cache of the cache nodes. As an example, the functions implemented by the cache nodes in a distributed cache system are integrated into a computing device that is a storage node in a distributed storage system. That is to say, in addition to implementing the functions of the storage node, the computing device also implements the functions of the cache nodes. In this case, part of the cache in the independent device is used as the cache of the cache nodes.
[0056] As an example, refer to Figure 2 , Figure 2 which shows another schematic diagram of a shared cache scenario. As shown in Figure 2 shown, Figure 2 shows j computing nodes accessing the cache system 20 in a distributed cache system 20, namely computing node 1, computing node 2, computing node 3, …, and computing node j, where j is a positive integer. And the cache system 20 includes 3 cache nodes, namely cache node 1, cache node 2, and cache node 3. Among them, the high-value data of any one of the j computing nodes is scattered and cached in the caches of multiple cache nodes in the cache system 20. In contrast, the cache of a single cache node in the cache system 20 is the shared cache of the j computing nodes.
[0057] The shared cache can also be a cache pool composed of the caches of multiple nodes. Among them, the cache pool is a cache pool composed of the caches of each node in the multiple nodes. In this case, the multiple entities that initiate IO requests to the shared cache are the multiple computing nodes accessing the cache pool. That is, the cache pool is used to cache the data of the multiple computing nodes.
[0058] Among them, the nodes that provide caches for the cache pool can be nodes in any network system, and there is no limitation on this. For example, if the nodes that provide caches for the cache pool are storage nodes in a storage system, the multiple computing nodes accessing the cache pool are multiple client devices of the storage system. For another example, if the nodes that provide caches for the cache pool are multiple servers in a server cluster, the multiple computing nodes accessing the cache pool are multiple client devices of the server cluster, or multiple applications accessing the server cluster are running among the multiple computing nodes accessing the cache pool, and there is no limitation on this.
[0059] As an example, refer to Figure 3 , Figure 3 which shows yet another schematic diagram of a shared cache scenario. As shown in Figure 3As shown in the figure, Node 1, Node 2, and Node 3 are storage nodes in the storage system. The cache of Node 1 is Cache 1, the cache of Node 2 is Cache 2, and the cache of Node 3 is Cache 3. Then, a part of Cache 1, a part of Cache 2, and a part of Cache 3 can form Cache Pool 30. Furthermore, the data of Computing Nodes 1, 2, 3, …, j in the storage system can be cached in Cache Pool 30, and Cache Pool 30 is the shared cache for these j computing nodes.
[0060] 3), Read Hit, Write Hit
[0061] Since the size of the cache is generally limited and much smaller than the size of the external storage, the amount of data that can be stored in the cache is very limited.
[0062] In the scenario of reading data, the IO request initiated by the entity for reading data carries the logical address of the data to be read. That is, the logical address carried in the IO request for reading data is the logical address to be accessed by this IO request.
[0063] When there is data in the cache corresponding to the logical address carried by the IO request, the IO request hits in the cache.
[0064] When there is no data in the cache corresponding to the logical address carried by the IO request, it means that the IO request for reading data does not hit the data to be read in the cache, that is, a cache miss occurs. Furthermore, the cache miss can trigger reading the data to be read requested by the IO request from the backend of the cache (such as the external storage), and caching the data read from the backend in the cache.
[0065] It can be seen that when the IO request hits in the cache, there is no need to read the data to be read from the backend, thus improving the response speed of the entity.
[0066] In the scenario of writing data, the IO request initiated by the entity for writing data carries the data to be written and the logical address for storing the data to be written. That is, the logical address carried in the IO request for writing data is the logical address to be accessed by this IO request.
[0067] When there is a physical address in the cache corresponding to the logical address carried by the IO request, it means that the logical address has been written with data before this IO request and cached in the cache, but has not been flushed from the cache to the external storage yet. In this case, it is said that the IO request for writing data hits the logical address for storing the data to be written in the cache, simply referred to as the IO request hitting in the cache for writing. Furthermore, the cache can update the data written to this logical address in the cache based on the data to be written carried by this IO request, and flush the updated data to the external storage subsequently.
[0068] When the physical address corresponding to the logical address carried by the IO request does not exist in the cache, it means that the logical address has not been written with data before this IO request, or the logical address has been written with data before this IO request and the written data has been flushed from the cache to the external memory. In this case, it is said that the IO request for writing data misses the logical address for storing the data to be written in the cache, simply referred to as the IO request not hitting the write in the cache. Furthermore, the cache allocates a corresponding physical address for the logical address carried in the IO request received this time, and writes the data to be written carried in this IO request to this physical address, thereby achieving caching of the data to be written.
[0069] It can be seen that when the IO request hits the write in the cache, it can update the data already written in the logical address in the cache, thereby reducing the number of times the cache flushes data to the external memory, and further saving the bandwidth between the cache and the external memory.
[0070] 4) Eviction algorithm
[0071] In the scenario of reading data, when the IO request misses the data to be read in the cache, data needs to be read from the backend of the cache (such as the external memory) into the cache. For example, for the tertiary cache in the CPU, when a cache miss occurs in the tertiary cache of the CPU, data needs to be read from the memory of the computing device where the CPU is located into the tertiary cache of the CPU. Another example is for the memory. When a cache miss occurs in the memory, data needs to be read from the external memory of the device where the memory is located into the memory.
[0072] If the free space in the cache is not enough to store the data read from the backend, it is necessary to evict the existing data in the cache (such as deleting some or all of the data, or marking some or all of the data as invalid, etc.), so as to provide storage space for the newly read data from the backend. Among them, the algorithm used to determine the data that needs to be evicted among the existing data in the cache is the eviction algorithm.
[0073] Generally, the eviction algorithm is designed based on the access pattern (such as the frequency of accessing data, etc.) when the IO request accesses data. Thus, by applying the eviction algorithm in the cache, high-value data can be retained in the cache for as long as possible, while low-value data is evicted. This can improve the hit rate of the IO request for reading data hitting the read in the cache, thereby improving the cache performance.
[0074] In the data writing scenario, the data cached in the cache is periodically flushed to the external storage, or when the data cached in the cache exceeds a certain quantity, the currently cached data is flushed to the external storage. In this case, some data with a relatively low subsequent update frequency can be flushed to the external storage, while the data with a relatively high subsequent update frequency is retained in the cache. In this way, the hit rate of subsequent IO requests in the cache write hit can be increased, thereby reducing the number of times the cache flushes data to the external storage. Here, the algorithm used to determine which data in the cache needs to be flushed to the external storage is the eviction algorithm.
[0075] 5) Cache Performance
[0076] Cache performance can generally be evaluated by the hit rate or the miss rate.
[0077] In the scenario where the read and write caches are separated (i.e., the read cache and the write cache are isolated from each other physically or through software), the hit rate of read hit is: the ratio of the number of times an IO request is read-hit in the cache within a period of time to the total number of all IO read requests during this period. The miss rate in the read data scenario: the ratio of the number of times an IO request experiences a cache miss in the cache within a period of time to the total number of all IO read requests during this period. The hit rate of write hit is: the ratio of the number of times an IO request is write-hit in the cache within a period of time to the total number of all IO write requests during this period. The miss rate in the write data scenario is: the ratio of the number of times an IO request fails to be write-hit in the cache within a period of time to the total number of all IO write requests during this period.
[0078] In the scenario where the read and write caches are integrated (i.e., the read data and the write data share the cache space), the hit rate is: the ratio of the sum of the number of times an IO request is read-hit in the cache and the number of times an IO request is write-hit in the cache within a period of time to the total number of all IO requests during this period. The miss rate is: the ratio of the number of times an IO request experiences a cache miss in the cache and the number of times an IO request fails to be write-hit in the cache within a period of time to the total number of all IO requests during this period.
[0079] It can be understood that cache performance is related not only to the eviction algorithm applied in the cache but also to the size of the cache itself. Therefore, in practice, cache performance is generally characterized by the cache miss rate curve (MRC) or the cache hit rate curve (HRC). Among them, the MRC is the curve of the correspondence between the cache size and the miss rate, used to describe the miss rate of IO requests in the cache when the cache is of different sizes under a certain eviction algorithm. The HRC is the curve of the correspondence between the cache size and the hit rate, used to describe the hit rate of IO requests in the cache when the cache is of different sizes under a certain eviction algorithm.
[0080] 6) Reuse Distance
[0081] For an I / O request initiated by an entity, the reuse distance of the I / O request is used to indicate the number of different logical addresses accessed by other I / O requests during the interval between two consecutive accesses to the logical address carried by the I / O request.
[0082] Among them, the reuse distance of the I / O request can be characterized by the number of different logical addresses. Specifically, for any one of the multiple I / O requests initiated by an entity within a period of time (for example, the first I / O request), the reuse distance of the first I / O request is: among the multiple I / O requests initiated by the entity, the number of different logical addresses accessed by the I / O requests located between the first I / O request and another I / O request in time sequence. Among them, the other I / O request is the I / O request that accessed the logical address carried by the first I / O request previously in time sequence.
[0083] As an example, assume that the I / O requests initiated by an entity within a period of time include 10 I / O requests, and the logical addresses accessed by these 10 I / O requests in time sequence are (a, b, c, d, a, d, a, c, b, a). Among them, each letter represents a logical address.
[0084] Then, for the first I / O request among the above 10 I / O requests, the logical address carried by this I / O request is a. Since there is no I / O request that accessed the logical address a before the first I / O request among the above 10 I / O requests. Therefore, the reuse distance of the first I / O request is usually defaulted to infinity (symbol: ∞). Similarly, for the second I / O request, the third I / O request, and the fourth I / O request among the above 10 I / O requests, the reuse distances of the second I / O request, the third I / O request, and the fourth I / O request are all infinity.
[0085] For the fifth I / O request among the above 10 I / O requests, the logical address carried by this I / O request is a. Since there is an I / O request that accessed the logical address a before the fifth I / O request among the above 10 I / O requests, and the previous I / O request that accessed the logical address a is the first I / O request among the above 10 I / O requests, and the number of different logical addresses accessed by the I / O requests located between the fifth I / O request and the first I / O request in time sequence is 3 (including the logical address b accessed by the second I / O request, the logical address c accessed by the third I / O request, and the logical address d accessed by the fourth I / O request). Therefore, the reuse distance of the fifth I / O request is 3. Similarly, for the sixth I / O request among the above 10 I / O requests, the reuse distance of the sixth I / O request is 1. And, for the seventh I / O request among the above 10 I / O requests, the reuse distance of the seventh I / O request is 1.
[0086] For the 8th I / O request among the above 10 I / O requests, the logical address carried by this I / O request is c. Since there is an I / O request accessing the logical address c before the 8th I / O request among the above 10 I / O requests, and the previous I / O request accessing the logical address c is the 3rd I / O request among the above 10 I / O requests, and the number of different logical addresses accessed by the I / O requests temporally between the 8th I / O request and the 3rd I / O request is 2 (including the logical address d accessed by the 4th I / O request and the 6th I / O request, and the logical address a accessed by the 5th I / O request and the 7th I / O request). Therefore, the reuse distance of the 8th I / O request is 2. Similarly, for the 9th I / O request among the above 10 I / O requests, the reuse distance of the 9th I / O request is 3. And for the 10th I / O request among the above 10 I / O requests, the reuse distance of the 10th I / O request is 2.
[0087] 7), Reuse time
[0088] The reuse time is used to indicate the time interval between two adjacent accesses to the same logical address. Therefore, the reuse time can be called the reuse time of the logical address.
[0089] Generally, the reuse time of a logical address can be characterized by the number of I / O requests. Specifically, for any one I / O request (such as the first I / O request) among multiple I / O requests initiated by an entity within a period of time, the reuse time of the logical address carried by the first I / O request is: the number of I / O requests temporally between the first I / O request and another I / O request among the multiple I / O requests initiated by the entity. Among them, the other I / O request is the previous I / O request that accessed the logical address carried by the first I / O request in time sequence.
[0090] As an example, assume that the I / O requests initiated by an entity within a period of time include 10 I / O requests, and the logical addresses accessed by these 10 I / O requests in time sequence are (a, b, d, c, b, d, a, a, c, d). Among them, each letter represents a logical address.
[0091] Then for the 1st I / O request among the above 10 I / O requests, the logical address accessed by this I / O request is a, and there is no I / O request accessing the logical address a before the 1st I / O request among the above 10 I / O requests. Therefore, the reuse time of the logical address a accessed by the 1st I / O request is usually defaulted to infinity. Similarly, for the 2nd I / O request, the 3rd I / O request, and the 4th I / O request among the above 10 I / O requests, the reuse times of the logical address b accessed by the 2nd I / O request, the logical address d accessed by the 3rd I / O request, and the logical address c accessed by the 4th I / O request are all infinity.
[0092] For the 5th I / O request among the above 10 I / O requests, the logical address accessed by this I / O request is b. Since there is an I / O request accessing the logical address b among the above 10 I / O requests before the 5th I / O request, and the previous I / O request accessing the logical address b is the 2nd I / O request among the above 10 I / O requests, and there are 2 I / O requests (including the 3rd I / O request and the 4th I / O request) between the 5th I / O request and the 2nd I / O request in terms of timing. Therefore, the reuse time of the logical address b accessed by the 5th I / O request is 2. Similarly, for the 6th I / O request among the above 10 requests, the reuse time of the logical address d accessed by the 6th I / O request is 2. For the 7th I / O request among the above 10 requests, the reuse time of the logical address a accessed by the 7th I / O request is 5. For the 8th I / O request among the above 10 requests, the reuse time of the logical address a accessed by the 8th I / O request is 0. For the 9th I / O request among the above 10 requests, the reuse time of the logical address c accessed by the 9th I / O request is 4. For the 10th I / O request among the above 10 requests, the reuse time of the logical address d accessed by the 10th I / O request is 3.
[0093] 8), Other terms
[0094] In the embodiments of the present application, the terms "first" and "second" do not represent an order relationship, but are used to distinguish different objects. The first, second, etc. mentioned in the following documents are also used to distinguish different packets, etc., and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features.
[0095] It should also be understood that in each embodiment of the present application, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0096] To avoid cache contention, the shared cache can usually be partitioned so that the data requested by different types of I / O requests is cached in different partitions of the cache. Among them, different types of I / O requests are, for example, I / O requests initiated by different entities, and there is no limitation on this. However, currently, when partitioning the cache in the shared cache, the partition size is usually artificially specified based on the memory access pattern (such as the frequency of accessing data) when the I / O request accesses the memory within a period of time, or a simple heuristic strategy (such as the priority of different types of I / O requests) is used to determine the partition size of the cache. And the partitions of the cache divided in this way can generally only ensure the cache performance within a period of time, but cannot ensure the long-term cache performance. In addition, the same cache eviction algorithm is used for the entire shared cache, resulting in poor performance of the shared cache.
[0097] Based on this, an embodiment of the present application provides a method for managing a shared cache. The method first determines the access characteristics of the I / O requests of each category in multiple categories for accessing the shared cache, and determines the partition size and replacement algorithm of the I / O requests of each category in the shared cache according to the access characteristics of the I / O requests of each category in the multiple categories determined and the hit rate of the shared cache. Furthermore, the determined partition size and replacement algorithm of the I / O requests of each category in the shared cache are applied to the shared cache. Among them, the access characteristics of the I / O requests of each category are the relationship between the hit rate of the I / O requests of each category in multiple categories under multiple replacement algorithms and the cache size. Through this method, the cache size and replacement algorithm can be set for the I / O requests of each category according to the access characteristics of the I / O requests of each category, thereby improving the performance of the shared cache.
[0098] In addition, when the above method is executed periodically, it is possible to timely adjust the partition size and replacement algorithm of the I / O requests of each category in the shared cache according to the access characteristics of the I / O requests of each category in multiple categories determined periodically, so as to continuously ensure the cache performance of the shared cache in the time dimension.
[0099] An embodiment of the present application also provides a management device for a shared cache. The management device is applied to a computing device, and the computing device can manage the shared cache by executing the method provided by the embodiment of the present application. For a detailed description of the shared cache, reference can be made to the description in the above terms and will not be repeated here. As an example, the computing device can be any computing device such as a general computer, a laptop computer, a tablet computer, a mobile phone, a vehicle-mounted terminal, etc.
[0100] Optionally, the above computing device can be any computing device including a shared cache. Exemplarily, the computing device is a computing device with Figure 1 the shown CPU. Again exemplarily, the computing device can be a server or a cache node including a shared cache (such as Figure 2 the shown cache node). Also exemplarily, the computing device can be any node including a shared cache as shown in Figure 3 etc., and is not limited thereto. It can be understood that when the computing device can be any node including a shared cache as shown in Figure 3 the node can obtain the data (such as I / O requests) required in executing the method provided by the embodiment of the present application through interaction with other nodes including a shared cache in Figure 3 and execute the method described below in the embodiment of the present application based on the obtained data.
[0101] Optionally, the above computing device can also be a computing device connected and communicating with a node including a shared cache. As an example, when the node including a shared cache is Figure 3The node that provides caching for the cache pool is shown. The computing device can be an independent node independent of the node that provides caching for the cache pool, such as a management node, and this is not limited. In this case, the computing device can obtain the data (such as IO requests) required to execute the method provided in the embodiments of the present application through interaction with the node that provides caching for the cache pool, and execute the method described below in the embodiments of the present application based on the obtained data.
[0102] Reference Figure 4 , Figure 4 FIG. shows a schematic hardware structure diagram of a computing device provided by an embodiment of the present application. As Figure 4 shown, the computing device 40 includes a processor 401, a memory 402, a communication interface 403, and a bus 404. The processor 401, the memory 402, and the communication interface 403 are connected through the bus 404.
[0103] The processor 401 is the control center of the computing device 40. It can be a general-purpose CPU. The processor 401 may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, graphics processing units (GPUs), neural processing units (NPUs), tensor processing units (TPUs), artificial intelligent chips, data processing units (DPUs), etc.
[0104] As an example, the processor 401 includes one or more CPUs, such as Figure 4 the CPUs 0 and 1 shown in. In addition, the present application does not limit the number of processor cores in each processor.
[0105] The memory 402 is used to store program instructions or data to be accessed by application processes. The processor 401 can implement the shared cache management method provided by the embodiments of the present application by executing the program instructions in the memory 402.
[0106] The memory 402 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Among them, the volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static RAM (SRAM), DRAM, synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM). The non-volatile memory may be a storage class memory (SCM), a solid state drive (SSD), a hard disk drive (HDD), etc. Among them, the storage class memory may be, for example, a non-volatile memory (NVM), a phase-change memory (PCM), a persistent memory, etc.
[0107] In a possible implementation, the memory 402 exists independently of the processor 401. The memory 402 is connected to the processor 401 through a bus 404 and is used to store data, instructions, or program codes. When the processor 401 calls and executes the instructions or program codes stored in the memory 402, the management method of the shared cache provided by the embodiments of the present application can be implemented.
[0108] In another possible implementation, the memory 402 and the processor 401 are integrated together.
[0109] The communication interface 403 is used for the computing device 40 to communicate with other devices (such as Figure 2 or Figure 3The computing nodes shown in [figure] are connected through a communication network, which can be an Ethernet network, a radio access network (RAN), a wireless local area network (WLAN), etc. The communication interface 403 includes a receiving unit for receiving data / messages and a sending unit for sending data / messages.
[0110] The bus 404 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, a Peripheral Component Interconnect Express (PCIe) bus, a Compute Express Link (CXL) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity in representation, Figure 4 it is only represented by a thick line in [figure], but it does not mean that there is only one bus or one type of bus.
[0111] It should be noted that, Figure 4 the structure shown in [figure] does not constitute a limitation on the computing device 40. Except for Figure 4 the components shown, the computing device 40 includes more or fewer components than Figure 4 those shown, or combines some components, or has a different component layout.
[0112] Next, with reference to the accompanying drawings, the method for managing a shared cache provided in the embodiments of the present application will be described in detail.
[0113] In the embodiments of the present application, the shared cache is used to cache the data requested by the IO requests initiated by multiple entities (such as reading or writing). In the following text, it is described by taking the IO requests initiated by the multiple entities corresponding to K categories and the shared cache corresponding to N replacement algorithms as an example. That is, the IO requests initiated by the multiple entities include IO requests of K categories, and the shared cache is preconfigured with N replacement algorithms. Among them, K and N are respectively integers greater than 1.
[0114] It should also be understood that each IO request carries a logical address, which is the logical address to be accessed by the IO request. For example, the IO request reads the data to be read stored in the logical address it carries, or the IO request writes the data to be written to the logical address it carries. Therefore, in the following text of the embodiments of the present application, the logical address carried by the IO request and the logical address accessed by the IO request can be used interchangeably.
[0115] Embodiment 1
[0116] In the scenario where the read and write caches of the shared cache are separated, the management method of the shared cache provided by the embodiments of the present application can be used to manage the read cache in the shared cache or to manage the write cache in the shared cache. Among them, the separation of the read and write caches of the shared cache means that the read cache and the write cache in the shared cache are isolated physically or by software. The read cache is used to cache the data requested to be read by the IO requests initiated by the entity. The write cache is used to cache the data written by the IO requests initiated by the entity.
[0117] Reference Figure 5 , Figure 5 shows a schematic flowchart of a management method for a shared cache provided by an embodiment of the present application. Optionally, this method can be applied to Figure 1 the CPU shown in Figure 2 or applied to Figure 3 the cache node shown in Figure 4 or applied to
[0118] S101. Obtain the IO requests initiated by multiple entities and determine the category of each IO request. Among them, the IO requests initiated by multiple entities include IO requests of K categories.
[0119] Among them, since the process of the computing device executing the management method of the shared memory provided by the embodiments of the present application is parallel to the process of the IO requests initiated by the entity accessing the shared cache. Therefore, during the process of the IO requests initiated by multiple entities accessing the shared cache, the computing device obtains a copy of the IO requests initiated by these multiple entities for accessing the shared cache. Thus, the computing device obtains the IO requests initiated by multiple entities, and these IO requests are IO requests including K categories.
[0120] As an example, in combination with Figure 2 Reference Figure 6 , Figure 6 shows a schematic diagram of an IO request initiated by an entity accessing a shared cache. As Figure 6As shown in the figure, the cache node 2 includes a shared cache, and j computing nodes (including computing node 1, computing node 2, computing node 3, …, and computing node j) can all access the shared cache of the cache node 2. Moreover, the shared cache in the cache node 2 also communicates with the backend storage of the cache node 2 (the backend storage can be located inside the cache node 2 or outside the cache node 2). In this way, in the scenario of writing data, the shared cache can periodically flush the data written into the shared cache by the j computing nodes through IO requests into the backend storage. In the scenario of reading data, the shared cache can pre-load the data in the backend storage into its own storage space so that the j computing nodes can subsequently read this data through IO requests. In addition, when the j computing nodes access the shared cache through an IO request for reading data and a cache miss occurs in the shared cache, the shared cache can load the data to be read from the backend storage so that the j computing nodes can read the data to be read. In the embodiment of the present application, during the process in which the computing device accesses the shared cache through the IO requests of the j computing nodes, the computing device obtains a copy of the IO requests initiated by the j computing nodes for accessing the shared cache, so as to execute the method provided in the embodiment of the present application.
[0121] It should be noted that when the method provided in the embodiment of the present application is used to manage the read cache in the shared cache, the IO request obtained by the computing device in S101 is an IO read request for reading data. When the method provided in the embodiment of the present application is used to manage the write cache in the shared cache, the IO request obtained by the computing device in S101 is an IO write request for writing data.
[0122] Furthermore, the computing device determines the category of each obtained IO request.
[0123] In a possible implementation, the IO request initiated by the entity carries a category tag indicating the category to which the IO request belongs. In this way, the computing device determines the category of each IO request according to the category tag carried by each IO request in the obtained IO requests.
[0124] Optionally, the category tag carried in the IO request may be a category tag added to the IO request when the entity initiates the IO request. In this case, the category tag carried in the IO request can be used to indicate the entity that initiates the IO request. That is to say, the categories of the IO requests are divided according to the entities that initiate the IO requests. In this way, the IO requests of K categories correspond to K entities, and the IO requests initiated by the same entity have the same category tag, that is, the IO requests initiated by the same entity belong to the same category. Among them, the entity may be different applications, different processing cores, or different computing nodes, client devices, etc. that access the shared cache, and this is not limited thereto.
[0125] In another possible implementation, the computing device can determine the category of each obtained IO request through a classifier. The classifier is used to classify IO requests into K categories of IO requests, and specifically used to classify IO requests carrying logical addresses with relatively high similarity into one category. Among them, the classifier is pre-generated based on the characteristics of a certain number of IO requests accessing logical addresses. For a detailed description of how the computing device pre-generates the classifier based on the characteristics of a certain number of IO requests accessing logical addresses, refer to the description below, which will not be elaborated here.
[0126] Specifically, the computing device sequentially inputs the logical address carried by each obtained IO request into the classifier, and the classifier can output a category identifier indicating the category of each IO request.
[0127] Optionally, after the computing device determines the category of the IO request through the classifier, it adds a category identifier indicating its category to the IO request.
[0128] S102. Determine the access characteristics of IO requests in each of the K categories accessing the shared cache. The access characteristics are the relationship between the hit rate and the cache size of IO requests in each of the K categories under N replacement algorithms.
[0129] Among them, the N replacement algorithms are the replacement algorithms pre-configured for the shared cache, that is, the shared cache corresponds to N replacement algorithms.
[0130] Among them, the computing device presets X cache sizes for IO requests in each category. For an IO request in a certain category (such as the first category of IO requests), the X cache sizes are the X cache sizes preset for the cache corresponding to the first category of IO requests. The cache corresponding to the first category of IO requests is the partition in the shared cache used to cache the data requested to be read by the first category of IO requests, and the cache size of the cache corresponding to the first category of IO requests is the size of the partition in the shared cache used to cache the data requested to be read by the first category of IO requests. It can be understood that the cache sizes preset by the computing device for IO requests in each category are smaller than the size of the shared cache.
[0131] It should be understood that this application embodiment does not specifically limit the rule of the computing device presetting X cache sizes for IO requests in each category, nor the size interval between the X cache sizes.
[0132] As an example, the computing device presets 3 cache sizes for the first category of IO requests. Assuming the size of the shared cache is 4M, the 3 cache sizes preset by the computing device for the first category of IO requests can be 1M, 2M, and 3M.
[0133] Optionally, the number of cache sizes preset for the IO requests of each of the K categories by the computing device may be X, that is, the number of cache sizes preset for the IO requests of each of the K categories by the computing device is the same. Of course, the number of cache sizes preset for the IO requests of each of the K categories by the computing device may also be different. For example, the number of cache sizes preset for the first category of IO requests among the K categories by the computing device is 3, and the number of cache sizes preset for the second category of IO requests among the K categories by the computing device is 4, and so on. Among them, the second category of IO requests is any category of IO requests other than the first category among the K categories.
[0134] For simplicity of description, in the following embodiments of the present application, an example is given in which the number of cache sizes preset for the IO requests of each of the K categories by the computing device is X.
[0135] Optionally, the relationship between the hit rate and the cache size of the IO requests of each of the K categories characterized by the above access characteristics under N replacement algorithms can be characterized by the HRC of the shared cache for the IO requests of each of the K categories under N replacement algorithms. Alternatively, the relationship between the hit rate and the cache size of the IO requests of each of the K categories characterized by the above access characteristics under N replacement algorithms can be characterized by the MRC of the shared cache for the IO requests of each of the K categories under N replacement algorithms. The embodiments of the present application do not limit this.
[0136] Specifically, after the computing device determines the hit rate of the IO requests of each category in caches of different sizes under N replacement algorithms based on the obtained IO requests of each of the K categories, that is, obtains the relationship between the hit rate and the cache size of the IO requests of each of the K categories under N replacement algorithms, that is, obtains the access characteristics of the IO requests of each of the K categories accessing the shared cache.
[0137] Among them, the process by which the computing device determines the access characteristics of the IO requests of each of the K categories accessing the shared cache based on the obtained IO requests of each of the K categories can be implemented by the following possible implementation manners.
[0138] In the first possible implementation manner, the computing device simulates in the shared cache to obtain the hit rate of the IO requests of each category in caches of different sizes under N replacement algorithms, so as to obtain the relationship between the hit rate and the cache size of the IO requests of each category, for example, obtain the HRC or MRC of the IO requests of each category under N replacement algorithms.
[0139] Specifically, the computing device can scale and simulate the hit rates of the sample I / Os of each category under N eviction algorithms in a shared cache at different cache sizes, so as to obtain the relationship between the hit rates of the sample I / Os of each category and the cache size under the N eviction algorithms. For example, obtain the HRC or MRC of the sample I / Os of each category under the N eviction algorithms respectively. Furthermore, the computing device determines the relationship between the hit rates of the sample I / Os of each category and the cache size under the N eviction algorithms as the relationship between the hit rates of the I / O requests of each category and the cache size under the N eviction algorithms. For example, the computing device determines the HRC of the sample I / Os of each category under the N eviction algorithms as the HRC of the I / O requests of each category under the N eviction algorithms respectively. Or, the computing device determines the MRC of the sample I / Os of each category under the N eviction algorithms as the MRC of the I / O requests of each category under the N eviction algorithms respectively.
[0140] Among them, the sample I / O of each category is an I / O request sampled by the computing device based on the I / O requests of each category in the K categories obtained in S101. The specific sampling process is referred to the following description and will not be elaborated here.
[0141] Among them, for the detailed description of the process in which the computing device scales and simulates the hit rates of the sample I / Os of each category under N eviction algorithms in a shared cache at different cache sizes to obtain the relationship between the hit rates of the sample I / Os of each category and the cache size under the N eviction algorithms, reference can be made to the following description and will not be elaborated here.
[0142] In the second possible implementation manner, the computing device determines the hit rates of the I / O requests of each category in a shared cache at different cache sizes under N eviction algorithms according to the reuse distance of each I / O request in the I / O requests of each category and different cache sizes, so as to obtain the relationship between the hit rates of the I / O requests of each category and the cache size under the N eviction algorithms.
[0143] For example, for the first category of I / O requests among the I / O requests of K categories and the first eviction algorithm among the N eviction algorithms, the computing device can determine the hit rate of the first category of I / O requests when applying the first eviction algorithm in a shared cache at different cache sizes according to the reuse distance of each I / O request in the first category of I / O requests and different cache sizes, so as to obtain the relationship between the hit rate of the first category of I / O requests and the cache size under the first eviction algorithm. Among them, the first eviction algorithm is any one of the N eviction algorithms.
[0144] Optionally, to save the computing resources and efficiency of the computing device, the computing device can determine the hit rates of the sample I / Os of each category in caches of different sizes under N elimination algorithms according to the reuse distance of each sample I / O in the sample I / O of each category and different cache sizes, so as to obtain the relationship between the hit rate and the cache size of the sample I / O of each category under N elimination algorithms. Furthermore, the computing device determines the relationship between the hit rate and the cache size of the sample I / O of each category under N elimination algorithms as the relationship between the hit rate and the cache size of the I / O requests of each category under N elimination algorithms. Wherein, the sample I / O of each category is the I / O request sampled by the computing device based on the I / O requests of each category among the K categories obtained in S101.
[0145] For example, for the first category of sample I / O and the first elimination algorithm, the computing device can determine the hit rate of the first category of sample I / O when applying the first elimination algorithm in caches of different sizes according to the reuse distance of each sample I / O in the first category of sample I / O and different cache sizes, so as to obtain the relationship between the hit rate and the cache size of the first category of sample I / O under the first elimination algorithm. Then, the computing device approximately determines the relationship between the hit rate and the cache size of the first category of sample I / O under the first elimination algorithm as the relationship between the hit rate and the cache size of the first category of I / O requests under the first elimination algorithm. For the detailed description of the reuse distance of the sample I / O, reference can be made to the detailed description of the reuse distance in the above terms and will not be elaborated here.
[0146] For the detailed description of how the computing device determines the hit rate of the first category of sample I / O when applying the first elimination algorithm in caches of different sizes according to the reuse distance of each sample I / O in the first category of sample I / O and different cache sizes, so as to obtain the relationship between the hit rate and the cache size of the first category of sample I / O under the first elimination algorithm, reference can be made to the following description and will not be elaborated here.
[0147] It can be seen that for the I / O requests of one category, under one elimination algorithm, the computing device can determine a set of relationships between different cache sizes and hit rates, that is, one HRC or MRC. Furthermore, for the I / O requests of one category, under N elimination algorithms, the computing device can determine N sets of relationships between cache sizes and hit rates, that is, N HRCs or MRCs. Furthermore, for the I / O requests of K categories, under N elimination algorithms, the computing device can determine N×K sets of relationships between cache sizes and hit rates, that is, N×K HRCs or MRCs.
[0148] It should be noted that when the computing device determines the HRC or MRC of the sample I / O of each category under N elimination algorithms, it can be obtained by adopting the first possible implementation method described above, or by adopting the second possible implementation method described above. Of course, for the HRC or MRC of the sample I / O of some categories of I / O requests under some elimination algorithms, the computing device can obtain them by adopting the first possible implementation method described above, while for the HRC or MRC of the sample I / O of other categories of I / O requests under other elimination algorithms, the computing device can adopt the second possible implementation method described above to determine. The embodiments of the present application do not limit this.
[0149] Next, a simple description will be given of the process in which the computing device samples the I / O requests of each category among the K categories obtained in S101 to obtain the sample I / O of each category.
[0150] Specifically, the computing device samples the I / O requests of each category among the K categories obtained based on a preset sampling condition to obtain the sample I / O of each category. Among them, the logical address carried by the sample I / O satisfies the preset sampling condition.
[0151] Among them, the preset sampling condition is a sampling condition designed based on a preset sampling rate. The preset sampling rate is preset by the computing device. For example, the preset sampling rate is 0.01, that is, one sample I / O is sampled from 100 I / O requests. It should be noted that when the computing device samples the I / O requests of each category according to the preset sampling condition designed according to the preset sampling rate, the sampling rate of sampling the I / O requests of each category can be the preset sampling rate.
[0152] As an example, if the preset sampling rate is L, the sample I / O that satisfies the preset sampling condition satisfies: the remainder of the hash value of the logical address carried by the sample I / O divided by the coefficient A is less than or equal to the product of the preset sampling rate L and the coefficient A. Among them, the hash value of the logical address can be obtained by hashing the logical address based on any hash algorithm, and the embodiments of the present application do not limit this. The coefficient A is a preset coefficient, and the embodiments of the present application do not limit the value of A.
[0153] It should be noted that in an IO request of a certain category, the preset sampling condition can ensure that: when the computing device samples a sample IO that meets the preset sampling condition in the IO requests of this category, all IO requests in the IO requests of this category with the same logical address carried by the sample IO can be sampled as sample IOs by the computing device. In other words, the preset sampling condition can ensure that: when the computing device samples a sample IO that meets the preset sampling condition in the IO requests of this category, all IO requests in the IO requests of this category that access the same logical address as the sample IO can be sampled as sample IOs by the computing device. Therefore, any sampling condition that can ensure the above sampling purpose should be within the protection scope of the embodiments of the present application.
[0154] Specifically, for any category of IO requests (such as the first category of IO requests) among the K categories, the computing device determines whether each IO request in the first category of IO requests meets the preset sampling condition. If it meets, the computing device determines the IO request that meets the preset sampling condition as a sample IO. In this way, the computing device can sample multiple sample IOs at a preset sampling rate from the obtained first category of IO requests.
[0155] As an example, for any IO request (such as the first IO request) in the first category of IO requests, the computing device first determines the hash value of the logical address carried in the first IO request, and further determines whether the remainder Y obtained by dividing the hash value by the coefficient A is less than or equal to the product P of the preset sampling rate L and the coefficient A. When Y is less than or equal to P, the computing device samples the first IO request, that is, determines the first IO request as a sample IO of the first category of IO requests.
[0156] It can be understood that in practice, since the IO requests initiated by the entity access the shared cache in real time, during this process, after the computing device obtains and determines the category of each IO request, it immediately determines whether the IO request meets the preset sampling condition to determine whether to sample it.
[0157] S103. Determine the partition size and replacement algorithm of the IO requests of each category in the shared cache according to the access characteristics of the IO requests of each category among the K categories and the hit rate of the shared cache.
[0158] Specifically, based on the relationship between the hit ratios of the I / O requests of each of the K categories determined by the computing device in S102 and the cache size under each of the N eviction algorithms (i.e., the access characteristics of the I / O requests of each of the K categories), the computing device can determine the X hit ratios corresponding to the X cache sizes of the I / O requests of each category under each eviction algorithm. In this way, based on the X hit ratios corresponding to the X cache sizes of the I / O requests of each category under each eviction algorithm, the computing device can determine the hit ratios of the I / O requests of the K categories in the shared cache under different combinations. Here, for the detailed description of the X cache sizes, reference can be made to the above text and will not be elaborated here.
[0159] Among them, different combinations refer to the combinations of different cache sizes and different eviction algorithms corresponding to the I / O requests of each category. For example, if the computing device presets X cache sizes and the shared cache is preconfigured with N eviction algorithms, then the combinations corresponding to the I / O requests of one category include X×N combinations.
[0160] It should be understood that based on the relationship between the hit ratios of the I / O requests of each of the K categories determined by the computing device in S102 and the cache size under each of the N eviction algorithms (i.e., the access characteristics of the I / O requests of each of the K categories), it can be known that for the I / O requests of one category under one combination, there is a corresponding hit ratio of the I / O requests of this category under this combination. Therefore, based on the hit ratio of the I / O requests of each category under any one combination, the computing device can obtain a hit ratio of the I / O requests of the K categories in the shared cache.
[0161] Specifically, for the first category of I / O requests among the I / O requests of the K categories, the computing device can determine the hit ratio of the first category of I / O requests in the shared cache under the first combination based on the hit ratio of the first category of I / O requests under any one combination (such as the first combination) and the proportion of the first category of I / O requests among the I / O requests of the K categories. Exemplarily, the computing device performs a multiplication operation on the hit ratio of the first category of I / O requests under the first combination and the proportion of the first category of I / O requests among the I / O requests of the K categories, so as to obtain the hit ratio of the first category of I / O requests in the shared cache under the first combination. For example, assume that the hit ratio of the first category of I / O requests under the first combination is Z1, and the proportion of the first category of I / O requests among the I / O requests of the K categories is R1, then the hit ratio K1 of the first category of I / O requests in the shared cache under the first combination = Z1×R1.
[0162] Among them, the proportion of the first type of IO requests in the K types of IO requests is determined by the computing device according to the number of IO requests in the first type of IO requests and the total number of the K types of IO requests. Optionally, after the computing device obtains the K types of IO requests initiated by the entity and determines the category of each IO request, it can determine the proportion of each type of IO request in the K types of IO requests according to the total number of the obtained IO requests and the number of IO requests of each category.
[0163] Similarly, the computing device can determine the hit rate of each type of IO request in the shared cache under any combination. Furthermore, the computing device sums up the K hit rates of the K types of IO requests in the shared cache respectively, and then the hit rate of the K types of IO requests in the shared cache can be obtained. Among them, the K hit rates used for summation are the hit rates of the K types of IO requests in the shared cache under any one combination respectively. For example, taking K = 2 as an example, assuming that the hit rate of the first type of IO requests in the shared cache under the first combination is K1, and the hit rate of the second type of IO requests in the shared cache under the second combination is K2, then the hit rate of the K types of IO requests in the shared cache is: K1 + K2.
[0164] It should be noted that when the computing device sums up the K hit rates of the K types of IO requests in the shared cache respectively to obtain the hit rate of the K types of IO requests in the shared cache, the sum of the cache sizes of the K types of IO requests is less than or equal to the size of the shared cache.
[0165] Furthermore, for the K types of IO requests, the computing device can obtain the (X × N) K hit rates of the K types of IO requests in the shared cache according to the X × N hit rates of each type of IO request under the X × N combinations respectively.
[0166] As an example, taking the value of K as 2, the value of N as 2, and the value of X as 2, in this case, the IO requests obtained by the computing device include the first type of IO request and the second type of IO request, and the proportion of the first type of IO request in the two types of IO requests is R1, and the proportion of the second type of IO request in the two types of IO requests is R2. The eviction algorithms for shared cache provisioning include the first eviction algorithm and the second eviction algorithm. The cache sizes preset by the computing device for each type of IO request include cache size 1 and cache size 2. Moreover, there are 4 combinations corresponding to each type of IO request, and these 4 combinations include: combination 1 composed of the first eviction algorithm and cache size 1, combination 2 composed of the second eviction algorithm and cache size 1, combination 3 composed of the first eviction algorithm and cache size 2, and combination 4 composed of the second eviction algorithm and cache size 2. In this way, on the premise that the sum of the cache sizes of the first type of IO request and the second type of IO request is less than or equal to the shared cache, the computing device can calculate the (X×N) K (i.e., 4 2 = 16) hit rates of the two types of IO requests in the shared cache, which are respectively:
[0167] The hit rate 1 of the two types of IO requests in the shared cache is: the product of the hit rate of the first type of IO request under combination 1 and R1 (i.e., the hit rate of the first type of IO request in the shared cache under combination 1) + the product of the hit rate of the second type of IO request under combination 1 and R2 (i.e., the hit rate of the second type of IO request in the shared cache under combination 1);
[0168] The hit rate 2 of the two types of IO requests in the shared cache is: the product of the hit rate of the first type of IO request under combination 1 and R1 (i.e., the hit rate of the first type of IO request in the shared cache under combination 1) + the product of the hit rate of the second type of IO request under combination 2 and R2 (i.e., the hit rate of the second type of IO request in the shared cache under combination 2);
[0169] The hit rate 3 of the two types of IO requests in the shared cache is: the product of the hit rate of the first type of IO request under combination 1 and R1 (i.e., the hit rate of the first type of IO request in the shared cache under combination 1) + the product of the hit rate of the second type of IO request under combination 3 and R2 (i.e., the hit rate of the second type of IO request in the shared cache under combination 3);
[0170] The hit rate 4 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 1 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 1)+the product of the hit rate of the second - type IO requests under combination 4 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 4);
[0171] The hit rate 5 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 2 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 2)+the product of the hit rate of the second - type IO requests under combination 1 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 1);
[0172] The hit rate 6 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 2 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 2)+the product of the hit rate of the second - type IO requests under combination 2 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 2);
[0173] The hit rate 7 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 2 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 2)+the product of the hit rate of the second - type IO requests under combination 3 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 3);
[0174] The hit rate 8 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 2 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 2)+the product of the hit rate of the second - type IO requests under combination 4 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 4);
[0175] The hit rate 9 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 3 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 3)+the product of the hit rate of the second - type IO requests under combination 1 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 1);
[0176] The hit rate 10 of the IO requests of 2 categories in the shared cache is: the product of the hit rate of the first - type IO requests under combination 3 and R1 (i.e., the hit rate of the first - type IO requests in the shared cache under combination 3)+the product of the hit rate of the second - type IO requests under combination 2 and R2 (i.e., the hit rate of the second - type IO requests in the shared cache under combination 2);
[0177] The hit rate of the IO requests of 2 categories in the shared cache 11 is: the product of the hit rate of the first category of IO requests under combination 3 and R1 (i.e., the hit rate of the first category of IO requests in the shared cache under combination 3) + the product of the hit rate of the second category of IO requests under combination 3 and R2 (i.e., the hit rate of the second category of IO requests in the shared cache under combination 3);
[0178] The hit rate of the IO requests of 2 categories in the shared cache 12 is: the product of the hit rate of the first category of IO requests under combination 3 and R1 (i.e., the hit rate of the first category of IO requests in the shared cache under combination 3) + the product of the hit rate of the second category of IO requests under combination 4 and R2 (i.e., the hit rate of the second category of IO requests in the shared cache under combination 4);
[0179] The hit rate of the IO requests of 2 categories in the shared cache 13 is: the product of the hit rate of the first category of IO requests under combination 4 and R1 (i.e., the hit rate of the first category of IO requests in the shared cache under combination 4) + the product of the hit rate of the second category of IO requests under combination 1 and R2 (i.e., the hit rate of the second category of IO requests in the shared cache under combination 1);
[0180] The hit rate of the IO requests of 2 categories in the shared cache 14 is: the product of the hit rate of the first category of IO requests under combination 4 and R1 (i.e., the hit rate of the first category of IO requests in the shared cache under combination 4) + the product of the hit rate of the second category of IO requests under combination 2 and R2 (i.e., the hit rate of the second category of IO requests in the shared cache under combination 2);
[0181] The hit rate of the IO requests of 2 categories in the shared cache 15 is: the product of the hit rate of the first category of IO requests under combination 4 and R1 (i.e., the hit rate of the first category of IO requests in the shared cache under combination 4) + the product of the hit rate of the second category of IO requests under combination 3 and R2 (i.e., the hit rate of the second category of IO requests in the shared cache under combination 3);
[0182] The hit rate of the IO requests of 2 categories in the shared cache 16 is: the product of the hit rate of the first category of IO requests under combination 4 and R1 (i.e., the hit rate of the first category of IO requests in the shared cache under combination 4) + the product of the hit rate of the second category of IO requests under combination 4 and R2 (i.e., the hit rate of the second category of IO requests in the shared cache under combination 4).
[0183] Furthermore, the computing device determines the maximum hit rate of the IO requests of K categories in the shared cache, and determines the cache size indicated by the combination corresponding to each category when obtaining the maximum hit rate as the partition size of the IO requests of each category in the shared cache, and determines the replacement algorithm indicated by the combination corresponding to each category when obtaining the maximum hit rate as the replacement algorithm of the IO requests of each category in the shared cache.
[0184] For example, assume that the hit rate of the above two categories of IO requests in the shared cache is 15, which is the maximum hit rate of the above two categories of IO requests in the shared cache. Then, the computing device determines the cache size 2 indicated by combination 4 corresponding to the first category of IO requests when obtaining the maximum hit rate as the partition size of the first category of IO requests in the shared cache, and determines the second replacement algorithm indicated by combination 4 corresponding to the first category of IO requests when obtaining the maximum hit rate as the replacement algorithm of the first category of IO requests in the shared cache. Moreover, the computing device determines the cache size 2 indicated by combination 3 corresponding to the second category of IO requests when obtaining the maximum hit rate as the partition size of the second category of IO requests in the shared cache, and determines the first replacement algorithm indicated by combination 3 corresponding to the second category of IO requests when obtaining the maximum hit rate as the replacement algorithm of the second category of IO requests in the shared cache.
[0185] Optionally, in a possible implementation, the computing device can determine the maximum hit rate of K categories of IO requests in the shared cache under different combinations by solving the following formula (2) under the constraint of formula (1) according to the X hit rates corresponding to the X cache sizes determined for each category of IO requests under each replacement algorithm.
[0186]
[0187] Wherein, m represents a preset cache size, and the number of m is X preset by the computing device. M represents the size of the shared cache. Therefore, the value of m satisfies m ∈ [0, M]. K represents the number of categories of IO requests, or it can be understood that K represents the number of partitions in the shared cache. It should be understood that one category of IO requests corresponds to one partition of the shared cache, that is, one partition of the shared cache is used to cache the data of one category of IO requests. i represents the i-th partition among the K partitions, or it can be understood that i represents the i-th category of IO requests among the K categories of IO requests. m i represents the cache size corresponding to the i-th partition in the shared cache. p represents a replacement algorithm, and {P1, P2, …, P N} represents N replacement algorithms. p i represents the replacement algorithm configured for the i-th partition. p i takes any one of {P1, P2, …, P N}. Ri represents the proportion of the i-th category of IO requests in all IO requests.
[0188] Furthermore, formula (1) represents that the sum of the K cache sizes of K categories of IO requests in the shared cache is less than or equal to the size of the shared cache.
[0189] In formula (2), Indicates under the eviction algorithm p i when the hit rate of the i-th category of IO requests with a cache size of m i where m i and p i are a combination as described above. Indicates the hit rate of the i-th category of IO requests in the shared cache under the combination formed by m i and p i Furthermore, in formula (2) indicates the hit rate of the IO requests of K categories in the shared cache under different combinations. argmax in formula (2) means to take the maximum value, that is, argmax means to take the maximum hit rate of the IO requests of K categories in the shared cache under different combinations.
[0190] Thus, by solving formula (2) under the constraint of formula (1), the computing device can determine the maximum hit rate of the IO requests of K categories in the shared cache under different combinations, and determine the cache size indicated by the combination corresponding to the IO requests of each category when obtaining the maximum hit rate as the partition size of the IO requests of each category in the shared cache, and determine the eviction algorithm indicated by the combination corresponding to the IO requests of each category when obtaining the maximum hit rate as the eviction algorithm of the IO requests of each category in the shared cache.
[0191] In another possible implementation, the computing device can determine the maximum hit rate of the IO requests of K categories in the shared cache under different combinations by solving the following formula (3) under the constraint of formula (1) according to the X hit rates corresponding to the X cache sizes determined for each category of IO requests under each eviction algorithm.
[0192]
[0193] where the detailed descriptions of m, M, K, i, m i , p, {P1, P2,..., P N}, p i , and Ri can be referred to the descriptions in the previous possible implementation and will not be elaborated here.
[0194] In formula (3), indicates the miss rate of the i-th category of IO requests with a cache size of m i under the eviction algorithm p i where m i and p i are a combination as described above. Indicates i under m iThe miss rate of the i-th category of IO requests in the shared cache under the composed combination. Furthermore, in formula (3), can represent the miss rates of the IO requests of K categories in the shared cache under different combinations. argmin in formula (3) represents taking the minimum value, that is, argmin represents taking the minimum miss rate of the IO requests of K categories in the shared cache under different combinations.
[0195] Thus, by solving formula (3) under the constraint of formula (1), the computing device can determine the minimum miss rate of the IO requests of K categories in the shared cache under different combinations. Correspondingly, the computing device determines the maximum hit rate of the IO requests of K categories in the shared cache under different combinations. In this way, the computing device can determine the cache size indicated by the combination corresponding to the IO requests of each category when obtaining the minimum miss rate as the partition size of the IO requests of each category in the shared cache, and determine the replacement algorithm indicated by the combination corresponding to the IO requests of each category when obtaining the minimum miss rate as the replacement algorithm of the IO requests of each category in the shared cache.
[0196] S104. Configure the cache size of the IO requests of each category in the shared cache as the partition size of the IO requests of each category in the shared cache determined above, and configure the replacement algorithm of the IO requests of each category in the shared cache as the replacement algorithm of the IO requests of each category in the shared cache determined above.
[0197] Specifically, after the computing device determines the partition size and replacement algorithm of the IO requests of each category in the shared cache, it configures the cache size of the IO requests of each category in the shared cache as the partition size of the IO requests of each category in the shared cache determined above, and configures the replacement algorithm of the IO requests of each category in the shared cache as the replacement algorithm of the IO requests of each category in the shared cache determined above.
[0198] Furthermore, after configuring the cache size and replacement algorithm of the IO requests of each category in the shared cache, different categories of IO requests initiated by the entity can access the shared cache under this configuration.
[0199] Specifically, in the read data scenario, for any category of IO requests (such as the first category of IO requests) initiated by the entity, during the process of the first category of IO requests accessing the shared cache under this configuration, when the shared cache monitors that the data size cached by the first category of IO requests in the shared cache exceeds the replacement threshold, the data cached by the first category of IO requests in the shared cache is replaced according to the replacement algorithm configured for the first category of IO requests in the shared cache.
[0200] In the data writing scenario, for any category of IO requests (such as the first category of IO requests) initiated by an entity, during the process of the first category of IO requests accessing the shared cache under this configuration, one possible implementation is that when the shared cache monitors that the data size cached by the first category of IO requests in the shared cache exceeds the eviction threshold, the data cached by the first category of IO requests in the shared cache is evicted according to the eviction algorithm configured for the first category of IO requests in the shared cache. Another possible implementation is that the shared cache periodically evicts the data cached by the first category of IO requests in the shared cache according to the eviction algorithm.
[0201] In summary, through the management method of the shared cache described in S101 - S104, it is possible to determine the corresponding partition size and cache eviction algorithm for each category of IO requests in the shared cache according to the access characteristics of each category of IO requests, so as to improve the cache performance of each category of IO requests, and further improve the cache performance in the entire shared cache.
[0202] When the method described in S101 - S104 is used to manage the read cache in the shared cache, based on the partition size and eviction algorithm determined for each category of IO requests by the method provided in the embodiments of the present application, it is possible to improve the read hit rate of the IO read requests in the shared cache, thereby improving the data reading efficiency. When the method described in S101 - S104 is used to manage the write cache in the shared cache, based on the partition size and eviction algorithm determined for each category of IO requests by the method provided in the embodiments of the present application, it is possible to improve the write hit rate of the IO write requests in the shared cache, and further save the number of disk flushes from the shared cache to the backend storage unit, that is, save the data transmission bandwidth between the shared cache and the backend storage unit.
[0203] In some embodiments, the method described in S101 - S104 above can be executed periodically. In this case, the IO requests obtained by the computing device in S101 are the IO requests initiated by the entity obtained by the computing device within one period (such as the current period). Furthermore, the computing device executes S102 - S103 based on the obtained IO requests within one period, so as to determine the partition size and eviction algorithm of each category of IO requests in the shared cache. Then in S104, the computing device can configure the shared cache periodically based on the partition size and eviction algorithm of each category of IO requests determined for each period. In other words, the computing device can adjust the partition size of each partition in the shared cache and its corresponding eviction algorithm periodically based on the partition size and eviction algorithm of each category of IO requests determined for each period.
[0204] It should be noted that when the method described in S101 - S104 is executed for the first time, the shared cache can be configured according to the partition size and replacement algorithm specified by the user, or the shared cache is a pooled cache, and the embodiments of the present application do not limit this.
[0205] It can be seen that when the method described in S101 - S104 is executed periodically, according to the cache size and replacement algorithm of the IO requests of each category determined periodically by the computing device in the shared cache, the shared cache can periodically adjust the cache size and replacement algorithm of the IO requests of each category in the shared cache, thereby achieving the hit rate of the IO requests in the shared cache guaranteed in the time domain, that is, guaranteeing the cache performance of the shared cache in the time domain.
[0206] Next, taking the first - type sample IO sampled from the first - type IO requests and taking the first replacement algorithm among the N replacement algorithms as an example, the process of "the computing device calculates the hit rates of the sample IOs of each category in different - sized caches under the scaled - down simulation of N replacement algorithms in the shared cache, so as to obtain the relationship between the hit rates and cache sizes of the sample IOs of each category under the N replacement algorithms" in S102 will be described by way of example.
[0207] Among them, the descriptions of the first - type IO requests, the first - type sample IOs, and the first replacement algorithm can refer to the relevant descriptions in the above text and will not be elaborated here.
[0208] As an example, assume that the number of preset cache sizes for the first - type IO requests by the computing device is 3 (i.e., X takes the value of 3), and they are 1M, 2M, and 3M respectively, and the preset sampling rate L for sampling by the computing device in the first - type IO requests is 0.01. Then, when the computing device determines the hit rates of the first - type sample IOs applying the first replacement algorithm in different - sized caches, the corresponding cache sizes are 0.01M (i.e., 1M×0.01), 0.02M (2M×0.01), and 0.03M (3M×0.01). In this way, the computing device applies for caches with sizes of 0.01M, 0.02M, and 0.03M (hereinafter referred to as simulated caches) in the shared cache for the first - type sample IOs in the shared cache, respectively, to simulate the hit rates of the first - type sample IOs applying the first replacement algorithm in these three cache spaces, so as to obtain the hit rate of the first - type sample IOs applying the first replacement algorithm in the cache with a size of 0.01, the hit rate of the first - type sample IOs applying the first replacement algorithm in the cache with a size of 0.02M, and the hit rate of the first - type sample IOs applying the first replacement algorithm in the cache with a size of 0.03M, and further be able to obtain the relationship between the hit rate and cache size of the first - type sample IOs under the first replacement algorithm, that is, obtain the HRC of the first - type sample IOs under the first replacement algorithm.
[0209] Taking the example that the computing device applies for a simulated cache 1 with a cache size of 0.01M in the shared cache for the first - type sample IOs, the computing device can successively instruct each sample IO in the first - type sample IOs to access the simulated cache 1 and count the number of sample IOs that hit.
[0210] Specifically, when the simulation starts, the computing device can first instruct the first sample IO in the first - type sample IOs to access the simulated cache 1. Since the simulated cache 1 is empty at this time, the computing device caches the logical address of the first sample IO in the simulated cache 1. Then, the computing device instructs the second sample IO in the first - type sample IOs to access the simulated cache 1. If the logical address of the second sample IO is the same as that of the first sample IO, it means that the second sample IO hits in the simulated cache 1; if the logical address of the second sample IO is different from that of the first sample IO, it means that the second sample IO misses in the simulated cache 1. At this time, the computing device caches the logical address of the second IO sample in the simulated cache 1. And so on, the computing device instructs each sample IO in the first - type sample IOs to access the simulated cache 1 successively and counts the number of sample IOs that hit. It should be noted that the computing device also monitors the size of the logical addresses stored in the simulated cache 1 during the simulation process. When the size of the logical addresses stored in the simulated cache 1 exceeds a certain threshold, a part of the logical addresses in the simulated cache 1 is eliminated through the first elimination algorithm (for example, deleting a part or setting a part to invalid).
[0211] Furthermore, based on the total number of IOs of the first - type sample IOs and the number of sample IOs that hit, the computing device calculates the hit rate of the first - type sample IOs in the simulated cache 1. Similarly, the computing device simulates and counts the number of sample IOs that hit in the simulated cache 2 with a cache size of 0.02M for the first - type sample IOs, and based on the total number of IOs of the first - type sample IOs and the number of sample IOs that hit, calculates the hit rate of the first - type sample IOs in the simulated cache 2. Also, the computing device simulates and counts the number of sample IOs that hit in the simulated cache 3 with a cache size of 0.03M for the first - type sample IOs, and based on the total number of IOs of the first - type sample IOs and the number of sample IOs that hit, calculates the hit rate of the first - type sample IOs in the simulated cache 3.
[0212] In this way, if the computing device presets X cache sizes for the first - type IO requests, the computing device can determine the X hit rates corresponding to the X cache sizes when applying the first elimination algorithm for the first - type sample IOs. Thus, the computing device obtains the relationship between the hit rate and the cache size of the first - type sample IOs under the first elimination algorithm, that is, obtains the HRC of the first - type sample IOs under the first elimination algorithm.
[0213] Optionally, the computing device can also first determine the number of unhit sample I / Os based on the total number of I / Os of the first type of sample I / O and the number of hit sample I / Os, and then calculate the miss rate of the first type of sample I / O in the simulation cache 1 according to the total number of I / Os of the first type of sample I / O and the number of unhit sample I / Os. Similarly, the computing device simulates and counts the number of hit sample I / Os of the first type of sample I / O in the simulation cache 2, and first determines the number of unhit sample I / Os based on the total number of I / Os of the first type of sample I / O and the number of hit sample I / Os, and then calculates the miss rate of the first type of sample I / O in the simulation cache 2 according to the total number of I / Os of the first type of sample I / O and the number of unhit sample I / Os. And, the computing device simulates and counts the number of hit sample I / Os of the first type of sample I / O in the simulation cache 3, and first determines the number of unhit sample I / Os based on the total number of I / Os of the first type of sample I / O and the number of hit sample I / Os, and then calculates the miss rate of the first type of sample I / O in the simulation cache 3 according to the total number of I / Os of the first type of sample I / O and the number of unhit sample I / Os.
[0214] In this way, if the computing device presets X cache sizes for the first type of I / O requests, the computing device can determine the X miss rates corresponding to the X cache sizes when applying the first replacement algorithm for the first type of sample I / O. Thus, the computing device obtains the relationship between the miss rate and the cache size of the first type of sample I / O under the first replacement algorithm, that is, obtains the MRC of the first type of sample I / O under the first replacement algorithm.
[0215] It should be noted that since the simulation cache applied by the computing device for the sample I / O is empty at the beginning of each simulation, it is equivalent to the state after the cache is restarted after a power-off (i.e., cache cold start). Therefore, the computing device can appropriately increase the hit rate in the HRC curve after determining the HRC of the sample I / O of each category, or appropriately reduce the miss rate in the MRC curve after determining the MRC of the sample I / O of each category, so as to compensate for the cache misses caused by cache cold start. It should be understood that what the embodiments of the present application actually want to simulate is the hit rate of I / O requests when there is data cached in the cache, and the hit rate simulated when the cache is initially empty is usually lower than the hit rate when there is data cached in the cache. Therefore, the embodiments of the present application compensate for the hit rate in the simulated HRC or the miss rate in the MRC.
[0216] Among them, the specific compensation value for cold start compensation of the hit rate in the simulated HRC or the miss rate in the MRC in the embodiments of the present application can be determined by the sample IOs of each category used to simulate the HRC or MRC. For example, in the embodiments of the present application, when the first category of sample IOs start to be simulated in the applied simulated cache, the number of sample IOs that are not hit within a preset duration can be used as the specific compensation value for the hit rate compensation in the simulated HRC or the miss rate in the MRC. Here, the embodiments of the present application do not make specific limitations on the preset duration, nor on the specific method for determining the compensation value.
[0217] It should also be noted that when the IO requests initiated by the entity access the shared cache, a data prefetching mechanism can be adopted. That is, the entity will preload the data that may be accessed in a future period of time into the shared cache. Therefore, when the subsequent IO requests initiated by the entity access the prefetched data, it will surely hit. Therefore, the hit rate of the IO requests under the prefetching mechanism is higher than that of the IO requests without the prefetching mechanism. Therefore, after calculating the HRC of the sample IOs of each category, the computing device can appropriately reduce the hit rate in the HRC curve, or, after calculating the MRC of the sample IOs of each category, appropriately increase the miss rate in the MRC curve, so as to reduce the impact of the prefetching mechanism on the hit rate of the IO requests.
[0218] Among them, the specific value for reducing the hit rate in the simulated HRC in the embodiments of the present application (or the specific value for increasing the miss rate in the simulated MRC in the embodiments of the present application) can be the number of sample IOs that access the prefetched data among the sample IOs of each category used to simulate the HRC or MRC. It should be understood that when the IO requests initiated by the entity access the prefetched data, the entity will mark a prefetch identifier in the IO request to indicate that the data accessed by the IO request is prefetched. Therefore, the computing device only needs to count the number of sample IOs that access the prefetched data among the sample IOs of each category to determine the specific value for reducing the hit rate in the simulated HRC (or the specific value for increasing the miss rate in the simulated MRC in the embodiments of the present application).
[0219] Similarly, the computing device can simulate the relationship between the hit rate and the cache size of the first category of sample IOs under N replacement algorithms respectively, and, simulate the relationship between the hit rate and the cache size of the sample IOs of each category under N replacement algorithms respectively.
[0220] Further, the computing device determines the relationship between the hit rate and the cache size of each category of sample I / O simulated under N eviction algorithms as the relationship between the hit rate and the cache size of each category of I / O requests under N eviction algorithms. For example, the computing device determines the HRC of each category of sample I / O under N eviction algorithms as the HRC of each category of I / O requests under N eviction algorithms. Alternatively, the computing device determines the MRC of each category of sample I / O under N eviction algorithms as the MRC of each category of I / O requests under N eviction algorithms.
[0221] Next, the detailed process of "the computing device determines the hit rate of the first category of sample I / O when applying the first eviction algorithm in caches of different sizes according to the reuse distance of each sample I / O in the first category of sample I / O and different cache sizes, so as to obtain the relationship between the hit rate and the cache size of the first category of sample I / O under the first eviction algorithm" in S102 will be described.
[0222] Specifically, the computing device can determine the reuse distance of each sample I / O in the first category of sample I / O based on the logical address accessed by each sample I / O in the first category of sample I / O. Then, in the first category of sample I / O, the computing device counts the number of sample I / Os with the same reuse distance in the reuse distance of each sample I / O, that is, counts the frequency of each reuse distance.
[0223] As an example, assume that the first category of sample I / O includes 10 sample I / Os, and the logical addresses accessed by these 10 sample I / Os in sequence are (a, b, c, d, a, d, a, c, b, a), and the computing device counts the reuse distances of each of these 10 sample I / Os as (∞, ∞, ∞, ∞, 3, 1, 1, 2, 3, 2). Then, the computing device counts the number of sample I / Os with the same reuse distance based on the reuse distance of each of these 10 sample I / Os. Among them, the number (i.e., frequency) of reuse distances of ∞ is 4, the number (i.e., frequency) of reuse distances of 1 is 2, the number (i.e., frequency) of reuse distances of 2 is 2, and the number (i.e., frequency) of reuse records of 3 is 2.
[0224] Further, the computing device determines the hit rate of the first category of sample I / O at different cache sizes based on the preset rule and the counted number of sample I / Os with the same reuse distance in the first category of sample I / O, so as to obtain the relationship between the hit rate and the cache size of the first category of sample I / O under the first eviction algorithm. Among them, the preset rule is designed based on the eviction algorithm, one eviction algorithm corresponds to one preset rule, and N eviction algorithms correspond to N preset rules. The specific design of the preset rule corresponding to the eviction algorithm in the embodiments of the present application is not limited.
[0225] Taking the first elimination algorithm as the least recently used (LRU) algorithm as an example, for the first type of sample I / O, the preset rule corresponding to the LRU can be: determining the number of reuse distances less than a preset value as the hit count of the first type of sample I / O when the cache size is the size of the data accessed by a preset number of sample I / Os. Here, the computing device can pre-determine the value of the preset number corresponding to each cache size based on the value of each cache size among the preset X cache sizes and the size of the data accessed by a single sample I / O. For example, if one of the X cache sizes is 1M and the size of a single sample I / O is 4K, then the value of the preset number = 1M / 4K = 256. In this way, the computing device can determine the X hit counts corresponding to the X cache sizes when applying the first elimination algorithm to the first type of sample I / O. The description of the X cache sizes can refer to the above description and will not be elaborated here.
[0226] Optionally, in the embodiments of the present application, it can be defaulted that the size of the data accessed by each sample I / O is the same, for example, all are 4K, and this is not limited. In this case, the computing device pre-sets the size of the data requested to be accessed by a single sample I / O. Optionally, the computing device can also determine the average size of the data requested to be accessed by each sample I / O according to the number of sample I / Os included in the first type of sample I / O and the total size of the data requested to be accessed by the first type of sample I / O, and use this average size as the size of the data requested to be accessed by a single sample I / O. The embodiments of the present application are not limited thereto.
[0227] As an example, assume that the first type of sample I / O includes 10 sample I / Os, and based on these 10 sample I / Os, the number of reuse distances counted as ∞ is 4, the number of reuse distances of 1 is 2, the number of reuse distances of 2 is 2, and the number of reuse distances of 3 is 2.
[0228] Then, if the value of the preset number is 1, the computing device determines that there is no reuse distance less than 1. Therefore, the computing device determines that when the cache size is the size of the data accessed by 1 sample I / O, the hit count of the first type of sample I / O is 0.
[0229] If the value of the preset number is 2, the computing device can determine the number 2 of reuse distances less than 2 (i.e., the number 2 of reuse distances of 1) as the hit count of the first type of sample I / O when the cache size is the size of the data accessed by 2 sample I / Os.
[0230] If the value of the preset number is 3, the computing device may determine the number 4 of reuse distances less than 3 (i.e., the sum of the number 2 of reuse distance 1 and the number 2 of reuse distance 2 (i.e., 2 + 2)) as the hit count of the first type of sample I / O when the cache size is the size of the data requested to be accessed by 3 sample I / Os.
[0231] If the value of the preset number is 4, the computing device may determine the number 6 of reuse distances less than 4 (i.e., the sum of the number 2 of reuse distance 1, the number 2 of reuse distance 2, and the number 2 of reuse distance 3 (i.e., 2 + 2 + 2)) as the hit count of the first type of sample I / O when the cache size is the size of the data requested to be accessed by 4 sample I / Os.
[0232] Furthermore, based on the values of different preset numbers, the preset rules corresponding to the LRU algorithm, and the counted number of the same reuse distances in the first type of sample I / O, the computing device may determine the hit counts of the first type of sample I / O under the LRU algorithm for different cache sizes. Further, based on the determined hit counts and the number of the first type of sample I / O, the computing device may calculate the hit rates of the first type of sample I / O under the LRU algorithm for different cache sizes. In this way, for the X cache sizes preset by the computing device for the first type of I / O requests, the computing device determines the X hit rates corresponding to the X cache sizes when the first type of sample I / O applies the LRU algorithm. Thus, the computing device obtains the relationship between the hit rate of the first type of sample I / O under the LRU algorithm and the cache size, that is, obtains the HRC of the first type of sample I / O under the LRU algorithm.
[0233] Optionally, after determining the hit counts of the first type of sample I / O under the LRU algorithm for different cache sizes, the computing device may determine the miss counts of the first type of sample I / O under the LRU algorithm for different cache sizes based on the total number of the first type of sample I / O and the hit counts of the first type of sample I / O under the LRU algorithm for different cache sizes. Furthermore, the computing device calculates the miss rates of the first type of sample I / O under the LRU algorithm for different cache sizes according to the total number of the first type of sample I / O and the miss counts of the first type of sample I / O under the LRU algorithm for different cache sizes. In this way, for the X cache sizes preset by the computing device for the first type of I / O requests, the computing device determines the X hit rates corresponding to the X cache sizes when the first type of sample I / O applies the LRU algorithm. Thus, the computing device obtains the relationship between the miss rate of the first type of sample I / O under the LRU algorithm and the cache size, that is, obtains the MRC of the first type of sample I / O under the LRU algorithm.
[0234] Similarly, based on the values of different preset values, the N preset rules corresponding to the N elimination algorithms, and the number of the same reuse distances in the sample IOs of each category, the computing device can respectively determine the hit times (or miss times) of the sample IOs of each category under the N elimination algorithms at different cache sizes. Further, based on the determined hit times (or miss times) and the number of sample IOs of each category, the computing device can calculate the hit rate (or miss rate) of the sample IOs of each category under the N elimination algorithms at different cache sizes. In this way, the computing device obtains the relationship between the hit rate and the cache size (or the relationship between the miss rate and the cache size) of the sample IOs of each category under the N elimination algorithms respectively, that is, obtains the HRC (MRC) of the sample IOs of each category under the N elimination algorithms respectively.
[0235] Next, the process of "the computing device pre-generates a classifier based on the characteristics of accessing logical addresses by a certain number of IO requests" in S101 will be described.
[0236] Specifically, before S101, the method provided by the embodiments of the present application further includes Figure 7 a method for generating a classifier as shown. As Figure 7 shown, before S101, the method provided by the embodiments of the present application further includes S101a - S101d.
[0237] S101a: Obtain multiple IO requests initiated by multiple entities.
[0238] Among them, the description of the computing device obtaining the multiple IO requests can refer to the relevant description of obtaining the IO requests initiated by the entity in S101, and will not be elaborated here.
[0239] Optionally, when the method described in S101 - S104 is executed periodically, taking the period for executing the method described in S101 - S104 as the first period as an example, in a possible implementation manner, the multiple IO requests obtained by the computing device in S101a are the multiple IO requests obtained by the computing device in the first time period of the first period. In this case, the K categories of IO requests obtained by the computing device in S101 are the K categories of IO requests obtained by the computing device in the second time period of the first period. Wherein, the first time period is a time period starting from the start time of the first period and having a preset duration within the first period. The second time period is the remaining time period within the first period except the first time period. Here, the embodiments of the present application do not limit the specific values of the duration of the first period and the preset duration.
[0240] As an example, assume that the first period is 1 hour, the start time of the first period is 10:00, and the preset duration is 10 minutes. Then the first time period is the time period from 10:00 to 10:10, and the second time period is the time period from 10:11 to 11:00. In this case, the multiple IO requests obtained by the computing device in S101a are the multiple IO requests obtained by the computing device within the time period from 10:00 to 10:10, and the K types of IO requests obtained by the computing device in S101 are the K types of IO requests obtained by the computing device within the time period from 10:11 to 11:00.
[0241] In another possible implementation, the multiple IO requests obtained by the computing device in S101a are the first preset number of IO requests in terms of time sequence among all the IO requests obtained by the computing device within the first period. In this case, the K types of IO requests obtained by the computing device in S101 are the IO requests other than the preset number of IO requests among all the IO requests obtained by the computing device within the first period. Herein, the specific value of the preset number is not limited in the embodiments of the present application.
[0242] As an example, assume that the value of the preset number is 1000, and all the IO requests obtained by the computing device within the first period include 10,000 IO requests. Then the multiple IO requests obtained by the computing device in S101a are the first 1000 IO requests in terms of time sequence among the 10,000 IO requests, that is, the 1st IO request to the 1000th IO request in terms of time sequence among the 10,000 IO requests. In this case, the K types of IO requests obtained by the computing device in S101 are 9000 IO requests other than the 1st IO request to the 1000th IO request among the 10,000 IO requests, that is, the 1001st IO request to the 10,000th IO request in terms of time sequence among the 10,000 IO requests.
[0243] S101b. Extract the features of the logical addresses accessed by the above multiple IO requests.
[0244] Among them, the features of the logical addresses accessed by the multiple IO requests include the access frequency of the multiple IO requests accessing the same logical address, and / or, include the reuse time of the logical addresses accessed by the multiple IO requests.
[0245] Optionally, after obtaining the above multiple IO requests, for the logical address (such as the first logical address) accessed by any one of the multiple IO requests, the computing device counts the number of IO requests (such as the first number) accessing the first logical address among the multiple IO requests. It can be understood that the first number is the access frequency of the multiple IO requests accessing the first logical address. Furthermore, the computing device determines the determined first number as the frequency feature of each IO request accessing the first logical address among the multiple IO requests.
[0246] Similarly, the computing device can determine the frequency characteristics of each of the multiple IO requests described above.
[0247] Optionally, during the process of obtaining the multiple IO requests, the computing device can determine the reuse time of the logical address accessed by each IO request as soon as an IO request is obtained. Alternatively, the computing device can also determine the reuse time of the logical address accessed by each of the multiple IO requests after obtaining the multiple IO requests. For a detailed description of the reuse time, reference can be made to the description in the above terms and will not be elaborated here.
[0248] It can be understood that among the multiple IO requests described above, there may be at least two IO requests that access the same logical address (such as the first logical address), and for these at least two IO requests, the computing device determines different reuse times for the first logical address. That is to say, for one logical address, the computing device can determine multiple different reuse times. Exemplarily, for 10 IO requests that sequentially access logical addresses (a, b, d, c, b, d, a, a, c, d) in time sequence, the reuse time of the logical address d accessed by the 3rd IO request among these 10 requests is infinity, the reuse time of the logical address d accessed by the 6th IO request is 2, and the reuse time of the logical address d accessed by the 10th IO request is 3.
[0249] In this case, after determining the reuse time of the logical address accessed by each IO request, the computing device can select one reuse time from different reuse times of the same logical address and use the selected reuse time as the reuse time feature of each IO request that accesses the logical address.
[0250] Optionally, the computing device can arbitrarily select one reuse time from different reuse times of the same logical address and use the selected reuse time as the reuse time feature of each IO request that accesses the logical address. Among them, the reuse time arbitrarily selected by the computing device is a non-infinite reuse time. Alternatively, the computing device can calculate the average value (or round up / down after calculating the average value) of the multiple different reuse times determined for the same logical address and use the calculated average value (or the rounded value after calculating the average value) as the reuse time feature of each IO request that accesses the logical address. This application does not make any restrictions in this regard. Among them, the multiple different reuse times used for calculating the average value are the multiple different reuse times except for the reuse time with a value of infinity.
[0251] As an example, for 10 IO requests that sequentially access logical addresses (a, b, d, c, b, d, a, a, c, d) in time sequence, since the reuse time of the logical address d accessed by the 3rd IO request among these 10 requests is infinite, the reuse time of the logical address d accessed by the 6th IO request is 2, and the reuse time of the logical address d accessed by the 10th IO request is 3. Then, the computing device can use 2 or 3 as the reuse time feature of each IO request accessing the logical address d. Alternatively, the computing device takes the average of the reuse time of 2 for the logical address d accessed by the 6th IO request and the reuse time of 3 for the logical address d accessed by the 10th IO request and rounds up (i.e., rounds up [(2 + 3) / 2]), and uses the rounded-up value of 3 as the reuse time feature of each IO request accessing the logical address d.
[0252] S101c. According to the characteristics of the above-mentioned multiple IO requests accessing logical addresses, divide the multiple IO requests into K categories of IO requests.
[0253] Specifically, the computing device divides the multiple IO requests into K categories of IO requests according to the characteristics of the above-mentioned multiple IO requests accessing logical addresses and at least one feature threshold. Among them, the feature threshold includes a frequency threshold and / or a reuse time threshold.
[0254] In a possible case, when the characteristics of the above-mentioned multiple IO requests accessing logical addresses include the access frequency of the multiple IO requests accessing the same logical address, the computing device divides the multiple IO requests into K categories of IO requests according to the frequency feature of each IO request among the multiple IO requests and K - 1 frequency thresholds. It should be noted that the logical addresses accessed by the IO requests belonging to the same category have a high degree of similarity.
[0255] Optionally, the K - 1 frequency thresholds can be user - preset frequency thresholds. Or, the K - 1 frequency thresholds are that the computing device divides the multiple frequencies determined in S101b into K frequency ranges according to a first preset algorithm, and the K - 1 critical frequencies between the K frequency ranges are the K - 1 frequency thresholds. Among them, the first preset algorithm can be any classification algorithm, and the embodiments of the present application do not make limitations in this regard. Among them, the logical addresses accessed by the IO requests belonging to a frequency range have a high degree of similarity.
[0256] In another possible case, when the characteristics of the logical addresses accessed by the multiple IO requests include the reuse time of the logical addresses accessed by the multiple IO requests, the computing device divides the multiple IO requests into K categories of IO requests according to the reuse time characteristics of each IO request in the multiple IO requests and K-1 reuse time thresholds. It should be noted that the logical addresses accessed by the IO requests belonging to the same category have a high similarity.
[0257] Optionally, the K-1 reuse time thresholds may be user preset reuse time thresholds. Alternatively, the K-1 reuse time thresholds are obtained by dividing the multiple reuse times determined by S101b into K reuse time ranges according to a second preset algorithm, and the K-1 critical reuse times between the K reuse time ranges are the K-1 reuse time thresholds. The second preset algorithm may be any classification algorithm, which is not limited in comparison with the embodiments of the present application. The logical addresses accessed by IO requests whose reuse time characteristics belong to a reuse time range have a high degree of similarity.
[0258] In another possible scenario, when the characteristics of the multiple IO requests accessing the logical address include the access frequency of the multiple IO requests accessing the same logical address, and the reuse time of the logical address accessed by the multiple IO requests, the computing device divides the multiple IO requests into K categories of IO requests according to the frequency characteristics of each IO request in the multiple IO requests, p frequency thresholds, the reuse time characteristics of each IO request in the multiple IO requests, and q reuse time thresholds. Wherein, p and q are both positive integers, and p+q≤K-1. Wherein, the detailed description of the frequency threshold and the reuse time threshold can be referred to the above description and will not be repeated here.
[0259] Optionally, for multiple IO requests obtained by the computing device in S101a, the computing device may first divide the multiple IO requests into q+1 categories based on the reuse time characteristics of each IO request in the multiple IO requests and q reuse time thresholds. Furthermore, for at least one category of IO requests in the aforementioned q+1 categories, the computing device divides each category of IO requests in the at least one category into p+1 categories of IO requests based on the frequency characteristics of each IO request in the at least one category and p frequency thresholds, thereby achieving the division of the multiple IO requests into K categories of IO requests. It should be noted that the logical addresses accessed by IO requests belonging to the same category have a high degree of similarity.
[0260] As an example, take the values of p and q as 1 and the value of K as 3. Figure 8For multiple IO requests obtained by the computing device in S101a, the computing device may first divide the multiple IO requests into a first category of IO requests and a second category of IO requests according to the reuse time characteristics of each IO request in the multiple IO requests and a reuse time threshold (such as threshold 1). The first category of IO requests are IO requests with a reuse time characteristic less than threshold 1, and the second category of IO requests are IO requests with a reuse time characteristic greater than threshold 1. The embodiment of the present application does not limit the case where the reuse time characteristic is equal to threshold 1. For example, the IO request with a reuse time characteristic equal to threshold 1 may be an IO request of the first category, or an IO request of the second category.
[0261] Furthermore, for any one of the IO requests of the first category or the IO requests of the second category, such as the IO requests of the second category, the computing device divides the IO requests of the second category into IO requests of the third category and IO requests of the fourth category according to the frequency characteristics of each IO request in the IO requests of the second category and a frequency threshold (such as threshold 2). The IO requests of the third category are IO requests with a frequency characteristic less than threshold 2, and the IO requests of the fourth category are IO requests with a frequency characteristic greater than threshold 2. The embodiment of the present application does not limit the case where the frequency characteristic is equal to threshold 2. For example, the IO request with a frequency characteristic equal to threshold 2 can be an IO request of the third category, or an IO request of the fourth category.
[0262] In this way, the multiple IO requests acquired by the computing device in S101a are divided into IO requests of the first category, IO requests of the third category, and IO requests of the fourth category.
[0263] Optionally, for multiple IO requests obtained by the computing device in S101a, the computing device may first divide the multiple IO requests into p+1 categories based on the frequency characteristics of each IO request in the multiple IO requests and p frequency thresholds. Furthermore, for at least one category of IO requests in the aforementioned p+1 categories, the computing device divides the IO requests in each category in the at least one category into q+1 categories based on the reuse time characteristics of each IO request in the multiple IO requests and q reuse time thresholds, thereby achieving the division of the multiple IO requests into K categories of IO requests. It should be noted that the logical addresses accessed by IO requests belonging to the same category have a high degree of similarity.
[0264] S101d. Generate a classifier according to the multiple IO requests divided into K categories.
[0265] After dividing the multiple IO requests obtained above into K categories of IO requests, the computing device marks a category identifier for each IO request, where the category identifier is used to indicate the category to which the IO request belongs.
[0266] For example, the computing device divides the multiple obtained IO requests into two categories of IO requests (including the IO requests of category 1 and the IO requests of category 2). In this way, the computing device marks each IO request in category 1 with an identifier 1 indicating category 1, and marks each IO request in category 2 with an identifier 2 indicating category 2.
[0267] Furthermore, the computing device generates a classifier according to the sample parameters of each IO request in the multiple IO requests and a third preset algorithm. Among them, the sample parameters of an IO request include the category identifier of the IO request and the logical address carried by the IO request. Among them, the third preset algorithm can be any supervised learning classification algorithm, and the embodiments of the present application do not limit this.
[0268] Optionally, the third preset algorithm can be the k-Nearest-Neighbor (KNN) algorithm, or the KNN+prototype algorithm. Among them, the KNN+prototype algorithm is simpler than the KNN algorithm. Furthermore, the computing device can input the sample parameters of each IO request in the multiple IO requests into KNN (or KNN+prototype), so as to obtain a classifier. In this way, when the classifier receives the logical address of the IO request to be classified, it can first determine the sample parameters where the logical address with the greatest similarity to this logical address is located, and output the category identifier in the sample parameters to indicate the category of the IO request to be classified.
[0269] For example, in S101, the computing device can input the logical address 1 of the obtained IO request 1 into the classifier generated above. In this way, the classifier determines that the logical address with the greatest similarity to the logical address 1 is the logical address 2. Furthermore, the classifier outputs the category identifier 1 of the sample parameter 1 where the logical address 2 is located to indicate the category of the IO request 1.
[0270] In this way, based on the method described in S101a-S101d, the computing device can generate a classifier for classifying IO requests. When the methods described in S101-S104 are executed periodically, the classifier generated by the computing device based on a part of the IO requests in the current period can be used to classify another part of the IO requests in the current period (for example, divided into K categories), so as to determine the partition size and replacement algorithm of each category of IO requests in the shared cache according to the access characteristics of the K categories of IO requests accessing the shared cache after classification.
[0271] In addition, the classifier generated by the computing device during the current cycle can be used to classify the IO requests initiated by an entity in the next cycle of the current cycle, so that the classified IO requests access the corresponding cache. Therefore, based on the methods described in S101a - S101d, the entity initiating the IO request does not need to mark the category tag for classification for the IO request, so that no additional resource overhead is generated for the entity initiating the IO request. Moreover, since the method provided in the embodiments of the present application does not intrude into the upper layer of the cache (i.e., the entity initiating the IO request), the method provided in the embodiments of the present application can be applied to a general cache system for diverse customers without ecological support.
[0272] Embodiment 2
[0273] The management method of the shared cache provided in the embodiments of the present application can also be applied to the scenario of read - write cache fusion in the shared cache. Among them, the read - write cache fusion in the shared cache means that in the shared cache, the read cache and the write cache are not distinguished, and the IO requests for reading data and the IO requests for writing data share the shared cache.
[0274] In the scenario of read - write cache fusion in the shared cache, referring to Figure 9 , Figure 9 shows a schematic flowchart of another management method of the shared cache provided in the embodiments of the present application. Optionally, this method can be applied to the CPU shown in Figure 1 , or applied to the cache node shown in Figure 2 , or applied to the node shown in Figure 3 . This method can be executed by a computing device with the hardware structure shown in Figure 4 , and this method includes the following steps.
[0275] S201. Obtain the IO requests initiated by multiple entities and determine the category of each IO request. Among them, the IO requests initiated by multiple entities include IO requests of K categories, and each category of IO requests in the K categories includes an IO read request and an IO write request.
[0276] For the detailed description of the computing device obtaining the IO requests initiated by multiple entities and determining the category of each IO request, reference can be made to the description of S101, which will not be elaborated here.
[0277] It should be noted that in S101, the IO request obtained by the computing device is either an IO read request for reading data or an IO write request for writing data. While in S201, the IO requests obtained by the computing device include both an IO read request for reading data and an IO write request for writing data.
[0278] S202. Determine the read access characteristics of the IO read requests of each of the K categories accessing the shared cache, and determine the write access characteristics of the IO write requests of each of the K categories accessing the shared cache.
[0279] Among them, the read access characteristic is the relationship between the read hit rate and the cache size of the IO read requests of each of the K categories under N replacement algorithms, and the write access characteristic is the relationship between the write hit rate and the cache size of the IO write requests of each of the K categories under N replacement algorithms. Here, the read hit rate refers to the hit rate of the IO read requests when reading in the cache. The write hit rate refers to the hit rate of the IO write requests when writing in the cache.
[0280] Specifically, the computing device determines the read access characteristics of the IO read requests of each of the K categories accessing the shared cache according to the obtained IO read requests of each of the K categories. And, the computing device determines the write access characteristics of the IO write requests of each of the K categories accessing the shared cache according to the obtained IO write requests of each of the K categories.
[0281] Among them, the detailed description of how the computing device determines the read access characteristics of the IO read requests of each of the K categories according to the IO read requests of each of the K categories, and the detailed description of how to determine the write access characteristics of the IO write requests of each of the K categories according to the IO write requests of each of the K categories, can all refer to the relevant description of "the computing device determines the access characteristics of the IO requests of each of the K categories accessing the shared cache based on the obtained IO requests of each of the K categories" in S102, and will not be elaborated here.
[0282] S203. Determine the partition size and replacement algorithm of the IO requests of each category in the shared cache according to the read access characteristics of the IO read requests of each of the K categories, the write access characteristics of the IO write requests of each of the K categories, and the hit rate of the shared cache.
[0283] Specifically, the detailed description of how the computing device determines the partition size and replacement algorithm of the IO requests of each category in the shared cache according to the read access characteristics of the IO read requests of each of the K categories, the write access characteristics of the IO write requests of each of the K categories, and the hit rate of the shared cache, can refer to the description of S103.
[0284] It should be noted that for "the computing device can obtain the hit rate of the IO requests of K categories in the shared cache according to the hit rate of the IO requests of each category under any combination in S103", in S203, the computing device can obtain the comprehensive hit rate of the IO requests of K categories in the shared cache according to the read hit rate of the IO read requests of each category under any combination, the read hit weight coefficient, the write hit rate of the IO write requests of each category under any combination, and the write hit weight coefficient.
[0285] Among them, the read hit weight coefficient is used to characterize the influence degree of the read hit rate of the IO read requests of each category in the shared cache on the comprehensive hit rate of the shared cache, and the write hit weight coefficient is used to characterize the influence degree of the write hit rate of the IO write requests of each category in the shared cache on the comprehensive hit rate of the shared cache. Among them, the comprehensive hit rate of the shared cache is the hit rate of the IO requests of K categories in the shared cache. In this embodiment of the application, the specific values of the read hit weight coefficient and the write hit weight coefficient are not limited. For example, both the read hit weight coefficient and the write hit weight coefficient are preset weight coefficients. For another example, the read hit weight coefficient and / or the write hit weight coefficient can be determined based on the access rules of the entity to the shared cache in a recent period of time, or the read hit weight coefficient and / or the write hit weight coefficient can be determined based on the predicted access rules of the entity to the shared cache in a future period of time.
[0286] Specifically, for the IO requests of any one category among the K categories (such as the first category of IO requests), the computing device can determine the hit rate of the first category of IO requests in the shared cache under the first combination according to the hit rate of the IO read requests (the first category of IO read requests) in the first category of IO requests under any combination (such as the first combination) and the proportion of the first category of IO read requests in the IO requests of K categories. For example, the computing device performs a product operation on the hit rate of the first category of IO read requests under the first combination and the proportion of the first category of IO read requests in the IO requests of K categories, so as to obtain the hit rate of the first category of IO read requests in the shared cache under the first combination.
[0287] Similarly, for the IO write requests (hereinafter referred to as the first category of IO write requests) in the first category of IO requests, the computing device can determine the hit rate of the first category of IO write requests in the shared cache under the first combination according to the hit rate of the first category of IO write requests under any combination (such as the first combination) and the proportion of the first category of IO write requests in the IO requests of K categories. For example, the computing device performs a product operation on the hit rate of the first category of IO write requests under the first combination and the proportion of the first category of IO write requests in the IO requests of K categories, so as to obtain the hit rate of the first category of IO write requests in the shared cache under the first combination.
[0288] Among them, for the detailed description of the proportion of the first type of IO read requests among the K types of IO requests and the proportion of the first type of IO write requests among the K types of IO requests, reference can be made to the relevant description in S103, which will not be elaborated here.
[0289] Furthermore, the computing device determines the hit rate of the first type of IO requests in the shared cache under the first combination (such as the first hit rate), the read hit weight coefficient, the hit rate of the first type of IO write requests in the shared cache under the first combination (such as the second hit rate), and the write hit weight coefficient, and then determines the hit rate of the first type of IO requests in the shared cache under the first combination. For example, the computing device performs a summation operation on the product of the first hit rate and the read hit weight coefficient and the product of the second hit rate and the write hit weight coefficient, so as to obtain the hit rate of the first type of IO requests in the shared cache under the first combination. As an example, assuming that the read hit weight coefficient is W1 and the write hit weight coefficient is W2, the hit rate of the first type of IO requests in the shared cache under the first combination is: the first hit rate × W1 + the second hit rate × W2.
[0290] Similarly, the computing device can determine the hit rate of each type of IO requests in the shared cache under any combination. Furthermore, the computing device sums up the K hit rates of the K types of IO requests in the shared cache respectively, and then can obtain the hit rate of the K types of IO requests in the shared cache. Further, for the K types of IO requests, based on the X×N hit rates of each type of IO requests under the X×N combinations, the computing device can obtain the (X×N) K hit rates of the K types of IO requests in the shared cache. Further, the computing device determines the maximum hit rate of the K types of IO requests in the shared cache, and determines the cache size indicated by the combination corresponding to each type when the maximum hit rate is obtained as the partition size of each type of IO requests in the shared cache, and determines the replacement algorithm indicated by the combination corresponding to each type when the maximum hit rate is obtained as the replacement algorithm of each type of IO requests in the shared cache. For the specific description, reference can be made to the description in S103, which will not be elaborated here.
[0291] It should be noted that when the computing device sums up the K hit rates of the K types of IO requests in the shared cache respectively to obtain the hit rate of the K types of IO requests in the shared cache, the sum of the cache sizes of the K types of IO requests is less than or equal to the size of the shared cache.
[0292] Then, the computing device executes S104.
[0293] In this way, through the management method of the shared cache described in S201-S104, in the scenario of the fusion of read and write caches in the shared cache, it is possible to determine and configure an appropriate cache size and replacement algorithm for each category of IO requests, thereby improving the hit rate of IO requests in the shared cache and further enhancing the overall cache performance of the shared cache.
[0294] In some embodiments, the method described in S201-S104 above can be executed periodically. For relevant descriptions, reference can be made to the relevant descriptions in Embodiment 1, which will not be elaborated here. By periodically executing S201-S104, the shared cache can regularly adjust the cache size and replacement algorithm of each category of IO requests in the shared cache according to the cache size and replacement algorithm of each category of IO requests determined periodically by the computing device, thereby achieving the guarantee of the hit rate of IO requests in the shared cache in the time domain, that is, ensuring the cache performance of the shared cache in the time domain.
[0295] Embodiment 3
[0296] In the scenario of the fusion of read and write caches in the shared cache, refer to Figure 10 , Figure 10 shows a flowchart of another management method of the shared cache provided by the embodiment of the present application. Optionally, this method can be applied to the Figure 1 shown CPU, or applied to the Figure 2 shown cache node, or applied to the Figure 3 shown node. And it can be executed by a computing device having the Figure 4 shown hardware structure. The computing device first executes S201, and then, the computing device executes S302.
[0297] S302. Determine the access characteristics of each category of IO requests accessing the shared cache according to each category of IO read requests, read hit weight coefficients, each category of IO write requests, and write hit weight coefficients among K categories.
[0298] Among them, for the detailed descriptions of the read hit weight coefficient and the write hit weight coefficient, reference can be made to the above text, which will not be elaborated here.
[0299] Specifically, the computing device first determines the read hit rates of the IO read requests of each of the K categories under N replacement algorithms in caches of different sizes according to the obtained IO read requests of each of the K categories. And the computing device determines the write hit rates of the IO write requests of each of the K categories under N replacement algorithms in caches of different sizes according to the obtained IO write requests of each of the K categories. For example, taking the first type of IO requests among the IO requests of the K categories as an example, the computing device can determine the read hit rates of the first type of IO read requests under N replacement algorithms in caches of different sizes according to the obtained first type of IO read requests. And the computing device determines the write hit rates of the first type of IO write requests under N replacement algorithms in caches of different sizes according to the obtained first type of IO write requests.
[0300] Among them, the descriptions of IO read requests, read hit rates, IO write requests, write hit rates, the first type of IO read requests, and the first type of IO write requests can all refer to the relevant descriptions above and will not be elaborated here.
[0301] Among them, the detailed description of how the computing device determines the read hit rates of the IO read requests of each of the K categories under N replacement algorithms in caches of different sizes according to the obtained IO read requests of each of the K categories, and the detailed description of how the computing device determines the write hit rates of the IO write requests of each of the K categories under N replacement algorithms in caches of different sizes according to the obtained IO write requests of each of the K categories can both refer to the relevant description of determining the hit rates of the IO requests of each category under N replacement algorithms in caches of different sizes in S102 and will not be elaborated here.
[0302] Furthermore, for the first type of IO requests among the IO requests of the K categories, under the same replacement algorithm and the same cache size, the computing device performs an addition operation on the product of the read hit rate and the read hit weight coefficient of the first type of IO read requests, and the product of the write hit rate and the write hit weight coefficient of the first type of IO write requests, so as to obtain the hit rate of the first type of IO requests under this replacement algorithm and cache size.
[0303] Similarly, the computing device can determine the hit rates of the IO requests of each category under N replacement algorithms in caches of different sizes. In this way, the relationship between the hit rates of the IO requests of each of the K categories and the cache size under N replacement algorithms is obtained, that is, the access characteristics of the IO requests of each of the K categories accessing the shared cache are obtained.
[0304] Then, the computing device executes S103 - S104.
[0305] In this way, through Figure 10The described method for managing a shared cache can determine and configure an appropriate cache size and replacement algorithm for each category of I / O requests in the scenario of the integration of read and write caches in the shared cache, thereby improving the hit rate of I / O requests in the shared cache and further enhancing the overall cache performance of the shared cache.
[0306] In some embodiments, the above Figure 10 described method can be executed periodically. For relevant descriptions, reference can be made to the relevant descriptions in Embodiment 1 and will not be elaborated here. By executing Figure 10 the described method periodically, the shared cache can regularly adjust the cache size and replacement algorithm of each category of I / O requests in the shared cache according to the cache size and replacement algorithm of each category of I / O requests determined periodically by the computing device, thereby achieving the guarantee of the hit rate of I / O requests in the shared cache in the time domain, that is, guaranteeing the cache performance of the shared cache in the time domain.
[0307] The above mainly introduces the solution provided in the embodiments of the present application from the perspective of the method.
[0308] To implement the above functions, as Figure 11 shown, Figure 11 FIG. shows a schematic structural diagram of a management device 110 for a shared cache provided in an embodiment of the present application. The management device 110 is used to execute the above-described method for managing a shared cache, for example, for executing Figure 5 , Figure 7 , Figure 9 or Figure 10 the method shown. Among them, the management device 110 may include a determination unit 111 and a configuration unit 112.
[0309] The determination unit 111 is used to determine the access characteristics of each category of I / O requests accessing the shared cache among K categories, and is used to determine the partition size and replacement algorithm of each category of I / O requests in the shared cache according to the access characteristics of the I / O requests of K categories and the hit rate of the shared cache. Among them, the access characteristics are the relationship between the hit rate and the cache size of each category of I / O requests among K categories under N replacement algorithms. The configuration unit 112 is used to configure the cache size of each category of I / O requests in the shared cache as the determined partition size of each category of I / O requests in the shared cache, and to configure the replacement algorithm of each category of I / O requests in the shared cache as the determined replacement algorithm of each category of I / O requests in the shared cache.
[0310] As an example, in combination with Figure 5 , the determination unit 111 can be used to execute S102 and S103, and the configuration unit 112 can be used to execute S104. In combination with Figure 9, the determination unit 111 can be used to execute S202 and S203, and the configuration unit 112 can be used to execute S104. Combined with Figure 10 , the determination unit 111 can be used to execute S302 and S103, and the configuration unit 112 can be used to execute S104.
[0311] Optionally, the management device 110 further includes: an emulation unit 113, configured to, for a first eviction algorithm among N eviction algorithms, emulate the hit rate of the IO requests of each of the K categories in different-sized caches when applying the first eviction algorithm in a shared cache, so as to obtain the relationship between the hit rate and the cache size. The first eviction algorithm is any one of the N eviction algorithms.
[0312] Optionally, the determination unit 111 is specifically configured to: for a first eviction algorithm among N eviction algorithms, determine the hit rate of the first type of IO requests in different-sized caches when applying the first eviction algorithm according to the reuse distance of each IO request in the first type of IO requests and different cache sizes, so as to obtain the relationship between the hit rate and the cache size. The first eviction algorithm is any one of the N eviction algorithms, and the first type of IO requests is the IO requests of any one of the K categories.
[0313] Optionally, the determination unit 111 is further specifically configured to: determine the hit rate of the IO requests of the K categories in the shared cache for each combination according to the X hit rates corresponding to the X cache sizes determined for the IO requests of each category under each eviction algorithm, and determine the cache size corresponding to the IO requests of each category when the hit rate of the IO requests of the K categories in the shared cache is the maximum as the partition size of the IO requests of each category in the shared cache, and determine the eviction algorithm corresponding to the IO requests of each category when the hit rate of the IO requests of the K categories in the shared cache is the maximum as the eviction algorithm of the IO requests of each category in the shared cache. For any one of the K categories of IO requests, the X cache sizes and the N eviction algorithms form X*N combinations, and each combination includes a cache size and an eviction algorithm; the X cache sizes are X cache sizes preset for the cache corresponding to the IO requests of each category.
[0314] As an example, combined with Figure 5 , the determination unit 111 can be used to execute S103. Combined with Figure 9 , the determination unit 111 can be used to execute S203.
[0315] Optionally, the management device 110 further includes: an obtaining unit 114, configured to obtain a plurality of I / O requests before determining the access characteristics of the I / O requests of each of the K categories accessing the shared cache; and a classification unit 115, configured to divide the I / O requests into K categories according to the characteristics of the addresses of the data accessed by the plurality of I / O requests or according to the category tags carried in the plurality of I / O requests.
[0316] As an example, in combination with Figure 5 , the obtaining unit 114 and the classification unit 115 may be used to execute S101.
[0317] Optionally, if the above-mentioned shared cache is the LLC of the CPU in the computing device, then the above-mentioned plurality of I / O requests are I / O requests initiated by multiple processing cores in the CPU.
[0318] Optionally, if the above-mentioned shared cache is the cache in the cache node, then the above-mentioned plurality of I / O requests are I / O requests initiated by multiple computing nodes accessing the cache node.
[0319] Optionally, if the above-mentioned shared cache is a cache pool composed of caches in multiple nodes, then the above-mentioned plurality of I / O requests are I / O requests initiated by multiple computing nodes accessing the cache pool.
[0320] Optionally, the above-mentioned access characteristics are characterized by the HRC or MRC of the I / O request.
[0321] Optionally, the determining unit 111 is further configured to periodically determine the access characteristics of the I / O requests of each of the K categories accessing the shared cache. For the access characteristics of the I / O requests of each of the K categories accessing the shared cache determined in the first period, the determining unit 111 is specifically configured to determine the partition size and the eviction algorithm of the I / O requests of each category in the shared cache in the first period according to the access characteristics of the I / O requests of the K categories determined in the first period and the hit rate of the shared cache, where the first period is any period for determining the access characteristics of the I / O requests of each of the K categories accessing the shared cache.
[0322] For the specific descriptions of the above optional manners, reference may be made to the foregoing method embodiments, which will not be elaborated herein. In addition, for any explanation of the management device 110 provided above and the description of the beneficial effects, reference may be made to the corresponding method embodiments above, which will not be elaborated.
[0323] As an example, in combination with Figure 4 , the functions implemented by the determining unit 111, the configuration unit 112, the simulation unit 113, and the classification unit 115 in the management device 110 may be executed by the processor 401 in Figure 4 and the program code in the memory 402 in Figure 4 . The functions implemented by the obtaining unit 114 may be throughFigure 4 is implemented by the communication interface 403 in
[0324] Those skilled in the art should easily realize that, for the units and algorithm steps of each example described in combination with the embodiments disclosed herein, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described function for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0325] It should be noted that Figure 11 the division of modules in is illustrative, merely a logical function division, and there can be other division methods in actual implementation. For example, two or more functions can also be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software function modules.
[0326] This application embodiment also provides a computer program product and a computer-readable storage medium for storing the computer program product. The computer program product can include one or more program instructions, and when the one or more program instructions are run by one or more processors, it can provide the functions or partial functions described above for Figure 5 , Figure 7 , Figure 9 or Figure 10 described. Therefore, for example, one or more features of S101 - S104 in Figure 5 can be borne by one or more instructions in the computer program product.
[0327] In some examples, such as for the management device of the shared cache for executing Figure 5 , Figure 7 , Figure 9 or Figure 10 the described method can be configured to provide various operations, functions, or actions in response to one or more program instructions stored in a computer-readable storage medium.
[0328] This application embodiment also provides a computing device, which can be used to execute the method described above in Figure 5 , Figure 7 , Figure 9 or Figure 10 to realize the management of the shared cache.
[0329] Optionally, the computing device can be the same device as the device including the shared cache, or the computing device is a device connected and communicating with the device including the shared cache. For detailed description, reference can be made toFigure 4 The related description of the computing device will not be elaborated here.
[0330] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are executed on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more integrated media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0331] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for managing a shared cache, characterized in that, The shared cache is used to cache data requested by multiple input / output (IO) requests. The multiple IO requests include K categories of IO requests, and the shared cache includes N replacement algorithms. The method includes: Determine the access characteristics of the IO requests of each category among the K categories accessing the shared cache. The access characteristics are the relationships between the hit ratios of the IO requests of each category under the N replacement algorithms and the cache size. Determine the partition size of the IO requests of each category in the shared cache and the replacement algorithm of the IO requests of each category in the shared cache according to the access characteristics of the IO requests of the K categories and the hit ratio of the shared cache. The cache size of the IO requests of each category in the shared cache is the determined partition size of the IO requests of each category in the shared cache, and the replacement algorithm of the IO requests of each category in the shared cache is the determined replacement algorithm of the IO requests of each category in the shared cache.
2. The method according to claim 1, wherein The determining the access characteristics of the IO requests of each category among the K categories accessing the shared cache includes: For the first replacement algorithm among the N replacement algorithms, simulate the hit ratios of the IO requests of each category among the K categories applying the first replacement algorithm in caches of different sizes in the shared cache to obtain the relationships between the hit ratios and the cache size. Wherein, the first replacement algorithm is any one of the N replacement algorithms.
3. The method according to claim 1, characterized in that The determining the access characteristics of the IO requests of each category among the K categories accessing the shared cache includes: For the first replacement algorithm among the N replacement algorithms, determine the hit ratios of the first category of IO requests applying the first replacement algorithm in caches of different sizes according to the reuse distances of each IO request in the first category of IO requests and different cache sizes to obtain the relationships between the hit ratios and the cache size. Wherein, the first replacement algorithm is any one of the N replacement algorithms, and the first category of IO requests is any one of the K categories of IO requests.
4. The method according to any one of claims 1 to 3, characterized in that The determining the partition size of the IO requests of each category in the shared cache and the replacement algorithm of the IO requests of each category in the shared cache according to the access characteristics of the IO requests of the K categories and the hit ratio of the shared cache includes: Determine the hit ratios of the IO requests of the K categories in the shared cache for each combination according to the X hit ratios corresponding to the X cache sizes determined for the IO requests of each category under each replacement algorithm. Wherein, for any one of the K categories of IO requests, the X cache sizes and the N replacement algorithms form X*N combinations, and each combination includes a cache size and a replacement algorithm. The X cache sizes are X cache sizes preset for the cache corresponding to the IO requests of each category. Determine the cache size corresponding to the I / O requests of each category when the hit rate of the I / O requests of the K categories in the shared cache is the highest as the partition size of the I / O requests of each category in the shared cache, and determine the replacement algorithm corresponding to the I / O requests of each category when the hit rate of the I / O requests of the K categories in the shared cache is the highest as the replacement algorithm of the I / O requests of each category in the shared cache.
5. The method according to claim 1, characterized in that Before determining the access characteristics of the I / O requests of each category in the K categories accessing the shared cache, the method further includes: Obtain the multiple I / O requests; Divide the I / O requests into the K categories according to the characteristics of the addresses of the data accessed by the multiple I / O requests or according to the category tags carried in the multiple I / O requests.
6. The method according to claim 1, characterized in that, The shared cache is the last-level cache (LLC) of the central processing unit (CPU) in the computing device, and the multiple I / O requests are the I / O requests initiated by multiple processing cores in the CPU.
7. The method according to claim 1, characterized in that The shared cache is the cache in the cache node, and the multiple I / O requests are the I / O requests initiated by multiple computing nodes accessing the cache node.
8. The method according to claim 1, characterized in that, The shared cache is a cache pool composed of caches in multiple nodes, and the multiple I / O requests are the I / O requests initiated by multiple computing nodes accessing the cache pool.
9. The method according to claim 1, wherein The access characteristics are characterized by the hit rate curve (HRC) or the miss rate curve (MRC) of the I / O requests.
10. The method according to claim 1, wherein The determination of the access characteristics of the I / O requests of each category in the K categories accessing the shared cache includes: Periodically determine the access characteristics of the I / O requests of each category in the K categories accessing the shared cache; For the access characteristics of the I / O requests of each category in the K categories accessing the shared cache determined in the first period, the determination of the partition size and the replacement algorithm of the I / O requests of each category in the shared cache according to the access characteristics of the I / O requests of the K categories and the hit rate of the shared cache includes: Determine the partition size and the replacement algorithm of the I / O requests of each category in the shared cache in the first period according to the access characteristics of the I / O requests of the K categories determined in the first period and the hit rate of the shared cache; the first period is any period for determining the access characteristics of the I / O requests of each category in the K categories accessing the shared cache.
11. A management device for a shared cache, characterized in that, The shared cache is used to cache the data requested by multiple input / output (I / O) requests, the multiple I / O requests include I / O requests of K categories, and the shared cache includes N replacement algorithms; the device includes: A determination unit, configured to determine the access characteristics of the I / O requests of each category in the K categories accessing the shared cache, where the access characteristics are the relationship between the hit rate and the cache size of the I / O requests of each category in the K categories under the N replacement algorithms respectively; and, configured to determine the partition size of the I / O requests of each category in the shared cache and the replacement algorithm of the I / O requests of each category in the shared cache according to the access characteristics of the I / O requests of the K categories and the hit rate of the shared cache. The cache size of the IO requests for each category in the shared cache is the partition size of the IO requests for each category determined in the shared cache, and the cache eviction algorithm of the IO requests for each category in the shared cache is the cache eviction algorithm of the IO requests for each category determined in the shared cache.
12. The device according to claim 11, characterized in that, The device further includes: A simulation unit, configured to, for a first cache eviction algorithm among the N cache eviction algorithms, simulate the hit rate of the IO requests for each of the K categories in different-sized caches when applying the first cache eviction algorithm in the shared cache, so as to obtain the relationship between the hit rate and the cache size; wherein the first cache eviction algorithm is any one of the N cache eviction algorithms.
13. The device according to claim 11, characterized in that, The determination unit is specifically configured to: For a first cache eviction algorithm among the N cache eviction algorithms, according to the reuse distance of each IO request in the first type of IO requests and different cache sizes, determine the hit rate of the first type of IO requests when applying the first cache eviction algorithm in different-sized caches, so as to obtain the relationship between the hit rate and the cache size; wherein the first cache eviction algorithm is any one of the N cache eviction algorithms, and the first type of IO requests is the IO requests of any one of the K categories.
14. The device according to any one of claims 11-13, characterized in that The determination unit is further specifically configured to: According to the X hit rates corresponding to the X cache sizes determined for the IO requests for each category under each cache eviction algorithm, determine the hit rate of the IO requests for the K categories in the shared cache for each combination; wherein, for any one of the K categories of IO requests, the X cache sizes and the N cache eviction algorithms form X*N combinations, and each combination includes a cache size and a cache eviction algorithm; the X cache sizes are X cache sizes preset for the cache corresponding to the IO requests for each category. Determine the cache size corresponding to the IO requests for each category when the hit rate of the IO requests for the K categories in the shared cache is the largest as the partition size of the IO requests for each category in the shared cache, and determine the cache eviction algorithm corresponding to the IO requests for each category when the hit rate of the IO requests for the K categories in the shared cache is the largest as the cache eviction algorithm of the IO requests for each category in the shared cache.
15. The device according to claim 11, characterized in that, The device further includes: An acquisition unit, configured to acquire the plurality of IO requests before determining the access characteristics of the IO requests for each of the K categories accessing the shared cache. A classification unit, configured to classify the IO requests into the K categories according to the characteristics of the addresses of the data accessed by the plurality of IO requests or according to the category tags carried in the plurality of IO requests.
16. The device according to claim 11, characterized in that, The shared cache is the last-level cache (LLC) of the central processing unit (CPU) in a computing device, and the plurality of IO requests are IO requests initiated by a plurality of processing cores in the CPU.
17. The device according to claim 11, characterized in that, The shared cache is the cache in a cache node, and the plurality of IO requests are IO requests initiated by a plurality of computing nodes accessing the cache node.
18. The device according to claim 11, characterized in that, The shared cache is a cache pool composed of caches in multiple nodes, and the multiple IO requests are IO requests initiated by multiple computing nodes accessing the cache pool.
19. The device according to claim 11, characterized in that, The access characteristics are characterized by the hit rate curve HRC or the miss rate curve MRC of the IO requests.
20. The apparatus according to claim 11, wherein the determining unit is further configured to periodically determine the access characteristics of the IO requests of each of the K categories accessing the shared cache; For the access characteristics of the IO requests of each of the K categories accessing the shared cache determined in the first period, the determining unit is specifically configured to determine the partition size and the replacement algorithm of the IO requests of each category in the shared cache in the first period according to the access characteristics of the IO requests of the K categories determined in the first period and the hit rate of the shared cache; The first period is any period for determining the access characteristics of the IO requests of each of the K categories accessing the shared cache.
21. A computing device, characterized in that, The computing device is used to manage the shared cache, and the computing device includes: a memory, one or more processors, and the one or more processors are configured to read program instructions stored in the memory to execute the method according to any one of claims 1-10.
22. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes program instructions, and when the program instructions run on a computer or a processor, the computer or the processor is caused to execute the method according to any one of claims 1-10.
Citation Information
Patent Citations
Handling cache write-back and cache eviction for cache coherence
CN104520824A
Predicting and optimizing I / O performance characteristics in a multi-level caching system
US8112586B1