A cache loading method, a storage device

CN122489452BActive Publication Date: 2026-09-25LANGCHAO ELECTRONIC INFORMATION IND CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610959606.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25
Estimated Expiration
2046-06-30

AI Technical Summary

Technical Problem

[0003]本发明提供了一种缓存加载方法、存储设备,以解决存储系统的缓存预取策略无法适应深度学习模型的数据访问模式的技术问题

Benefits of technology

[0007]本发明的有益效果在于:本发明首先可根据客户端下发的读访问请求,检测连续数据访问长度,并根据检测结果更新最大连续访问长度;其中,读访问请求由客户端中运行的深度学习模型产生。维护最大连续访问长度的目的在于:深度学习模型在运行时,通常会连续读取数据,但不同模型的连续读取长度不同。因此,维护最大连续访问长度,可自动检测出不同深度学习模型的连续读取长度。进而,当检测到未命中缓存的目标读访问请求时,根据目标读访问请求的首地址确定数据加载首地址,并根据最大连续访问长度确定数据加载长度;根据数据加载首地址和数据加载长度,将位于磁盘设备中的数据加载至缓存。这样,本发明可以根据深度学习模型的数据访问模式,设置更为合适的缓存加载方法,可提升存储设备在应对深度学习场景时的缓存加载效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122489452B_ABST
    Figure CN122489452B_ABST
Patent Text Reader

Abstract

The application provides a cache loading method and a storage device, and relates to the technical field of storage. The method is applied to the storage device and can include the following steps: detecting a continuous data access length according to a read access request issued by a client, and updating a maximum continuous access length according to a detection result; wherein the read access request is generated by a deep learning model running in the client; when a target read access request that does not hit the cache is detected, determining a data loading start address according to a start address of the target read access request, and determining a data loading length according to the maximum continuous access length; and loading data located in a disk device to the cache according to the data loading start address and the data loading length. The cache preloading can be adapted to the continuous access mode of the deep learning model, and the maximum continuous access length can be updated to adapt to different continuous access lengths of different deep learning models, so that the cache preloading effect of the storage device in the deep learning scenario can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of storage technology, and in particular to a cache loading method and a storage device. Background Technology

[0002] In related technologies, the cache prefetching strategies of storage systems mainly rely on the principles of temporal and spatial locality of data access, employing LRU (Least Recently Used) and LFU (Least Frequently Used) strategies. However, these strategies are not suitable for the data access patterns of deep learning, resulting in low cache hit rates and poor caching performance. Summary of the Invention

[0003] This invention provides a cache loading method and a storage device to solve the technical problem that the cache prefetching strategy of the storage system cannot adapt to the data access pattern of the deep learning model.

[0004] This invention provides a cache loading method applied to a storage device, the method comprising: The system detects the length of continuous data access based on the read access request sent by the client, and updates the maximum length of continuous access based on the detection result. The read access request is generated by a deep learning model running in the client. When a target read access request that misses the cache is detected, the starting address of the data loading is determined based on the starting address of the target read access request, and the length of data loading is determined based on the maximum length of continuous access. Based on the starting address and length of the data to be loaded, the data located on the disk device is loaded into the cache.

[0005] The present invention also provides a storage device, comprising: Memory, used to store computer programs; The processor is used to implement the cache loading method described above when executing computer programs.

[0006] This invention provides a cache loading method applied to a storage device. The method includes: detecting the length of continuous data access based on a read access request issued by a client, and updating the maximum continuous access length based on the detection result; wherein the read access request is generated by a deep learning model running in the client; when a target read access request that misses the cache is detected, determining the starting address of data loading based on the starting address of the target read access request, and determining the data loading length based on the maximum continuous access length; and loading the data located on the disk device into the cache based on the starting address of the data loading request and the data loading length.

[0007] The beneficial effects of this invention are as follows: First, it can detect the continuous data access length based on read access requests issued by the client, and update the maximum continuous access length based on the detection results; wherein, the read access requests are generated by the deep learning model running in the client. The purpose of maintaining the maximum continuous access length is that: deep learning models typically read data continuously during runtime, but the continuous read length varies between different models. Therefore, maintaining the maximum continuous access length can automatically detect the continuous read length of different deep learning models. Furthermore, when a target read access request that misses the cache is detected, the starting address of the data loading is determined based on the starting address of the target read access request, and the data loading length is determined based on the maximum continuous access length; based on the starting address of the data loading and the data loading length, the data located on the disk device is loaded into the cache. In this way, this invention can set a more suitable cache loading method according to the data access pattern of the deep learning model, which can improve the cache loading effect of the storage device when dealing with deep learning scenarios.

[0008] The present invention also provides a storage device that has the above-mentioned beneficial effects. Attached Figure Description

[0009] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart of a cache loading method provided in an embodiment of the present invention; Figure 2 A structural block diagram of a sampling record linked list provided in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating in-volume cache resource management as provided in an embodiment of the present invention; Figure 4 A schematic diagram of the data processing flow provided in an embodiment of the present invention; Figure 5 This is a structural block diagram of a cache loading device provided in an embodiment of the present invention. Detailed Implementation

[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0012] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0013] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0014] In related technologies, the cache prefetching strategies of storage systems mainly rely on the principles of temporal and spatial locality of data access, employing LRU (Least Recently Used) and LFU (Least Frequently Used) strategies. However, these strategies are not suitable for the data access patterns of deep learning, resulting in low cache hit rates and poor caching.

[0015] For example, in training deep learning models, the models typically access data in batches according to a certain data access length. Furthermore, the data access length varies randomly across different models; for example, it could be 128k, 256k, or 4k. Due to these characteristics, traditional strategies suffer from low cache hit rates and long I / O paths, leading to storage performance limitations that cannot meet the training requirements of deep learning models.

[0016] In view of this, in order to address the technical problem of how to solve the problem that the cache prefetch strategy of the storage system cannot adapt to the data access mode of the deep learning model, the present invention can provide a cache loading method that can adapt to the different consecutive access lengths of different deep learning models by updating the maximum consecutive access length, thereby improving the cache preloading effect of the storage device in the deep learning scenario.

[0017] It should be noted that this method can be executed by storage devices. Storage devices refer to those that provide data access services to clients and manage the backend disk devices and the data on them, typically consisting of one or more controllers. Controllers can be configured with caches, which are storage spaces within the storage device used for temporary data storage with access speeds higher than those on the disk devices. The front end of the storage device can connect to the server (i.e., the client) via technologies such as FC and RDMA (roce), while the back end can connect to the disk devices via technologies such as SAS and NVME. Furthermore, storage devices can set up storage volumes based on disk devices. A storage volume is a logical storage unit that, from the client's perspective, is equivalent to an independent storage space; data access on different storage volumes is independent of each other.

[0018] When a client application generates a read I / O request, the request is first sent to one of the controllers on the storage device, such as controller A, via the front-end card (e.g., FC). Controller A identifies the read request and searches its cache based on the volume to which the requested data belongs and its address within that volume. If the data is found, it is directly returned to the client, completing one host read I / O request, known as a cache hit. If the data is not in the cache, the cache module further passes the read I / O request to lower-level modules until the data is read from the back-end storage disk. After processing through the entire I / O stack, the data is returned to the client, completing one host read I / O request. Therefore, storage caching plays a crucial role in storage I / O performance.

[0019] For easier understanding, please refer to Figure 1 , Figure 1 A flowchart of a cache loading method provided in an embodiment of the present invention, the method being applied to a storage device, may include: S11. Detect the length of continuous data access based on the read access request sent by the client, and update the maximum length of continuous access based on the detection result; wherein, the read access request is generated by the deep learning model running in the client.

[0020] It should be noted that this embodiment is not limited to a specific deep learning model. The model can be, for example, an image processing model, a language processing model, etc., and the data read by the model can be images, text, etc.

[0021] In this step, a deep learning model runs on the client, and this model can send read access requests to the storage device through the client to read the required data. Since different models use different data access lengths when performing continuous data access, the underlying storage device cannot set a fixed continuous data access length to accommodate different models. Therefore, this embodiment detects the continuous data access length by monitoring the read access requests sent by the client, and updates the maximum continuous access length based on this length. This maximum continuous access length is then considered as the continuous data access length used by the deep learning model running on the client, enabling cache preloading.

[0022] Specifically, this embodiment uses sampling records to continuously record the continuous data access information of each read access request. The sampling record may include the access end address and the recorded access length. The access end address refers to the last address of the data accessed by the read access request, and the recorded access length records the cumulative length of continuous access by one or more read access requests. When a read access request is received from the client, this embodiment matches the access start address in the read access request with the access end address in the sampling record. If the match is successful, it indicates that the read access request is continuous with previous historical read access requests, and the access end address and recorded access length in the sampling record can be updated. If the match fails, it indicates that the read access request is not continuous with previous read access requests. In this case, to facilitate subsequent matching, a new sampling record can be created, and the newly created sampling record is used to record the access end address and recorded access length of the current read access request.

[0023] The following describes how to use the sampling records. In one implementation, detecting the continuous data access length based on the read access request sent by the client, and updating the maximum continuous access length based on the detection result, may include: S1111. Upon receiving a read access request, the starting address of the read access request is matched with the ending address of the access in the pre-set sampling record; wherein, the sampling record contains the ending address of the access and the recorded access length.

[0024] In this step, the storage device is pre-configured with sampling records. Each sampling record contains an access end address (last_lba) and a recorded access length (current_max_length). When a read access request is received, the storage device can compare the access start address of the read access request with the access end address in the sampling record to determine whether the current access is consecutive in address with a previously recorded access.

[0025] S1112. When a match is successful, update the access end address in the successfully matched sample record to the access end address of the read access request, and add the data access length of the read access request to the recorded access length of the sample record.

[0026] In this step, a successful match indicates that the current read access request is address-continuous with the historical access recorded in the sampling record. Therefore, the access end address of the sampling record can be updated to the access end address of the current read access request, and the data access length of the current read access request can be added to the recorded access length of the sampling record, so that the sampling record can continuously track the cumulative length of a continuous access.

[0027] S1113. When a match fails, set a new sampling record and use the new sampling record to save the access end address and data access length corresponding to the read access request.

[0028] In this step, a failed match indicates that the current read access request is not contiguous in address with the historical accesses recorded in the existing sampling records. At this time, a new sampling record is set, and the new sampling record is used to save the end address and data access length corresponding to the current read access request, as the starting record of a new contiguous access.

[0029] S1114. After completing the recording of the read access request, update the maximum continuous access length according to the recorded access length in each of the current sampled records.

[0030] In this step, after completing the above updates or settings, the storage device can update the maintained maximum consecutive access length based on the recorded access lengths in each current sampling record. Specifically, the maximum consecutive access length can be the largest among the recorded access lengths in each current sampling record.

[0031] As can be seen, by setting sampling records, this embodiment can quickly and effectively determine the maximum continuous access length of the deep learning model, thereby facilitating cache preloading.

[0032] Furthermore, to avoid excessive sampling records, this embodiment can also set a limited number of sampling records and manage them using a linked list structure to achieve cyclical use of the sampling records, thereby saving resources and quickly iterating to find the optimal data access length. In one possible scenario... Please refer to... Figure 2 , Figure 2 This is a structural block diagram of a sampling record linked list provided in an embodiment of the present invention. Figure 2In this model, the sampling record linked list (i.e., the first linked list) contains a first number of sampling records. This first number is determined by a preset multiple of the preset concurrency level. The preset concurrency level is the minimum concurrency level among all hardware units in the storage link, where the storage link is the link between the client and the storage device. For example, a typical storage link consists of a server block device, server multipath control software, server FC card, storage FC card, and storage front-end concurrent resource control. Each module on the path has a maximum concurrency level or maximum queue depth. The preset concurrency level is the minimum of the maximum concurrency levels of all modules on that path. The preset multiple is an adjustable value that can be adjusted according to the access characteristics of different models (such as the number of accesses per access batch in the model). The default value can be n=5. The first linked list can be set for each storage volume. For example, if the maximum concurrency level of the storage path is 16, then each storage volume can have 80 sampling records.

[0033] The specific uses of the first linked list are described below. In one possible scenario, the length of consecutive data accesses by the client is detected based on the read access request issued by the client, and the maximum consecutive access length is updated based on the detection result. This could include: S1121. Upon receiving a read access request, the starting address of the read access request is matched with the ending address of the access in the pre-set sampling record; wherein, the sampling record contains the ending address of the access and the recorded access length.

[0034] S1122. When the access start address matches the access end address in the sampling record, update the access end address in the matching sampling record to the access end address of the read access request, add the data access length of the read access request to the recorded access length of the sampling record, and adjust the matching sampling record to the first position of the first linked list.

[0035] In this embodiment, when the access start address successfully matches the access end address in the sampling record, the sampling record can be updated, and the sampling record can be adjusted to the first position of the first linked list. In this way, by adjusting the position of the sampling records, the sampling record that has not been updated for the longest time can be gradually moved to the end of the first linked list.

[0036] S1123. When the access start address fails to match the access end address in the sampling record, select the sampling record at the end of the first linked list as the new sampling record.

[0037] S1124. Update the access end address of the new sampled record to the access end address of the read access request, update the recorded access length of the selected new sampled record to the data access length of the read access request, and move the new sampled record to the first head of the first linked list.

[0038] In steps S1123~S1124, when the access start address fails to match the access end address in the sampling record, a new sampling record can be selected from the end of the first linked list, and the original data in the new sampling record can be overwritten using the access end address and data access length of the current read access request. Then, the new sampling record is moved back to the beginning of the first linked list to realize the reuse of the sampling record.

[0039] S1125. After completing the recording of the read access request, update the maximum continuous access length according to the recorded access length in each of the current sampled records.

[0040] As can be seen in this embodiment, the storage device can use sampling records to track the cumulative length of continuous access, and collect the length of continuous data access with a fixed number of sampling records managed by a linked list. This can avoid uncontrollable time interval factors and quickly iterate to find the optimal length of continuous data access, further improving the detection accuracy and adaptability of the maximum continuous access length.

[0041] S12. When a target read access request that misses the cache is detected, the starting address of the data load is determined based on the starting address of the target read access request, and the data load length is determined based on the maximum continuous access length.

[0042] In this step, when the storage device detects that the data requested by a certain read access request (i.e., the target read access request) is not cached, the storage device can determine the starting address of subsequent data loading based on the starting address of the target read access request, i.e., determine the starting address of data loading; on the other hand, it can determine the length of this data loading based on the maintained maximum consecutive access length, i.e., determine the data loading length. Specifically, since the maximum consecutive access length reflects the consecutive read length of the current deep learning model, it can be used as the data loading length so that cache loading automatically matches the actual access length of the model.

[0043] Of course, in the early stages of operation, the storage device may not have recorded the effective maximum consecutive access length. Therefore, when determining the data loading length, the larger of the data access length of the target read access request and the maximum consecutive access length can be used as the data loading length.

[0044] Based on this, determining the data loading length according to the maximum continuous access length can include: S121. Use the larger of the data access length of the target read access request and the maximum continuous access length as the data loading length.

[0045] In this step, the larger value between the data access length of the target read access request and the maximum continuous access length is taken as the data loading length. This ensures that the data loading length is not less than the data access length required by the target read access request itself, and can also cover the continuous reading length of the current deep learning model, thereby achieving effective cache preloading.

[0046] S13. Load the data located on the disk device into the cache according to the data loading start address and data loading length.

[0047] In this step, the storage device preloads the corresponding data from the disk device into a cache, starting from the data load address and according to the data load length. Thus, when a client subsequently initiates a read access request for that data segment, it can directly hit the cache and receive the data from the cache, without needing to access the backend disk device again.

[0048] As can be seen, since the data loading length can be determined based on the predetermined maximum continuous data access length, and the maximum continuous data access length can be effectively adapted to the data access length used by the model, this embodiment can ensure that the data preloaded into the cache better meets the usage requirements of the model, thereby achieving a better cache preloading effect.

[0049] Based on the above embodiments, the present invention first detects the continuous data access length according to the read access request issued by the client, and updates the maximum continuous access length according to the detection result; wherein, the read access request is generated by the deep learning model running in the client. The purpose of maintaining the maximum continuous access length is that: when a deep learning model is running, it usually reads data continuously, but the continuous read length is different for different models. Therefore, maintaining the maximum continuous access length can automatically detect the continuous read length of different deep learning models. Furthermore, when a target read access request that misses the cache is detected, the starting address of the data loading is determined according to the starting address of the target read access request, and the data loading length is determined according to the maximum continuous access length; according to the starting address of the data loading and the data loading length, the data located on the disk device is loaded into the cache. In this way, the present invention can set a more suitable cache loading method according to the data access pattern of the deep learning model, which can improve the cache loading effect of the storage device when dealing with deep learning scenarios.

[0050] Based on the above embodiments, the maximum consecutive access length record can also be recorded across different storage volumes. The following describes a method for recording the maximum consecutive access length separately for different storage volumes. In one possible scenario, this method may further include: S21. Based on the read access requests sent by the client to each storage volume, detect the length of continuous data access in each storage volume, and update the maximum length of continuous access in each storage volume according to the detection results; wherein, the read access requests are generated by the deep learning model running in the client.

[0051] S22. When a target read access request that misses the cache is detected, determine the target storage volume corresponding to the target read access request, and determine the data loading length based on the maximum continuous access length of the target storage volume.

[0052] S23. Based on the data loading start address and data loading length, load the data located in the disk device into the cache area corresponding to the target storage volume.

[0053] In steps S21 to S23, since different storage volumes may carry data access for one or more deep learning models, the continuous data access length can be detected separately for each storage volume and the maximum continuous access length of each storage volume can be maintained separately. This ensures that the detection process of each storage volume does not affect each other, thereby more accurately matching the data access length on each storage volume.

[0054] Furthermore, each storage volume can be configured with a corresponding cache area to cache the data required by that storage volume. To ensure efficient use of cache resources, each storage volume can also utilize various linked list structures for effective management of cache resources.

[0055] The organization and eviction of cached resources within each storage volume's cache region are further explained below. Please refer to [link / reference needed]. Figure 3 , Figure 3 This is a schematic diagram illustrating in-volume cache resource management as provided in an embodiment of the present invention. Figure 3 In this system, the cache resources in the cache regions corresponding to each storage volume are recorded in the second, third, and fourth linked lists, respectively. The second linked list records free cache resources, the third linked list records cache resources with cached data that have not yet been accessed, and the fourth linked list records cache resources with cached data that have already been accessed. A cache resource refers to a unit of cache space used for caching data.

[0056] In other words, idle cache resources can be placed in the second linked list. When an idle cache resource is used to cache data, and the cached data has not yet been accessed by the client, the cache resource can be moved to the third linked list. When the cached data is accessed by the client, the corresponding cache resource can be moved to the fourth linked list. The usage of these three linked lists is described below.

[0057] In one scenario, loading data located on a disk device into the cache area corresponding to the target storage volume may include: S231. When there is a cache resource in the second linked list of the target storage volume, use the cache resource in the second linked list to cache the data to be loaded, and adjust the cache resource to the first position of the third linked list.

[0058] In this step, when there are idle cache resources, the idle cache resources can be used first to cache the data to be loaded, and then transferred to the third linked list to indicate that the cache resource has cached data but has not yet been accessed.

[0059] In one case, this method may also include: S31. When a read access request hits data that already exists in the cache area corresponding to the target storage volume, the cache resource containing the hit data is moved to the first position of the fourth linked list.

[0060] In this step, when cached data in the cache is hit by a read access request, the cache resource containing the hit data can be transferred to the fourth linked list, indicating that the data cached by the cache resource has been accessed.

[0061] In one scenario, loading data located on a disk device into the cache area corresponding to the target storage volume also includes: S41. When there is no cached resource in the second linked list and there is a cached resource in the fourth linked list, select the cached resource at the end of the fourth linked list to cache the data to be loaded, and adjust the selected cached resource to the first position of the third linked list.

[0062] S42. When there are no cached resources in the second linked list and no cached resources in the fourth linked list, select the cached resource at the end of the third linked list to cache the data to be loaded, and adjust the selected cached resource to the first position of the third linked list.

[0063] In steps S41-S42, when there are no free cache resources, the cache resource at the end of the fourth linked list is preferentially selected to cache the data to be loaded. Specifically, the fourth linked list records cached data that has already been accessed, and its probability of being accessed again in the short term is low, so it can be released and reused first; while the end of the fourth linked list corresponds to the cache resource that has not been accessed for the longest time, which conforms to the principle of least recently used. When there are no cache resources in the fourth linked list, the cache resource at the end of the third linked list is selected to cache the data to be loaded. After the reused cache resources cache new data to be loaded, they are all adjusted to the first position of the third linked list.

[0064] As can be seen in this embodiment, the storage device manages the cache resources in each storage volume cache area according to the categories of idle, cached but not accessed, and cached and accessed, respectively, using the second linked list, the third linked list, and the fourth linked list. When the cache resources are insufficient, the accessed cache resources are reused first. Compared with the method of uniformly eviction, the device can more accurately retain cache data that still has access value and reduce the probability of valid data being mistakenly evicted.

[0065] Based on the above embodiments, the allocation method of cache resources will be described below. In one possible case, this method may further include: S51. During storage device initialization, cache resources are evenly distributed to the cache area of ​​each storage volume according to the total number of storage volumes.

[0066] In this step, since the storage device does not yet have access statistics for each storage volume during the initialization phase, the cache resources can be evenly distributed to the cache area of ​​each storage volume according to the total number of storage volumes, providing initial cache resources for each storage volume.

[0067] S52. Obtain the cache hit rate, access bandwidth, and total cache resources occupied by each storage volume.

[0068] In this step, the storage device can obtain the cache hit rate, access bandwidth, and the total amount of cache resources currently occupied by each storage volume, which will serve as the basis for determining the cache resource adjustment method later. Specifically, the access bandwidth can be calculated separately for each storage volume within a unit of time.

[0069] S53. Determine the scheduling priority of each storage volume based on its cache hit rate, access bandwidth, and total cache resources used, and sort the storage volumes according to their scheduling priorities.

[0070] In this step, the storage device can determine the scheduling priority of each storage volume by comprehensively considering its cache hit rate, access bandwidth, and total cache resource usage. The scheduling priority represents the benefit gained from allocating cache resources to the corresponding storage volume. Subsequently, the storage volumes can be sorted according to their scheduling priorities. In one feasible implementation, the higher the cache hit rate, the greater the access bandwidth, and the smaller the total cache resource usage, the higher the scheduling priority of the corresponding storage volume, and the larger the scheduling priority value, the higher the priority can be, thus allowing them to be sorted in descending order.

[0071] It should be noted that this invention does not limit the specific method for determining scheduling priority. For example, in one case, scheduling priority can be determined as follows: Scheduling priority = cache hit rate Total bandwidth / volume usage of cache resources.

[0072] S54. When a cache adjustment operation is triggered, the cache resources occupied by each storage volume are adjusted according to the sorting results.

[0073] In this step, when the triggering conditions for the cache adjustment operation are met, the storage device adjusts the cache resources occupied by each storage volume according to the sorting results, so that the storage volume with higher scheduling priority occupies more cache resources. In one feasible implementation, the cache adjustment operation can be triggered when the average cache hit rate of each storage volume is lower than a preset threshold (e.g., 50%). The cache adjustment operation can be performed periodically, and the adjustment period can be set to no less than seconds to avoid excessively frequent adjustments that could cause cache jitter. This invention does not limit this.

[0074] The following describes the specific methods for adjusting the cache resources used by each storage volume based on the sorting results. In one possible scenario, adjusting the cache resources used by each storage volume based on the sorting results may include: S5421. For the first storage volume that is in the last preset ratio in the sorting result, set the release ratio of each first storage volume according to the sorting position, and release the cache resources of each first storage volume according to the release ratio to obtain the mobile cache resources.

[0075] The first preset ratio can be, for example, the first 30%. The specific release ratio can be determined according to the following formula: Percent = sorting percentage – 70%.

[0076] S5422. Divide the motor cache resources into the first motor cache resources and the second motor cache resources according to the second preset ratio.

[0077] The second preset ratio can be, for example, 80%.

[0078] S5423. For the second storage volume that ranks in the top third of the sorting results, set the allocation ratio of each second storage volume according to the sorting position, and allocate the first mobile cache resource to the second storage volume according to the allocation ratio.

[0079] The third preset ratio, for example, can be the bottom 30%. The allocation ratio can be determined as follows: 1. Calculate the total number of volumes in the first three preset proportions, and denote it as x; 2. Let X = 1 + 2 + 3 + ... + (x - 1) + (x), where X is the accumulated value of id; 3. Let the sorting number of the volume itself be id, then the range of values ​​for id is [1, x]. 4. The amount of cached resources obtained for each volume is (x – id + 1) / X N.

[0080] S5424. Select a third storage volume from the storage volumes with no available cache resources, and allocate the second mobile cache resources evenly to the third storage volume.

[0081] In this embodiment, the cache resources of the first storage volume, which is ranked lower, are released according to their ranking position to form a flexible cache resource. The first flexible cache resource is allocated to the second storage volume, which is ranked higher, according to its ranking position. Then, a third storage volume is selected from the storage volumes with no available cache resources, and the second flexible cache resource is allocated evenly among them, allowing these storage volumes to re-participate in the allocation of cache resources. It should be noted that this embodiment does not limit the specific values ​​of the preset ratio and preset threshold; they can be set according to actual application requirements.

[0082] The following section provides a complete description of the cache loading method based on specific examples. This invention proposes a data caching and prefetching strategy specifically for the data access patterns of deep learning. Utilizing these patterns, data is pre-fetched from the backend disk into the data cache. When the frontend accesses data, this improves the cache hit rate and shortens the real-time data access path, thereby effectively enhancing the overall data access efficiency of the storage device.

[0083] Please refer to Figure 4 , Figure 4 This is a schematic diagram of a data processing flow provided in an embodiment of the present invention. For example... Figure 4 As shown, the main data flow of this invention is divided into five modules. First, the detection module monitors real-time data access, mainly including front-end data access bandwidth, continuous data access length, overall cache resource usage, and cache hit rate, and feeds this information back to the control module. The control module controls the prefetching module to prefetch data based on the continuous data access length, data access address, and access bandwidth. Based on the overall cache resource usage, the control module controls the eviction module to evict some cached data and reclaim cache resources. Based on the overall business situation, the control module controls the allocation module to allocate cache resources. The working principle of each module is explained in detail below.

[0084] I. Detection Module: The detection module is mainly responsible for collecting information on the current state of the cache and storage system. This will be explained below according to the types of information: 1. Collection of continuous data access length: A. Monitoring of continuous access length is collected by storage volume (equivalent to an independent hard drive from the server's perspective), meaning that monitoring data from different storage volumes do not interfere with each other.

[0085] B. The sampling resources for monitoring data should be set to n times the maximum concurrency. Here, maximum concurrency refers to the maximum concurrency value along the I / O path from the server to the storage. A typical path might be server block device, server multipath control software, server FC card, storage FC card, and storage front-end concurrent resource control. Each module on the path has a maximum concurrency or maximum queue depth, and the maximum concurrency is the minimum of the maximum concurrency values ​​of all modules on that path. n times is an adjustable value that can be adjusted according to the access characteristics of different models (such as the number of accesses in each access batch). A default value of n=5 is recommended. For example, if the maximum concurrency is 16, then the sampling resources for each storage volume are 80.

[0086] C. For example Figure 2 As shown, the sampled resource contains only two records: the maximum address of the last sampled I / O access (last_lba) and the length of the recorded consecutive data accesses. The sampled resource is managed using a linked list. In addition, each storage volume also needs to maintain a current_max_length entry to indicate the current maximum access length.

[0087] D. When a new IO arrives, first obtain its starting address, then compare it with last_lba in the sampled resource list. If the comparison is successful, update the value of last_lba to the maximum data access address of the new IO, and add the data access length of the new IO to length. Compare the accumulated length with current_max_length. If length is larger, update current_max_length with length. Finally, move the sampled resource that successfully matched to the head of the list. If the comparison fails, remove the sampled resource from the tail of the list, update last_lba to the maximum access address of the new IO, update length to the data access length of the new IO, and similarly try to update current_max_length with the new length. Finally, insert the sampled resource into the head of the list.

[0088] E. The cache control module periodically reads current_max_length to know the maximum continuous data access length in real time.

[0089] Here, by setting a fixed number of sampling resources to collect the continuous data access length from the monitoring client, compared to using a fixed time interval, the uncontrollable factors of the time interval can be naturally avoided. Furthermore, the reasonable setting of the number of sampling resources, the resource LRU usage mode, and the periodic reading by the control module can quickly iterate to the optimal data access length. This is within the information retrieval period interval.

[0090] 2. Access bandwidth statistics are still calculated per second based on storage volume.

[0091] 3. Cache hit rate monitoring: A. Calculate cache hit rate by storage volume; B. Data blocks in the cache that are hit by client I / O will be marked with a hit flag by the detection module and moved to the head of the hit data block linked list; C. The formula for calculating cache hit rate is: The amount of data that was hit in the cache during IO reads / the total amount of data read during IO reads 100%; D. The cache hit rate update cycle is the same as the current_max_length reading cycle of the cache control module. After each reading, this value is reset to zero and the statistics are repeated.

[0092] 4. Overall monitoring of cache resource usage: The system monitors two aspects of cache usage: firstly, the total amount of cache resources and their usage; and secondly, the maximum amount of cache resources that can be used for each volume and their usage.

[0093] II. Cache Allocation Module: 1. Cache allocation during initialization: During system initialization, cache resources are evenly distributed according to the total number of storage volumes.

[0094] 2. Dynamic adjustment of cache allocation: A. The cache control module obtains the cache hit rate and front-end bandwidth of the storage volume from the detection module; B. Cache allocation follows a greedy algorithm and profit-seeking principle, and the storage volumes are sorted according to the following formula: Cache hit rate Total bandwidth / volume usage of cache resources.

[0095] (1) When the average cache hit rate is below 50%, start dynamic allocation of cache resources. Here, 50% is a reference value. (2) Dynamically adjust periodically, with a period of at least seconds, such as 5 seconds, to avoid excessively frequent adjustments that could cause cache jitter and lead to performance instability; C. After sorting, the percentage of cached resources released for the remaining 30% of volumes is calculated using the following formula: Percent = sort percentage – 70%; Once the cache resources are exhausted, the storage volume will no longer participate in the sorting process, and the released cache resources will be used as spare resources for the entire cache pool.

[0096] D. Volumes 30% to 70% after sorting do not require adjustment; E. The top 30% of volumes are allocated 80% of the total available resources in the cache pool, denoted as N, and distributed according to the following principles: (1) Calculate the total number of the top 30% of the papers, and let it be x; (2) Let X = 1 + 2 + 3 + ... + (x - 1) + (x), where X is the accumulated value of id; (3) Let the sorting number of the volume itself be id, then the range of values ​​for id is [1, x]; (4) The amount of cached resources obtained for each volume is (x – id + 1) / X N; F. From the volumes that have exhausted their cache resources, randomly select a portion of the volumes and evenly distribute the remaining 20% ​​of the cache resources in the cache pool so that they can participate in the allocation of cache resources again.

[0097] III. Cache eviction: When the cache control module learns from the detection module that the cache resources of a certain volume are 100% used, it needs to initiate the cache eviction mechanism. The volume-level cache resource management mode is as follows: Figure 3 As shown.

[0098] 1. In-volume cache resource management is divided into 3 linked lists: Free list, cached resources that have not been used by data in this volume; A valid data linked list, where the cached resources store the data of this volume, and the data has not been accessed by IO; The linked list was hit, the cached resource stores the data of this volume, and the data has been accessed by IO; 2. If there are still free cache resources in the volume, that is, the cache resource utilization rate is not 100%, then the free cache resources will be used first. 3. If the cache resource utilization rate has reached 100%, the cache resources in the hit list will be used to cache the data first. That is, the cache resources at the end of the hit list will be taken to cache the data (following the LRU principle), and then the resource will be moved to the head of the valid data list. 4. If the hit list is empty, then the cached resource at the tail of the valid data list is taken to cache the data (following the LRU principle), and then the resource is moved to the head of the valid data list.

[0099] IV. Cache prefetching: When the detection module detects an I / O that misses the cache, it should notify the cache control module to start cache prefetching.

[0100] 1. The cache prefetch module starts the prefetch process after receiving the prefetch notification; 2. Determine the starting address of the prefetched data based on the starting address of the cache miss I / O, and take the larger value between the cache miss I / O data size and current_max_length as the data prefetch size; 3. Based on the cache resource usage principles determined in the cache eviction mechanism, obtain a certain amount of cache resource prefetch data according to the prefetch size.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0102] Please refer to Figure 5 , Figure 5 This is a structural block diagram of a cache loading device provided in an embodiment of the present invention. The device is applied to a storage device and may include: The detection module 501 is used to detect the continuous data access length of the deep learning model running in the client according to the read access request sent by the client, and update the maximum continuous access length according to the detection result. The determination module 502 is used to determine the starting address of data loading based on the starting address of the target read access request when a target read access request that misses the cache is detected, and to determine the data loading length based on the maximum continuous access length. Loading module 503 is used to load data located on the disk device into the cache according to the data loading start address and data loading length.

[0103] Optionally, the detection module 501 includes: The matching submodule is used to match the starting address of the read access request with the ending address of the access in the pre-set sampling record when a read access request is received; the sampling record contains the ending address of the access and the recorded access length. The first recording submodule is used to update the access end address in the successfully matched sampled record to the access end address of the read access request when a match is successful, and to add the data access length of the read access request to the recorded access length of the sampled record. The second recording submodule is used to set a new sampling record when the matching fails, and to use the new sampling record to save the access end address and data access length corresponding to the read access request; The update submodule is used to update the maximum continuous access length based on the recorded access length in each of the current sampled records after completing the read access request for the record.

[0104] Optionally, the first linked list contains a first number of sampling records, the first number being determined according to a preset multiple of a preset concurrency level, the preset concurrency level being the minimum concurrency level among the hardware units in the storage link, and the storage link being the link between the client and the storage device.

[0105] Optionally, the detection module 501 further includes: The first adjustment submodule is used to adjust the successfully matched sampling record to the first position of the first linked list when the access start address matches the access end address in the sampling record.

[0106] Optionally, the detection module 501 further includes: The second adjustment submodule is used to select the sampling record at the end of the first linked list as the new sampling record; update the access end address of the new sampling record to the access end address of the read access request; update the recorded access length of the selected new sampling record to the data access length of the read access request; and move the new sampling record to the beginning of the first linked list.

[0107] Optionally, module 502 includes: The data loading length determination submodule is used to determine the larger of the data access length of the target read access request and the maximum continuous access length as the data loading length.

[0108] Optionally, the detection module 501 is also used for: Based on the read access requests sent by the client to each storage volume, detect the length of continuous data access in each storage volume; The determining module 502 is also used to: determine the target storage volume corresponding to the target read access request, and determine the data loading length based on the maximum continuous access length of the target storage volume; The loading module 503 is also used to load data located on the disk device into the cache area corresponding to the target storage volume according to the data loading start address and the data loading length.

[0109] Optionally, the device further includes: The cache resource allocation module is used to evenly distribute cache resources to the cache area of ​​each storage volume according to the total number of storage volumes during storage device initialization. The cache usage statistics module is used to obtain the cache hit rate, access bandwidth, and total cache resources used by each storage volume. The sorting module is used to determine the scheduling priority of each storage volume based on its cache hit rate, access bandwidth, and total cache resources used, and to sort the storage volumes according to the scheduling priority. The reallocation module is used to adjust the cache resources occupied by each storage volume according to the sorting results when a cache adjustment operation is triggered.

[0110] Optionally, the device further includes: The cache adjustment trigger module is used to trigger a cache adjustment operation when the average cache hit rate of each storage volume is lower than a preset threshold.

[0111] Optionally, the reallocation module is used for: For the first storage volume that is at the last preset ratio in the sorting results, set the release ratio of each first storage volume according to the sorting position, and release the cache resources of each first storage volume according to the release ratio to obtain the mobile cache resources. The mobile cache resources are divided into the first mobile cache resources and the second mobile cache resources according to the second preset ratio; For the second storage volume that ranks in the top three of the sorting results, the allocation ratio of each second storage volume is set according to the sorting position, and the first mobile cache resources are allocated to the second storage volume according to the allocation ratio. Select a third storage volume from the storage volumes with no available cache resources, and allocate the second mobile cache resources evenly to the third storage volume.

[0112] Optionally, the cache resources in the cache area corresponding to each storage volume are recorded in the second linked list, the third linked list, and the fourth linked list, respectively; wherein, the second linked list records idle cache resources, the third linked list records cache resources with cached data that have not been accessed, and the fourth linked list records cache resources with cached data that have been accessed.

[0113] Optionally, module 503 is loaded, including: The third adjustment submodule is used to cache the data to be loaded using the cached resources in the second linked list when there are cached resources in the second linked list of the target storage volume, and to adjust the cached resources to the first position of the third linked list.

[0114] Optionally, it also includes: The fourth adjustment submodule is used to adjust the cache resource containing the hit data to the first position of the fourth linked list when a read access request hits data that already exists in the cache area corresponding to the target storage volume.

[0115] Optionally, loading data located on the disk device into the cache area corresponding to the target storage volume further includes: The fifth adjustment submodule is used to select the cached resource at the end of the fourth linked list to cache the data to be loaded when there is no cached resource in the second linked list and there is a cached resource in the fourth linked list, and then adjust the selected cached resource to the beginning of the third linked list. The sixth adjustment submodule is used to select the cached resource at the end of the third linked list to cache the data to be loaded when there is no cached resource in the second linked list and no cached resource in the fourth linked list, and then adjust the selected cached resource to the beginning of the third linked list.

[0116] For a description of the features in the embodiment corresponding to the cache loading device, please refer to the relevant description of the embodiment corresponding to the cache loading method, which will not be repeated here.

[0117] Embodiments of the present invention also provide a storage device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above-described cache loading method embodiments.

[0118] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described cache loading method embodiments at runtime.

[0119] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0120] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described cache loading method embodiments.

[0121] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described cache loading method embodiments.

[0122] Any of the components, modules, units, parts, methods, and operations described herein can be implemented using software, firmware, hardware (e.g., fixed logic circuitry), manual processing, or any combination thereof. Alternatively or additionally, any functionality described herein can be performed at least in part by one or more hardware logic components, such as, but not limited to, a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), an application-specific standard product (ASSP), a system-on-a-chip (SoC), a complex programmable logic device (CPLD), a microprocessor (MCU), etc. The terms "system," "computing device," or "apparatus" as used herein encompass various means, devices, and machines for processing data, including, for example, one or more programmable processors, computers, SoCs, or combinations thereof. The apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or one or more combinations thereof. The aforementioned computer program (also known as a program, software, software application, app, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, object, or other unit suitable for a computing environment.

[0123] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0124] The cache loading method and storage device provided by this invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. It should be noted that those skilled in the art can make several improvements and modifications to this invention without departing from the principles of this invention, and these improvements and modifications also fall within the protection scope of this invention.

Claims

1. A cache loading method, characterized in that, Applied to a storage device, the method includes: The maximum continuous access length is detected based on the read access request issued by the client, and the maximum continuous access length is updated based on the detection result; wherein, the read access request is generated by the deep learning model running in the client; When a target read access request that misses the cache is detected, the starting address of the data loading is determined based on the starting address of the target read access request, and the data loading length is determined based on the maximum continuous access length. Based on the data loading start address and the data loading length, the data located on the disk device is loaded into the cache; The first linked list contains a first number of sampled records; the step of detecting the continuous data access length based on the read access request issued by the client, and updating the maximum continuous access length based on the detection result, includes: Upon receiving the read access request, the starting address of the read access request is matched with the ending address of the access in the pre-set sampling record; wherein, the sampling record contains the ending address of the access and the recorded access length; When a match is successful, the access end address in the successfully matched sample record is updated to the access end address of the read access request, the data access length of the read access request is added to the recorded access length of the sample record, and the successfully matched sample record is adjusted to the first position of the first linked list. When a match fails, the sampled record at the end of the first linked list is selected as a new sampled record, and the new sampled record is used to save the access end address and data access length corresponding to the read access request, and the new sampled record is moved to the beginning of the first linked list. After completing the recording of the read access request, the maximum consecutive access length is updated according to the recorded access length in each of the current sampled records.

2. The cache loading method according to claim 1, characterized in that, The first quantity is determined according to a preset multiple of a preset concurrency level, where the preset concurrency level is the minimum concurrency level among the hardware units in the storage link, and the storage link is the link between the client and the storage device.

3. The cache loading method according to claim 1, characterized in that, The newly sampled record is used to store the access end address and data access length corresponding to the read access request, including: Update the access end address of the new sampled record to the access end address of the read access request; Update the recorded access length of the selected new sample record to the data access length of the read access request.

4. The cache loading method according to claim 1, characterized in that, Determining the data loading length based on the maximum continuous access length includes: The larger of the data access length of the target read access request and the maximum continuous access length is used as the data loading length.

5. The cache loading method according to any one of claims 1 to 4, characterized in that, The length of continuous data access is detected based on the read access request sent by the client, including: Based on the read access requests sent by the client to each storage volume, detect the length of continuous data access in each storage volume; Determining the data loading length based on the maximum continuous access length includes: Determine the target storage volume corresponding to the target read access request, and determine the data loading length based on the maximum continuous access length of the target storage volume; Based on the data load starting address and the data load length, the data located on the disk device is loaded into the cache, including: Based on the data loading start address and the data loading length, the data located on the disk device is loaded into the cache area corresponding to the target storage volume.

6. The cache loading method according to claim 5, characterized in that, Also includes: During storage device initialization, cache resources are evenly distributed to the cache area of ​​each storage volume according to the total number of storage volumes; Get the cache hit rate, access bandwidth, and total cache resources used by each storage volume; The scheduling priority of each storage volume is determined based on its cache hit rate, access bandwidth, and total cache resources used, and the storage volumes are sorted according to the scheduling priority. When a cache adjustment operation is triggered, the cache resources used by each storage volume are adjusted according to the sorting results.

7. The cache loading method according to claim 6, characterized in that, Also includes: The cache adjustment operation is triggered when the average cache hit rate of each storage volume is lower than a preset threshold.

8. The cache loading method according to claim 6, characterized in that, Adjust the cache resources used by each storage volume according to the sorting results, including: For the first storage volume that is in the last preset proportion in the sorting result, set the release ratio of each first storage volume according to the sorting position, and release the cache resources of each first storage volume according to the release ratio to obtain the mobile cache resources. The mobile cache resources are divided into first mobile cache resources and second mobile cache resources according to a second preset ratio; For the second storage volume that ranks in the top third of the sorting results, the allocation ratio of each second storage volume is set according to the sorting position, and the first mobile cache resource is allocated to the second storage volume according to the allocation ratio. A third storage volume is selected from the storage volumes with no available cache resources, and the second mobile cache resources are evenly allocated to the third storage volume.

9. The cache loading method according to claim 5, characterized in that, The cache resources in the cache area corresponding to each storage volume are recorded in the second linked list, the third linked list, and the fourth linked list, respectively; wherein, the second linked list records idle cache resources, the third linked list records cache resources with cached data that have not been accessed, and the fourth linked list records cache resources with cached data that have been accessed.

10. The cache loading method according to claim 9, characterized in that, Loading data located on the disk device into the cache area corresponding to the target storage volume includes: When a cached resource exists in the second linked list of the target storage volume, the cached resource in the second linked list is used to cache the data to be loaded, and the cached resource is moved to the first position of the third linked list.

11. The cache loading method according to claim 9, characterized in that, Also includes: When a read access request hits a cache region that already contains data, the cache resource containing the hit data is moved to the head of the fourth linked list.

12. The cache loading method according to claim 9, characterized in that, Loading data located on the disk device into the cache area corresponding to the target storage volume further includes: When there is no cached resource in the second linked list and there is a cached resource in the fourth linked list, the cached resource at the end of the fourth linked list is selected to cache the data to be loaded, and the selected cached resource is adjusted to the first position of the third linked list. When there are no cached resources in the second linked list and no cached resources in the fourth linked list, the cached resource at the end of the third linked list is selected to cache the data to be loaded, and the selected cached resource is adjusted to the first position of the third linked list.

13. A storage device, characterized in that, include: Memory, used to store computer programs; A processor for implementing the cache loading method as described in any one of claims 1 to 12 when executing the computer program.

Citation Information

Patent Citations

  • Hard disk pre-reading method and device, electronic equipment and storage medium

    CN120371731A