Data pre-reading method, storage medium and electronic equipment
By determining the target access identifier and access heat in the data read request and pre-reading the high-access heat sub-area, the data read delay problem during cache miss is solved, and the cache hit rate and storage system performance are improved.
Patent Information
- Application Number
- CN202511109190.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
In the existing technology, data reading takes a long time, mainly because when the cache misses, data needs to be read from storage devices such as disks, resulting in increased latency.
By determining the target access identifier based on the storage address of the data read request, pre-reading is performed using the access popularity information in the access information set, and sub-region data with high access popularity is added to the cache first, achieving more detailed partition management and control.
It improves the cache hit rate, reduces the time spent on data reading, and optimizes the performance of the storage system.
Smart Images

Figure CN120596038A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data pre-reading method, a storage medium, and an electronic device. Background Art
[0002] Currently, the data reading process follows a sequential response mode. That is, when an application on the server initiates a read input / output request, the request is first sent to the front end of the storage controller. The front end transmits the request to the storage device by using technologies such as Fibre Channel or remote direct memory access. After receiving the request, the storage controller will parse the data location information contained in the request and check whether the required data is in the cache. If the data is found in the cache, it is a "cache hit". The storage controller will quickly extract the data from the cache and send it directly back to the server to complete the read request. If the data is not found in the cache, it is a "cache miss". Then, it is necessary to further read the data from storage devices such as disks, which increases the time consumption of data reading.
[0003] That is, there is a technical problem in the prior art that it takes a long time to read data. Summary of the Invention
[0004] The present application provides a data pre-reading method, a storage medium, and an electronic device to at least solve the technical problem in the prior art that it takes a long time to read data.
[0005] The present application provides a data pre-reading method, comprising: determining a target access identifier based on a storage address included in a current data read request, the target access identifier being used to indicate a current sub-region among multiple storage sub-regions included in a current disk, the current sub-region being used to store a first data object matching the current data read request; determining an access information set matching the target access identifier, the access information set including multiple access request description information corresponding to the multiple storage sub-regions, the access request description information being used to indicate access heat of the corresponding storage sub-regions; and adding at least one second data object read in the current sub-region to a cache when the access information set includes current request description information corresponding to the current sub-region and a comparison result between the current access heat indicated by the current request description information and a reference access heat meets a pre-reading condition.
[0006] The present application also provides a data pre-reading device, including: a first determination unit, used to determine a target access identifier based on a storage address included in a current data read request, the target access identifier being used to indicate a current sub-area among multiple storage sub-areas included in the current disk, the current sub-area being used to store a first data object matching the current data read request; a second determination unit, used to determine an access information set matching the target access identifier, the access information set including multiple access request description information corresponding to the multiple storage sub-areas, respectively, the access request description information being used to indicate an access heat of the corresponding storage sub-area; a pre-reading unit, used to include current request description information corresponding to the current sub-area in the access information set, and if a comparison result between the current access heat indicated by the current request description information and the reference access heat meets a pre-reading condition, add at least one second data object read in the current sub-area to the cache.
[0007] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned data pre-reading methods when executing the computer program.
[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned data pre-reading methods are implemented.
[0009] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned data pre-reading methods when executed by a processor.
[0010] Through the present application, a target access identifier is determined based on the storage address included in the current data read request, the target access identifier is used to indicate the current sub-area among the multiple storage sub-areas included in the current disk, and the current sub-area is used to store the first data object matching the current data read request, by performing access control and management based on more detailed partitions rather than relying solely on the hierarchy of the entire disk or volume; an access information set matching the target access identifier is determined, the access information set includes multiple access request description information corresponding to the multiple storage sub-areas, and the access request description information is used to indicate the access heat of the corresponding storage sub-area; when the access information set includes the current request description information corresponding to the current sub-area, and the comparison result between the current access heat indicated by the current request description information and the reference access heat meets the pre-read condition, at least one second data object read in the current sub-area is added to the cache, and the sub-area with high access heat is effectively pre-read into the cache, thereby improving the cache hit rate and solving the technical problem of long data reading time in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A hardware structure block diagram of a server device for a data pre-reading method provided in an embodiment of the present application;
[0013] Figure 2 A flowchart of a data pre-reading method provided in an embodiment of the present application;
[0014] Figure 3 An architectural diagram of a data pre-reading method provided in an embodiment of the present application;
[0015] Figure 4 A structural diagram of a prefetcher provided in an embodiment of the present application;
[0016] Figure 5 is a structural diagram of a data pre-reading device according to an embodiment of the present application;
[0017] Figure 6 It is a structural diagram of a data pre-reading electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0020] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0021] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure diagram of a server device of a data pre-reading method according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The above-mentioned server device may also include a transmission device 106 and an input and output device 108 for communication functions. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0022] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data pre-reading method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to a server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0023] Transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communication provider of the server device. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] An embodiment of the present application provides a data pre-reading method, and the method is described in detail in conjunction with the execution flow of the data pre-reading method.
[0025] In order to more clearly understand the technical solutions provided by the embodiments of the present application, the key terms involved in the embodiments of the present application are first introduced here:
[0026] Data input and output (IO).
[0027] Remote Direct Memory Access (RDMA).
[0028] Cache (CA).
[0029] Redundant Arrays of Independent Disks (RAID).
[0030] As an optional implementation, Figure 2 As shown, the data pre-reading method includes:
[0031] S202: Determine a target access identifier based on a storage address included in a current data read request, where the target access identifier indicates a current sub-region among a plurality of storage sub-regions included in a current disk, and the current sub-region is used to store a first data object matching the current data read request.
[0032] S204: Determine an access information set that matches the target access identifier, where the access information set includes multiple access request description information corresponding to multiple storage sub-regions, where the access request description information indicates access popularity of the corresponding storage sub-regions.
[0033] S206. When the access information set includes current request description information corresponding to the current sub-area, and the comparison result between the current access heat indicated by the current request description information and the reference access heat meets the pre-reading condition, at least one second data object read in the current sub-area is added to the cache.
[0034] In step S202, a target access identifier is determined according to the storage address included in the current data read request, where the target access identifier is used to indicate a current sub-region among multiple storage sub-regions included in the current disk, and the current sub-region is used to store a first data object matching the current data read request;
[0035] It should be noted that the storage capacity in the storage system is managed at the volume granularity. Each volume is equivalent to an available disk. The cache module divides any storage volume into multiple storage sub-areas, namely regions, at a granularity of 32MB. For example, a volume with a capacity of 100G can be divided into 100×1024 / 32=3200 storage sub-areas (regions). The prefetching business of these regions is dispatched to different cores and prefetchers for processing. Each core is configured with a prefetcher to achieve higher concurrency, such as Figure 3 As shown, a volume with volume_ID 1 is divided into multiple regions at a granularity of 32MB. Data pre-reading operations are evenly distributed to different processing cores and prefetchers through a specific algorithm. That is, by mapping core_ID = f(region_ID), a schematic diagram of further data reading is performed. For example, core 1 corresponds to prefetcher 1, core 2 corresponds to prefetcher 2, core 3 corresponds to prefetcher 3, and core 4 corresponds to prefetcher 4.
[0036] As an optional implementation, the target access identifier may be determined according to the storage address included in the current data read request using the following formula: target access identifier region_id=lba / 32MB, where lba is the starting address of the host read IO request received by the storage end.
[0037] In the above step S204, an access information set matching the target access identifier is determined, wherein the access information set includes a plurality of access request description information corresponding to the plurality of storage sub-regions, and the access request description information is used to indicate the access popularity of the corresponding storage sub-regions;
[0038] As an optional implementation, the access information set may be a database or data structure containing access information (such as frequency, most recent access time, etc.) of all storage sub-regions. The access request description information records a detailed description of the historical access situation of each storage sub-region, including but not limited to indicators such as the number of accesses, access interval, length of read data, and step length. The access heat may be a quantitative indicator derived from the access request description information, which may be used to determine the frequency with which a storage sub-region is accessed. The access heat may specifically be a hot description, which records the volume_id and region_id (i.e., region_id) described therein. During a search, the volume_id and region_id of the server-side read IO request are compared with the volume_id and region_id recorded in the hotdescription in the prefetch bucket.
[0039] Further, in step S206, when the access information set includes current request description information corresponding to the current sub-area, and the comparison result between the current access heat indicated by the current request description information and the reference access heat meets the pre-reading condition, at least one second data object read in the current sub-area is added to the cache.
[0040] Optionally, the current access popularity stores a read_count variable, which is used to record the number of times the region is read by the server. Each time a server read IO request belonging to the region arrives, the value of the variable is incremented by 1.
[0041] The core processor currently processing the region read request records the total number of read requests for the region as core_region_read_count, and the number of all host-side read IO requests processed by the core is recorded as core_read_count. When a new server-side read IO request arrives, the variable is incremented by 1, and the difference between core_read_count and core_region_read_count is calculated to obtain the first difference value of the reference access heat indication mentioned above.
[0042] The ratio of the read_count in the current access heat and the first difference indicated by the reference access heat (ie, core_read_count-core_region_read_count) is further calculated to determine whether it is greater than a target threshold. If the threshold is met, it is determined that the pre-read condition is met.
[0043] Therefore, the next data block or area that may be accessed in the current sub-area can be predicted based on the access frequency, access time, access order and data block size in the current request description information, and at least one second data object in the predicted area can be added to the cache, for example, determining that the data object of the target type in the area is the above-mentioned second data object.
[0044] Through the above-mentioned implementation mode recorded in the present application, a target access identifier is determined according to the storage address included in the current data read request, the target access identifier is used to indicate the current sub-area among the multiple storage sub-areas included in the current disk, and the current sub-area is used to store the first data object matching the current data read request, by performing access control and management based on more detailed partitions, rather than relying solely on the hierarchy of the entire disk or volume; an access information set matching the target access identifier is determined, the access information set includes multiple access request description information corresponding to the multiple storage sub-areas, and the access request description information is used to indicate the access heat of the corresponding storage sub-area; when the access information set includes the current request description information corresponding to the current sub-area, and the comparison result between the current access heat indicated by the current request description information and the reference access heat meets the pre-read condition, at least one second data object read in the current sub-area is added to the cache, and the sub-area with high access heat is effectively pre-read into the cache, thereby improving the cache hit rate and solving the technical problem of long data reading time in the prior art.
[0045] In an optional implementation, the comparison result between the current access popularity and the reference access popularity satisfies the pre-reading condition, including:
[0046] S1, determining a first request quantity of data read requests for data in a plurality of storage sub-regions processed by a target core, and a second request quantity of data read requests for data in a current sub-region processed by the target core, wherein the target core is used to process the current data read request at a current moment;
[0047] S2, determining a difference between the first number of requests and the second number of requests as a first difference value of a reference access heat indication;
[0048] S3, determining a prefetch weight value according to a ratio of the total number of visits to the current sub-region indicated by the current access heat and the first difference;
[0049] S4: When the pre-fetch weight value is greater than the target threshold, determine whether the comparison result between the current access heat and the reference access heat satisfies the pre-fetch condition.
[0050] The above steps S1-S3 are fully described using a complete implementation method:
[0051] For step S1, the number of all host-side read IO requests from the core processor currently processing the current data read request (i.e., the current region read request) may be recorded as the first request count (core_read_count). When a new server-side read IO request arrives, this variable is incremented by 1. The total number of reads from the core processor currently processing the region read request for the region is recorded as the second request count (core_region_read_count).
[0052] Corresponding to the above step S2, the difference between core_read_count and core_region_read_count is calculated to obtain the first difference value of the reference access heat indication.
[0053] In the above step S3, the total number of visits to the current sub-region indicated by the current access heat is read_count. The variable records the number of times the region is read by the server. Each time a server-side read IO request belonging to the region arrives, its value is incremented by 1.
[0054] Furthermore, the prefetch weight can be calculated according to the following formula:
[0055] weight=read_count×W / (core_read_count-core_region_read_conut), where W is a constant that can be set based on experience.
[0056] In step S4, it is determined whether the prefetch weight is greater than the target threshold X. If the threshold condition is met, it is determined that the comparison result between the current access heat and the reference access heat meets the pre-reading condition, and a prefetch process is initiated to prefetch data into the cache.
[0057] By counting the total number of requests processed by the target core for all storage sub-areas and the number of requests for specific sub-areas, we can analyze the core's processing load and the access frequency of sub-areas. Furthermore, by calculating the prefetch weight value, we can more accurately determine which sub-area data is worth pre-reading into the cache, avoiding the resource waste that may be caused by blind pre-reading. When the pre-fetch weight value exceeds the threshold, it means that the access popularity of the current sub-area has reached the preset hot spot standard. At this time, pre-reading can significantly improve the cache hit rate, thereby optimizing the performance of the storage system.
[0058] In an optional implementation, determining whether a comparison result between the current access popularity and the reference access popularity satisfies a pre-reading condition further includes:
[0059] S1, determining a current access time of a current data read request indicated by a current access heat, and a high-frequency reading time period indicated by a reference access heat;
[0060] S2: When the current access time is within the high-frequency reading time period, determine that the pre-reading condition is met.
[0061] In the above S1-S2, the current access time of the current data read request indicated by the current access heat and the high-frequency reading time period indicated by the reference access heat are determined. If the current access time is within the high-frequency reading time period, it is determined that the pre-reading condition is met.
[0062] Optionally, you can analyze the timestamps of historical read requests to perform periodic statistics, time series analysis, etc., to determine the high-frequency read time period, and compare the current access time with the high-frequency read time period identified by the system. If the current access time falls within the high-frequency read time period, it means that the system may currently be in a period with high access demand for a specific region, and the pre-read condition is considered to be met.
[0063] For example, analyzing historical read requests reveals that region_123 is frequently accessed between 9:00 AM and 11:00 AM daily, a high-frequency read period. If a read request for region_123 is received at 10:30 AM, its timestamp indicates that the request time falls within the identified high-frequency read period. Therefore, the system initiates a prefetch operation, fetching data from region_123 that may be accessed by subsequent requests into the cache.
[0064] It can be understood that in this embodiment, after determining that the prefetch weight value meets the target threshold condition, the storage system can intelligently decide the prefetch operation based on the time pattern, thereby improving the cache hit rate and reducing the time consumption of data reading.
[0065] In an optional implementation, determining whether a comparison result between the current access popularity and the reference access popularity satisfies a pre-reading condition further includes:
[0066] S1, determining a current field position of a first data object indicated by a current access heat in a current sub-region, and a reference field position of a historical data object in a historical data read request indicated by a reference access heat in the current sub-region;
[0067] S2: When the current field position and the reference field position meet the target position condition, determine that the pre-reading condition is met.
[0068] In the above S1-S2, the current field position of the first data object indicated by the current access heat in the current sub-area and the reference field position of the historical data object in the historical data read request indicated by the reference access heat in the current sub-area are determined. When the current field position and the reference field position meet the target position condition, it is determined that the pre-reading condition is met.
[0069] Optionally, the reference field position can be determined by analyzing access patterns, such as continuous access, jump access, fixed-interval access, etc. When a new read request arrives, the current field position of the requested first data object in the region is calculated, and the current field position is compared with the reference field position to determine whether the two satisfy a certain specific position relationship, that is, the target position condition.
[0070] Assume that historical analysis shows that read requests for region_456 tend to access data between LBA 32MB and 96MB, that is, the starting point of the data access tendency is located at the 32nd megabyte from the starting point of the storage device, and the end point of the data access tendency is located at the 96th megabyte from the starting point of the storage device. When a new request points to data at LBA 64MB, the system will recognize that the current field position (that is, 64MB) and the historically frequently accessed reference field position (32MB to 96MB) conform to a certain continuous access pattern (such as jump access). Based on this, it can be determined that the pre-read condition is met, so the data adjacent to LBA 64MB is read into the cache to improve the cache hit rate of subsequent accesses.
[0071] It can be understood that in this embodiment, after determining that the prefetch weight value meets the target threshold condition, the storage system can intelligently decide the prefetch operation based on the data access location pattern, thereby improving the cache hit rate and reducing the time consumption of data reading.
[0072] In an optional implementation, adding at least one second data object read in the current sub-region to the cache includes:
[0073] S1, determining the data reading length of the first data object;
[0074] S2, determining the average of the historical average data read length indicated by the reference access heat and the data read length as the reference data read length;
[0075] S3, determining the product of the pre-fetch weight value and the reference data read length as the target data read length;
[0076] S4, determining a current data read start address indicated by a current data read request and a reference data read start address indicated by a historical data read request, wherein the historical data read request is generated before the current moment;
[0077] S5, determining a target step length according to the current data read start address, the reference data read start address, and a reference step length indicated by the reference access heat, wherein the reference step length is used to indicate the interval between the data read start addresses of two consecutive historical data access requests;
[0078] S6, determining the target data reading start address according to the sum of the target step size, the current data reading start address and the data reading length;
[0079] S7, determining at least one second data object to be read in the current sub-region according to the target data reading start address and the target data reading length;
[0080] S8: Add at least one second data object read in the current sub-region to the cache.
[0081] The above steps S1-S8 are described below using an optional implementation method:
[0082] In steps S1 - S3 , the data reading length of the pre-fetched data is determined. Specifically, first, in step S1 , the data reading length io_size of the first data object is determined.
[0083] In step S2 , the reference data read length is determined to be (read_size+io_size) / 2, where read_size is the average read size of the recorded server-side read IO requests.
[0084] In step S3 , the target data read length=prefetch weight value×reference data read length.
[0085] In steps S4-S6, the data reading start address of the pre-fetched data is determined. Specifically: in step S4, the current data reading start address lba_io and the reference data reading start address lba_last are determined, which may be, for example, the start address of the last server-side read io request.
[0086] In step S5, the target step length can be calculated using the following formula:
[0087] Target step size = reference step size × a + (1-a) × (|lba_io–lba_last|). The step size variable is used to record the starting address interval between two server-side read I / O requests. The value of a is [0, 1] and is adjustable.
[0088] In step S6, the target data reading start address can be calculated by the following formula:
[0089] The starting address of the pre-fetched data is = the current data read starting address + the target data read length + the target step length.
[0090] Optionally, in steps S7-S8, determining at least one second data object to be read in the current sub-region according to the target data read start address and the target data read length; and adding the at least one second data object read in the current sub-region to the cache can be specifically implemented by the following steps:
[0091] Update the available resource count: Figure 4 The prefetch resource pool in maintains a variable available_count to record the total number of available resources. Each time the prefetch process is initiated, the count is incremented by 1.
[0092] Request a prefetch IO resource from the prefetch resource pool, fill in the prefetch start address and size information, and submit it to the lower-level module of the IO stack to read data from the disk. After the lower-level module returns the data, it saves this part of the data in the cache.
[0093] Update the available resource count: The prefetch resource pool maintains a variable available_count to record the total number of available resources. After the prefetch process is completed, the count is decremented by 1, indicating that the data has been placed in the cache.
[0094] Through the above implementation method, the target data reading length is dynamically calculated based on the pre-fetch weight value and the reference data reading length, which means that the amount of pre-read data can be automatically adjusted according to the access popularity and historical reading scale, which will not cause waste of cache space and can effectively cover potential large-scale data reading needs. In addition, the determination of the target step length is based on the historical access step length and the current access position, which helps to learn the user data access pattern and predict the next access location, so as to pre-read data more accurately, realize dynamic adjustment of the pre-read data amount and position, prevent excessive resource consumption, and improve the cache hit rate.
[0095] In an optional implementation, determining the access information set matching the target access identifier includes:
[0096] S1, determining the target resource pool identifier that matches the current data read request based on the disk identifier and the target access identifier of the current disk;
[0097] S2, determining an access information set matching the target access identifier from the target resource pool corresponding to the target resource pool identifier, wherein the target resource pool includes access information storage sub-areas corresponding to multiple data read requests respectively, and the access information storage sub-areas store the access information set.
[0098] Optionally, in the above step S1, a target resource pool identifier matching the current data read request is determined according to the disk identifier of the current disk and the target access identifier.
[0099] It is understandable that in a storage system, data and access information may be allocated to different resource pools, each of which may be responsible for a specific set of data or access patterns; a resource pool identifier is jointly determined by the disk identifier and the target access identifier, and the resource pool to be queried is then quickly located.
[0100] For example, the target resource pool identifier can be calculated using the following formula:
[0101] Target access identifier region_id=lba / 32MB;
[0102] Each core is configured with a prefetcher. The core ID is the ID corresponding to the prefetcher (resource pool). core_id = f(region_id) = (region_id + lun_id + N) % core_total, where lba is the starting address of the host read I / O request received by the storage end, lun_id is the storage volume ID (that is, the disk ID of the current disk) to which the host read I / O request belongs, core_total is the total number of cores available for cache prefetching in the storage system, and N is a pressure balancing adjustable variable.
[0103] Further in the above step S2, an access information set matching the target access identifier is determined from the target resource pool corresponding to the target resource pool identifier, wherein the target resource pool includes an access information storage sub-area corresponding to each of multiple data read requests, and the access information storage sub-area stores the access information set.
[0104] Specifically, the prefetch bucket to which the IO belongs can be found according to the region and lun_id to which the IO belongs according to the following algorithm: hash_id=(region_id+lun_id)%hash_total, where hash_total is the total number of prefetch buckets.
[0105] It should be noted that the pre-fetch bucket stores the data corresponding to the request. Figure 4 The hot description in the prefetcher structure diagram shown above records the volume_id and region_id it describes, as well as the access information set. During the search, the volume_id and region_id of the server-side read IO request are compared with the volume_id and region_id recorded in the hotdescription in the prefetch bucket to determine the access information set that matches the target access identifier.
[0106] In this embodiment, the volume is first divided into different regions, and the region_id is involved in the algorithm for dispatching services to different cores, so that the pre-fetched services in each volume can be distributed to all cores in a relatively even manner; and the volume_id (lun_id) is involved in the algorithm for dispatching services to different cores. When the hot data in many volumes of the storage service are located in the same or similar locations, due to the different volume_ids, their services can still be evenly distributed to all cores.
[0107] The role of the N value is to adjust the distribution of prefetch services so that the resources of each prefetcher are fully utilized when uneven distribution of prefetch services still exists after applying the above two mechanisms, resulting in serious consumption of some prefetcher resources and widespread idleness of some prefetcher resources.
[0108] In an optional implementation, determining a target resource pool identifier that matches the current data read request according to the disk identifier of the current disk and the target access identifier includes:
[0109] S1, determining a pressure balancing parameter value according to the number of currently available prefetch resources, wherein the prefetch resources are resources used to perform data pre-reading operations;
[0110] S2, determining the sum of the target access identifier, the disk identifier, and the pressure balance parameter value;
[0111] S3, taking the remainder of the total number of resource pools according to the summation result, and determining the value of the remainder as the target resource pool identifier that matches the current data read request.
[0112] Optionally, in the above step S1 , the pressure balancing parameter value is determined according to the number of currently available pre-fetch resources, wherein the pre-fetch resources are resources used to perform data pre-reading operations.
[0113] For example, the average value and standard deviation of the available resource counts of all pre-readers can be calculated. If the available resource count of the current pre-reader is less than the average value, and the difference between any two pre-readers exceeds the threshold Z, or the standard deviation exceeds the threshold Y, then the updated pressure equalization parameter value N=N+1, where the thresholds Y and Z are both constants and can be adjusted based on experience.
[0114] In the above steps S2-S3, the sum of the target access identifier, disk identifier and pressure balance parameter value is determined; the remainder of the total number of resource pools is taken according to the summation result, and the value of the remainder result is determined as the target resource pool identifier that matches the current data read request.
[0115] Target access identifier region_id=lba / 32MB;
[0116] The target resource pool identifier core_id = f(region_id) = (region_id + lun_id + N) % core_total, where lba is the starting address of the host read IO request received by the storage end, lun_id is the storage volume _id (that is, the disk identifier of the current disk) to which the host read IO request received by the storage end belongs, core_total is the total number of cores available for cache prefetching in the storage system. Each core is configured with a prefetcher, and the core identifier is the identifier corresponding to the prefetcher (resource pool). N is an adjustable variable for pressure balancing.
[0117] Through the above-mentioned implementation method recorded in this application, the system can dynamically adjust the pressure balancing parameters according to the number of currently available pre-fetch resources. The above-mentioned calculation method increases the randomness of allocation to specific resource pools, avoids the concentration of requests in a few resource pools, and causes local overload, and ensures that pre-read operations can be evenly distributed among all available resource pools.
[0118] In an optional implementation, the root determines, from the target resource pool corresponding to the target resource pool identifier, a set of access information that matches the target access identifier, including:
[0119] S1, determining the sum of the target access identifier and the disk identifier;
[0120] S2, taking the remainder of the total number of access information storage sub-regions in the target resource pool according to the summation result, and determining the value of the remainder as the target region identifier indicating the access information storage sub-region;
[0121] S3: Determine an access information set in the target access information storage sub-area indicated by the target area identifier.
[0122] The above S1-S3 are described with a complete implementation method:
[0123] The following algorithm can be used to find the prefetch bucket to which the IO belongs based on its region_id (target access identifier) and lun_id (disk identifier):
[0124] region_id=lba / 32MB, where lba is the starting address of the host read IO request received by the storage end;
[0125] Target region ID (prefetch bucket ID) hash_id = (region_id + lun_id) % hash_total, where lun_id is the storage volume ID (i.e., disk ID) to which the host read IO request received by the storage end belongs, and hash_total is the total number of prefetch buckets in the target resource pool.
[0126] It should be noted that the prefetch bucket stores the heat descriptor corresponding to the request, such as Figure 4 As shown, the hot description records the volume_id and region_id it describes, as well as the access information set. During the search, the volume_id and region_id to which the server-side read IO request belongs are compared with the volume_id and region_id recorded in the hot description in the prefetch bucket, thereby determining the access information set that matches the target access identifier.
[0127] The data pre-reading method described in this application is described below using a complete implementation method:
[0128] Prefetch service distribution module: Figure 3 As shown in the figure, the storage capacity in the storage system is managed at the volume granularity. The cache module divides any storage volume into multiple regions at a granularity of 32MB (it can also be any other granularity, without limitation, 32M is used as an example). For example, a volume with a capacity of 100GB can be divided into 100×1024 / 32=3200 regions. The prefetch services of these regions are dispatched to different cores and prefetchers according to the following algorithm. Each core is configured with a corresponding prefetcher to achieve higher concurrency.
[0129] region_id=lba / 32MB;
[0130] core_id=f(region_id)=(region_id+lun_id+N)%core_total, where lba is the starting address of the host read IO request received by the storage end, lun_id is the storage volume ID to which the host read IO request received by the storage end belongs, core_total is the total number of cores available for cache prefetching in the storage system, and N is a pressure balancing adjustable variable.
[0131] This application first divides the volume into different regions, and involves region_id in the algorithm of dispatching services to different cores, so that the pre-fetched services in each volume can be distributed to all cores in a relatively even manner.
[0132] Including volume_id (equivalent to lun_id) in the algorithm for distributing services to different cores can ensure that evenly distributed services can be distributed to all cores when hot data in multiple volumes of a storage service are located in the same or similar locations due to the different volume_ids.
[0133] The significance of the N value lies in that, even after applying the above two mechanisms, if uneven distribution of prefetch services still occurs, resulting in severe consumption of some prefetcher resources and widespread idleness of some prefetcher resources, this value can be adjusted to adjust the distribution of prefetch services and fully utilize the resources of each prefetcher.
[0134] The prefetcher algorithm and resource module executes: S1, the prefetcher maintains a core_read_count variable to record the number of all host-side read IO requests processed by the core. When a new server-side read IO request arrives, the variable is incremented by 1.
[0135] S2: When a new server-side read IO request arrives, the prefetch bucket to which the IO belongs is first found according to the region and lun_id to which the IO belongs according to the following algorithm: hash_id = (region_id + lun_id) % hash_total, where hash_total is the total number of prefetch buckets.
[0136] S3, searches from the pre-fetch bucket (hash bucket) to check whether there is a corresponding heat descriptor, that is, Figure 4 The hot description in the prefetch bucket records the volume_id and region_id described in the hot description. When searching, the volume_id and region_id of the server-side read IO request are compared with the volume_id and region_id recorded in the hot description in the prefetch bucket.
[0137] S4-1: If the hot descriptor corresponding to the current server-side read IO request is found in the prefetch bucket (hash bucket), the information is directly updated in the hot description.
[0138] S4-2, if the heat descriptor corresponding to the current server-side read IO request is not found in the prefetch bucket (hash bucket), then Figure 4 Take the tail heat descriptor from the heat chain shown in , remove it from the current prefetch bucket, update the volume_id and region_id recorded in it to the volume_id and region_id of the current server-side read IO request, and record the current value of core_read_count, recorded as core_region_read_count, reset the remaining data to zero, and then insert it into the prefetch bucket to which it belongs according to the currently recorded volume_id and region_id.
[0139] S5, region read count update: The hot description maintains a read_count variable, which is used to record the number of times the region is read by the server. Each time a server-side read IO request belonging to the region arrives, its value is incremented by 1.
[0140] S6, step record update: hot description maintains a step_read variable, which is used to record the starting lba interval between two server-side read IO requests. The update algorithm is as follows:
[0141] step_read=step_read×a+(1-a)×(|lba_io–lba_last|), where a is in the range [0,1] and can be adjusted according to actual conditions. lba_io is the starting address of this server-side read IO request, and lba_last is the starting address of the last server-side read IO request. When a region is read for the first time, step_read is 0 and is not updated. Only lba_last is updated.
[0142] S7, read length update: The hot description maintains a read_size variable to record the average read size of the server-side read IO request. The update algorithm is as follows:
[0143] read_size=(read_size+io_size) / 2, where io_size is the size of the data requested by the server for this read IO request.
[0144] S8, prefetch weight calculation: Calculate the prefetch weight according to the following algorithm:
[0145] weight = read_count × W / (core_read_count - core_region_read_count), where W is a constant that can be set based on experience.
[0146] S9, if the prefetch weight > X, initiate the prefetch process to prefetch data into the cache. The prefetched data size is size = weight × read_size; the starting address of the prefetched data is lba = lba_io + io_size + step_read. The above X is also a constant and can be set based on experience.
[0147] S10, remove the hot description from the current position of the heat chain and add it to the head of the heat chain. This step and step S4-2 together implement the least recently used algorithm.
[0148] Backend prefetch module: This module uses the resources each time a prefetch process is initiated. Each time the prefetch algorithm and resource module initiate a prefetch process through prefetch algorithm calculation, the following operations are performed:
[0149] S1, update the available resource count: the pre-fetch resource pool maintains a variable available_count to record the total number of available resources. Each time the pre-fetch process is initiated, the count is incremented by 1.
[0150] S2, apply for a prefetch io resource from the prefetch resource pool, fill in the prefetch start address and size and other information, and then submit it to the lower module of the IO stack (such as Figure 3 Data input and output stack, other modules such as disk redundant array, thin stack, etc.) to store disk read data.
[0151] S3: After the lower-level module returns the data, it saves the data in the cache.
[0152] S4, update the available resource count: the pre-fetch resource pool maintains a variable available_count to record the total number of available resources. After the pre-fetch process is completed, the count is decremented by 1.
[0153] S5, calculate the average and standard deviation of the available resource counts of all prefetchers. If the available resource count of this prefetcher is less than the average and the difference exceeds the threshold Z, or the standard deviation exceeds the threshold Y, update the adjustable variable N=N+1, where the thresholds Y and Z are both constants and can be adjusted based on experience.
[0154] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0155] According to another aspect of the embodiment of the present application, a data pre-reading device for implementing the above-mentioned data pre-reading method is also provided. Figure 5 As shown, the device includes:
[0156] A first determining unit 502 determines a target access identifier based on a storage address included in the current data read request, where the target access identifier is used to indicate a current sub-region among multiple storage sub-regions included in the current disk, and the current sub-region is used to store a first data object matching the current data read request;
[0157] A second determining unit 504 determines an access information set matching the target access identifier, wherein the access information set includes a plurality of access request description information corresponding to the plurality of storage sub-regions, the access request description information indicating access popularity of the corresponding storage sub-regions;
[0158] The pre-reading unit 506 adds at least one second data object read in the current sub-area to the cache when the access information set includes current request description information corresponding to the current sub-area and the comparison result between the current access heat indicated by the current request description information and the reference access heat meets the pre-reading condition.
[0159] Optionally, the above-mentioned pre-reading unit 506 includes: a third determination unit, used to determine the first request number of data read requests for data in multiple storage sub-areas processed by the target core, and the second request number of data read requests for data in the current sub-area processed by the target core, wherein the target core is used to process the current data read request at the current moment; determining the difference between the first request number and the second request number as the first difference indicated by the reference access heat; determining the pre-fetch weight value based on the ratio of the total number of accesses to the current sub-area indicated by the current access heat to the first difference; when the pre-fetch weight value is greater than the target threshold, determining that the comparison result between the current access heat and the reference access heat meets the pre-read condition.
[0160] Optionally, the above-mentioned third determination unit is used to determine the current access time of the current data read request indicated by the current access heat, and the high-frequency reading time period indicated by the reference access heat; when the current access time is within the high-frequency reading time period, it is determined that the pre-reading condition is met.
[0161] Optionally, the above-mentioned third determination unit is also used to determine the current field position of the first data object indicated by the current access heat in the current sub-area, and the reference field position of the historical data object in the historical data reading request indicated by the reference access heat in the current sub-area; when the current field position and the reference field position meet the target position condition, it is determined that the pre-reading condition is met.
[0162] Optionally, the above-mentioned pre-reading unit 506 is also used to determine the data read length of the first data object; determine the historical average data read length indicated by the reference access heat and the average of the data read length as the reference data read length; determine the product of the pre-fetch weight value and the reference data read length as the target data read length; determine the current data read start address indicated by the current data read request, and the reference data read start address indicated by the historical data read request, wherein the historical data read request is generated before the current moment; determine the target step length according to the current data read start address, the reference data read start address and the reference step length indicated by the reference access heat, wherein the reference step length is used to indicate the interval between the data read start addresses of two consecutive historical data access requests; determine the target data read start address according to the target step length, the current data read start address and the sum of the data read length; determine at least one second data object to be read in the current sub-area according to the target data read start address and the target data read length; and add at least one second data object read in the current sub-area to the cache.
[0163] Optionally, the above-mentioned second determination unit 504 also includes a fourth determination module, which is used to determine the target resource pool identifier that matches the current data read request based on the disk identifier and the target access identifier of the current disk; determine the access information set that matches the target access identifier from the target resource pool corresponding to the target resource pool identifier, wherein the target resource pool includes an access information storage sub-area corresponding to each of the multiple data read requests, and the access information storage sub-area stores the access information set.
[0164] Optionally, the above-mentioned fourth determination module is used to determine the pressure balancing parameter value based on the number of currently available pre-fetch resources, wherein the pre-fetch resources are resources used to perform data pre-reading operations; determine the sum of the target access identifier, the disk identifier and the pressure balancing parameter value; take the remainder of the total number of resource pools according to the sum result, and determine the value of the remainder result as the target resource pool identifier that matches the current data read request.
[0165] Optionally, the above-mentioned fourth determination module is also used to determine the sum of the target access identifier and the disk identifier; take the remainder of the total number of access information storage sub-areas in the target resource pool according to the sum result, and determine the value of the remainder as the target area identifier indicating the access information storage sub-area; and determine the access information set in the target access information storage sub-area indicated by the target area identifier.
[0166] For the description of the features in the embodiment corresponding to the data pre-reading device, reference can be made to the relevant description of the embodiment corresponding to the data pre-reading method, which will not be repeated here.
[0167] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned data pre-reading method embodiments.
[0168] The electronic device used in this application can be a terminal device or a server. This embodiment takes the electronic device as a mobile phone or a computer as an example. Figure 6 As shown, the electronic device includes a memory 602 and a processor 604. The memory 602 stores a computer program, and the processor 604 is configured to execute the steps in any of the above method embodiments through the computer program.
[0169] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0170] Alternatively, those skilled in the art will appreciate that Figure 6 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 6 It does not limit the structure of the above electronic device. For example, the electronic device may also include Figure 6 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 6 Different configurations shown.
[0171] Among them, the memory 602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the data pre-reading method and device in the embodiment of the present application. The processor 604 executes various functional applications by running the software programs and modules stored in the memory 602, that is, realizing the above-mentioned data pre-reading method. The memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 602 may further include a memory remotely located relative to the processor 604, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 602 can be used specifically but not limited to store information such as signals and data. As an example, such as Figure 6As shown, the memory 602 may include, but is not limited to, the first determination unit 502, the second determination unit 504, and the pre-reading unit 506 in the data pre-reading device. In addition, it may also include, but is not limited to, other module units in the data pre-reading device, which will not be repeated in this example.
[0172] Optionally, the transmission device 606 is configured to receive or transmit data via a network. Specific examples of the aforementioned network may include wired networks and wireless networks. In one embodiment, the transmission device 606 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to enable communication with the Internet or a local area network. In one embodiment, the transmission device 606 is a radio frequency (RF) module configured to communicate with the Internet wirelessly.
[0173] In addition, the electronic device further includes: a display 608; and a connection bus 610 for connecting various module components in the electronic device.
[0174] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes via network communication. The nodes may form a point-to-point network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the point-to-point network.
[0175] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned data pre-reading method embodiments when running.
[0176] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0177] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned data pre-reading method embodiments are implemented.
[0178] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned data pre-reading method embodiments are implemented.
[0179] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0180] The above is a detailed introduction to a data pre-reading method, storage medium, and electronic device provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A data pre-reading method, characterized in that: include: Determining a target access identifier according to a storage address included in a current data read request, where the target access identifier is used to indicate a current sub-area among a plurality of storage sub-areas included in a current disk, the current sub-area being used to store a first data object matching the current data read request; Determining an access information set that matches the target access identifier, wherein the access information set includes a plurality of access request description information corresponding to the plurality of storage sub-regions, respectively, and the access request description information is used to indicate access popularity of the corresponding storage sub-regions; When the access information set includes current request description information corresponding to the current sub-area, and the comparison result between the current access heat indicated by the current request description information and the reference access heat meets the pre-reading condition, at least one second data object read in the current sub-area is added to the cache.
2. The method according to claim 1, characterized in that The comparison result between the current access heat and the reference access heat satisfies the pre-reading conditions, including: determining a first request quantity of data read requests for data in the plurality of storage sub-regions processed by a target core, and a second request quantity of data read requests for data in the current sub-region processed by the target core, wherein the target core is currently used to process the current data read request; determining a difference between the first number of requests and the second number of requests as a first difference of the reference access heat indication; determining a prefetch weight value according to a ratio of a total number of visits to the current sub-region indicated by the current access heat to the first difference; In a case where the pre-fetch weight value is greater than a target threshold, it is determined that the comparison result between the current access heat and the reference access heat satisfies the pre-reading condition.
3. The method according to claim 2, characterized in that Determining whether the comparison result between the current access popularity and the reference access popularity satisfies the pre-reading condition further includes: Determine a current access time of the current data read request indicated by the current access heat, and a high-frequency reading time period indicated by the reference access heat; In a case where the current access time is within the high-frequency reading time period, it is determined that the pre-reading condition is satisfied.
4. The method according to claim 2, characterized in that Determining whether the comparison result between the current access popularity and the reference access popularity satisfies the pre-reading condition further includes: Determine that the current access heat indicates a current field position of the first data object in the current sub-region, and the reference access heat indicates a reference field position of a historical data object in the historical data read request in the current sub-region; In a case where the current field position and the reference field position satisfy a target position condition, it is determined that the pre-reading condition is satisfied.
5. The method according to claim 2, characterized in that The adding the at least one second data object read from the current sub-region to the cache includes: Determining a data read length of the first data object; Determine the average of the historical average data read length indicated by the reference access heat and the data read length as a reference data read length; Determine a target data read length as a product of the pre-fetch weight value and the reference data read length; Determining a current data read start address indicated by the current data read request and a reference data read start address indicated by a historical data read request, wherein the historical data read request was generated before the current moment; determining a target step length according to the current data read start address, the reference data read start address, and a reference step length indicated by the reference access heat, wherein the reference step length is used to indicate an interval between data read start addresses of two consecutive historical data access requests; Determine a target data reading start address according to a sum of the target step size, the current data reading start address, and the data reading length; Determining, in the current sub-area according to the target data reading start address and the target data reading length, at least one second data object to be read; Adding at least one second data object read from the current sub-region to the cache.
6. The method according to claim 1, wherein The determining of the access information set matching the target access identifier includes: Determining a target resource pool identifier that matches the current data read request according to the disk identifier of the current disk and the target access identifier; The access information set matching the target access identifier is determined from the target resource pool corresponding to the target resource pool identifier, wherein the target resource pool includes access information storage sub-areas corresponding to multiple data read requests respectively, and the access information set is stored in the access information storage sub-areas.
7. The method according to claim 6, characterized in that The determining, according to the disk identifier of the current disk and the target access identifier, a target resource pool identifier that matches the current data read request includes: Determining a pressure balancing parameter value according to the number of currently available prefetch resources, wherein the prefetch resources are resources used to perform a data pre-reading operation; Determine a sum of the target access identifier, the disk identifier, and the pressure equalization parameter value; The remainder of the total number of the resource pools is taken according to the summation result, and the value of the remainder result is determined as the target resource pool identifier that matches the current data read request.
8. The method according to claim 6, characterized in that The determining, from the target resource pool corresponding to the target resource pool identifier, the access information set matching the target access identifier comprises: Determine a sum of the target access identifier and the disk identifier; Taking the remainder of the total number of the access information storage sub-areas in the target resource pool according to the summation result, and determining the value of the remainder as the target area identifier indicating the access information storage sub-area; The access information set is determined in the target access information storage sub-area indicated by the target area identifier.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the data pre-reading method according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data pre-reading method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Data prefetching method and device
CN112199304A
Data access method and device, storage medium and electronic equipment
CN117032596A
Data prefetching method and device
CN119557240A
Cache management method, electronic device, storage medium and program product
CN119620961A
Cited By
A cache loading method, a storage device
CN122489452A