Data Processing Method, Storage Medium, Electronic Device and Program Product

By using access history records to predict future data access in SSDs, the method enhances pre-read accuracy and hit rates, improving system responsiveness and reducing delays.

CN119902720BActive Publication Date: 2025-07-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510379918.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

In the prior art, the read-out mechanism of the solid-state drive cannot accurately predict the data that may be accessed in the future, resulting in a low read-out hit rate.

Method used

By maintaining the access history table of data blocks and data pages, recording access frequency and weights, dynamically adjusting the read-out position selection, and priority loading of high-weight data blocks and data pages into the cache.

Benefits of technology

Significantly improve the read preview hit rate, reduce invalid read preview operations, and improve system response speed and overall performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902720B_ABST
    Figure CN119902720B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, a storage medium, an electronic device, and a program product, which relate to the field of computer technologies. The method includes: first, extracting data access information from a solid-state drive to determine a currently accessed data block; then, obtaining an access history record table corresponding to the currently accessed data block, where the access history record table includes a data block access history record table and a data page access history record table; determining target prefetch data based on the weights of multiple data blocks and multiple data pages in the access history record table; and if it is detected that the position of the target prefetch data is not in the prefetch position linked list, loading the target prefetch data into the cache of the solid-state drive. Compared with the current existing technologies, by maintaining the access history record table and recording the access frequencies and weights of each data block and data page, the present application can more accurately predict the next most likely accessed data block and data page, thereby improving the prefetch hit rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data processing method, a storage medium, an electronic device, and a program product. Background Art

[0002] With the widespread use of solid state drives (SSDs) in computer systems, their performance has become one of the important indicators for measuring the overall performance of the system. By using the prefetching technology to load in advance the data that may be accessed in the future into the storage medium of the solid state drive, the performance of the solid state drive can be further improved.

[0003] Currently, the prefetching mechanisms in related technologies usually perform actual reading and prefetching operations based on a simple proportional strategy, and directly use the subsequent data pages of the actually read data as prefetch data. However, this method cannot accurately predict the prefetch data that may be accessed in the future, thereby resulting in a low prefetch hit rate. Summary of the Invention

[0004] The present disclosure provides a data processing method, a storage medium, an electronic device, and a program product. Its main purpose is to solve the problem that the related technologies cannot accurately predict the prefetch data that may be accessed in the future, thereby resulting in a low prefetch hit rate.

[0005] In a first aspect, the present application provides a data processing method, including:

[0006] Extracting data access information from a solid state drive to determine the currently accessed data block;

[0007] Obtaining an access history record table corresponding to the currently accessed data block, where the access history record table includes a data block access history record table and a data page access history record table. The data block access history record table is used to record the accessed data blocks and the weights of the accessed data blocks, and the data page access history record table is used to record the accessed data pages and the weights of the accessed data pages. The weights are used to reflect the probabilities of the data blocks and data pages in the solid state drive being accessed;

[0008] Determining target prefetch data based on the weights of multiple data blocks and the weights of multiple data pages in the access history record table;

[0009] If it is detected that the position of the target prefetch data is not in the prefetch position linked list, then loading the target prefetch data into the cache of the solid state drive.

[0010] In a second aspect, the present application provides a data processing device, including:

[0011] An extraction module configured to extract data access information from a solid state drive to determine the currently accessed data block;

[0012] An acquisition module, configured to acquire an access history record table corresponding to a currently accessed data block. The access history record table includes a data block access history record table and a data page access history record table. The data block access history record table is used to record the accessed data blocks and the weights of the accessed data blocks. The data page access history record table is used to record the accessed data pages and the weights of the accessed data pages. The weights are used to reflect the access probabilities of the data blocks and data pages in the solid-state drive.

[0013] A determination module, configured to determine target prefetch data based on the weights of multiple data blocks and the weights of multiple data pages in the access history record table.

[0014] A loading module, configured to load the target prefetch data into the cache of the solid-state drive if it is detected that the location of the target prefetch data is not in the prefetch location linked list.

[0015] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect is implemented.

[0016] In a fourth aspect, the present application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, the method of the first aspect is implemented.

[0017] In a fifth aspect, the present application provides a computer program product, on which a computer program is stored. When the computer program is executed by a processor, the method of the first aspect is implemented.

[0018] The data processing method, storage medium, electronic device, and program product provided by the present disclosure, wherein the method includes: First, extract data access information from a solid-state drive to determine the currently accessed data block; obtain the access history record table corresponding to the currently accessed data block, and the access history record table includes a data block access history record table and a data page access history record table. The data block access history record table is used to record the accessed data blocks and the weights of the accessed data blocks, and the data page access history record table is used to record the accessed data pages and the weights of the accessed data pages. The weight is used to reflect the probability of the data blocks and data pages in the solid-state drive being accessed; based on the weights of multiple data blocks and multiple data pages in the access history record table, determine the target prefetch data; if it is detected that the position of the target prefetch data is not in the prefetch position linked list, then load the target prefetch data into the cache of the solid-state drive. Compared with the current existing technologies, the present application can more accurately predict the next most likely accessed data blocks and data pages by maintaining the data block access history record table and the data page access history record table, recording the access frequencies and weights of each data block and data page, and then determine the target prefetch data. This dynamic adjustment mechanism based on the access pattern can significantly improve the hit rate of prefetching and reduce invalid prefetch operations. By preloading the target prefetch data into the cache of the solid-state drive in advance, when the user actually accesses this data, it can be directly read from the cache, thereby significantly reducing the data access latency and improving the system response speed.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 The flowchart shows a data processing method provided by an embodiment of the present application;

[0022] Figure 2 The flowchart shows another data processing method provided by an embodiment of the present application;

[0023] Figure 3 The flowchart shows an example process provided by an embodiment of the present application;

[0024] Figure 4Shows a schematic flowchart of an example provided by an embodiment of the present application;

[0025] Figure 5 Shows a schematic flowchart of an example provided by an embodiment of the present application;

[0026] Figure 6 Shows a schematic flowchart of an example provided by an embodiment of the present application;

[0027] Figure 7 Shows a schematic structural diagram of a data processing device provided by an embodiment of the present application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0029] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0030] Currently, traditional prefetching algorithms generally only allocate the number of actual reads and prefetches proportionally. For example, 10 actual reads are accompanied by 1 prefetch, and this prefetch is generally the subsequent data page of the actual read. This mechanism has problems such as low hit rate and limited size of the cache prefetch location buffer. This is mainly because traditional prefetching algorithms lack the ability to dynamically learn and adjust to the actual access pattern, resulting in inaccurate selection of the prefetch location. Therefore, designing a prefetching algorithm that can adaptively learn according to the actual access pattern to improve the prefetch hit efficiency as much as possible under the premise of limited memory size has become an urgent problem to be solved.

[0031] To improve the technical problem in the related art that the prefetch data that may be accessed in the future cannot be accurately predicted, resulting in a low prefetch hit rate. This embodiment provides a data processing method, as Figure 1 shown, this method includes the following steps:

[0032] Step 101, extract data access information from the solid-state drive and determine the currently accessed data block.

[0033] Exemplarily, system tools or dedicated software can be used to obtain the access logs of the SSD or collect the access records of data blocks. For example, the identifier of the currently accessed data block and related access information can be parsed from the access logs, so as to accurately track which data blocks are being accessed, providing an important basis for subsequent cache management, prefetching strategy optimization, and overall performance improvement.

[0034] Step 102: Obtain the access history record table corresponding to the currently accessed data block.

[0035] Among them, the access history record table includes a data block access history record table and a data page access history record table. The data block access history record table (Blk_AccessHistory) is used to record the accessed data blocks and the weights of the accessed data blocks (recorded in the weight table Blk_Weights), and the data page access history record table (Pg_AccessHistory) is used to record the accessed data pages and the weights of the accessed data pages (recorded in the weight table Pg_Weights). The weights are used to reflect the access probabilities of data blocks and data pages in the solid-state drive.

[0036] Exemplarily, the data block access history record table is a table or data structure that records data block access information. It usually contains the unique identifier (Identifier, ID) of each data block and the timestamp (Timestamp) of its most recent or multiple accesses, which can be used to track which data blocks have been accessed and the specific time points when these accesses occurred, helping to identify frequently accessed data blocks and thus providing a basis for cache management; the structure of the data page access history record table is similar to that of the data block access history record table and is used to record access information at the finer-grained data page level within the data block.

[0037] In some examples, as the system running time increases, the access history record table and the weight table are continuously updated, enabling the prefetching strategy to be adaptively adjusted to adapt to new access patterns. This adaptability allows the system to maintain high performance when facing dynamically changing data access requirements. By introducing the weight mechanism, the system can dynamically adjust the prefetching strategy according to the access frequency and access pattern, prefetching those data blocks and data pages that are more likely to be accessed first, thereby improving the prefetching hit rate.

[0038] Step 103: Determine the target prefetch data based on the weights of multiple data blocks and multiple data pages in the access history record table.

[0039] Exemplarily, the top N high-weight data blocks are selected from the data block access history record table according to the weights, and then for each selected data block, the top M high-weight data pages are similarly filtered according to the weights from the corresponding data page access history record table. These selected data pages are identified as target prefetch data. By analyzing and sorting the access weights of the data blocks and data pages, the data most likely to be accessed can be accurately selected as the target prefetch data, thereby improving the system performance and response speed.

[0040] Step 104: If it is detected that the position of the target prefetch data is not in the prefetch position linked list, then load the target prefetch data into the cache of the solid-state drive.

[0041] In some examples, the prefetch position linked list is used to store the data position information predicted by the system to be possibly accessed in the future, including identifiers such as data block IDs and data page IDs. According to the prefetch position linked list, it is possible to effectively track which data has been prefetched, its prefetch order, and status, ensuring that these data can be quickly located when actually needed. Inserting the position of the target prefetch data into the prefetch position linked list can optimize the cache usage and improve the prefetch hit rate.

[0042] Exemplarily, in this embodiment, an access history record table is maintained through an adaptive learning mechanism to track the access frequency of each data block, and a weight is defined for each data block to reflect the probability of its being accessed, so as to dynamically adjust the selection of the prefetch position according to the actual access pattern. During the prefetch process, the data block with the highest weight is preferentially selected for prefetching, ensuring that the system can intelligently predict and load the data most likely to be accessed, and improving the prefetch hit rate and data access efficiency.

[0043] In some embodiments, through the above data processing flow, the selection of the prefetch position is dynamically adjusted according to the actual access pattern, so as to maintain a high prefetch hit rate. Specifically, the algorithm statistically analyzes the access history records of the data blocks and data pages, dynamically adjusts the weights according to the access frequency, and identifies the data blocks and data pages most likely to be accessed. These high-weight data are pre-loaded into the cache in advance through the adaptive learning algorithm to accelerate the response speed of subsequent access requests. The method of this embodiment not only optimizes the data access efficiency, but also significantly improves the overall performance of the SSD, enabling the system to reduce the waiting time while increasing the speed and response ability of data processing. The method of this embodiment effectively enhances the performance of the storage system and the user experience by intelligently predicting and prefetching data.

[0044] Compared with the related technologies, in this embodiment, data access information is first extracted from the solid-state drive to determine the currently accessed data block; an access history record table corresponding to the currently accessed data block is obtained. The access history record table includes a data block access history record table and a data page access history record table. The data block access history record table is used to record the accessed data blocks and the weights of the accessed data blocks, and the data page access history record table is used to record the accessed data pages and the weights of the accessed data pages. The weights are used to reflect the probabilities of the data blocks and data pages in the solid-state drive being accessed; based on the weights of multiple data blocks and multiple data pages in the access history record table, target prefetch data is determined; if it is detected that the position of the target prefetch data is not in the prefetch position linked list, the target prefetch data is loaded into the cache of the solid-state drive. Compared with the current existing technologies, in this embodiment, by maintaining the data block access history record table and the data page access history record table, recording the access frequencies and weights of each data block and data page, the next most likely accessed data block and data page can be predicted more accurately, and then the target prefetch data can be determined. This dynamic adjustment mechanism based on the access pattern can significantly improve the hit rate of prefetching and reduce invalid prefetch operations. By preloading the target prefetch data into the cache of the solid-state drive in advance, when the user actually accesses this data, it can be directly read from the cache, thus significantly reducing the data access latency and improving the system response speed.

[0045] Further, as a refinement and extension of the above embodiment, to specifically illustrate the data processing process, optionally, step 103 may specifically further include: screening out the first preset number of data blocks from high to low according to the weights of multiple data blocks in the data block access history record table; according to the weights of multiple data pages in the data page access history record table, for multiple data blocks among the first preset number of data blocks, screening out the second preset number of data pages from high to low; determining the data in the second preset number of data pages as the target prefetch data.

[0046] Exemplarily, according to the weights of each data block in the data block access history record table, the first preset number (for example, the first 30) of data blocks are screened out from high to low; then, for these selected data blocks, according to the data page weights in their corresponding data page access history record tables, the second preset number (for example, the first 30 under each data block) of data pages are further screened out from high to low. Finally, the data in these selected data pages is determined as the target prefetch data to optimize the data access efficiency and ensure that the prefetch operation focuses on the data most likely to be accessed, thereby improving the overall system performance and reducing unnecessary read overhead.

[0047] In some embodiments, before performing the specific prefetch operation, it is also necessary to check the prefetch linked list to ensure that the prefetch operation is only performed on data blocks and data pages that have not been read, thereby avoiding redundant prefetching of data that already exists in the cache, improving the overall reading efficiency, and reducing unnecessary resource consumption. This can intelligently optimize the data access pattern and enhance the system performance.

[0048] Optionally, the method of this embodiment may further include: in response to a read request for a data block, updating the identifier and access timestamp of the first data block to the access history record table; updating the weight of the first data block based on the number of times the first data block is accessed, and updating the weight of the first data page corresponding to the first data block based on the number of times the first data page is accessed.

[0049] In some examples, when a read request is received, the system first records the identifier (block_id) of the currently accessed data block and its access timestamp (timestamp) into the initialized data block access history table (Blk_AccessHistory[block_id] = timestamp). Then, update the weight of this data block. The initial weight is set to 1. Whenever this data block is accessed again, its weight is incremented by a certain score (e.g., +1). Specifically, it is determined whether to increase the existing weight or set a new weight by checking whether the data block exists in the weight table. In addition, it is necessary to find the data page access table corresponding to the data block identifier, record the current data page, and update the weight of the data page in the same way to track the access frequency and importance of the data page. This process helps to optimize the data access and storage strategy and improve the read operation efficiency.

[0050] Optionally, the method of this embodiment may further include: if it is detected that the position of the target prefetch data is in the prefetch position linked list, increasing the prefetch hit count of the position of the target prefetch data.

[0051] Exemplarily, if it is detected that the position of the target prefetch data is in the prefetch position linked list, it indicates that the prefetch has been performed on this position before. At this time, the prefetch hit count should be increased to record the access frequency of the prefetch data. Ensure that frequently accessed data can be kept in the cache, thereby improving the cache utilization rate and prefetch hit rate, and ultimately enhancing the system response speed and overall performance. The system can optimize the cache management by updating the prefetch hit count.

[0052] Optionally, before determining the target prefetch data based on the weights of multiple data blocks and multiple data pages in the access history record table, the method of this embodiment may specifically further include: querying the access history record table to obtain the last access time of the currently accessed data block; if the difference between the last access time and the current time is greater than a preset interval period, stop obtaining the target prefetch data; correspondingly, step 103 may specifically further include: if the difference between the last access time and the current time is less than or equal to the preset interval period, determine the target prefetch data based on the weights of multiple data blocks and multiple data pages in the access history record table.

[0053] Exemplarily, a timer can be created and set to trigger the prefetch process periodically every 5 seconds. At each trigger, first check the data block access history record table to confirm whether the most recent access occurred within 5 seconds. If it exceeds 5 seconds, skip the current prefetch election process to avoid invalid prefetching.

[0054] Optionally, step 104 may specifically further include: inserting the target prefetch data into the prefetch position linked list and adding the target prefetch data to the prefetch buffer table.

[0055] Exemplarily, insert the specific positions of these target prefetch data (such as data block ID and data page ID) into the prefetch position linked list to record their prefetch order and status. At the same time, add the actual data content or pointers pointing to these data to the prefetch buffer to ensure that these potentially accessed data can be quickly obtained directly from the cache, reducing subsequent access latency, not only optimizing the cache usage efficiency, but also improving the prefetch hit rate and the overall performance of the system.

[0056] In some examples, through the SSD prefetch algorithm based on adaptive learning, under the condition of limited memory size, it is possible to dynamically adjust the prefetch strategy, improve the prefetch hit efficiency, reduce invalid read operations, and thus significantly enhance the overall disk performance and user interaction experience. This method effectively optimizes the cache usage and enhances the data processing speed and the overall response ability of the system.

[0057] As a refinement of this embodiment, before loading the target prefetch data into the cache of the solid-state drive, the space of the prefetch position buffer can be allocated in the following ways but not limited to, such as Figure 3 as shown Figure 3 is a flowchart of a data processing method provided by an embodiment of the present disclosure, including:

[0058] Step 201, determine whether there is enough space in the cache of the solid-state drive to store the target prefetch data.

[0059] Exemplarily, in the cache management policy, first initialize the pre-read buffer (PreReadBuffer), set the total cache capacity (TotalCapacity) to the maximum cache size (MAX_CACHE_SIZE) (for example, set it to 16K * 30 * 30, which means each data page size is 16KB, and for each of the first 30 data blocks selected for pre-reading, select the first 30 data pages again), initialize the currently used cache size (CurrentSize) to 0, and initialize the pre-read position linked list (PreReadAddr) to optimize the use of cache space, improve the data pre-reading hit rate and system performance.

[0060] Step 202: If there is not enough space in the cache, sort the nodes in the pre-read position linked list according to the access frequency, and delete the nodes with an access frequency lower than the preset threshold.

[0061] Optionally, if there is enough space in the cache, load the target pre-read data into the cache of the solid-state drive.

[0062] Exemplarily, when the adaptive learning algorithm determines the upcoming pre-read position, first check whether this position already exists in the pre-read position linked list. If it exists, increase the pre-read hit count (access frequency) of this position (Pre_cnt++); if it does not exist, initiate a pre-read request. After the pre-read data is returned, before putting it into the cache, the system will first check whether there is enough space currently. If there is enough space, the new node will be inserted into the pre-read position linked list, and at the same time set its pre-read hit count Pre_cnt = 1, and add the pre-read data to the pre-read buffer. If the space is insufficient, sort the existing nodes according to the access frequency, and delete the node with the lowest access frequency to free up space, and then add the new pre-read data and its node to the pre-read buffer and the pre-read position linked list. In this way, by dynamically managing the data in the cache, it is ensured that the data with high access frequency is retained, thereby improving the overall performance and pre-reading efficiency of the system.

[0063] Exemplarily, in this embodiment, through the intelligent cache management policy, on the premise of limited memory size, reasonably allocate the space of the pre-read position buffer and release space for the data blocks that are more likely to be accessed by regularly cleaning the data blocks with low access frequency, maintaining the freshness and relevance of the cache content. At the same time, an optimization algorithm is adopted to ensure the efficient operation of the cache management policy, ensuring that the data blocks with high access frequency are retained, thereby improving the overall cache utilization rate and system performance.

[0064] Optionally, the method of this embodiment may specifically further include: regularly cleaning the access history record table and / or the pre-read position linked list based on a preset time interval.

[0065] Exemplarily, by performing regular cleaning, low-frequency or outdated data records can be deleted, thereby freeing up valuable storage resources and enabling the cache to accommodate more active or useful data. Removing invalid or inefficient cache entries helps maintain the freshness and relevance of the cache content. This ensures that the data retained in the cache is the most likely to be accessed again, thereby improving the cache hit rate, reducing unnecessary disk operations, and accelerating data access speed.

[0066] Optionally, perform regular cleaning on the access history record table based on a preset time interval. Specifically, it may include: obtaining the current time and the last access time of each data block in the access history record table; in the access history record table, deleting the unaccessed data blocks whose difference between the last access time and the current time exceeds the preset time threshold.

[0067] In some examples, obtain the current time and the last access time of each data block in the access history record table. Then, set a preset time threshold for determining whether a data block is a data block that has not been accessed for a long time. Next, check the difference between the last access time of each data block in the access history record table and the current time. If this time difference exceeds the preset time threshold, the data block is considered to have not been accessed for a long time and is deleted from the access history record table and the weight table. This ensures that only those data blocks that are still active or have been recently accessed are retained in the system, thereby optimizing storage efficiency and improving system performance. This not only reduces unnecessary memory occupancy but also enables subsequent access and prefetch operations to focus more on valid data, improving the overall data processing speed and response efficiency.

[0068] Exemplarily, to keep the sizes of the access history record table and the weight table appropriate, it is necessary to regularly remove the data blocks that have not been accessed for a long time. First, obtain the current time current_time and set a time threshold (e.g., 60 seconds). Then, traverse each data block ID and its corresponding access timestamp in AccessHistory. If the difference between the current time and the access timestamp of a data block exceeds the set threshold, add the data block ID to the to - remove list. And delete the corresponding entries from AccessHistory and Weights according to the to - remove list to ensure that only active data block information is retained, thereby optimizing storage efficiency and system performance. Remove the corresponding data blocks from AccessHistory and Weights according to the to - remove list, clean up expired data, and maintain the validity of the data and the health of the storage structure.

[0069] Optionally, the prefetch position linked list is periodically cleared based on a preset time interval. Specifically, it may include: based on the preset time interval, counting the access frequencies of each data block in the prefetch position linked list; in the prefetch position linked list, deleting the second data blocks with access frequencies lower than a preset frequency threshold and the nodes corresponding to the second data blocks.

[0070] Exemplarily, to optimize cache usage and ensure that data blocks with greater potential to be accessed can obtain space, a timer can be set to periodically clear the prefetch position linked list. Specifically, it can be checked once every 10 seconds, and the prefetch nodes with access frequencies lower than 5 times are removed, so as to free up space for data blocks more likely to be accessed. This process not only maintains the freshness and relevance of the cache content, but also improves the overall performance and prefetch efficiency of the system.

[0071] Optionally, the method of this embodiment may specifically further include: in response to an erase operation request, querying in the access history record table; if there is a third data block to be erased in the access history record table, deleting the record of the third data block and the data page access table corresponding to the third data block.

[0072] In some examples, when an erase operation occurs, the access history of related data blocks needs to be cleared to maintain data consistency and the neatness of the storage structure. The specific steps are as follows: First, query the data block access history record table to check whether there is a record (i.e., block_id) of the data block to be erased. If the record of this data block exists in the data block access history record table, delete this record, and at the same time, also delete the data page access history record related to this data block, that is, remove the corresponding entry from the data page access history record table. This can ensure that the access history record is updated and cleared in a timely manner after the erase operation, and maintain the consistency of system data and the neatness of the storage structure.

[0073] Optionally, the method of this embodiment may specifically further include: in response to an erase operation request, querying in the prefetch position linked list; if there is a node corresponding to a fourth data block to be erased in the prefetch position linked list, deleting the record of the fourth data block and the prefetch buffer table corresponding to the fourth data block.

[0074] In some examples, when an erase operation occurs, query the prefetch position linked list. If there is a node corresponding to the block_id, delete the record of the block_id and its data in the prefetch buffer table to clear useless data and free up cache space. The cache management strategy effectively improves the cache utilization rate and prefetch hit rate on the premise of a limited memory size by comprehensively applying access frequency analysis, cache space allocation, and periodic cleaning mechanisms. Specifically, this strategy can dynamically adjust cache space allocation, preferentially retain data blocks with high access frequencies, and timely clean data blocks with low access frequencies, thereby significantly improving the overall performance of the SSD.

[0075] In some embodiments, as Figure 3 shown, when a read request of the Flash Translation Layer (FTL) is received, a thread is first created (creat) to update the access history of the read request. Then, the data block access history table Blk_AccessHistory is queried based on the block_id. If the block_id does not exist, the current block_id node is inserted, and its timestamp and weight Blk_Weights are initialized; if it exists, the timestamp of the node is updated and its weight is increased. Subsequently, the data page access history table Pg_AccessHistory is queried based on the identifier page_id of the data page. If the page_id does not exist, the current page_id node is inserted, and its timestamp and weight Pg_Weights are initialized; if it exists, the timestamp of the node is updated and its weight is increased. This ensures that the access history of relevant data blocks and data pages can be updated in a timely manner after each access, providing accurate data support for subsequent prefetching strategies.

[0076] In some embodiments, as Figure 4 shown, a timer timer is started every 5 seconds to obtain the current time. Each time the prefetching process is triggered, the data block access history table is first checked to confirm whether the most recent access occurred within 5 seconds. If it exceeds 5 seconds, the current prefetching election process is skipped to avoid invalid prefetching. Then, the timestamps of the data blocks are sorted (sort), and it is checked whether there are new access records within the most recent 60 seconds. If not, the next round of loop continues; if so, the top 30 data blocks with the highest weights are selected as the target prefetch blocks according to the Blk_Weights. Then, for each selected data block, the top 30 data pages with the highest weights are selected as the target prefetch pages according to the Pg_Weights. Subsequently, a prefetch operation is initiated for the selected high-weight data blocks and data pages, and the returned data is saved. At the same time, the timestamps of all data blocks are traversed, and the data blocks that have not been accessed for more than 60 seconds and their corresponding data page access history tables are deleted. This ensures that the system can dynamically adjust the prefetching strategy, improving the prefetching hit rate and overall performance.

[0077] In some embodiments, as Figure 5As shown, when the adaptive learning prefetch instruction is issued, first check whether the target prefetch data is already in the prefetch location linked list PreReadAddr. If it exists, increment the prefetch hit count Pre_cnt of this data and return; if it does not exist, it is processed by the back end (BE) and the prefetch data is returned. Next, check whether the total capacity TotalCapacity of the prefetch buffer is full (TotalCapacity = 0). If it is not full, insert the new data into the prefetch location linked list PreReadAddr and store the prefetch data in the prefetch buffer PreReadBuffer. If the prefetch buffer is full, find the data node with the least prefetch hit count in the prefetch location linked list, remove the min node, and free the corresponding prefetch buffer space. Subsequently, insert the new prefetch data into the prefetch location linked list and store it in the prefetch buffer, and at the same time save the currently prefetch returned data to the prefetch buffer to ensure the effective utilization of the cache space.

[0078] In some embodiments, as Figure 6 shown, when the host sends a read operation request, first process the data access information (host op message) read request through the FTL and create a thread to update the access history of the read request. Next, check whether the block_id of the target data block is in the data block access history table Blk_AccessHistory. If it exists, delete the node and its corresponding data page access history table Pg_AccessHistory. Then, check whether this block_id is in the prefetch location linked list PreReadAddr. If it exists, delete the node and free the corresponding prefetch buffer space. Finally, find the corresponding data buffer from the prefetch location linked list, obtain the data and return it to the host; if it does not exist, send a read request message to the back end (BE), process it according to the normal non-volatile memory (Non-volatile memory NANDflash memory, Nand) read process, and return the data buff to the host.

[0079] Compared with the current existing technologies, in this embodiment, by maintaining the data block access history record table and the data page access history record table, recording the access frequency and weight of each data block and data page, it is possible to more accurately predict the next most likely data block and data page to be accessed, and then determine the target prefetch data. This dynamic adjustment mechanism based on the access pattern can significantly improve the prefetch hit rate and reduce invalid prefetch operations. By preloading the target prefetch data into the cache of the solid-state drive in advance, when the user actually accesses this data, it can be directly read from the cache, thus significantly reducing the data access latency and improving the system response speed. By introducing an adaptive learning mechanism and an intelligent cache management strategy, on the premise of a limited memory size, the prefetch hit efficiency is significantly improved, thereby enhancing the overall SSD performance. This algorithm dynamically learns the actual access pattern, selects the data blocks most likely to be accessed for prefetching, effectively improving the prefetch hit rate; at the same time, it uses the intelligent cache management strategy to reasonably allocate the space of the prefetch location buffer, and avoids invalid prefetching by regularly clearing the data blocks with low access frequency, further optimizing the cache utilization rate. These features work together to ensure that the system can maximize the cache efficiency and resource utilization while maintaining efficient data access.

[0080] An embodiment of the present application also provides a data processing device, as Figure 1 a specific implementation of the method shown in Figure 7 shown, the device includes: an extraction module 31, an acquisition module 32, a determination module 33, and a loading module 34.

[0081] The extraction module 31 is configured to extract data access information from the solid-state drive and determine the currently accessed data block;

[0082] The acquisition module 32 is configured to acquire the access history record table corresponding to the currently accessed data block. The access history record table includes a data block access history record table and a data page access history record table. The data block access history record table is used to record the accessed data blocks and the weights of the accessed data blocks, and the data page access history record table is used to record the accessed data pages and the weights of the accessed data pages. The weights are used to reflect the probabilities of the data blocks and data pages in the solid-state drive being accessed;

[0083] The determination module 33 is configured to determine the target prefetch data based on the weights of multiple data blocks and multiple data pages in the access history record table;

[0084] The loading module 34 is configured to, if it is detected that the location of the target prefetch data is not in the prefetch location linked list, load the target prefetch data into the cache of the solid-state drive.

[0085] In some examples of this embodiment, the determining module 33 is specifically configured to screen out the first preset number of data blocks from high to low according to the weights of multiple data blocks in the data block access history record table; according to the weights of multiple data pages in the data page access history record table, screen out the second preset number of data pages from high to low for multiple data blocks among the first preset number of data blocks; and determine the data in the second preset number of data pages as the target prefetch data.

[0086] In some examples of this embodiment, the determining module 33 is further specifically configured to, in response to a read request for a data block, update the identifier and access timestamp of the first data block to the access history record table; update the weight of the first data block based on the number of times the first data block is accessed, and update the weight of the first data page corresponding to the first data block based on the number of times the first data page is accessed.

[0087] In some examples of this embodiment, the loading module 34 is further specifically configured to, if it is detected that the position of the target prefetch data is in the prefetch position linked list, increase the prefetch hit count of the position of the target prefetch data.

[0088] In some examples of this embodiment, the loading module 34 is further specifically configured to determine whether there is enough space in the cache of the solid-state drive to store the target prefetch data; if there is not enough space in the cache, sort the nodes in the prefetch position linked list according to the access frequency and delete the nodes with an access frequency lower than the preset threshold; if there is enough space in the cache, load the target prefetch data into the cache of the solid-state drive.

[0089] In some examples of this embodiment, the loading module 34 is further specifically configured to insert the target prefetch data into the prefetch position linked list and add the target prefetch data to the prefetch buffer table.

[0090] In some examples of this embodiment, the loading module 34 is further specifically configured to periodically clean the access history record table and / or the prefetch position linked list based on a preset time interval.

[0091] In some examples of this embodiment, the loading module 34 is further specifically configured to obtain the current time and the last access time of each data block in the access history record table; in the access history record table, delete the unaccessed data blocks whose difference between the last access time and the current time exceeds the preset time threshold.

[0092] In some examples of this embodiment, the loading module 34 is further specifically configured to, based on a preset time interval, count the access frequency of each data block in the prefetch position linked list; in the prefetch position linked list, delete the second data block and the corresponding node with an access frequency lower than the preset frequency threshold.

[0093] In some examples of this embodiment, the loading module 34 is further specifically configured to query in the access history record table in response to a wipe operation request; if there is a third data block to be wiped in the access history record table, delete the record of the third data block and the data page access table corresponding to the third data block.

[0094] In some examples of this embodiment, the loading module 34 is further specifically configured to query in the prefetch location linked list in response to a wipe operation request; if there is a node corresponding to a fourth data block to be wiped in the prefetch location linked list, delete the record of the fourth data block and the prefetch buffer table corresponding to the fourth data block.

[0095] In some examples of this embodiment, by querying the access history record table, obtain the last access time of the currently accessed data block; if the difference between the last access time and the current time is greater than a preset interval period, stop obtaining the target prefetch data; correspondingly, the determination module 33 is further specifically configured to, if the difference between the last access time and the current time is less than or equal to the preset interval period, determine the target prefetch data based on the weights of multiple data blocks and the weights of multiple data pages in the access history record table.

[0096] It should be noted that for other corresponding descriptions of each functional unit involved in the data processing device provided in this embodiment, reference can be made to Figure 1 the corresponding description in, which will not be elaborated here.

[0097] Based on the above as Figure 1 and Figure 2 shown in the method, correspondingly, this embodiment also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method shown in the above as Figure 1 and Figure 2 shown is implemented.

[0098] Based on the above as Figure 1 and Figure 2 shown in the method, correspondingly, this embodiment also provides a computer program product, on which a computer program is stored, and when the computer program is executed by a processor, the method shown in the above as Figure 1 and Figure 2 shown is implemented.

[0099] Based on such an understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.), and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of various implementation scenarios of this application.

[0100] Based on the above as Figure 1 andFigure 2 The method shown, and Figure 7 The virtual device embodiments shown. To achieve the above object, embodiments of the present application further provide an electronic device, such as a personal computer or a server. The device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the method shown in Figure 1 and Figure 2 shown above.

[0101] In some embodiments, the above-mentioned physical device may further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, sensors, an audio circuit, a WI-FI module, and so on. The user interface may include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. Optionally, the user interface may further include a USB interface, a card reader interface, etc. The network interface may include a standard wired interface, a wireless interface (such as a WI-FI interface), etc. in some embodiments.

[0102] Those skilled in the art can understand that the structure of the physical device provided in this embodiment does not limit the physical device, and it may include more or fewer components, or combine some components, or arrange different components.

[0103] The storage medium may further include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned physical device, and supports the operation of information processing programs and other software and / or programs. The network communication module is used to implement communication between components inside the storage medium, and communication between other hardware and software in the information processing physical device.

[0104] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary general hardware platform, or can also be implemented by hardware. By applying the solution of this embodiment, compared with the current existing technologies, this embodiment improves the overall performance of the SSD by integrating an adaptive learning mechanism and an intelligent cache management strategy. The adaptive learning mechanism designs a set of algorithms that can dynamically adjust the selection of prefetch positions according to the actual access pattern. By maintaining an access history record table of data blocks and their data pages, and introducing a weight mechanism to reflect the access probability, it preferentially selects data blocks and data pages with higher weights for prefetching. At the same time, when a data block is erased, the relevant access records are cleared and the node is removed. The intelligent cache management strategy, on the premise of a limited memory size, reasonably allocates the space of the prefetch buffer, regularly clears data blocks with low access frequencies to free up space for data blocks more likely to be accessed, and improves the cache utilization rate. Finally, these algorithms are optimized and integrated into the firmware of the SSD to ensure its efficient operation and compatibility with other functions, becoming the core function to improve the performance of the SSD. This systematic method not only enhances the data access efficiency, but also significantly improves the overall performance and response speed of the storage system.

[0105] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without more limitations, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0106] The above are only specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to these embodiments herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A data processing method, characterized in that, Including: Extract data access information from a solid-state drive to determine the currently accessed data block; Obtain an access history record table corresponding to the currently accessed data block. The access history record table includes a data block access history record table and a data page access history record table. The data block access history record table is used to record the accessed data blocks and the weights of the accessed data blocks. The data page access history record table is used to record the accessed data pages and the weights of the accessed data pages. The weights are used to reflect the probabilities of the data blocks and data pages in the solid-state drive being accessed; Based on the weights of multiple data blocks and multiple data pages in the access history record table, determine target prefetch data, including: screening out a first preset number of data blocks from high to low according to the weights of multiple data blocks in the data block access history record table; according to the weights of multiple data pages in the data page access history record table, for multiple data blocks among the first preset number of data blocks, screening out a second preset number of data pages from high to low; determining the data in the second preset number of data pages as the target prefetch data; by querying the access history record table, obtain the last access time of the currently accessed data block; if the difference between the last access time and the current time is greater than a preset interval period, stop obtaining the target prefetch data; if the difference between the last access time and the current time is less than or equal to the preset interval period, determine the target prefetch data based on the weights of multiple data blocks and multiple data pages in the access history record table; If it is detected that the position of the target prefetch data is not in the prefetch position linked list, then load the target prefetch data into the cache of the solid-state drive, including: determining whether there is enough space in the cache of the solid-state drive to store the target prefetch data; if there is not enough space in the cache, sort the nodes in the prefetch position linked list according to the access frequency and delete the nodes with an access frequency lower than a preset threshold; if there is enough space in the cache, load the target prefetch data into the cache of the solid-state drive.

2. The method according to claim 1, wherein The method further includes: In response to a read request for a first data block, update the identifier and access timestamp of the first data block to the access history record table; Update the weight of the first data block based on the number of times the first data block is accessed, and update the weight of the first data page corresponding to the first data block based on the number of times the first data page is accessed.

3. The method according to claim 1, characterized in that, The method further includes: If it is detected that the position of the target prefetch data is in the prefetch position linked list, then increase the prefetch hit count of the position of the target prefetch data.

4. The method according to claim 1, characterized in that, The loading the target prefetch data into the cache of the solid-state drive includes: Insert the target prefetch data into the prefetch position linked list and add the target prefetch data to the prefetch buffer table.

5. The method according to claim 1, characterized in that, The method further includes: Regularly clean the access history record table and / or the prefetch position linked list based on a preset time interval.

6. The method according to claim 5, wherein Regularly clean the access history record table based on a preset time interval, including: Obtain the current time and the last access time of each data block in the access history record table; In the access history record table, delete the unaccessed data blocks whose difference between the last access time and the current time exceeds a preset time threshold.

7. The method according to claim 5, wherein Regularly clean the prefetch location linked list based on a preset time interval, including: Based on a preset time interval, count the access frequencies of the data blocks in the prefetch location linked list; In the prefetch location linked list, delete the second data blocks with access frequencies lower than a preset frequency threshold and the nodes corresponding to the second data blocks.

8. The method according to claim 1, characterized in that, The method further includes: In response to a wipe operation request, query in the access history record table; If there is a third data block to be wiped in the access history record table, delete the record of the third data block and the data page access table corresponding to the third data block.

9. The method according to claim 1, wherein The method further includes: In response to a wipe operation request, query in the prefetch location linked list; If there is a node corresponding to a fourth data block to be wiped in the prefetch location linked list, delete the record of the fourth data block and the prefetch buffer table corresponding to the fourth data block.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

11. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein, When the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

12. A computer program product having a computer program stored thereon, characterized in that, When the computer program product is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Data pre-reading method and device

    CN117807040A

  • Software configuration method and device

    CN119415143A