A data prefetching method, device, electronic device and storage medium

By independently training the data prefetch algorithm for each process and forming an algorithm set, filtering the appropriate data prefetch algorithm to determine the target data prefetch strategy, the problem of inaccurate data prefetch results caused by the hardware prefetcher relying on a fixed memory access mode is solved, and more efficient data prefetch resource utilization is achieved.

CN119883954BActive Publication Date: 2025-06-17SHANDONG BOSUAN ZHIXIN INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510362116.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-17
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The existing hardware prefetcher relies on a fixed memory access mode, which reduces the accuracy of data prefetch results during process switching and wastes data prefetch resources of computing devices.

Method used

By obtaining the memory fetch information of all currently running processes, the data prefetch algorithm is independently trained for each process, forming an algorithm set, and filtering the appropriate data prefetch algorithm from the algorithm center based on the target process to which the current memory fetch request belongs to, to determine the target data prefetch strategy.

Benefits of technology

It improves the accuracy of data prefetch results, optimizes the utilization rate of data prefetch resources of computing devices, and adapts to the differences in memory access modes of different processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119883954B_ABST
    Figure CN119883954B_ABST
Patent Text Reader

Abstract

The present application discloses a data prefetching method, apparatus, electronic device, and storage medium, relating to the field of computer technologies. After obtaining the memory access information of all currently running processes, a corresponding data prefetching algorithm is independently trained for each process to form an algorithm set. When performing data prefetching, a suitable algorithm is selected from the algorithm set according to the target process to which the current memory access request belongs. Therefore, it is possible to solve the technical problem in the prior art that the hardware prefetcher depends on a fixed memory access pattern, and when the memory access pattern changes due to process switching, the accuracy of its data prefetching result decreases, and achieve the technical effect of improving the accuracy of the data prefetching result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and in particular, to a data prefetching method, apparatus, electronic device, and storage medium. Background Art

[0002] With the improvement of the performance of computing devices such as servers, the storage wall problem has become increasingly prominent. The memory access speed lags far behind the processor computing speed, resulting in memory access latency when data is missing. Therefore, how to perform prefetching of memory access data to reduce memory access latency has become a key research content.

[0003] In related technologies, a hardware prefetcher is usually used for data prefetching. However, the hardware prefetcher depends on a fixed memory access pattern. When the memory access pattern changes due to process switching, the accuracy of its data prefetching result decreases, wasting the data prefetching resources of the computing device. Summary of the Invention

[0004] This application provides a data prefetching method, apparatus, electronic device, and storage medium to at least solve the problem in related technologies that the accuracy of data prefetching results is reduced and the data prefetching resources of the computing device are wasted.

[0005] This application provides a data prefetching method, including:

[0006] Obtaining the memory access information of all currently running processes;

[0007] For any process, training the data prefetching algorithm corresponding to the process according to the memory access information of the process to obtain a data prefetching algorithm set;

[0008] When the current cache state of the preset cache meets the preset prefetch condition, obtaining the current memory access request and determining the target process to which the current memory access request belongs;

[0009] Filtering the target data prefetching algorithm from the data prefetching algorithm set according to the target process;

[0010] Based on the target data prefetching algorithm, determining the target data prefetching policy according to the memory access address information of the current memory access request.

[0011] This application also provides a data prefetching apparatus, including:

[0012] An obtaining module, configured to obtain the memory access information of all currently running processes;

[0013] A training module, configured to, for any process, train the data prefetching algorithm corresponding to the process according to the memory access information of the process to obtain a data prefetching algorithm set;

[0014] A determination module, configured to obtain a current memory access request and determine a target process to which the current memory access request belongs when a current cache state of a preset cache meets a preset prefetch condition;

[0015] A screening module, configured to screen a target data prefetch algorithm from a data prefetch algorithm set according to the target process;

[0016] A prefetch module, configured to determine a target data prefetch policy based on the target data prefetch algorithm according to memory access address information of the current memory access request.

[0017] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any one of the above data prefetch methods when executing the computer program.

[0018] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any one of the above data prefetch methods are implemented.

[0019] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of any one of the above data prefetch methods are implemented.

[0020] Through this application, after obtaining memory access information of all currently running processes, a corresponding data prefetch algorithm is independently trained for each process to form an algorithm set. When performing data prefetch, a suitable algorithm is screened from the algorithm set according to the target process to which the current memory access request belongs. Therefore, the technical problem in the prior art that the hardware prefetcher depends on a fixed memory access pattern and the accuracy of its data prefetch result decreases when the memory access pattern changes due to process switching can be solved, and the technical effect of improving the accuracy of the data prefetch result can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic structural diagram of a data prefetch system based on the embodiments of this application;

[0023] Figure 2 It is a schematic flowchart of a data prefetch method provided by the embodiments of this application;

[0024] Figure 3 It is a schematic structural diagram of an exemplary data prefetch system provided by the embodiments of this application;

[0025] Figure 4 Schematic flowchart of an exemplary data prefetching method provided by an embodiment of the present application;

[0026] Figure 5 Schematic structural diagram of another exemplary data prefetching system provided by an embodiment of the present application;

[0027] Figure 6 Schematic structural diagram of a prefetching and launching unit provided by an embodiment of the present application;

[0028] Figure 7 Schematic structural diagram of a data prefetching device provided by an embodiment of the present application;

[0029] Figure 8 Schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0031] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0032] In modern multi-core high-performance processors, resources such as the storage capacity allocated to cores are extremely tight, and the memory access speed lags far behind the computing speed of the processor. The bottleneck effect caused by the "memory wall" problem has become an important factor restricting the development of high-performance processors. The common approach of modern processors is to add a hidden small-capacity and fast cache between the CPU and the main memory. Although the cache can provide high-speed storage access for the CPU, when the data required by the program is missing in the cache, it is necessary to obtain data from a lower-level storage structure, resulting in a dozens-fold increase in access latency, which seriously affects the performance of computing.

[0033] Data prefetching technology refers to, before a program accesses data, based on the principle of locality of program memory access, predicting subsequent memory accesses by learning the past memory access patterns, and loading the required data into the cache in advance, so as to reduce cache misses and hide the memory latency problem. This technology does not necessarily bring benefits to the program, because the prediction of the data required by the future processor consumes additional storage bandwidth. If data blocks that will not be accessed in the future are prefetched, it will not only cause additional communication pressure, but may also replace useful data blocks in the cache, resulting in cache pollution and causing cache misses that would not have occurred otherwise. Therefore, it is necessary to judge the effectiveness of the data prefetching technology adopted.

[0034] Currently, the technical trend of high-performance processors generally adopts a multi-core and multi-threaded structure. The concurrent execution of multiple processes brings a complex computing environment and increased computing load. The memory access pattern characteristics and data ranges between different processes may be different, which will affect the locality of program memory access. The current prefetchers often analyze and prefetch for specific memory access patterns. Different application programs have different memory access patterns and thus are suitable for different types of prefetching algorithms. A single and static prefetching algorithm cannot have excellent prefetching results for each memory access pattern.

[0035] Existing research has combined multiple prefetchers and issued prefetch addresses according to the priority order. For example, the Slim-AMPM prefetcher combines the Access mapping pattern matching (AMPM) prefetcher with the (Delta-correlation-prefetcher-table, DCPT) prefetcher, and preferentially uses the prefetch addresses issued by the DCPT prefetcher. The AMPM prefetcher learns the memory access patterns that the DCPT prefetcher cannot learn and issues prefetch addresses. However, when a process switch occurs in this fusion prefetcher that does not perceive the process, the memory access pattern learned by the hardware prefetcher for the previous process may not be applicable to the new process. Especially when concurrent execution causes the process to switch repeatedly, it is more difficult to determine the memory access pattern of the program. Blindly continuing to use the prefetch training results of different processes will generate a large amount of redundant prefetches, wasting the bandwidth and power consumption of the processor accessing the on-chip cache.

[0036] Embodiments of the present application are provided to solve the above technical problems, and provide a data prefetching method, device, electronic device, and storage medium. The method includes: obtaining memory access information of all currently running processes; for any process, training a data prefetching algorithm corresponding to the process according to the memory access information of the process to obtain a data prefetching algorithm set; when the current cache state of a preset cache meets a preset prefetching condition, obtaining a current memory access request and determining a target process to which the current memory access request belongs; screening a target data prefetching algorithm from the data prefetching algorithm set according to the target process; and determining a target data prefetching policy based on the target data prefetching algorithm according to the memory access address information of the current memory access request. The method provided by the above solution, after obtaining the memory access information of all currently running processes, independently trains a corresponding data prefetching algorithm for each process to form an algorithm set, and when performing data prefetching, screens a suitable algorithm from the algorithm set according to the target process to which the current memory access request belongs, improving the accuracy of the data prefetching result and laying a foundation for the utilization rate of the data prefetching resources of the computing device.

[0037] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0038] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the data prefetching method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0039] First, the structure of the data prefetching system on which the present application is based is described:

[0040] The data prefetching method, device, electronic device, and storage medium provided by the embodiments of the present application are applicable to data prefetching for computing devices such as servers. As Figure 1 shown, it is a schematic structural diagram of the data prefetching system based on the embodiments of the present application, mainly including a host side, a buffer, and a data prefetching device. Among them, the host side is used to read service data from the buffer, and the data prefetching device is used to pre-train data prefetching algorithms suitable for different processes. When data prefetching is required, a suitable algorithm is screened according to the target process to which the current memory access request belongs, and data prefetching is performed for the buffer based on the screened algorithm, so as to cache the service data that the host side needs to read into the buffer before the host side accesses the buffer.

[0041] Embodiments of the present application provide a data prefetching method for prefetching service data to be accessed by a host-side process from a main memory to a buffer, so that the host side can quickly read service data from the buffer. The execution subject of the embodiments of the present application is an electronic device, such as a server, a desktop computer, a laptop computer, a tablet computer, and other electronic devices that can be used for data prefetching.

[0042] AsFigure 2 As shown in the figure, it is a schematic flowchart of the data prefetching method provided by the embodiment of the present application. The method includes:

[0043] Step 201, obtain the memory access information of all currently running processes.

[0044] It should be noted that multiple applications are deployed on the host side. When an application starts, several processes related to the current business are created and start running. Among them, the memory access information of the process at least includes memory access address, time, frequency, etc.

[0045] Step 202, for any process, train the data prefetching algorithm corresponding to the process according to the memory access information of the process to obtain a data prefetching algorithm set.

[0046] Among them, the data prefetching algorithms include Block-based Prefetcher (abbreviated as BOP), Machine Learning-based Prefetcher (abbreviated as MLOP), Spatial Prefetcher (abbreviated as SPP), Temporal Prefetcher, Spatio-Temporal Prefetcher (abbreviated as STeMs), Access Mapping Pattern Matching Prefetcher (abbreviated as AMPM), Delta Correlation Prediction Table Prefetcher (abbreviated as Delta Correlation Prediction Table Prefetcher), and Triage.

[0047] Step 203, when the current cache state of the preset cache meets the preset prefetch condition, obtain the current memory access request and determine the target process to which the current memory access request belongs.

[0048] Among them, the preset cache is the above-mentioned cache, which is a pre-set high-speed cache area for storing data, used to reduce the latency of the host-side CPU accessing the memory. The current memory access request is the memory data access request initiated by the host side to the preset cache.

[0049] Specifically, the state of the preset cache can be continuously monitored. When the cache state meets the preset prefetch condition, obtain the memory access request at this time and determine which process the request belongs to, so as to prepare for screening a suitable prefetching algorithm later.

[0050] Step 204, screen the target data prefetching algorithm from the data prefetching algorithm set according to the target process.

[0051] Specifically, according to the target process to which the current memory access request belongs, a target data prefetching algorithm suitable for the process can be selected from the data prefetching algorithm set obtained from previous training, so as to ensure the matching degree between the target data prefetching algorithm and the target process and improve the accuracy of the prefetching result.

[0052] Step 205: Based on the target data prefetching algorithm, determine a target data prefetching policy according to the memory access address information of the current memory access request.

[0053] Among them, the target data prefetching policy at least includes a data prefetching address.

[0054] Specifically, based on the target data prefetching algorithm, according to the memory access address information of the current memory access request, the memory access address information of the memory access requests that the target process will initiate subsequently can be predicted to obtain a data prefetching address, and then the target data prefetching policy can be determined.

[0055] Based on the above embodiments, as an implementable manner, in one embodiment, for any process, according to the memory access information of the process, train the data prefetching algorithm corresponding to the process to obtain a data prefetching algorithm set, including:

[0056] Step 2021: Obtain an initial algorithm set;

[0057] Step 2022: For any process, train each initial algorithm in the initial algorithm set according to the memory access information of the process;

[0058] Step 2023: When the preset training period is reached, determine the data prefetching algorithm corresponding to the process according to the prefetching accuracy of each trained algorithm;

[0059] Step 2024: Construct a data prefetching algorithm set according to the data prefetching algorithms corresponding to each process.

[0060] It should be noted that the embodiments of the present application realize the training and selection of personalized data prefetching algorithms for different processes, can fully adapt to the memory access pattern differences of different processes, improve the accuracy and efficiency of data prefetching, reduce invalid prefetching, reduce the bandwidth occupation and cache pollution of the system, and thus improve the overall performance of the system. Especially in a multi-core and multi-thread environment where different processes run concurrently, it can better meet the memory access requirements of each process and improve the response speed and throughput of the system.

[0061] Specifically, for each process running in the system, each algorithm in the initial algorithm set is trained using its memory access information. By inputting the memory access information of the process into the algorithm, the algorithm learns the memory access pattern and rules of the process, etc., to adjust the internal parameters of the algorithm so that the algorithm can better adapt to the memory access behavior of the process. When the training reaches this preset training cycle, the prefetch accuracy of each trained algorithm is evaluated. The prefetch accuracy can be measured by comparing the matching degree between the prefetch address predicted by the algorithm and the address actually accessed by the process. Select the algorithm with the highest prefetch accuracy as the data prefetch algorithm corresponding to this process to ensure that the most accurate prefetch service can be provided for this process in subsequent operations. Finally, collect the most suitable data prefetch algorithms determined for each process to form a new set, that is, the data prefetch algorithm set. The data prefetch algorithm set contains the optimal prefetch algorithms for all processes. When performing data prefetch later, according to the process to which the current memory access request belongs, the corresponding prefetch algorithm can be quickly found from this algorithm set.

[0062] Specifically, in one embodiment, for any process, according to the memory access address information represented by the memory access information of this process, the memory access rule of this process can be determined; according to the memory access rule of the process, each initial algorithm in the initial algorithm set is trained to optimize the configuration parameters of each initial algorithm.

[0063] It should be noted that the memory access information of the process covers various data when the process accesses the memory, such as memory access address information, etc. The memory access address information records the specific memory addresses accessed by the process. The specific memory addresses accessed by the process generally do not appear randomly, but contain the internal rules of the process's memory access. For example, when processing an array, the process may access the memory addresses in sequence according to the order of the array elements, showing a sequential access rule; or when performing matrix operations, it may access addresses in a specific step size in a jumping manner.

[0064] Specifically, after determining the memory access rule of the process, apply these rules to the training process of the initial algorithm. By adjusting the configuration parameters of the algorithm, the algorithm can better adapt to the memory access pattern of the process.

[0065] Specifically, in one embodiment, during the training process of each initial algorithm in the initial algorithm set, the actual memory access address of the process can be obtained; when the preset training cycle is reached, the training prefetch results of each trained algorithm are obtained; for any trained algorithm, according to the difference between the training prefetch result of this trained algorithm and the actual memory access address of the process, the prefetch accuracy of this trained algorithm is determined; the trained algorithms with prefetch accuracy reaching the preset threshold are used as candidate algorithms; the candidate algorithm with the highest prefetch accuracy is used as the data prefetch algorithm corresponding to this process.

[0066] Among them, the actual memory access address reflects the real data access requirements of the process.

[0067] Specifically, for each trained algorithm, by comparing the difference between its training prefetch result and the actual memory access address of the process, the prefetch accuracy of the algorithm is determined. Among them, during the training process, the trained algorithm will generate multiple data prefetch addresses. By comparing the number of matches between these data prefetch addresses and the actual memory access address, that is, according to how many data prefetch addresses match (are consistent with) the actual memory access address, the prefetch accuracy of the algorithm is determined. After determining the prefetch accuracies of multiple trained algorithms, the trained algorithms whose prefetch accuracies reach the preset threshold are used as candidate algorithms, and the one with the highest prefetch accuracy is selected from the candidate algorithms as the data prefetch algorithm corresponding to this process, that is, the data prefetch algorithm of the process is the trained algorithm whose prefetch accuracy reaches the preset threshold and has the highest prefetch accuracy.

[0068] Specifically, in one embodiment, for any process, if the prefetch accuracies of all the trained algorithms of this process still do not reach the preset threshold after being trained for multiple preset training cycles, then this process is regarded as a discarded process; the training of the data prefetch algorithm for the discarded process is abandoned.

[0069] Specifically, during the training process of the data prefetch algorithm, the system uses each algorithm in the initial algorithm set to train for each process, and evaluates the prefetch accuracy of the trained algorithm when the preset training cycle is reached. When the prefetch accuracies of all the trained algorithms of a certain process still cannot reach the preset threshold even after being trained for multiple preset training cycles, it means that these algorithms are difficult to adapt to the memory access pattern of this process, and it is difficult to achieve ideal results by continuing to train these algorithms. In this case, this process is marked as a discarded process, and the training work of its data prefetch algorithm is stopped, and no more computing resources and time will be invested in trying to optimize the prefetch algorithm for this process, reducing unnecessary training rounds and computing amounts, enabling the system to complete the training and optimization of the data prefetch algorithms for other processes faster and accelerating the construction process of the entire prefetch algorithm set.

[0070] Correspondingly, in one embodiment, if the prefetch accuracy of a trained algorithm reaches the preset threshold after a short period of training, the training of this algorithm can also be terminated, thereby reducing the training cycle of the algorithm.

[0071] Specifically, in one embodiment, when any process enters a preset interruption state, the process is regarded as an interrupted process; the training of the interrupted process is paused, and the trained algorithm set of the interrupted process and the training prefetch results generated by each trained algorithm in the trained algorithm set during the training process are packaged to obtain a training interruption data packet of the interrupted process; the training interruption data packet is stored in a preset temporary storage area (process temporary storage area); when any interrupted process enters a preset active state, the training interruption data packet corresponding to the interrupted process is extracted from the preset temporary storage area; based on the training interruption data packet, the trained algorithm set of the interrupted process is continued to be trained.

[0072] It should be noted that a process includes multiple states, such as a ready state, a running state, a blocked state, and a completed state, etc. In the embodiments of the present application, a process that enters the completed state (preset interruption state) can be regarded as an interrupted process. Among them, when a process enters the preset interruption state, the process terminates running, and the system will recycle the resources occupied by the process.

[0073] Among them, as Figure 3 shown, it is a schematic structural diagram of an exemplary data prefetch system provided by the embodiments of the present application. The data prefetch system includes a memory management unit (MMU) and a multi-level cache (preset cache). The memory management unit includes a page table address space and a prefetch configuration unit. A prefetch algorithm unit, a prefetch training unit, a process temporary storage unit, and a prefetch emission unit are provided between the memory management unit and the multi-level cache. Among them, the prefetch algorithm unit is an integration unit of each data prefetch algorithm, and is used to generate a data prefetch policy; the prefetch training unit is used to train the initial algorithm in the prefetch algorithm unit according to the memory access information of the process. The prefetch emission unit is used to send the prefetch instruction corresponding to the data prefetch policy to the multi-level cache after the algorithm is trained and the data prefetch policy is determined, so as to prefetch the data that may be accessed by the process in advance, thereby improving the access efficiency of the system. The prefetch emission unit is also used to monitor the cache state and adjust the prefetch degree, and send a request to the multi-level cache according to the received prefetch address; the process temporary storage unit includes two parts in the lower-level cache (process temporary storage area) of the memory management unit and the multi-level cache, and is used to temporarily store unfinished processes (training interruption data packets of interrupted processes), and a separate completion space is set to record the situation of completed processes (interrupted processes) for subsequent rapid training; the page table address space is used to convert the page table base address and the running state into a virtual page table address; the prefetch configuration unit is used to configure the parameters and rules related to prefetching, such as recording the process information of completed (interrupted processes) and non-prefetched (abandoned processes), and the initial configuration information of the prefetch training unit for the new page table base address. Among them, when any interrupted process enters the preset active state, the initial configuration information of its data prefetch algorithm is determined according to the training interruption data packet of the interrupted process, so as to continue training the trained algorithm set of the interrupted process on the basis of the training interruption data packet to improve the training efficiency of the algorithm.

[0074] Specifically, in one embodiment, when any process becomes a completed process or a process to be replaced, it is determined that the process enters a preset interrupt state.

[0075] Among them, when a new process is generated and the number of parallel processes reaches a preset parallel threshold, the least used process among the original processes is used as the process to be replaced.

[0076] Specifically, when a new process is generated and the number of currently running parallel processes has reached the preset parallel threshold, where the preset parallel threshold is the maximum number of processes that can run in parallel simultaneously set by the system in advance, the system makes a decision to select one of the original processes for replacement to make room for the new process. Specifically, the least used process among the original processes can be selected as the process to be replaced to avoid excessive competition for system resources, so that system resources can be more reasonably allocated to each valuable process and improve the overall operation efficiency of the system.

[0077] Specifically, in one embodiment, the page table base address of the interrupted process is obtained in the preset page table base address register; the page table base address is converted into a virtual page table address; wherein, the address length of the virtual page table address is less than the address length of the page table base address; the virtual page table address is added to the training interrupt data packet; when the virtual page table address of the new process matches the virtual page table address of any interrupted process, it is determined that the interrupted process enters a preset active state.

[0078] It should be noted that in the operating system of a computing device, each process has its own page table, which is used to implement the conversion from virtual address to physical address. The page table base address register is a register specifically used to store the page table base address of a process. When a process enters the preset interrupt state, the system obtains the page table base address of the interrupted process in the preset page table base address register. Since the page table base address is relatively long and contains a lot of information, in order to facilitate subsequent processing and storage, the page table base address is converted into a virtual page table address to reduce the amount of data, improve the processing efficiency, and also facilitate storage and transmission in the training interrupt data packet.

[0079] Among them, multi-core and multi-threading may result in a large number of concurrent processes, and each process may use several page table base addresses. Recording a large number of page table base addresses, running states, and corresponding information on the chip requires a large amount of storage resources. Therefore, in the embodiments of the present application, a separate page table address space is set up. By introducing an additional indirect level, any relevant page table base address is converted into a continuous address (virtual page table address) in a new structure that is only visible to the data prefetch device, so that the page table base addresses can appear in sequence and the bit width can be reduced, thereby reducing the storage space.

[0080] Specifically, when a new process is generated, the virtual page table address of the new process is compared with the virtual page table address of the previously interrupted process in the training interruption data packet. If it is found that the virtual page table address of the new process matches the virtual page table address of a certain interrupted process, this means that there is a certain correlation between the new process and the interrupted process in terms of memory access mode or resource usage, or the new process is a restarted interrupted process. At this time, it is determined that the interrupted process enters a preset active state.

[0081] Based on the above embodiments, as an implementable manner, in one embodiment, the method further includes:

[0082] Step 301, when the cache hit rate characterized by the current cache state of the preset cache is lower than the preset hit rate threshold, it is determined that the current cache state of the preset cache meets the preset prefetch condition.

[0083] It should be noted that the cache hit rate refers to the ratio of the number of times of successfully obtaining data in the preset cache to the total number of data accesses, which reflects the effective utilization degree of the cache. The preset hit rate threshold is a standard value set according to system performance requirements and expectations. For example, assume that the preset hit rate threshold is set to 80%, which means that the system expects that in an ideal situation, 80% of the data accesses can be directly obtained from the cache without accessing the slower main memory.

[0084] Specifically, when the cache hit rate characterized by the current cache state of the preset cache is lower than the preset hit rate threshold, it indicates that the preset cache fails to effectively meet the data access requirements, and a large number of data accesses need to be obtained from the main memory, resulting in an increase in access latency and affecting system performance. At this time, the system determines that the current cache state meets the preset prefetch condition, that is, it is considered necessary to start the data prefetch operation to replace the original data in the preset cache to improve the cache hit rate.

[0085] Correspondingly, in one embodiment, it can also be determined that the current cache state of the preset cache meets the preset prefetch condition when the remaining cache space characterized by the current cache state of the preset cache is not lower than the preset remaining cache space threshold.

[0086] It should be noted that the preset remaining cache space not being lower than the preset remaining cache space threshold indicates that the preset cache has enough free space to store the prefetch data. Therefore, it is determined that the current cache state of the preset cache meets the preset prefetch condition. In this case, performing the data prefetch operation will not cause the newly prefetched data to be unable to be stored due to insufficient cache space, nor will it frequently trigger cache replacement due to tight cache space, ensuring the stability and performance of the preset cache.

[0087] Accordingly, in one embodiment, when the current cache state of the preset cache indicates that the current memory access request misses the cache, it is determined that the current cache state of the preset cache meets the preset prefetch condition.

[0088] It should be noted that the current memory access request missing the cache means that the preset cache has a memory access miss. The current memory access request missing also indicates that subsequent related data may not be in the cache either. If no prefetch measure is taken, subsequent memory accesses may still face high latency problems. Therefore, it is determined that the current cache state meets the preset prefetch condition to perform data anticipation, making subsequent memory access requests more likely to hit the cache, thereby effectively reducing access latency and improving the overall performance of the system.

[0089] Based on the above embodiment, as an implementable manner, in one embodiment, obtaining the current memory access request and determining the target process to which the current memory access request belongs includes:

[0090] Step 2031, after obtaining the current memory access request, analyze the current memory access request to determine the page table base address of the current memory access request;

[0091] Step 2032, according to the page table base address, determine the target process to which the current memory access request belongs.

[0092] It should be noted that in the operating system, each process has an independent page table. The page table base address is the unique identifier and is the key information of a process's page table. After obtaining the current memory access request, the page table base address corresponding to the current memory access request can be obtained in the page table base address register. Through the page table base address, the corresponding page table can be found, and this page table belongs to a specific process. After obtaining the page table base address of the current memory access request, according to the mapping relationship between the page table and the process, the target process corresponding to the page table base address can be searched, that is, the target process to which the current memory access request belongs is determined.

[0093] Based on the above embodiment, as an implementable manner, in one embodiment, before determining the target data prefetch policy according to the memory access address information of the current memory access request based on the target data prefetch algorithm, the method further includes:

[0094] Step 401, according to the cache hit rate and cache replacement rate represented by the current cache state of the preset cache, determine the prefetch degree of the target data prefetch algorithm.

[0095] Among them, the target data prefetch algorithm is negatively correlated with both the cache hit rate and the cache replacement rate. The prefetch degree represents the amount of data prefetched by the target data prefetch algorithm in one prefetch operation. The higher the prefetch degree, the larger the amount of prefetched data; the lower the prefetch degree, the smaller the amount of prefetched data.

[0096] It should be noted that the cache hit rate represents the ratio of the number of times a process successfully retrieves data from a preset cache to the total number of data accesses. The higher the cache hit rate, the higher the utilization efficiency of the cache. The cache replacement rate represents the frequency at which data in the preset cache is replaced per unit time. A higher cache replacement rate means that the data in the preset cache is updated frequently.

[0097] Specifically, when the cache hit rate is high, it indicates that the current cache utilization is good. At this time, the prefetch degree of the target data prefetch algorithm can be appropriately reduced to reduce unnecessary prefetch operations and avoid wasting system resources. On the contrary, when the cache hit rate is low, it means that there is insufficient data in the cache, and the prefetch degree needs to be increased to prefetch more data into the cache to improve the hit rate of subsequent data accesses. For the cache replacement rate, when the cache replacement rate is high, it means that the data in the cache is updated frequently and there is an over-prefetching situation. At this time, the prefetch degree should be reduced, the amount of prefetch data should be decreased, and the pressure of frequent replacement of data in the cache should be reduced. When the cache replacement rate is low, the prefetch degree can be appropriately increased to increase the amount of prefetch data to make full use of the cache space and improve the data access efficiency.

[0098] Based on the above embodiments, as an implementable manner, in one embodiment, based on the target data prefetch algorithm, according to the memory access address information of the current memory access request, a target data prefetch policy is determined, including:

[0099] Step 2051, based on the target data prefetch algorithm, according to the memory access address information of the current memory access request, determine data prefetch addresses under multiple different look-ahead degrees;

[0100] Step 2052, according to the data prefetch addresses under multiple different look-ahead degrees, determine the target data prefetch policy.

[0101] Among them, the target data prefetch policy at least includes data prefetch addresses under multiple different look-ahead degrees and the issuing order of each data prefetch address. The look-ahead degree represents the prediction degree or lead of future data accesses. For example, if the look-ahead degree is 10, it means that the currently predicted data prefetch address is the memory access address that the host may initiate access to within the next 10 clock cycles, and so on.

[0102] Exemplarily, such as Figure 4As shown in the figure, it is a schematic flowchart of an exemplary data prefetching method provided by an embodiment of the present application. When a memory access request with a cache miss occurs, first, it is determined whether the page table base address of the process to which the current memory access request belongs is recorded in the page table address space. If not, the process is determined to be a new process. At this time, the least recently used entry (the process to be replaced) is replaced into the process temporary storage area, and then the software configuration information corresponding to the process is imported. The software configuration information includes memory access address information, which is used to characterize key contents such as the memory access mode of the process and is used for subsequent simulated prefetching operations. If it exists, it is further determined whether the process is in the prefetch training unit. If not, it is determined that the process was previously paused as an interrupted process for training. Therefore, the training interruption data packet is extracted from the process temporary storage area to determine the data prefetch address. If it exists, the previously determined data prefetch address is directly obtained in the prefetch training unit. After obtaining the data prefetch address, by comparing the difference between the training prefetch result (data prefetch address) and the actual memory access address of the process, the prefetch accuracy is determined. When it is determined that the prefetch accuracy reaches the preset threshold, prefetch instructions are sent to multiple levels of caches simultaneously with different look-ahead degrees to prefetch data into the cache in advance, so as to improve the efficiency of subsequent data access.

[0103] Among them, when the corresponding process ends, the prefetch training unit will send the page table base address, the selected prefetch result and the hit situation to the completion table, and increase or decrease the score of the corresponding prefetch algorithm in the initial configuration register according to the situation, that is, continue to evaluate and maintain the score table. The initial configuration register can also be adjusted by software to assign values to the register, and software recommendation instructions are added to adjust and accelerate the prefetch algorithm. The score of the prefetch algorithm is positively correlated with its prefetch accuracy. That is, the method provided by the embodiment of the present application supports the regulation function of software on the prefetch unit. A software configuration port is set in the initial configuration register. Since the software layer has a better understanding of the characteristics of the memory access stream, in some cases, the characteristics of the instruction stream memory access can be obtained in advance. Therefore, the prefetch unit can be configured accordingly to reduce the training cycle and the initial hit times, so as to perform prefetch quickly and reduce useless losses.

[0104] Specifically, in one embodiment, the data prefetch addresses under multiple different look-ahead degrees can be sent to the high-level cache or the low-level cache according to the target data prefetch policy and the process priority of the current memory access request;

[0105] Among them, the preset cache is divided into two parts: a high-level cache and a low-level cache.

[0106] Specifically, the emitter can send the prefetch results (target data prefetch policy) of frequently used processes to the high-level cache (upper-level cache) and those of less frequently used processes to the low-level cache (lower-level cache) according to the information in the process staging area and the prefetch training unit, and adjust the prefetch degree accordingly based on the cache hit situation feedback by the monitor. Among them, the more frequently a process is used, the higher its process priority. Among them, the upper-level cache is closer to the host CPU, with a short data transmission path and faster access speed, and can respond to data requests from the host in an extremely short time; the lower-level cache is farther from the host CPU, with a long data transmission time and relatively slower access speed, but the cache capacity of the lower-level cache is larger than that of the upper-level cache.

[0107] Exemplarily, as Figure 5 shown, it is a schematic structural diagram of another exemplary data prefetch system provided by an embodiment of the present application. The prefetch training unit includes sub-units under different look-ahead degrees, and each sub-unit includes an address selector, a threshold checker, a prefetch register bank, and a status register. Among them, the address selector is used to select the data prefetch address with the best comprehensive accuracy and hit count and send it out; the threshold checker is used to check whether relevant parameters reach a preset threshold to judge the rationality of the prefetch operation; the prefetch register bank contains the same number of prefetch registers as the number of algorithms in the prefetch algorithm unit, and is used to record virtual page table addresses, prefetch information, and simulated prefetch hit situations, etc.; the status register records the status information during the prefetch training process, and two counters are set in the status register, one for recording the number of unused times and one for recording the training cycle. The sub-modules under different look-ahead degrees can implement different degrees of forward-looking prefetch training to adapt to different memory access patterns. The prefetch configuration unit includes a non-training queue, a completion table, and an initial configuration register. Among them, the non-training queue is used to cache abandoned processes, and the completion table is used to cache interrupted processes. When the same page table base address as the interrupted process is encountered again, the training cycle and the initial hit count are configured accordingly in the initial configuration register to reduce the training cycle of this process. The replacement register is used for process replacement. When the number of parallel page table base addresses is too large and there are insufficient internal entries in the prefetch training unit, the replacement register replaces the least used process into the process staging area according to the information in the unused counter, and records the correspondence between the virtual page table address and the unit storage address (predicted data prefetch address) for subsequent retrieval. The process staging area is configured in the lower-level cache of the preset cache to have a larger storage space, and is replaced into the prefetch training unit when the replaced process is used.

[0108] Specifically, in one embodiment, the memory access latency of the data prefetch address in the preset cache is obtained; when the memory access latency of the data prefetch address reaches the preset latency condition, the target data prefetch algorithm is updated to replace the target data prefetch algorithm corresponding to the target process to which the current memory access request belongs.

[0109] It should be noted that the memory access latency refers to the time elapsed from the request to access the data prefetch address to the actual acquisition of the corresponding data. A large memory access latency indicates that the prediction of the expected data address is late. The host has initiated a memory access request, but the prefetch data has not been successfully prefetched.

[0110] Specifically, when the memory access latency of the data prefetch address reaches the preset latency condition, it indicates that the current target data prefetch algorithm cannot well meet the data access requirements. At this time, the target data prefetch algorithm corresponding to the target process to which the current memory access request belongs can be replaced. Since different prefetch algorithms are applicable to different memory access modes, the algorithm with the second highest prefetch accuracy among the candidate algorithms of the target process can be used as the target data prefetch algorithm at this time. By replacing the algorithm, it is possible to find a prefetch strategy more suitable for the memory access characteristics of the current process, thereby reducing the memory access latency and improving the data access efficiency.

[0111] Among them, as Figure 6 shown, it is a schematic structural diagram of the prefetch emission unit provided by the embodiment of the present application. The prefetch emission unit includes a cache monitor, a prefetch degree regulator, a prefetch address buffer, a transmitter, and an emitted sequence. The data prefetch address and related information sent by the prefetch training unit are first temporarily stored in the prefetch address buffer, and then extracted according to the prefetch degree sent by the prefetch degree regulator and sent to the transmitter according to the look-ahead degree. The transmitter sends the prefetch address of the ongoing process to the high-level cache according to the process temporary storage area and the virtual page table address entry information in the prefetch training unit, sends the prefetch address of the process that has not been performed for a long time to the low-level cache, and stores the sent prefetch address in the emitted queue. The set cache monitor monitors the hit rate and replacement rate of the cache and feeds back the situation to the prefetch degree regulator, and the prefetch degree regulator adaptively adjusts the prefetch degree according to the set threshold. The data recorded in the emitted sequence will be sent to the prefetch address buffer and corresponding flag bits are set to filter duplicate requests for prefetch results with different look-aheads in the transmitter. At the same time, the cache monitor will also check whether the subsequent memory access requests corresponding to the prefetch address have been executed. For the delayed memory access requests, the corresponding counter is incremented by 1. When the value is reached, it is determined that the requests issued by the prefetch are generally delayed requests, that is, the memory access latency of the data prefetch address reaches the preset latency condition, and the subsequent operations are stopped and the second-best prefetch (using the algorithm with the second highest prefetch accuracy among the candidate algorithms of the target process as the target data prefetch algorithm) is used, and so on.

[0112] The data prefetching method provided by the embodiment of the present application includes: obtaining the memory access information of all currently running processes; for any process, training the data prefetching algorithm corresponding to the process according to the memory access information of the process to obtain a data prefetching algorithm set; when the current cache state of the preset cache meets the preset prefetching condition, obtaining the current memory access request and determining the target process to which the current memory access request belongs; screening the target data prefetching algorithm from the data prefetching algorithm set according to the target process; and determining the target data prefetching strategy based on the target data prefetching algorithm and according to the memory access address information of the current memory access request. In the method provided by the above solution, after obtaining the memory access information of all currently running processes, the corresponding data prefetching algorithm for each process is independently trained to form an algorithm set. When performing data prefetching, the appropriate algorithm is screened from the algorithm set according to the target process to which the current memory access request belongs, improving the accuracy of the data prefetching result and laying a foundation for the utilization rate of the data prefetching resources of the computing device. Moreover, the processes are distinguished by using the page table base address, and prefetching is performed based on the memory access characteristics of the same process, avoiding the destruction of the program memory access locality in the complex computing environment under multiple cores and multiple threads. The virtual-to-real conversion of the page table address is used to reduce the storage space occupied in multiple internal tables, enabling implementation in the limited on-chip space. The initial situation is configured using software and the completed information, and the processes that cannot effectively identify the memory access pattern are excluded, which is beneficial to accelerating the prefetching training speed and providing more prefetching opportunities. Filtering the prefetching requests with delays and low precision, avoiding excessive additional costs to the bandwidth, and alleviating the problem of limited storage bandwidth of the shared cache. Multiple prefetching requests are sent to caches at different levels according to different look-aheads and process usage frequencies, making full use of the prefetching opportunities, increasing the probability of finding the data obtained through prefetching in the cache for subsequent memory accesses, and improving the memory access performance.

[0113] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0114] The embodiment of the present application also provides a data prefetching device for executing the data prefetching method provided by the above embodiment.

[0115] As Figure 7 shown, it is a schematic structural diagram of the data prefetching device provided by the embodiment of the present application. The data prefetching device 70 includes: an obtaining module 701, a training module 702, a determining module 703, a screening module 704, and a prefetching module 705.

[0116] Among them, an acquisition module is configured to acquire the memory access information of all currently running processes; a training module is configured to, for any process, train the data prefetch algorithm corresponding to the process according to the memory access information of the process to obtain a data prefetch algorithm set; a determination module is configured to, when the current cache state of a preset cache meets a preset prefetch condition, acquire a current memory access request and determine the target process to which the current memory access request belongs; a screening module is configured to screen a target data prefetch algorithm from the data prefetch algorithm set according to the target process; and a prefetch module is configured to, based on the target data prefetch algorithm and according to the memory access address information of the current memory access request, determine a target data prefetch policy.

[0117] For the description of the features in the embodiments corresponding to the data prefetch device, reference may be made to the relevant descriptions in the embodiments corresponding to the data prefetch method, which will not be elaborated here one by one.

[0118] An embodiment of the present application further provides an electronic device, as Figure 8 shown, is a schematic structural diagram of the electronic device provided by the embodiment of the present application, including a processor 10 and a memory 20. A computer program is stored in the memory 20, and the processor 10 is configured to run the computer program to execute the steps in any one of the above data prefetch method embodiments.

[0119] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above data prefetch method embodiments when running.

[0120] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc, and other various media that can store computer programs.

[0121] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above data prefetch method embodiments.

[0122] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any one of the above data prefetch method embodiments.

[0123] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0124] The above has introduced in detail a data prefetching method, apparatus, electronic device, and storage medium provided by this application. Specific examples are used herein to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A data pre-fetching method, characterized in that: include: Get the memory access information of all currently running processes; For any of the processes, according to the memory access information of the process, a data prefetching algorithm corresponding to the process is trained to obtain a data prefetching algorithm set; When the current cache state of the preset cache meets the preset pre-fetch condition, obtaining the current memory access request, and determining the target process to which the current memory access request belongs; According to the target process, a target data prefetching algorithm is selected from the data prefetching algorithm set; Based on the target data pre-fetching algorithm, according to the memory access address information of the current memory access request, determining the target data pre-fetching strategy; The step of training a data prefetching algorithm corresponding to any of the processes according to the memory access information of the process to obtain a data prefetching algorithm set includes: Get the initial algorithm set; For any of the processes, training each initial algorithm in the initial algorithm set according to the memory access information of the process; When the preset training cycle is reached, the data prefetching algorithm corresponding to the process is determined according to the prefetching accuracy of each trained algorithm; According to the data prefetching algorithm corresponding to each process, a data prefetching algorithm set is constructed.

2. The data pre-fetching method according to claim 1, characterized in that: The step of training each initial algorithm in the initial algorithm set according to the memory access information of any of the processes comprises: For any of the processes, determining a memory access rule of the process according to memory access address information represented by memory access information of the process; According to the memory access rule of the process, each initial algorithm in the initial algorithm set is trained to optimize the configuration parameters of each initial algorithm.

3. The data pre-fetching method according to claim 1, characterized in that: When the preset training cycle is reached, the data prefetching algorithm corresponding to the process is determined according to the prefetching accuracy of each trained algorithm, including: In the process of training each initial algorithm in the initial algorithm set, obtaining the actual memory access address of the process; When the preset training cycle is reached, the training pre-fetch results of each post-training algorithm are obtained; For any of the trained algorithms, determining the prefetch accuracy of the trained algorithm according to a difference between a training prefetch result of the trained algorithm and an actual memory access address of the process; The trained algorithm whose pre-fetching accuracy reaches a preset threshold is used as a candidate algorithm; The selected algorithm with the highest pre-fetching accuracy is used as the data pre-fetching algorithm corresponding to the process.

4. The data pre-fetching method according to claim 3, characterized in that: The method further comprises: For any of the processes, if all trained algorithms of the process have been trained for multiple rounds of preset training cycles and the pre-fetching accuracy still does not reach the preset threshold, the process is regarded as an abandoned process; Abandoning the training of the data prefetch algorithm of the abandoned process.

5. The data pre-fetching method according to claim 3, characterized in that: The method further comprises: When any of the processes enters a preset interrupt state, the process is treated as an interrupt process; The training of the interruption process is suspended, and the training algorithm set of the interruption process and the training pre-fetch results generated by each post-trained algorithm in the training algorithm set during the training process are packaged and processed to obtain a training interruption data packet of the interruption process; Storing the training interruption data packet in a preset temporary storage area; When any of the interruption processes enters a preset active state, extracting a training interruption data packet corresponding to the interruption process in the preset temporary storage area; Based on the training interruption data packet, continue to train the post-training algorithm set of the interruption process.

6. The data pre-fetching method according to claim 5, characterized in that: The method further comprises: When any of the processes becomes a completed process or a process to be replaced, determining that the process enters a preset interrupt state; When a new process is generated and the number of parallel processes reaches a preset parallel threshold, the least used process among the original processes is used as the process to be replaced.

7. The data pre-fetching method according to claim 6, characterized in that: The method further comprises: Obtaining a page table base address of the interrupt process in a preset page table base address register; Converting the page table base address into a virtual page table address; wherein the address length of the virtual page table address is smaller than the address length of the page table base address; Adding the virtual page table address to the training interrupt data packet; When the virtual page table address of the new process matches the virtual page table address of any of the interrupted processes, it is determined that the interrupted process enters a preset active state.

8. The data pre-fetching method according to claim 1, characterized in that: The method further comprises: When the cache hit rate represented by the current cache state of the preset cache is lower than a preset hit rate threshold, it is determined that the current cache state of the preset cache meets the preset pre-fetch condition.

9. The data pre-fetching method according to claim 1, characterized in that: The method further comprises: When the remaining cache space represented by the current cache state of the preset cache is not less than a preset remaining cache space threshold, it is determined that the current cache state of the preset cache meets the preset pre-fetch condition.

10. The data pre-fetching method according to claim 1, characterized in that: The method further comprises: When the current cache state of the preset cache indicates that the current memory access request misses the cache, it is determined that the current cache state of the preset cache meets the preset pre-fetch condition.

11. The data pre-fetching method according to claim 1, characterized in that: The obtaining of the current memory access request and determining the target process to which the current memory access request belongs includes: After obtaining the current memory access request, analyzing the current memory access request to determine a page table base address of the current memory access request; The target process to which the current memory access request belongs is determined according to the page table base address.

12. The data pre-fetching method according to claim 1, characterized in that: Before determining the target data prefetching strategy based on the target data prefetching algorithm and according to the memory access address information of the current memory access request, the method further includes: Determining the prefetching degree of the target data prefetching algorithm according to a cache hit rate and a cache replacement rate represented by a current cache state of the preset cache; The target data pre-fetching algorithm is negatively correlated with both the cache hit rate and the cache replacement rate.

13. The data pre-fetching method according to claim 1, characterized in that: The method of determining a target data prefetching strategy based on the target data prefetching algorithm and according to the memory access address information of the current memory access request includes: Based on the target data prefetch algorithm, and according to the memory access address information of the current memory access request, determining a plurality of data prefetch addresses at different degrees of foresight; Determining a target data prefetching strategy according to the data prefetching addresses at the plurality of different foresight levels; The target data pre-fetching strategy at least includes the data pre-fetching addresses at the multiple different foresight levels and the order in which the data pre-fetching addresses are issued.

14. The data pre-fetching method according to claim 13, characterized in that: The method further comprises: According to the target data prefetching strategy and the process priority of the current memory access request, the data prefetching addresses at the multiple different look-ahead degrees are sent to the high-level cache or the low-level cache; The preset cache is divided into two parts: a high-level cache and a low-level cache.

15. The data pre-fetching method according to claim 14, characterized in that: The method further comprises: Obtaining a memory access delay of a data prefetch address in the preset cache; When the memory access delay of the data prefetch address reaches a preset delay condition, the target data prefetch algorithm is updated to replace the target data prefetch algorithm corresponding to the target process to which the current memory access request belongs.

16. A data pre-fetching device, characterized in that: include: The acquisition module is used to obtain the memory access information of all currently running processes; A training module, used for training a data prefetching algorithm corresponding to any of the processes according to the memory access information of the process, so as to obtain a data prefetching algorithm set; A determination module, configured to obtain a current memory access request and determine a target process to which the current memory access request belongs when a current cache state of the preset cache satisfies a preset pre-fetch condition; A screening module, used for screening a target data prefetching algorithm in the data prefetching algorithm set according to the target process; A prefetch module, configured to determine a target data prefetch strategy based on the target data prefetch algorithm and the memory access address information of the current memory access request; The training module is specifically used for: Get the initial algorithm set; For any of the processes, training each initial algorithm in the initial algorithm set according to the memory access information of the process; When the preset training cycle is reached, the data prefetching algorithm corresponding to the process is determined according to the prefetching accuracy of each trained algorithm; According to the data prefetching algorithm corresponding to each process, a data prefetching algorithm set is constructed.

17. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data prefetching method according to any one of claims 1 to 15 when executing the computer program.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the data pre-fetching method according to any one of claims 1 to 15.

19. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the data pre-fetching method according to any one of claims 1 to 15 are implemented.

Citation Information

Patent Citations

  • Data prefetching method and related equipment

    CN117519571A

  • Method for intelligently prefetching data of storage system

    CN119473144A