Data prefetching method based on reinforcement learning and related equipment

Through the data prefetching method based on reinforcement learning, the problem of low accuracy and timeliness of data prefetching in the prior art is solved, effective capture and pre-retrieval decision adjustment of complex memory access modes is realized, and memory performance is improved.

CN119938555AActive Publication Date: 2025-05-06湖南工商大学
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510429962.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-06
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

When faced with complex memory access modes, existing data prefetching technologies have low accuracy and timeliness, making it difficult to effectively utilize caches.

Method used

The data prefetching method based on reinforcement learning is adopted, and the prefetching step is calculated by obtaining the memory access page number of the target computer, and the reinforcement learning model is used to determine the memory prefetching action and reward function values, and the model is trained to finally realize data prefetching.

Benefits of technology

Improves the accuracy and timeliness of data prefetching, and can more effectively capture complex memory access patterns and adjust prefetch decisions, thereby improving memory performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938555A_ABST
    Figure CN119938555A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data prefetching, and provides a data prefetching method based on reinforcement learning and related equipment. The method comprises the following steps: acquiring memory access page numbers of a target computer at multiple moments; calculating a plurality of prefetch steps based on all memory access page numbers; according to all the memory access page numbers and all the prefetch steps, determining a memory prefetch action by using a reinforcement learning model, and calculating a reward function value corresponding to the memory prefetch action; training the reinforcement learning model according to the memory prefetching action and the reward function value to obtain a final reinforcement learning model; and obtaining a final memory prefetching action of the target computer at the current moment by utilizing the final reinforcement learning model, and performing data prefetching on the target computer according to the final memory prefetching action. According to the method, the accuracy and timeliness of data prefetching can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data pre-fetching, and in particular to a data pre-fetching method based on reinforcement learning and related equipment. Background Art

[0002] Memory is an integral part of computer architecture. It is responsible for storing processor executable code and computing data while applications are running. Its performance has a significant impact on the overall computing speed and stability of the computer. During program execution, computing units such as the central processing unit (CPU) need to constantly interact with memory to read instructions, load data, and complete data processing. However, there is a significant speed mismatch between the processor and memory. In addition, the improvement of processor performance far exceeds the improvement of memory performance, which leads to a serious computer performance bottleneck, the well-known "memory wall" problem.

[0003] There are currently many ways to improve and optimize memory performance: 1. Multi-level memory structure. By designing a multi-level cache between the computing unit and the memory, the cache closer to the computing unit has a faster access speed, and vice versa, the access speed is slower. This ensures that the computing unit always interacts with its nearest cache, and then stores the calculation results in the memory and disk through the intermediate cache; 2. Memory separation technology. The current computer design tightly binds the CPU and memory together in physical structure, thereby limiting the memory resources that the CPU can use and reducing the memory utilization. Therefore, the memory resources are pooled and the memory resources that the CPU can access are flexibly allocated, which greatly improves the memory performance and utilization; 3. Data prefetching technology. When the CPU needs to change the data in a certain block of memory, it takes a lot of time from issuing an instruction to returning the memory data of the instruction, which easily causes the computing resources to always be in a waiting state. By analyzing the memory data previously processed by the CPU to predict the memory unit that will be used later, and taking it to the cache closer to the CPU in advance, this problem can be greatly solved.

[0004] Data prefetching usually requires identifying the access patterns and memory access characteristics of different applications from historical data, so as to accurately predict and obtain the memory data needed later. Current applications have different memory access patterns, which are usually divided into two types: regular access and random access. Regular access: 1. Sequential access. Applications access memory addresses in sequence, such as increasing or decreasing regularly. This memory access pattern has good locality characteristics and can greatly improve the efficiency of data prefetching. 2. Jump access. Applications access memory addresses according to a certain pattern as a whole, but have jump access characteristics in time or space. For example, they will jump to another memory address that is far away at a certain point in time. This usually exists in matrix and array operations. Random access: Applications access memory address space in an irregular order, with poor locality. It is difficult for commonly used prefetching technologies to learn this access pattern, and most prefetching operations are invalid, resulting in ineffective use of cache.

[0005] Traditional data prefetching can achieve good prefetching effects when there is high spatial locality and temporal locality of memory access. Common ones include: 1. Stride prefetcher, which identifies the memory access strides of different streams by saving a prefetch cache window for different memory access sequences. However, this method of dividing the prefetch cache separately for each memory access sequence brings additional storage overhead; 2. Offset prefetcher, which configures a unified prefetch offset window for all memory access sequences. For example, if the offset variable is d, then the content that may be saved in this window is [d, 2d, 3d, ...], which saves storage resources and can more accurately identify the overall memory access characteristics; 3. Spatial prefetcher, which uses the feature that a sequence that has been accessed in the past may be accessed again to identify and cache the memory access page data with prefetching value, so as to timely re-attempt the memory access operations that may occur again. However, these traditional prefetchers can achieve good results when facing programs with relatively simple memory access patterns. But as the scale and complexity of current artificial intelligence and big data applications are growing exponentially, the more complex memory access patterns they bring about cause traditional prefetchers to have problems with low accuracy and timeliness of data prefetching. Summary of the invention

[0006] The present application provides a data pre-fetching method and related equipment based on reinforcement learning, which can solve the problems of low accuracy and timeliness of data pre-fetching.

[0007] In a first aspect, an embodiment of the present application provides a data pre-fetching method based on reinforcement learning, the data pre-fetching method comprising: Obtaining memory access page numbers of a target computer at multiple times; the memory access page number is the number of the memory address accessed by the target computer; Calculate multiple prefetch strides based on all memory access page numbers; the prefetch stride is used to describe the difference between the memory access page numbers of two adjacent moments; According to all memory access page numbers and all prefetch strides, a reinforcement learning model is used to determine the memory prefetch action and calculate the reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of data prefetches according to the target prefetch stride; The reinforcement learning model is trained according to the memory pre-fetch action and the reward function value to obtain the final reinforcement learning model; The final reinforcement learning model is used to obtain the final memory prefetch action of the target computer at the current moment, and data is prefetched for the target computer according to the final memory prefetch action.

[0008] Optionally, multiple prefetch strides are calculated based on all memory access page numbers, including: Determine a plurality of current memory access page numbers from all memory access page numbers; For each two adjacent current memory access page numbers, a difference between the two adjacent current memory access page numbers is calculated, and the calculated difference is used as the prefetch stride corresponding to the two adjacent current memory access page numbers.

[0009] Optionally, multiple current memory access page numbers are determined from all memory access page numbers, including: When the number of all memory access pages is greater than the preset number of pages, the previous N All memory access page numbers are used as the current memory access page number; N Indicates the preset number of page numbers; When the number of all memory access page numbers is less than or equal to the preset number of page numbers, all memory access page numbers are used as current memory access page numbers.

[0010] Optionally, a reinforcement learning model is used to determine memory prefetch actions based on all memory access page numbers and all prefetch strides, including: The memory access page numbers corresponding to all prefetch strides are used as the current environment state; Determines memory prefetch actions based on the current environment state and all prefetch strides.

[0011] Optionally, calculate the reward function value corresponding to the memory prefetch action, including: By formula: Calculate the reward function value : in, Indicates the impact of memory prefetching on memory access operations. Decision values ​​representing confidence: ; ; in, represents the weight factor of the delayed consistency reward, represents the maximum number of time steps to consider delay, k represents the index of future time steps, T(t+k) represents the time when memory access occurs in the future k time steps after the memory prefetch action is executed, and T(t) represents the time when the current memory access occurs. represents the confidence reward weight, Indicates the current environment status. Indicates memory prefetch action, represents information entropy, represents the maximum possible information entropy, The Q value entropy representing the memory prefetch action is: ; ; ; in, represents the normalized probability, and All action sets A subset of represents the entropy temperature control factor.

[0012] Optionally, the reinforcement learning model is trained according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model, including: The current environment state, memory prefetch action, reward function value, and environment state after executing the memory prefetch action are stored as a piece of data in the experience replay pool; Determine whether the amount of data in the experience replay pool has reached the preset amount; If the number of data in the experience replay pool reaches the preset number, a target data is extracted from all the data in the experience replay pool, and the target data is used to train the reinforcement learning model to obtain the trained reinforcement learning model. The number of iterations is increased by 1, and it is determined whether the number of iterations is greater than or equal to the preset number of iterations; If the number of iterations is greater than or equal to the preset number of iterations, the trained reinforcement learning model is used as the final reinforcement learning model; If the amount of data in the experience replay pool does not reach the preset number or the number of iterations is less than the preset number of iterations, the multiple memory access page numbers contained in the environment state after the memory prefetch action is executed are used as all memory access page numbers in the step of calculating multiple prefetch strides based on all memory access page numbers, and the step of calculating multiple prefetch strides based on all memory access page numbers is returned.

[0013] Optionally, the target data is used to train the reinforcement learning model to obtain a trained reinforcement learning model, including: The target data is used to update the Q-value function in the Q network in the reinforcement learning model, and the reinforcement learning model is updated according to the updated Q network to obtain a trained reinforcement learning model.

[0014] Optionally, the target data is used to update the Q-value function in the Q network in the reinforcement learning model, including: By formula: ; Update the Q value function; in, represents the Q-value function, represents the time distribution corresponding to the target data, , represents the tracking attenuation factor, represents the discount factor, Indicates the time distribution corresponding to the previous data of the target data, Indicates the current state of the environment in the target data, Indicates the memory prefetch action in the target data, represents the reward function value in the target data, and Both represent weight factors, and All action sets A subset of represents the maximum Q value that can be achieved in the next environment state of the target data, It means that at time step t+1, Uncertainty about the distribution of all possible actions a; Update the reinforcement learning model based on the updated Q network, including: By formula: Update the parameters in the reinforcement learning model; in, represents the parameters of the reinforcement learning model, represents a small discount factor, Represents the updated parameters of the Q network.

[0015] In a second aspect, an embodiment of the present application provides a data pre-fetching device based on reinforcement learning, comprising: A first acquisition module is used to acquire the memory access page number of the target computer at multiple times; the memory access page number is the number of the memory address accessed by the target computer; A calculation module, used for calculating a plurality of prefetch strides based on all memory access page numbers; the prefetch stride is used for describing the difference between the memory access page numbers at two corresponding adjacent moments; A determination module is used to determine a memory prefetch action based on all memory access page numbers and all prefetch strides using a reinforcement learning model, and calculate a reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of data prefetches according to the target prefetch stride; A training module is used to train the reinforcement learning model according to the memory pre-fetch action and the reward function value to obtain the final reinforcement learning model; The second acquisition module is used to use the final reinforcement learning model to acquire the final memory prefetch action of the target computer at the current moment, and prefetch data for the target computer according to the final memory prefetch action.

[0016] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned reinforcement learning-based data prefetching method when executing the above-mentioned computer program.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned reinforcement learning-based data prefetching method.

[0018] The above solution of the present application has the following beneficial effects: In an embodiment of the present application, by obtaining the memory access page number of the target computer at multiple times, then calculating multiple pre-fetch strides based on all memory access page numbers, and then determining the memory pre-fetch action based on all memory access page numbers and all pre-fetch strides using a reinforcement learning model, and calculating the reward function value corresponding to the memory pre-fetch action, then training the reinforcement learning model based on the memory pre-fetch action and the reward function value to obtain a final reinforcement learning model, and finally using the final reinforcement learning model to obtain the final memory pre-fetch action of the target computer at the current moment, and pre-fetching data for the target computer according to the final memory pre-fetch action. Among them, reinforcement learning has advantages and dynamic decision-making capabilities in high-dimensional nonlinear data representation, can learn valuable pre-fetch features from complex memory access conditions, capture the computer memory access mode at multiple times and adjust the pre-fetch decision, thereby improving the accuracy of the memory pre-fetch action, training the reinforcement learning model, can improve the performance of the reinforcement learning model, using the final reinforcement learning model to obtain the final memory pre-fetch action, and pre-fetching data for the target computer at the current moment, can effectively improve the accuracy and timeliness of data pre-fetching.

[0019] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 A flowchart of a data pre-fetching method based on reinforcement learning provided in one embodiment of the present application; Figure 2 A schematic diagram of prefetch stride provided in an embodiment of the present application; Figure 3 A schematic diagram of reinforcement learning provided in an embodiment of the present application; Figure 4 A schematic diagram of the structure of a data pre-fetching device based on reinforcement learning provided in one embodiment of the present application; Figure 5 A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0022] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0023] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0024] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0025] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0026] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0027] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0028] In response to the problems of low accuracy and timeliness of existing data prefetching, an embodiment of the present application provides a data prefetching method based on reinforcement learning. The data prefetching method obtains the memory access page numbers of the target computer at multiple times, and then calculates multiple prefetching steps based on all memory access page numbers. Then, based on all memory access page numbers and all prefetching steps, a reinforcement learning model is used to determine the memory prefetching action and calculate the reward function value corresponding to the memory prefetching action. Then, the reinforcement learning model is trained based on the memory prefetching action and the reward function value to obtain the final reinforcement learning model. Finally, the final reinforcement learning model is used to obtain the final memory prefetching action of the target computer at the current moment, and data is prefetched for the target computer based on the final memory prefetching action. Among them, reinforcement learning has advantages in high-dimensional nonlinear data representation and dynamic decision-making capabilities. It can learn valuable pre-fetch features from complex memory access conditions, capture the computer's memory access patterns at multiple times and adjust pre-fetch decisions, thereby improving the accuracy of memory pre-fetch actions. Training the reinforcement learning model can improve the performance of the reinforcement learning model. Using the final reinforcement learning model to obtain the final memory pre-fetch action and pre-fetch data for the target computer at the current moment can effectively improve the accuracy and timeliness of data pre-fetching.

[0029] Next, the data pre-fetching method based on reinforcement learning provided in this application is exemplified.

[0030] like Figure 1 As shown, the data pre-fetching method based on reinforcement learning provided by the present application includes the following steps: Step 11, obtaining the memory access page numbers of the target computer at multiple times.

[0031] The memory access page number is the number of the memory address accessed by the target computer (the memory address accessed at each moment may be different, and the memory access page number at each moment is the number of the memory address accessed by the target computer at that moment). The target computer is the computer that needs to pre-fetch data.

[0032] In some embodiments of the present application, the accessed memory address may be first obtained, and then the memory address may be shifted right to obtain the memory access page number of the memory address.

[0033] It should be noted that the number of bits by which the memory address is shifted right is determined by the number of bytes in the memory address itself.

[0034] Exemplarily, the memory address defaults to the cacheline granularity, that is, 64 bytes, and a page defaults to 4KB, that is, each memory address is shifted right by 12 bits to obtain the page number of the memory access. A record buffer can be maintained to store multiple memory addresses of the target computer, and when performing this step, the memory address is read from the record buffer.

[0035] It is worth mentioning that using memory access page numbers for data prefetching can greatly reduce the overhead caused by traditional data prefetchers storing complete memory addresses.

[0036] Step 12, calculate multiple prefetch strides based on all memory access page numbers.

[0037] The prefetch stride is used to describe the difference between two corresponding memory access page numbers, and the corresponding two memory access page numbers are two memory access page numbers required to calculate the prefetch stride.

[0038] In some embodiments of the present application, the above step of calculating multiple prefetch strides based on all memory access page numbers is specifically:

[0039] In the first step, multiple current memory access page numbers are determined from all memory access page numbers.

[0040] Specifically, when the number of all memory access page numbers is greater than the preset number of page numbers, the first N memory access page numbers are used as the current memory access page numbers in the order of all moments; N represents the preset number of page numbers.

[0041] When the number of all memory access page numbers is less than or equal to the preset number of page numbers, all memory access page numbers are used as current memory access page numbers.

[0042] Exemplarily, the memory access page numbers are 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 85, 87, 95, 96, 104, 105, and N is equal to 20. Then, all target memory access page numbers are 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87.

[0043] In the second step, for each two adjacent current memory access page numbers, the difference between the two adjacent current memory access page numbers is calculated, and the calculated difference is used as the prefetch stride corresponding to the two adjacent current memory access page numbers.

[0044] For example, if the first current memory access page number is 100 and the second current memory access page number is 150, the prefetch stride between the two is 50. Figure 2 As shown in the figure, the horizontal axis represents the execution time of data prefetching, and the vertical axis represents the inter-page stride (ie, prefetch stride).

[0045] Step 13: According to all memory access page numbers and all prefetch strides, a reinforcement learning model is used to determine the memory prefetch action, and the reward function value corresponding to the memory prefetch action is calculated.

[0046] The above-mentioned memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times data is prefetched according to the target prefetch stride.

[0047] In some embodiments of the present application, the steps of determining the memory prefetch action using the reinforcement learning model according to all memory access page numbers and all prefetch strides, and calculating the reward function value corresponding to the memory prefetch action include: In the first step, the memory access page numbers corresponding to all prefetch strides are taken as the current environment state.

[0048] Exemplarily, from step 12, it can be seen that the memory access page numbers corresponding to all prefetch strides are all current memory access page numbers, and all current memory access page numbers are 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87, respectively. Then the current environment state is 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87.

[0049] In the second step, the memory prefetch action is determined based on the current environment state and all prefetch strides.

[0050] For example, a strategy in reinforcement learning (such as a strategy in a deep Q-network (DQN)) can be used to select a target prefetch step from all prefetch steps and generate a prefetch degree. For example, a Q-value-based strategy is currently used to calculate a Q-value for each possible prefetch step using a Q-network trained by DQN, that is, is 15, is 10, is 2, then the prefetch stride 0->9 with the largest Q value is selected, and the model believes that the number of times the prefetch stride 9 appears in the next few memory accesses is 1, so the final prefetch degree is determined to be 1. If the current environment state is 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87, the target prefetch stride is 9, and the prefetch degree is 1, it means that the data in the memory area corresponding to the memory access page 96 needs to be prefetched based on the memory access page number 87. After executing this action, the environment state becomes 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 85, 87, 95, 96.

[0051] The third step is to calculate the reward function value corresponding to the memory prefetch action.

[0052] Specifically, through the formula: Calculate the reward function value : in, Indicates the impact of memory prefetching on memory access operations. Decision values ​​representing confidence: ; ; in, represents the weight factor of the delayed consistency reward, Indicates the maximum number of time steps to consider delays, controls the scope of consideration of future time steps, k indicates the index of future time steps, and represents the time span of delay impact. For example, for a prefetch action made at the tth time step, k=1 indicates the memory access of the next time step (i.e., t+1). If K=3, the impact of the current prefetch behavior on the next three time steps (i.e., t+1, t+2, t+3) is considered, while more distant time steps are ignored. A larger K can increase the consideration of the impact of long-term delays, while a smaller K focuses on the impact within a shorter time range; represents the time scaling factor, which is used to control the impact of delay on reward calculation. T(t+k) represents the time when memory access occurs in the next k time steps after the memory prefetch action is executed. T(t) represents the time when the current memory access occurs. For example, at time t=5, the memory access of the current page number 87 occurs at time T(5), and when the target prefetch stride is 9, the memory access of the target page number 96 will occur at time T(6) (i.e., the next time step after prefetch). represents a time step or time point in the current reinforcement learning environment, Represents the confidence reward weight, which is used to adjust the impact of confidence in the reward. Indicates the current environment status. Indicates memory prefetch action, represents information entropy, represents the maximum possible information entropy, The Q value entropy representing the memory prefetch action is: ; ; ; in, represents the normalized probability, and All action sets A subset of represents the entropy temperature control factor.

[0053] It can be understood that the process recorded in the above step 13 is the data processing process in the reinforcement learning model.

[0054] The following is an illustrative example of reinforcement learning.

[0055] like Figure 3 As shown in the figure, the agent (i.e., the reinforcement learning model) obtains actions through strategies, interacts with the environment, and transmits rewards and states to the agent. The agent again generates actions based on the states and obtains the next state and rewards after interacting with the environment.

[0056] It is worth mentioning that reinforcement learning has advantages in high-dimensional nonlinear data representation and dynamic decision-making capabilities. It can learn valuable pre-fetching features from complex memory access conditions, capture the computer's memory access patterns at multiple times and adjust pre-fetching decisions.

[0057] Step 14: Train the reinforcement learning model according to the memory pre-fetch action and the reward function value to obtain the final reinforcement learning model.

[0058] In some embodiments of the present application, the step of training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain the final reinforcement learning model includes: The first step is to store the current environment state, memory prefetch action, reward function value, and environment state after executing the memory prefetch action as a piece of data in the experience replay pool; The second step is to determine whether the amount of data in the experience replay pool has reached the preset amount.

[0059] If the number of data in the experience replay pool reaches the preset number, a target data is extracted from all the data in the experience replay pool, and the reinforcement learning model is trained using the target data to obtain the trained reinforcement learning model. The number of iterations is increased by 1, and it is determined whether the number of iterations is greater than or equal to the preset number of iterations.

[0060] If the number of iterations is greater than or equal to the preset number of iterations, the trained reinforcement learning model is used as the final reinforcement learning model.

[0061] If the amount of data in the experience replay pool does not reach the preset number or the number of iterations is less than the preset number of iterations, the multiple memory access page numbers contained in the environment state after the memory prefetch action is executed are used as all memory access page numbers in the step of calculating multiple prefetch strides based on all memory access page numbers, and the step of calculating multiple prefetch strides based on all memory access page numbers is returned.

[0062] It should be noted that the initial number of iterations is 0.

[0063] The above-mentioned steps of training the reinforcement learning model with the target data to obtain the trained reinforcement learning model are specifically: updating the Q-value function in the Q network in the reinforcement learning model with the target data, and updating the reinforcement learning model according to the updated Q network to obtain the trained reinforcement learning model.

[0064] By formula: ; Update the Q-value function.

[0065] in, represents the Q-value function, represents the time distribution corresponding to the target data, , represents the tracking attenuation factor, represents the discount factor, Indicates the time distribution corresponding to the previous data of the target data, Indicates the current state of the environment in the target data, Indicates the memory prefetch action in the target data, represents the reward function value in the target data, and Both represent weight factors, and All action sets A subset of represents the maximum Q value that can be achieved in the next environment state of the target data, It means that at time step t+1, The uncertainty of the distribution of all possible actions a, that is, the state The entropy of action selection under .

[0066] By formula: ; Update the parameters in the reinforcement learning model.

[0067] in, represents the parameters of the reinforcement learning model, represents a small discount factor that controls the update speed, Represents the updated parameters of the Q network.

[0068] Step 15, using the final reinforcement learning model to obtain the final memory prefetch action of the target computer at the current moment, and prefetching data for the target computer according to the final memory prefetch action.

[0069] Specifically, the memory access page numbers of the target computer at T moments are obtained, where the Tth moment is the moment before the current moment, and then multiple prefetch strides are calculated based on all memory access page numbers. According to all memory access page numbers and all prefetch strides, the final memory prefetch action is determined using the final reinforcement learning model.

[0070] Exemplarily, the current time is 9 o'clock, T=5, and T times may be 4 o'clock, 5 o'clock, 6 o'clock, 7 o'clock, and 8 o'clock. The multiple times in step 11 may be historical times before the current time, such as 2 o'clock, 3 o'clock, 4 o'clock, 5 o'clock, and 6 o'clock. When performing the above step of obtaining the memory access page number of the target computer at T times, data has not been pre-fetched from the target computer at the current time.

[0071] It is worth mentioning that reinforcement learning has advantages and dynamic decision-making capabilities in high-dimensional nonlinear data representation. It can learn valuable pre-fetch features from complex memory access conditions, capture the computer's memory access patterns at multiple times and adjust pre-fetch decisions, thereby improving the accuracy of memory pre-fetch actions. Training the reinforcement learning model can improve the performance of the reinforcement learning model. Using the final reinforcement learning model to obtain the final memory pre-fetch action and pre-fetch data for the target computer at the current moment can effectively improve the accuracy and timeliness of data pre-fetching.

[0072] In addition, the method of this application includes several key points: Key point 1, memory access data processing mechanism; Technical effect: Maintain a buffer to record the most recently accessed memory addresses, extract page numbers for continuous memory accesses, filter invalid information, and avoid storing complete memory addresses to reduce memory overhead.

[0073] Key point 2, inter-page memory stride filtering mechanism; Technical effect: By sampling the memory access sequence, only a preset number of multiple memory access page numbers are retained. This mechanism effectively avoids invalid prefetching of high-span random access patterns, reduces system bandwidth waste, and improves overall prefetching accuracy and cache hit rate.

[0074] Key point 3, deep reinforcement learning network learning technology; Technical effect: Construct a reinforcement learning model based on deep reinforcement learning, combine the advantages of reinforcement learning in high-dimensional nonlinear data representation with the dynamic decision-making ability of reinforcement learning, and learn valuable pre-fetch features from complex memory access conditions. Through multi-step reward mechanism, soft network update, and optimized pre-fetch strategy, it can accurately identify multiple memory access modes and adjust pre-fetch decisions in real time.

[0075] Key point 4, adaptive reward optimization mechanism; Technical effect: Design a multi-layer reward mechanism, and incorporate factors such as delay consistency and confidence reward into the overall evaluation. By dynamically adjusting the weight parameters in the reward function, the system performance is adaptively balanced to improve the long-term robustness and stability of the algorithm.

[0076] Key point 5, dynamic entropy regularization and exploration-utilization balance mechanism; Technical effect: Add a dynamic entropy regularization term to the Q-value function update, dynamically adjust the entropy temperature parameter according to the needs of exploration and utilization, encourage the model to explore more extensively in the early stages, and focus on high-value pre-fetching decisions in the later stages. This mechanism improves the robustness of the pre-fetching strategy and avoids premature convergence to a suboptimal solution.

[0077] Key point 6, multi-step prefetch mechanism based on latency awareness; Technical effect: Comprehensively analyze the impact of current memory access on the latency of multiple future accesses, optimize the time and space consistency of prefetch operations by increasing the prediction of future multi-step rewards, and reduce the overall system latency.

[0078] The following is an exemplary description of the data pre-fetching device based on reinforcement learning provided in the present application.

[0079] like Figure 4 As shown, the embodiment of the present application provides a data pre-fetching device based on reinforcement learning, and the data pre-fetching device 400 based on reinforcement learning includes: The first acquisition module 401 is used to acquire the memory access page number of the target computer at multiple times; the memory access page number is the number of the memory address accessed by the target computer; A calculation module 402 is used to calculate a plurality of prefetch strides based on all memory access page numbers; the prefetch stride is used to describe the difference between the memory access page numbers of two corresponding adjacent moments; A determination module 403 is used to determine a memory prefetch action using a reinforcement learning model according to all memory access page numbers and all prefetch strides, and calculate a reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times data is prefetched according to the target prefetch stride;

[0080] A training module 404 is used to train the reinforcement learning model according to the memory pre-fetch action and the reward function value to obtain a final reinforcement learning model; The second acquisition module 405 is used to acquire the final memory prefetch action of the target computer at the current moment by using the final reinforcement learning model, and prefetch data for the target computer according to the final memory prefetch action.

[0081] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0082] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0083] like Figure 5 As shown, an embodiment of the present application provides a terminal device. The terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 5 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps of any of the above-mentioned method embodiments when executing the computer program D102.

[0084] Specifically, when the processor D100 executes the computer program D102, it obtains the memory access page numbers of the target computer at multiple times, then calculates multiple pre-fetch steps based on all memory access page numbers, and then uses the reinforcement learning model to determine the memory pre-fetch action according to all memory access page numbers and all pre-fetch steps, and calculates the reward function value corresponding to the memory pre-fetch action, and then trains the reinforcement learning model according to the memory pre-fetch action and the reward function value to obtain the final reinforcement learning model, and finally uses the final reinforcement learning model to obtain the final memory pre-fetch action of the target computer at the current moment, and pre-fetches data for the target computer according to the final memory pre-fetch action. Among them, reinforcement learning has advantages and dynamic decision-making capabilities in high-dimensional nonlinear data representation, can learn valuable pre-fetch features from complex memory access conditions, capture the computer memory access mode at multiple times and adjust the pre-fetch decision, thereby improving the accuracy of the memory pre-fetch action, training the reinforcement learning model can improve the performance of the reinforcement learning model, using the final reinforcement learning model to obtain the final memory pre-fetch action, and pre-fetching data for the target computer at the current moment, which can effectively improve the accuracy and timeliness of data pre-fetching.

[0085] The processor D100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0086] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or is to be output.

[0087] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0088] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0089] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the data pre-fetching method device / terminal device based on reinforcement learning, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.

[0090] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0091] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0092] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data pre-fetching method based on reinforcement learning, characterized in that: include: Obtain the memory access page numbers of the target computer at multiple times; The memory access page number is the number of the memory address accessed by the target computer; Calculate multiple prefetch strides based on all memory access page numbers; The prefetch stride is used to describe the difference between the memory access page numbers of two corresponding adjacent moments; According to all memory access page numbers and all prefetch strides, a memory prefetch action is determined using a reinforcement learning model, and a reward function value corresponding to the memory prefetch action is calculated; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times data is prefetched according to the target prefetch stride; Training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model; The final reinforcement learning model is used to obtain a final memory prefetch action of the target computer at the current moment, and data is prefetched for the target computer according to the final memory prefetch action.

2. The data pre-fetching method according to claim 1, characterized in that: The calculating of a plurality of prefetch strides based on all memory access page numbers comprises: Determine a plurality of current memory access page numbers from all memory access page numbers; For each two adjacent current memory access page numbers, a difference between the two adjacent current memory access page numbers is calculated, and the calculated difference is used as the prefetch stride corresponding to the two adjacent current memory access page numbers.

3. The data pre-fetching method according to claim 2, characterized in that: The step of determining a plurality of current memory access page numbers from all memory access page numbers includes: When the number of all memory access page numbers is greater than the preset number of page numbers, the first N memory access page numbers are used as the current memory access page numbers in the order of all moments; N represents the preset number of page numbers; When the number of all memory access page numbers is less than or equal to the preset number of page numbers, all memory access page numbers are used as current memory access page numbers.

4. The data pre-fetching method according to claim 3, characterized in that: The memory prefetch action is determined by using a reinforcement learning model according to all memory access page numbers and all prefetch strides, including: The memory access page numbers corresponding to all prefetch strides are used as the current environment state; A memory prefetch action is determined according to the current environment state and all prefetch strides.

5. The data pre-fetching method according to claim 1, characterized in that: The calculating the reward function value corresponding to the memory prefetch action includes: By formula: Calculate the reward function value : in, Indicates the impact of the memory prefetch action on the memory access operation, Decision values ​​representing confidence: ; ; in, represents the weight factor of the delayed consistency reward, represents the maximum number of time steps to consider delay, k represents the index of future time steps, T(t+k) represents the time when memory access occurs in the future k time steps after the memory prefetch action is executed, and T(t) represents the time when the current memory access occurs. represents the confidence reward weight, Indicates the current environment status. Indicates memory prefetch action, represents information entropy, represents the maximum possible information entropy, The Q value entropy of the memory prefetch action is represented by: ; ; ; in, represents the normalized probability, and All action sets A subset of represents the entropy temperature control factor.

6. The data pre-fetching method according to claim 4, characterized in that: The step of training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model includes: The current environment state, the memory prefetch action, the reward function value, and the environment state after executing the memory prefetch action are stored as a piece of data in the experience replay pool; Determining whether the amount of data in the experience replay pool reaches a preset amount; If the number of data in the experience replay pool reaches a preset number, a target data is extracted from all the data in the experience replay pool, and the reinforcement learning model is trained using the target data to obtain a trained reinforcement learning model. The number of iterations is increased by 1, and it is determined whether the number of iterations is greater than or equal to the preset number of iterations; If the number of iterations is greater than or equal to the preset number of iterations, the trained reinforcement learning model is used as the final reinforcement learning model; If the amount of data in the experience replay pool does not reach the preset amount or the number of iterations is less than the preset number of iterations, the multiple memory access page numbers contained in the environment state after executing the memory prefetch action will be used as all the memory access page numbers in the step of calculating multiple prefetch strides based on all memory access page numbers, and the step of calculating multiple prefetch strides based on all memory access page numbers will be returned.

7. The data pre-fetching method according to claim 6, characterized in that: The step of training the reinforcement learning model using the target data to obtain a trained reinforcement learning model includes: The target data is used to update the Q value function in the Q network in the reinforcement learning model, and the reinforcement learning model is updated according to the updated Q network to obtain a trained reinforcement learning model.

8. The data pre-fetching method according to claim 7, characterized in that: The updating of the Q value function in the Q network in the reinforcement learning model by using the target data includes: By formula: Updating the Q value function; in, represents the Q-value function, represents the time distribution corresponding to the target data, , represents the tracking attenuation factor, represents the discount factor, Indicates the time distribution corresponding to the previous data of the target data, Indicates the current state of the environment in the target data, Indicates the memory prefetch action in the target data, represents the reward function value in the target data, and Both represent weight factors, and All action sets A subset of represents the maximum Q value that can be achieved in the next environment state of the target data, It means that at time step t+1, Uncertainty about the distribution of all possible actions a; The updating of the reinforcement learning model according to the updated Q network includes: By formula: ; Update the parameters in the reinforcement learning model; in, represents the parameters of the reinforcement learning model, represents a small discount factor, represents the updated parameters of the Q network.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the data pre-fetching method according to any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the data pre-fetching method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Coordinated prefetching in hierarchically cached processors

    CN103324585A

  • Reinforcement learning driven network map region clustering prefetching method

    CN106503238A

  • Data prefetching method and device

    CN114721974A

  • Data prefetching method and device, electronic equipment and readable storage medium

    CN118276946A

  • Methods and apparatus to prefetch memory objects

    US20040216097A1