A data prefetching method and related devices based on reinforcement learning

The reinforcement learning-based data pre-fetching method addresses the accuracy and timeliness issues in complex memory access patterns by training a model to determine optimal pre-fetch actions, improving data pre-fetching efficiency in modern computing systems.

CN119938555BActive Publication Date: 2025-07-15湖南工商大学
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429962.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-15
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

When faced with complex memory access modes, existing data prefetching technologies have low accuracy and timeliness, making it difficult to effectively utilize caches.

Method used

The data prefetching method based on reinforcement learning is adopted, and the prefetching step is calculated by obtaining the computer's memory access page number, and the memory prefetching action and reward function value is determined using the reinforcement learning model, and the final reinforcement learning model is trained to improve the accuracy and timeliness of data prefetching.

Benefits of technology

Improve the accuracy and timeliness of data prefetching, and can learn valuable prefetching features from complex memory access conditions, adjust prefetching decisions, and optimize memory performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938555B_ABST
    Figure CN119938555B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of data prefetching, and provides a data prefetching method and related devices based on reinforcement learning. The method includes: obtaining memory access page numbers of a target computer at multiple moments; calculating multiple prefetch strides based on all the memory access page numbers; determining a memory prefetch action by using a reinforcement learning model according to all the memory access page numbers and all the prefetch strides, and calculating a reward function value corresponding to the memory prefetch action; training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model; using the final reinforcement learning model to obtain a final memory prefetch action of the target computer at the current moment, and performing data prefetching on the target computer according to the final memory prefetch action. The method of this application can improve the accuracy and timeliness of data prefetching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data prefetching, and particularly to a data prefetching method and related devices based on reinforcement learning. Background Art

[0002] Memory is an indispensable part of the computer architecture, responsible for storing the executable code and operation data that can be processed by the processor during the operation of the application program. Its performance has a significant impact on the overall operation speed and stability of the computer. During program execution, the arithmetic unit (such as the central processing unit (CPU, Central Processing Unit)) needs to continuously interact with the memory to read instructions, load data, and complete data processing. However, there is a significant speed mismatch problem between the processor and the memory. Coupled with the fact that the progress of processor performance far exceeds the improvement of memory performance, this has led to a serious computer performance bottleneck, namely the well-known "memory wall" problem.

[0003] Currently, there are various ways to improve and optimize memory performance: 1. Multi-level memory structure. By designing multi-level caches between the arithmetic unit and the memory, the closer the cache is to the arithmetic unit, the faster the access speed, and vice versa. This ensures that the arithmetic unit always interacts with its nearest cache, and then stores the operation results in the memory and disk through the intermediate cache; 2. Memory separation technology. In the current computer design, the CPU and the memory are tightly bound together physically, which limits the memory resources that the CPU can use and reduces the memory utilization rate. Therefore, pooling the memory resources and flexibly allocating the memory resources that the CPU can access can greatly improve the memory performance and utilization rate; 3. Data prefetching technology. When the CPU needs to change the data in a certain memory block, it takes a long time for the memory data of the instruction to return from the time the instruction is issued. This easily causes the computing resources to always be in a waiting state. By analyzing the memory data processed by the CPU before to predict the memory units that will be used later and fetching them in advance to the cache closer to the CPU, this problem can be greatly solved.

[0004] Data prefetching usually needs to identify the access patterns and memory access characteristics of different applications from historical data, so as to accurately predict and obtain the subsequent required memory data. Current application programs have different memory access patterns, which are generally divided into two types: regular access and random access. Regular access: 1. Sequential access, where the application program accesses memory addresses in sequence, such as increasing or decreasing regularly. This memory access pattern has good locality characteristics and can greatly improve the efficiency of data prefetching; 2. Strided access, where the application program accesses memory addresses according to a certain rule as a whole, but has strided access characteristics in terms of time or space. For example, it will jump to another memory address far away at a certain time point, which usually exists in matrix and array operations. Random access: The application program accesses the memory address space in an irregular order, with poor locality. Common prefetching techniques are difficult to learn this access pattern, and most of the prefetching operations are ineffective, resulting in the inability to effectively utilize the cache.

[0005] Traditional data prefetching can achieve good prefetching effects when there is high memory access spatial locality and temporal locality. Common ones include: 1. Stride prefetchers, which save a prefetch cache window for different memory access sequences to identify the memory access strides of different streams. However, this method of separately dividing prefetch caches for each memory access sequence brings additional storage overhead; 2. Offset prefetchers, which configure a unified prefetch offset window for all memory access sequences. For example, if the offset variable is d, the content that the window may save is [d, 2d, 3d,...], saving storage resources and being able to more accurately identify the overall memory access characteristics; 3. Spatial prefetchers, which use the characteristic that the sequences accessed in the past may be accessed again to identify and cache the memory page data with prefetch value, so as to timely re-possibly perform the memory access operations that may appear again. However, these traditional prefetchers can achieve good results when facing programs with relatively simple memory access patterns. But as the scale and complexity of current artificial intelligence and big data applications grow exponentially, the more complex memory access patterns they bring make the traditional prefetchers have problems of low accuracy and timeliness in data prefetching. Summary of the Invention

[0006] This application provides a data prefetching method and related devices based on reinforcement learning, which can solve the problems of low accuracy and timeliness in data prefetching.

[0007] In a first aspect, an embodiment of this application provides a data prefetching method based on reinforcement learning. The data prefetching method includes:

[0008] Obtain the memory access page numbers of the target computer at multiple moments; the memory access page number is the number of the memory address accessed by the target computer;

[0009] Calculate multiple prefetch strides based on all memory access page numbers; the prefetch strides are used to describe the difference between the memory access page numbers at two adjacent corresponding moments.

[0010] Based on all memory access page numbers and all prefetch strides, use a reinforcement learning model to determine a memory prefetch action, and calculate the reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times of data prefetch according to the target prefetch stride.

[0011] Train the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model.

[0012] Use the final reinforcement learning model to obtain the final memory prefetch action of the target computer at the current moment, and perform data prefetch on the target computer according to the final memory prefetch action.

[0013] Optionally, calculating multiple prefetch strides based on all memory access page numbers includes:

[0014] Determine multiple current memory access page numbers from all memory access page numbers.

[0015] For each pair of current memory access page numbers corresponding to two adjacent moments respectively, calculate the difference between the current memory access page numbers corresponding to the two adjacent moments, and use the calculated difference as the prefetch stride corresponding to the current memory access page numbers corresponding to the two adjacent moments.

[0016] Optionally, determining multiple current memory access page numbers from all memory access page numbers includes:

[0017] When the number of all memory access page numbers is greater than the preset number of page numbers, in the order of all moments, take the first N memory access page numbers as the current memory access page numbers; N denotes the preset number of page numbers;

[0018] When the number of all memory access page numbers is less than or equal to the preset number of page numbers, take all memory access page numbers as the current memory access page numbers.

[0019] Optionally, using a reinforcement learning model to determine a memory prefetch action based on all memory access page numbers and all prefetch strides includes:

[0020] Take the memory access page numbers corresponding to all prefetch strides as the current environmental state;

[0021] Determine a memory prefetch action according to the current environmental state and all prefetch strides.

[0022] Optionally, calculating the reward function value corresponding to the memory prefetch action includes:

[0023] Through the formula:

[0024]

[0025] Calculate the reward function value :

[0026] wherein, represents the impact of the memory prefetch action on the memory access operation, represents the decision value of confidence:

[0027] ;

[0028] ;

[0029] wherein, represents the weight factor of the latency consistency reward, represents the maximum number of time steps considering latency, k represents the index of future time steps, T(t + k) represents the memory access occurrence time of the future k time steps after the memory prefetch action is executed, and T(t) represents the current memory access occurrence time, represents the confidence reward weight, represents the current environmental state, represents the memory prefetch action, represents the information entropy, represents the maximum possible information entropy, represents the Q-value entropy of the memory prefetch action:

[0030] ;

[0031] ;

[0032] ;

[0033] wherein, represents the normalized probability, and are both subsets of the action set ; represents the entropy temperature control factor.

[0034] Optionally, train the reinforcement learning model according to the memory prefetch action and the reward function value to obtain the final reinforcement learning model, including:

[0035] Store the current environmental state, memory prefetch action, reward function value, and environmental state after the memory prefetch action is executed as a piece of data in the experience replay pool;

[0036] Determine whether the number of data in the experience replay pool reaches the preset number;

[0037] If the number of data in the experience replay pool reaches the preset number, extract a target data from all the data in the experience replay pool, and use the target data to train the reinforcement learning model to obtain the trained reinforcement learning model. Increment the iteration count by 1 and determine whether the iteration count is greater than or equal to the preset iteration count.

[0038] If the iteration count is greater than or equal to the preset iteration count, use the trained reinforcement learning model as the final reinforcement learning model.

[0039] If the number of data in the experience replay pool does not reach the preset number or the iteration count is less than the preset iteration count, use the multiple memory access page numbers included in the environment state after performing the memory prefetch action as all the memory access page numbers in the step of calculating multiple prefetch strides based on all the memory access page numbers, and return the step of calculating multiple prefetch strides based on all the memory access page numbers.

[0040] Optionally, using the target data to train the reinforcement learning model to obtain the trained reinforcement learning model includes:

[0041] Update the Q-value function in the Q-network of the reinforcement learning model using the target data, and update the reinforcement learning model according to the updated Q-network to obtain the trained reinforcement learning model.

[0042] Optionally, updating the Q-value function in the Q-network of the reinforcement learning model using the target data includes:

[0043] Through the formula:

[0044] ;

[0045] Update the Q-value function;

[0046] Where represents the Q-value function, represents the time distribution corresponding to the target data, , represents the trace decay factor, represents the discount factor, represents the time distribution corresponding to the previous data of the target data, represents the current environment state in the target data, represents the memory prefetch action in the target data, represents the reward function value in the target data, and both represent weight factors, and are both subsets of the action set of Represents the maximum Q value achievable in the next environmental state of the target data, Indicates that at time step t+1, The uncertainty of the distribution of all possible actions a;

[0047] Update the reinforcement learning model according to the updated Q network, including:

[0048] Through the formula:

[0049]

[0050] Update the parameters in the reinforcement learning model;

[0051] Wherein, Represents the parameters of the reinforcement learning model, Represents the small discount factor, Represents the parameters of the updated Q network.

[0052] In a second aspect, an embodiment of the present application provides a data prefetching device based on reinforcement learning, including:

[0053] A first acquisition module, configured to acquire the memory access page numbers of the target computer at multiple moments; the memory access page number is the number of the memory address accessed by the target computer;

[0054] A calculation module, configured to calculate multiple prefetch strides based on all the memory access page numbers; the prefetch stride is used to describe the difference between the memory access page numbers of two adjacent moments;

[0055] A determination module, configured to determine a memory prefetch action according to all the memory access page numbers and all the prefetch strides by using a reinforcement learning model, and calculate a reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times of data prefetch according to the target prefetch stride;

[0056] A training module, configured to train the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model;

[0057] A second acquisition module, configured to obtain the final memory prefetch action of the target computer at the current moment by using the final reinforcement learning model, and perform data prefetch on the target computer according to the final memory prefetch action.

[0058] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the above-mentioned data prefetching method based on reinforcement learning is implemented.

[0059] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned data prefetching method based on reinforcement learning.

[0060] The above solution of the present application has the following beneficial effects:

[0061] In the embodiment of the present application, by obtaining the memory access page numbers of the target computer at multiple moments, then calculating multiple prefetch strides based on all the memory access page numbers, and then using a reinforcement learning model to determine the memory prefetch action according to all the memory access page numbers and all the prefetch strides, and calculating the reward function value corresponding to the memory prefetch action, and then training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model, and finally using the final reinforcement learning model to obtain the final memory prefetch action of the target computer at the current moment, and performing data prefetching on the target computer according to the final memory prefetch action. Among them, reinforcement learning has advantages in high-dimensional non-linear data representation and dynamic decision-making ability, can learn valuable prefetch features from complex memory access situations, capture the patterns of the computer accessing memory within multiple moments and adjust the prefetch decision, thereby improving the accuracy of the memory prefetch action. Training the reinforcement learning model can improve the performance of the reinforcement learning model. Using the final reinforcement learning model to obtain the final memory prefetch action and performing data prefetching on the target computer at the current moment can effectively improve the accuracy and timeliness of data prefetching.

[0062] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0064] Figure 1 is a flowchart of a data prefetching method based on reinforcement learning provided by an embodiment of the present application;

[0065] Figure 2 is a schematic diagram of a prefetch stride provided by an embodiment of the present application;

[0066] Figure 3 is a schematic diagram of reinforcement learning provided by an embodiment of the present application;

[0067] Figure 4Schematic structural diagram of a data prefetching device based on reinforcement learning provided by an embodiment of the present application;

[0068] Figure 5 Schematic structural diagram of a terminal device provided by an embodiment of the present application. Detailed implementation manners

[0069] In the following description, specific details such as specific system architectures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0070] It should be understood that when used in the specification and appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0071] It should also be understood that the term "and / or" as used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0072] As used in the specification and appended claims of the present application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" according to the context.

[0073] In addition, in the description of the specification and appended claims of the present application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0074] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear at different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0075] Aiming at the problems of low accuracy and timeliness of existing data prefetching, an embodiment of this application provides a data prefetching method based on reinforcement learning. This data prefetching method obtains the memory access page numbers of the target computer at multiple moments, then calculates multiple prefetch strides based on all the memory access page numbers, and then determines the memory prefetch action by using a reinforcement learning model according to all the memory access page numbers and all the prefetch strides, and calculates the reward function value corresponding to the memory prefetch action. Then, the reinforcement learning model is trained according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model. Finally, the final reinforcement learning model is used to obtain the final memory prefetch action of the target computer at the current moment, and data prefetching is performed on the target computer according to the final memory prefetch action. Among them, reinforcement learning has advantages in high-dimensional non-linear data representation and dynamic decision-making ability, can learn valuable prefetch features from complex memory access situations, capture the patterns of the computer accessing memory within multiple moments and adjust the prefetch decision, thereby improving the accuracy of the memory prefetch action. Training the reinforcement learning model can improve the performance of the reinforcement learning model. Using the final reinforcement learning model to obtain the final memory prefetch action and performing data prefetching on the target computer at the current moment can effectively improve the accuracy and timeliness of data prefetching.

[0076] Next, an exemplary description will be given of the data prefetching method based on reinforcement learning provided by this application.

[0077] As Figure 1 shown, the data prefetching method based on reinforcement learning provided by this application includes the following steps:

[0078] Step 11, obtain the memory access page numbers of the target computer at multiple moments.

[0079] The above memory access page numbers are the numbers of the memory addresses accessed by the target computer (the memory addresses accessed at each moment may be different, and the memory access page number at each moment is the number of the memory address accessed by the target computer at that moment). The above target computer is the computer that needs to perform data prefetching.

[0080] In some embodiments of the present application, the accessed memory address can be obtained first, and then the memory address can be shifted right to obtain the memory access page number of the memory address.

[0081] It should be noted that the number of bits by which the memory address is shifted right is determined according to the number of bytes of the memory address itself.

[0082] Exemplarily, the memory address is defaulted to the cacheline granularity, that is, 64 bytes, and a page is defaulted to 4KB. That is, by shifting each memory address right by 12 bits, the memory access page number can be obtained. A record buffer can be maintained to store multiple memory addresses of the target computer. When performing this step, the memory addresses are read from the record buffer.

[0083] It is worth mentioning that using the memory access page number for data prefetching can greatly reduce the overhead brought by storing the complete memory address in the traditional data prefetcher.

[0084] Step 12, calculate multiple prefetch strides based on all memory access page numbers.

[0085] The above prefetch stride is used to describe the difference between the corresponding two memory access page numbers, and the corresponding two memory access page numbers are the two memory access page numbers required to calculate the prefetch stride.

[0086] In some embodiments of the present application, the step of calculating multiple prefetch strides based on all memory access page numbers is specifically as follows:

[0087] The first step is to determine multiple current memory access page numbers from all memory access page numbers.

[0088] Specifically, when the number of all memory access page numbers is greater than the preset number of page numbers, according to the order of all times, the first N memory access page numbers are all used as the current memory access page numbers; N represents the preset number of page numbers.

[0089] When the number of all memory access page numbers is less than or equal to the preset number of page numbers, all memory access page numbers are used as the current memory access page numbers.

[0090] Exemplarily, the memory access page numbers are 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 85, 87, 95, 96, 104, 105 in sequence, N is equal to 20, then all target memory access page numbers are 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87 in sequence.

[0091] In the second step, for each pair of currently accessed memory page numbers corresponding to two adjacent moments, calculate the difference between the currently accessed memory page numbers corresponding to the two adjacent moments, and use the calculated difference as the prefetch stride corresponding to the currently accessed memory page numbers corresponding to the two adjacent moments.

[0092] Exemplarily, if the 1st currently accessed memory page number is 100 and the 2nd currently accessed memory page number is 150, then the prefetch stride between them is 50. The prefetch strides within a period of time are as Figure 2 shown. In the figure, the horizontal axis represents the execution time of data prefetching, and the vertical axis represents the inter-page stride (i.e., the prefetch stride).

[0093] Step 13: Based on all the accessed memory page numbers and all the prefetch strides, use a reinforcement learning model to determine the memory prefetch action and calculate the reward function value corresponding to the memory prefetch action.

[0094] The above-mentioned memory prefetch action includes a target prefetch stride and a prefetch degree. The prefetch degree is used to describe the number of times of data prefetching according to the target prefetch stride.

[0095] In some embodiments of the present application, the step of using a reinforcement learning model to determine the memory prefetch action based on all the accessed memory page numbers and all the prefetch strides and calculating the reward function value corresponding to the memory prefetch action includes:

[0096] In the first step, use the memory page numbers corresponding to all the prefetch strides as the current environmental state.

[0097] Exemplarily, as can be seen from step 12, the memory page numbers corresponding to all the prefetch strides are all the currently accessed memory page numbers. All the currently accessed memory page numbers are 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87 in sequence. Then the current environmental state is 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87.

[0098] In the second step, determine the memory prefetch action according to the current environmental state and all the prefetch strides.

[0099] Exemplarily, a policy in reinforcement learning (such as the policy in the Deep Q-Network (DQN)) can be used to select a target prefetch stride from all the prefetch strides according to the policy and generate a prefetch degree. For example, when currently using a Q-value-based policy, the Q-network trained using DQN calculates a Q-value for each possible prefetch stride, that is is 15, is 10, If it is 2, then select the prefetch stride with the largest Q value from 0 to 9. And if the model believes that the prefetch stride 9 will appear 1 time in the next few memory accesses, then determine that the final prefetch degree is 1. For example, if the current environmental state is 0, 1, 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 86, 87, the target prefetch stride is 9, and the prefetch degree is 1, it means that data in the memory area corresponding to memory access page 96 needs to be prefetched based on memory access page number 87. After executing this action, the environmental state becomes 9, 10, 19, 20, 28, 29, 38, 39, 47, 48, 57, 58, 66, 67, 76, 77, 85, 87, 95, 96.

[0100] Step 3: Calculate the reward function value corresponding to the memory prefetch action.

[0101] Specifically, through the formula:

[0102]

[0103] Calculate the reward function value :

[0104] Among them, represents the impact of the memory prefetch action on the memory access operation, represents the decision value of confidence:

[0105] ;

[0106] ;

[0107] Among them, represents the weight factor of the delayed consistency reward, represents the maximum number of time steps considering the delay, which controls the consideration range of future time steps. k represents the index of the future time step, representing the time span of the delay impact. For example, for a prefetch action made at the t-th time step, k = 1 represents the memory access in the next time step (i.e., t + 1). If K = 3, then consider the impact of the current prefetch behavior on the next 3 time steps (i.e., t + 1, t + 2, t + 3), and ignore more distant time steps. A larger K can increase the consideration of the long-term delay impact, while a smaller K focuses on the impact within a shorter time range; represents the time scaling factor, which is used to control the degree of influence of latency on reward calculation. T(t + k) represents the time of memory access in the next k time steps after performing the memory prefetch action, and T(t) represents the time of the current memory access. For example, at the moment t = 5, the memory access of the current page number 87 occurs at the moment T(5). When the target prefetch stride is 9, the memory access of the target page number 96 will occur at the moment T(6) (i.e., the next time step after prefetch). represents a certain time step or time point in the current reinforcement learning environment. represents the confidence reward weight, which is used to adjust the influence of confidence in the reward. represents the current environmental state. represents the memory prefetch action. represents the information entropy. represents the maximum possible information entropy. represents the Q - value entropy of the memory prefetch action:

[0108] ;

[0109] ;

[0110] ;

[0111] Among them, represents the normalized probability, and are both subsets of the action set of represents the entropy temperature control factor.

[0112] It can be understood that the process recorded in step 13 above is the data - processing process in the reinforcement learning model.

[0113] Next, a specific example is used to illustrate the reinforcement learning exemplarily.

[0114] As Figure 3 shown, the agent (i.e., the reinforcement learning model) obtains actions through the policy, interacts with the environment, and transmits the reward and state to the agent. The agent then generates actions again according to the state, and after interacting with the environment, obtains the next state and reward.

[0115] It is worth mentioning that reinforcement learning has advantages in high - dimensional non - linear data representation and dynamic decision - making ability. It can learn valuable prefetch features from complex memory access situations, capture the patterns of computer memory access in multiple moments, and adjust the prefetch decision.

[0116] Step 14, train the reinforcement learning model according to the memory prefetch action and the reward function value to obtain the final reinforcement learning model.

[0117] In some embodiments of the present application, the step of training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain the final reinforcement learning model includes:

[0118] First, store the current environmental state, the memory prefetch action, the reward function value, and the environmental state after performing the memory prefetch action as a piece of data in the experience replay pool;

[0119] Second, determine whether the number of data in the experience replay pool reaches a preset number.

[0120] If the number of data in the experience replay pool reaches the preset number, extract a target data from all the data in the experience replay pool, use the target data to train the reinforcement learning model to obtain the trained reinforcement learning model, increment the iteration count by 1, and determine whether the iteration count is greater than or equal to the preset iteration count.

[0121] If the iteration count is greater than or equal to the preset iteration count, use the trained reinforcement learning model as the final reinforcement learning model.

[0122] If the number of data in the experience replay pool does not reach the preset number or the iteration count is less than the preset iteration count, use the multiple memory access page numbers included in the environmental state after performing the memory prefetch action as all the memory access page numbers in the step of calculating multiple prefetch strides based on all memory access page numbers, and return to the step of calculating multiple prefetch strides based on all memory access page numbers.

[0123] It should be noted that the initial iteration count is 0.

[0124] The step of using the target data to train the reinforcement learning model to obtain the trained reinforcement learning model is specifically: updating the Q-value function in the Q-network of the reinforcement learning model using the target data, and updating the reinforcement learning model according to the updated Q-network to obtain the trained reinforcement learning model.

[0125] Through the formula:

[0126] ;

[0127] Update the Q-value function.

[0128] Wherein, represents the Q-value function, represents the time distribution corresponding to the target data, , represents the tracking decay factor, represents the discount factor, represents the time distribution corresponding to the previous data of the target data Represents the current environmental state in the target data, Represents the memory prefetch action in the target data, Represents the reward function value in the target data, and both represent weight factors, and are both subsets of the action set of Represents the maximum Q value achievable in the next environmental state of the target data, Represents at time step t + 1, the uncertainty of the distribution of all possible actions a, i.e., the entropy of the action selection under state of the action selection under state

[0129] Through the formula:

[0130] ;

[0131] Update the parameters in the reinforcement learning model.

[0132] Wherein, represents the parameters of the reinforcement learning model, represents a small discount factor for controlling the update speed, represents the parameters of the updated Q network.

[0133] Step 15: Use the final reinforcement learning model to obtain the final memory prefetch action of the target computer at the current moment, and perform data prefetch on the target computer according to the final memory prefetch action.

[0134] Specifically, obtain the memory access page numbers of the target computer at T moments, where the T-th moment is the moment before the current moment, and then calculate multiple prefetch strides based on all the memory access page numbers. Based on all the memory access page numbers and all the prefetch strides, use the final reinforcement learning model to determine the final memory prefetch action.

[0135] Exemplarily, the current moment is 9 o'clock, T = 5, and the T moments can be 4 o'clock, 5 o'clock, 6 o'clock, 7 o'clock, 8 o'clock. The multiple moments in step 11 can be historical moments before the current moment, such as the multiple moments are 2 o'clock, 3 o'clock, 4 o'clock, 5 o'clock, 6 o'clock. When performing the above step of obtaining the memory access page numbers of the target computer at T moments, data has not been prefetched from the target computer at the current moment.

[0136] It is worth mentioning that reinforcement learning has advantages in high-dimensional non-linear data representation and dynamic decision-making capabilities. It can learn valuable prefetch features from complex memory access situations, capture the patterns of computer memory access over multiple moments, and adjust prefetch decisions, thereby improving the accuracy of memory prefetch actions. Training the reinforcement learning model can improve the performance of the reinforcement learning model. Using the final reinforcement learning model to obtain the final memory prefetch action and perform data prefetch on the target computer at the current moment can effectively improve the accuracy and timeliness of data prefetch.

[0137] In addition, the method of this application includes multiple key points:

[0138] Key point 1, memory access data processing mechanism; Technical effect: Maintain a buffer that records the recently accessed memory addresses, extract the page numbers of consecutive memory accesses, filter out invalid information, and avoid storing complete memory addresses to reduce memory overhead.

[0139] Key point 2, inter-page memory stride filtering mechanism; Technical effect: By sampling the memory access sequence, only retain multiple memory access page numbers of a preset quantity. This mechanism effectively avoids invalid prefetching for high-stride random access patterns, reduces system bandwidth waste, and improves overall prefetch accuracy and cache hit rate.

[0140] Key point 3, deep reinforcement learning network learning technology; Technical effect: Construct a reinforcement learning model based on deep reinforcement learning, combine the advantages of reinforcement learning in high-dimensional non-linear data representation with the dynamic decision-making ability of reinforcement learning, and learn valuable prefetch features from complex memory access situations. Through a multi-step reward mechanism and soft network update, optimize the prefetch strategy, and accurately identify various memory access patterns and adjust prefetch decisions in real time.

[0141] Key point 4, adaptive reward optimization mechanism; Technical effect: Design a multi-layer reward mechanism, incorporate factors such as latency consistency and confidence reward into the overall evaluation. By dynamically adjusting the weight parameters in the reward function, adaptively balance system performance, and enhance the long-term robustness and stability of the algorithm.

[0142] Key point 5, dynamic entropy regularization and exploration-exploitation balance mechanism; Technical effect: Add a dynamic entropy regularization term to the Q-value function update, dynamically adjust the entropy temperature parameter according to the needs of exploration and exploitation, encourage the model to explore more widely in the early stage, and focus on high-value prefetch decisions in the later stage. This mechanism improves the robustness of the prefetch strategy and avoids premature convergence to suboptimal solutions.

[0143] Key Point 6, a multi-step prefetching mechanism based on latency awareness; Technical effect: Comprehensively analyze the impact of current memory access on future multi-access latency, optimize the temporal and spatial consistency of prefetch operations by increasing the prediction of future multi-step rewards, and reduce the overall system latency.

[0144] The data prefetching device based on reinforcement learning provided by the present application will be exemplarily described below.

[0145] As Figure 4 shown, an embodiment of the present application provides a data prefetching device based on reinforcement learning. The data prefetching device 400 based on reinforcement learning includes:

[0146] A first acquisition module 401, configured to acquire memory access page numbers of a target computer at multiple moments; the memory access page number is the number of the memory address accessed by the target computer;

[0147] A calculation module 402, configured to calculate multiple prefetch strides based on all memory access page numbers; the prefetch stride is used to describe the difference between the memory access page numbers corresponding to two adjacent moments;

[0148] A determination module 403, configured to determine a memory prefetch action by using a reinforcement learning model according to all memory access page numbers and all prefetch strides, and calculate a reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times of data prefetching according to the target prefetch stride;

[0149] A training module 404, configured to train the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model;

[0150] A second acquisition module 405, configured to use the final reinforcement learning model to acquire a final memory prefetch action of the target computer at the current moment, and perform data prefetching on the target computer according to the final memory prefetch action.

[0151] It should be noted that the information interaction, execution process, etc. between the above-mentioned device / units, due to being based on the same concept as the method embodiment of the present application, for their specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.

[0152] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0153] As Figure 5 shown, an embodiment of the present application provides a terminal device. The terminal device D10 in this embodiment includes: at least one processor D100 ( Figure 5 only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100. When the processor D100 executes the computer program D102, the steps in any of the foregoing method embodiments are implemented.

[0154] Specifically, when the processor D100 executes the computer program D102, it obtains the memory access page numbers of the target computer at multiple times, then calculates multiple prefetch strides based on all the memory access page numbers, and then determines the memory prefetch action by using the reinforcement learning model according to all the memory access page numbers and all the prefetch strides, and calculates the reward function value corresponding to the memory prefetch action. Then, according to the memory prefetch action and the reward function value, the reinforcement learning model is trained to obtain the final reinforcement learning model. Finally, the final reinforcement learning model is used to obtain the final memory prefetch action of the target computer at the current time, and data prefetch is performed on the target computer according to the final memory prefetch action. Among them, reinforcement learning has advantages in high-dimensional non-linear data representation and dynamic decision-making ability. It can learn valuable prefetch features from complex memory access situations, capture the patterns of the computer accessing memory within multiple times and adjust the prefetch decision, thereby improving the accuracy of the memory prefetch action. Training the reinforcement learning model can improve the performance of the reinforcement learning model. Using the final reinforcement learning model to obtain the final memory prefetch action and performing data prefetch on the target computer at the current time can effectively improve the accuracy and timeliness of data prefetch.

[0155] The so-called processor D100 may be a central processing unit (CPU), and the processor D100 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0156] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as the hard disk or memory of the terminal device D10. In some other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc., equipped on the terminal device D10. Further, the memory D101 may also include both the internal storage unit and the external storage device of the terminal device D10. The memory D101 is used to store an operating system, application programs, a boot loader, data, and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or is to be output.

[0157] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0158] An embodiment of the present application provides a computer program product, and when the computer program product runs on a terminal device, the terminal device can be caused to implement the steps in the above-mentioned method embodiments when executed.

[0159] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the device / terminal device for data prefetching method based on reinforcement learning, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk, or an optical disc, etc.

[0160] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0161] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0162] The above is the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle described in this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A data prefetching method based on reinforcement learning, characterized in that Including: Obtaining memory access page numbers of a target computer at multiple moments; The memory access page number is the number of the memory address accessed by the target computer; Calculating multiple prefetch strides based on all the memory access page numbers; The prefetch stride is used to describe the difference between the memory access page numbers at two adjacent moments; According to all the memory access page numbers and all the prefetch strides, using a reinforcement learning model to determine a memory prefetch action, and calculating a reward function value corresponding to the memory prefetch action; the memory prefetch action includes a target prefetch stride and a prefetch degree, and the prefetch degree is used to describe the number of times of data prefetch according to the target prefetch stride; Training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model; Using the final reinforcement learning model to obtain a final memory prefetch action of the target computer at the current moment, and performing data prefetch on the target computer according to the final memory prefetch action; Wherein, calculating the reward function value corresponding to the memory prefetch action includes: By the formula: Calculate the reward function value : Among them, indicates the impact of the memory prefetch action on the memory access operation, represents the decision value of the confidence level: ; ; Among them, represents the weight factor of the latency consistency reward, represents the maximum number of time steps considering latency, k represents the index of future time steps, T(t + k) represents the time of memory access in the future k time steps after performing the memory prefetch action, and T(t) represents the time of the current memory access, represents the confidence reward weight, represents the current environmental state, represents the memory prefetch action, represents the information entropy, represents the maximum possible information entropy, represents the time scaling factor, represents the time point in the current reinforcement learning environment, represents the Q - value entropy of the memory prefetch action: ; ; ; Among them, represents the normalized probability, and are both subsets of the action set , represents the entropy temperature control factor.

2. The data prefetching method according to claim 1, wherein The calculating multiple prefetch strides based on all the memory access page numbers includes: Determining multiple current memory access page numbers from all the memory access page numbers; For the current memory access page numbers corresponding to every two adjacent moments respectively, calculating the difference between the current memory access page numbers corresponding to the two adjacent moments, and taking the calculated difference as the prefetch stride corresponding to the current memory access page numbers corresponding to the two adjacent moments.

3. The data prefetching method according to claim 2, wherein The determining multiple current memory access page numbers from all the memory access page numbers includes: When the number of all the memory access page numbers is greater than a preset number of page numbers, taking the first N memory access page numbers as the current memory access page numbers in the order of all the moments; N represents the preset number of page numbers; When the number of all the memory access page numbers is less than or equal to the preset number of page numbers, taking all the memory access page numbers as the current memory access page numbers.

4. The data prefetching method according to claim 3, wherein The determining a memory prefetch action using a reinforcement learning model according to all the memory access page numbers and all the prefetch strides includes: Taking the memory access page numbers corresponding to all the prefetch strides as the current environmental state; Determining a memory prefetch action according to the current environmental state and all the prefetch strides.

5. The data prefetching method according to claim 4, wherein The training the reinforcement learning model according to the memory prefetch action and the reward function value to obtain a final reinforcement learning model includes: Storing the current environmental state, the memory prefetch action, the reward function value, and the environmental state after executing the memory prefetch action as a piece of data into an experience replay pool; Judging whether the number of data in the experience replay pool reaches a preset number; If the number of data in the experience replay pool reaches the preset number, extracting a target piece of data from all the data in the experience replay pool, training the reinforcement learning model using the target piece of data to obtain a trained reinforcement learning model, incrementing the iteration number by 1, and judging whether the iteration number is greater than or equal to a preset iteration number; If the iteration number is greater than or equal to the preset iteration number, taking the trained reinforcement learning model as the final reinforcement learning model; If the number of data in the experience replay pool does not reach the preset number or the number of iterations is less than the preset number of iterations, then the multiple memory access page numbers included in the environmental state after performing the memory prefetch action are used as the all memory access page numbers in the step of calculating multiple prefetch strides based on all memory access page numbers, and the step of calculating multiple prefetch strides based on all memory access page numbers is returned.

6. The data prefetching method according to claim 5, wherein The training of the reinforcement learning model with the target data to obtain a trained reinforcement learning model includes: Updating the Q-value function in the Q-network of the reinforcement learning model with the target data, and updating the reinforcement learning model according to the updated Q-network to obtain a trained reinforcement learning model.

7. The data prefetching method according to claim 6, wherein The updating of the Q-value function in the Q-network of the reinforcement learning model with the target data includes: Through the formula: Updating the Q-value function; Among them, represents the Q-value function, represents the time distribution corresponding to the target data, , represents the tracking decay factor, represents the discount factor, represents the time distribution corresponding to the previous data of the target data, represents the current environmental state in the target data, represents the memory prefetch action in the target data, represents the reward function value in the target data, and both represent weight factors, and are both subsets of the action set ; represents the maximum Q-value achievable in the next environmental state of the target data, represents at time step t + 1, the uncertainty of the distribution of all possible actions a; The updating of the reinforcement learning model according to the updated Q-network includes: Through the formula: ; Updating the parameters in the reinforcement learning model; Among them, represents the parameters of the reinforcement learning model, represents the small discount factor, represents the parameters of the updated Q-network.

8. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data prefetch method according to any one of claims 1 to 7.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the data prefetch method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Reinforcement learning driven network map region clustering prefetching method

    CN106503238A

  • Data prefetching method and device

    CN114721974A