Branch predictor based on reinforcement learning

Through a branch predictor based on reinforcement learning, the storage pool and processing module generate state vectors and dynamically update index values, the problems of low hardware resource utilization efficiency and stagnation of prediction performance are solved, and higher branch prediction accuracy and processor performance improvement are achieved.

CN120371400APending Publication Date: 2025-07-25TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510234127.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The hardware resource utilization efficiency of existing branch predictors is low, and the prediction performance improvement is stagnant, and it is impossible to effectively deal with the path uncertainty of complex branch instructions, resulting in wasted CPU cycles.

Method used

Using a branch predictor based on reinforcement learning, the first storage pool and the second storage pool are combined with the instruction processing module and the processor, and the state vector is generated by XOR processing, and the index value of the two-dimensional table is dynamically updated to improve prediction accuracy, and the hardware resource utilization is optimized using reinforcement learning algorithm.

Benefits of technology

It improves branch prediction accuracy, efficiently utilizes hardware resources, improves processor performance, is highly adaptable, can handle complex branch situations, and has a more efficient feedback and update mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371400A_ABST
    Figure CN120371400A_ABST
Patent Text Reader

Abstract

The invention discloses a branch predictor based on reinforcement learning. The branch predictor comprises a first storage pool, an instruction processing module, a second storage pool and a processor, the branch predictor queries the tuple of the received branch instruction in the first storage pool; when the corresponding tuple does not exist in the first storage pool, the instruction processing module carries out XOR processing on the branch instruction to obtain a state vector, and the state vector serves as an index value to be input into the second storage pool; the second storage pool queries the corresponding maximum index value of the state vector in the two-dimensional table as a predicted value and feeds the predicted value back to the processor; the processor calculates the predicted value in a subsequent period to obtain an actual execution value, gives a reward value corresponding to the execution value and feeds back the reward value to the second storage pool; the second storage pool updates the two-dimensional table by the execution memory value according to each branch instruction period; according to the method, the prediction strategy can be dynamically updated and adjusted according to program operation, and the prediction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of superscalar processing chips, and particularly relates to a branch predictor based on reinforcement learning. Background Art

[0002] Most modern high-performance processors use Instruction Level Parallelism (ILP), which enables the CPU to execute multiple instructions in one cycle to accelerate computing. However, one of the constraints of ILP is the large number of branch instructions. According to Haque et al., 15% to 25% of all instructions in a program are branch instructions. The three types of branch instructions are 1) jump, 2) call, and 3) return. Branch instructions can interrupt the sequential flow of instruction execution, making the path of branch execution uncertain. Especially during the parallel execution of instructions, a large number of CPU cycles may be wasted.

[0003] To alleviate the problems caused by branch instructions, a technique called Branch Prediction (BP) is used in modern architectures. The Branch Prediction Unit (BPU) is a key technology widely adopted by modern superscalar processors. The branch predictor can achieve a larger instruction window, which is crucial for instruction-level and memory-level parallelism. Under a huge instruction base, with the further enhancement of width and pipeline width, even a slight improvement in the prediction accuracy of the BPU will bring a significant gain to the overall performance of the processor. In the past, engineers usually used large hardware tables for branch prediction. However, with the end of Moore's Law, hardware resources have become a bottleneck. How to more efficiently utilize limited hardware resources has become an extremely important issue for current engineers. Nowadays, the branch predictors used in most processors are engineering variants of TAGE or perceptron, and gratifying achievements have been made in prediction performance. However, the problem that researchers cannot ignore is that the improvement of these two series of branch predictors has reached a standstill. Researchers need to find new ideas to continuously optimize the existing ones or design new branch predictors.

[0004] Machine learning is an important branch of artificial intelligence (AI). In recent years, machine learning has played a crucial role in the progress of the field of computer architecture. However, the rapid development of emerging technologies such as machine learning has led to an increasing demand for system performance. Researchers, system engineers, and computer architects need to continue to design more advanced hardware architectures to meet the computing requirements. Currently, researchers have conducted some studies on how to use technologies such as machine learning and deep learning to improve system performance, mainly focusing on areas such as memory replacement and CPU scheduling. Reinforcement learning is a subfield of machine learning that emphasizes how an agent acts based on the environment to achieve the maximum expected benefit, which is very suitable for the application scenario of branch prediction. Summary of the Invention

[0005] Aiming at the technical problem that the development of branch predictors relying only on large hardware table types has stagnated in the prior art, the present invention provides a branch predictor based on reinforcement learning, which can make full use of the matching hardware resources and accurately and quickly judge the prediction tasks.

[0006] To solve the problems of the prior art, the present invention adopts the following technical solutions:

[0007] A branch predictor based on reinforcement learning, the branch predictor includes: a first storage pool, an instruction processing module, a second storage pool, and a processor; wherein: the process of the branch predictor obtaining a prediction value by processing branch instructions includes:

[0008] The branch predictor queries the tuples in the first storage pool for the received branch instructions;

[0009] When there is no corresponding tuple in the first storage pool, the instruction processing module performs exclusive OR processing on the branch instructions to obtain a state vector, and inputs the state vector as an index value to the second storage pool;

[0010] The second storage pool queries the state vector in the two-dimensional table for its corresponding maximum index value and feeds it back to the processor as a prediction value;

[0011] The processor calculates the actual execution value from the prediction value in subsequent cycles and gives a corresponding reward value to the second storage pool;

[0012] The second storage pool updates the execution memory value for the two-dimensional table according to each branch instruction cycle;

[0013] The branch predictor selects the final index value in the two-dimensional table as the prediction value output.

[0014] Further, the second storage pool is composed of at least three or more parallel two-dimensional tables, and each two-dimensional table uses a state vector as an index to find the corresponding index value; wherein: each index value includes:

[0015] The value to jump in the S state, the value not to jump in the S state, and the confidence level; the two-dimensional table uses a 32-bit integer fixed-point data structure; the decimal point of the fixed-point data structure is between the 15th and 16th bits, the highest bit is the sign bit, 0 represents a positive number, 1 represents a negative number, the next 15 bits represent the integer part, and the lower 16 bits are the decimal part.

[0016] Further, the first storage pool is a translation lookaside buffer composed of a first-in, first-out queue. The first storage pool stores the execution memory values of historical branches, wherein: the branch predictor queries the tuples in the first storage pool for the received branch instructions, including;

[0017] If a corresponding tuple is found in the first storage pool, it is judged whether the predicted value in the tuple is the same as the actual execution value. When the predicted value is the same as the actual execution value, the previous predicted value is continued; otherwise, the actual execution value is used as the standard.

[0018] Further, the instruction processing module includes a program counter and a global history register.

[0019] Beneficial effects

[0020] Compared with the traditional technical solution, the beneficial effects brought by the present invention are:

[0021] 1. The branch prediction accuracy is improved. The branch predictor based on reinforcement learning can dynamically update and adjust the prediction strategy according to the program operation. The application of real numbers enables the branch predictor to capture the differences between different branches, facilitating more accurate predictions.

[0022] 2. Hardware resources are utilized more efficiently. When the table storage is limited to 64 KB, using the SPEC 2017 dataset workload, with Mispredictions Per Kilo Instructions, MPKI as the standard, the prediction accuracy of Q-BP is 11.38% higher than that of the gshare branch predictor under the condition of using the same historical length; at the same time, it is 9.1% higher than that of the neuron-based predictor.

[0023] 3. Strong scalability. In the reinforcement learning algorithm of Q-BP, the state vector ensures scalability. The selection of the state vector mainly includes control flow and data flow. The most basic Q-BP selects PC in the control flow and GHR in the data flow as the state vector. In fact, all data in the program runtime environment (such as data in memory, etc.) can be used as the state vector. In addition, this technology can be used in combination with other prediction technologies to further improve the prediction accuracy and processing ability.

[0024] 4. Strong adaptability. Compared with traditional branch predictors, Q-BP can better explore the program runtime environment and handle more complex branch situations.

[0025] 5. More advanced and efficient feedback update mechanism. As a natural agent, the branch predictor continuously adjusts the prediction strategy according to the prediction results to maintain high prediction accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Schematic diagram of the structure of a branch predictor based on reinforcement learning according to the present invention;

[0027] Figure 2 Schematic diagram of the second storage pool in a branch predictor based on reinforcement learning according to the present invention;

[0028] Figure 3 Fixed-point table of the second storage pool in a branch predictor based on reinforcement learning according to the present invention;

[0029] Figure 4 Flowchart of generating prediction values by a branch predictor based on reinforcement learning according to the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0030] The following is an explanation of the present invention with reference to the attached Figure 1 - attached Figure 4 The following description is made for the present invention:

[0031] As Figure 1 shown, the present invention provides a branch predictor based on reinforcement learning. The branch predictor 100 includes a first storage pool 101, an instruction processing module 102, a second storage pool 103, and a processor 104; wherein:

[0032] When the program starts running, the queue is empty. As the program runs, the processor 104 constructs the state value, prediction value, actual execution value, and obtained reward value of the executed branch instruction into a tuple, that is, an empty queue, and records it in the first storage pool 101.

[0033] When a new branch instruction arrives at the branch predictor, the branch predictor 100 will first search for the corresponding tuple in the first storage pool 101.

[0034] The first storage pool 101 consists of a translation lookaside buffer (TLB) formed by a first-in-first-out (FIFO) queue. The first storage pool 101 stores the execution memory values of historical branches, and each execution memory value stores the most recently executed state vector. If a corresponding tuple is found in the first storage pool 101, it is determined whether the predicted value and the actual execution value in the tuple are the same. When the predicted value and the actual execution value are the same, the previous predicted value is continued; otherwise, the actual execution value is used as the standard.

[0035] The instruction processing module 102 includes a program counter and a global history register. The instruction processing module performs an exclusive OR operation on the branch instruction through the program counter and the global history register to obtain a state vector, which is input as the control flow into the second storage pool 103.

[0036] The second storage pool 103 consists of n two-dimensional tables. The structure of each two-dimensional table is the same. The state vector given by the instruction processing module is used as the index value, and each index value can find three values, Q(S,T), Q(S,NT), and confidence, as Figure 2 shown. Among them, Q(S,T) is the Q value of jumping in state S, Q(S,NT) is the Q value of not jumping in state S, and confidence is the confidence level of whether the prediction is correct. Among them:

[0037] When the corresponding tuple is not found in the first storage pool 101, the second storage pool 103 is searched for the index value. The second storage pool 103 is the core of the entire branch predictor 100. It not only outputs the prediction result but also dynamically updates the prediction strategy according to the feedback of the program environment, which more prominently reflects the innovation of applying reinforcement learning to the processor.

[0038] The second storage pool 103 includes several two-dimensional tables. The two-dimensional tables are tables that index different parts of the state vector and are searched in parallel. Finally, the index value corresponding to the larger action is selected as the corresponding predicted value, and the predicted value is fed back to the processor. The CPU calculates the actual execution value of this branch in a certain cycle later, and then gives the corresponding reward value. Specifically, the reward value is 1 if the prediction is correct, otherwise it is -1. The second storage pool uses the Q-learning algorithm to update the corresponding index value indexed by the state vector.

[0039] During this request process, the processor 104 constructs an empty queue as a tuple and stores it in the first storage pool 101. If the empty queue is full, the earliest stored tuple is deleted. At this time, the memory execution value of the first storage pool 101 is updated in the two-dimensional table in the second storage pool. In addition, real numbers can represent more meanings than integers. The second storage pool 103 uses real numbers as index values. To solve the technical problem that real numbers are difficult to calculate at the hardware level, the present invention uses fixed-point data to construct the two-dimensional table. For specific operations, see Figure 3 The second storage pool of the present invention adopts multiple two-dimensional tables that can be searched in parallel. Each two-dimensional table uses a state value as an index to search for the corresponding index value, and finally, the final index value is selected according to the decision of the branch predictor as the basis for prediction.

[0040] The fixed-point of the present invention is a 32-bit integer fixed-point structure, that is, it is agreed that the decimal point is between the 15th and 16th bits. The highest bit is the sign bit, 0 represents a positive number, 1 represents a negative number, the next 15 bits represent the integer part, and the lower 16 bits are the decimal part.

[0041] As Figure 4 shown, in the process of actually generating the predicted value in the present invention; the search process is divided into five stages

[0042] In the first stage, an index vector is generated. The present invention uses the exclusive OR value of the program counter PC and the global history register GHR as the state vector; if other state values are added, only the eigenvalue needs to be added after the state vector.

[0043] In the second stage, the corresponding tuple is searched in the first storage pool queue.

[0044] In the third stage, the index values of the corresponding eigenvectors are searched in parallel in the n two-dimensional tables of the second storage pool.

[0045] In the fourth stage, the final index value is output as the predicted value.

[0046] Although the present invention has been described above, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many variations without departing from the purpose of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A branch predictor based on reinforcement learning, characterized in that, The branch predictor includes: a first storage pool, an instruction processing module, a second storage pool, and a processor; wherein: the process of the branch predictor obtaining a prediction value by processing a branch instruction includes: The branch predictor queries for tuples in the first storage pool for the received branch instruction; When there is no corresponding tuple in the first storage pool, the instruction processing module performs an exclusive OR operation on the branch instruction to obtain a status vector, and inputs the status vector as an index value to the second storage pool; The second storage pool queries the status vector in the two-dimensional table for its corresponding maximum index value and feeds it back to the processor as the prediction value; The processor calculates the actual execution value from the prediction value in a subsequent cycle, and gives a reward value corresponding to the actual execution value and feeds it back to the second storage pool; The second storage pool updates the two-dimensional table with the execution memory value for each branch instruction cycle; The branch predictor selects the final index value in the two-dimensional table as the prediction value output; 2. The branch predictor based on reinforcement learning according to claim 1, wherein The second storage pool is composed of at least 3 parallel two-dimensional tables, and each two-dimensional table uses the status vector as an index to find the corresponding index value; wherein: each index value includes: The value to jump in the S state, the value not to jump in the S state, and the confidence level; the two-dimensional table uses a 32-bit integer fixed-point data structure; the decimal point of the fixed-point data structure is between the 15th and 16th bits, the highest bit is the sign bit, 0 represents a positive number, 1 represents a negative number, the next 15 bits represent the integer part, and the lower 16 bits are the decimal part.

3. The branch predictor based on reinforcement learning according to claim 1, wherein The first storage pool is a translation lookaside buffer composed of a first-in, first-out queue. The first storage pool stores the execution memory values of historical branches. Wherein: the process of the branch predictor querying for tuples in the first storage pool for the received branch instruction includes; If a corresponding tuple is found in the first storage pool, then determine whether the prediction value and the actual execution value in the tuple are the same. When the prediction value and the actual execution value are the same, continue with the previous prediction value; otherwise, use the actual execution value as the standard.

4. The branch predictor based on reinforcement learning according to any one of claims 1-3, characterized in that, The instruction processing module includes a program counter and a global history register.