Hybrid branch predictor capable of dynamically selecting different prediction methods and application

By designing a hybrid branch predictor that combines global and local history predictors and optimizes the branch target predictor, the branch prediction problem of traditional microprocessors under limited hardware resources is solved, achieving higher accuracy and lower resource consumption, making it suitable for high-performance processor designs.

CN121680941APending Publication Date: 2026-03-17CHINA STATE SHIPBUILDING CORP LTD RESEARCH INSTITUTE 719 +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511488274.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Traditional microprocessor branch predictors suffer from poor adaptability, low storage efficiency, and high hardware resource overhead, making it difficult to improve branch prediction accuracy with limited hardware resources, especially in IoT devices with complex computational requirements.

Method used

A hybrid branch predictor is adopted, which combines the TAGE global history predictor and the Local local history predictor. The prediction strategy is dynamically selected through the Choice module. The branch target predictor is optimized by using a four-way grouped BTB and a space-efficient RAS, thereby improving the branch prediction accuracy and reducing hardware resource consumption.

Benefits of technology

It significantly improves branch prediction accuracy, optimizes storage efficiency, reduces hardware resource overhead, and is suitable for high-performance processor designs, especially improving computing efficiency in IoT devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121680941A_ABST
    Figure CN121680941A_ABST
Patent Text Reader

Abstract

The invention discloses a hybrid branch predictor capable of dynamically selecting different prediction methods and application, and relates to the technical field of high-performance processor design, the hybrid branch predictor mainly comprises a branch direction predictor and a branch target predictor; the branch direction predictor is used for predicting whether a branch instruction jumps or not; and the branch target predictor is used for predicting a specific jump target address when the branch direction is predicted to be jump. By implementing the hybrid branch predictor capable of dynamically selecting different prediction methods and the application provided by the invention, the branch prediction accuracy can be improved, and the hardware resource overhead can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of high-performance processor design, and more particularly to a hybrid branch predictor capable of dynamically selecting different prediction methods and applications. BACKGROUND

[0002] With the rapid development of the Internet of Things field, the demand for terminal nodes is growing explosively. As the core control unit of terminal node devices, microprocessors have become increasingly important. Traditional microprocessors based on ARM architecture not only require high licensing fees, but also prohibit any unauthorized modifications, which is less flexible. The RISC-V instruction set architecture has the advantages of simplicity, modularity, and open source, which has attracted the attention of major enterprises and researchers.

[0003] The increasing number of Internet of Things nodes has led to a large influx of data into the cloud. To alleviate the computing pressure on the cloud, microprocessors need to be optimized in terms of high performance under the premise of low area, and can execute complex communication algorithms faster and support stable communication between a large number of terminal nodes. The RISC-V instruction set architecture provides a technical breakthrough for the innovation of Internet of Things microprocessors, and its modular nature has given rise to specialized processors for multiple scenarios from edge computing to industrial control. To pursue higher Instructions Per Cycle (IPC), microprocessor architectures have gradually evolved towards parallelization solutions such as multi-issue structures and deep pipeline designs. However, there are dense branch instructions (accounting for 15-30%) in programs such as environmental perception and edge reasoning, which can cause a lot of pipeline flushing and limit the increase of IPC. On the other hand, due to the area limitations of microprocessors in the Internet of Things field, it is forced to use 3-5 stage pipeline structures, making it difficult to meet the increasingly complex computing demands of terminal devices.

[0004] To reduce the impact of branch instructions on microprocessor performance, modern microprocessors often use branch prediction technology for speculative execution. However, when branch prediction errors occur, all instructions in the pipeline need to be cleared. The deeper the pipeline, the more instruction cycles are lost, which severely affects the efficiency of instruction execution. Therefore, existing microprocessors have to use more complex branch prediction schemes to achieve higher branch prediction accuracy, which not only increases the delay of the critical path but also occupies more hardware resources. At the same time, in the limited hardware resources of Internet of Things microprocessors, it is impossible to reduce the impact of inter-stage hazards on pipeline efficiency through complex bypass networks, and shallow pipelines mean that more work needs to be done at each stage, resulting in increased performance loss of microprocessors due to each hazard.

[0005] Therefore, traditional branch predictors have the following drawbacks: poor adaptability to single prediction strategies: for example, the TAGE predictor relies on global historical information and has insufficient prediction accuracy for branch instructions that are sensitive to local history; low storage efficiency: the BTB structure suffers from decreased prediction accuracy due to storage conflicts in multi-target branch scenarios, and the traditional 32-bit target address storage space is large; inefficient recursive call processing: RAS has low space utilization in recursive scenarios and is prone to prediction errors due to stack overflow. Summary of the Invention

[0006] The purpose of this invention is to provide a hybrid branch predictor and its application that can dynamically select different prediction methods, thereby improving branch prediction accuracy and reducing hardware resource overhead.

[0007] This invention provides a hybrid branch predictor, including a branch direction predictor and a branch target predictor; the branch direction predictor is used to predict whether a branch instruction will result in a jump; the branch target predictor is used to predict the specific jump target address when the branch direction is predicted to be a jump.

[0008] The present invention also provides an application of the hybrid branch predictor as described above, applied to a microprocessor.

[0009] Implementing the hybrid branch predictor and its application that can dynamically select different prediction methods provided by this invention has the following beneficial effects: This invention integrates a TAGE global history predictor and a Local history predictor, dynamically selects prediction strategies through a Choice module, and utilizes a storage-optimized target predictor composed of a four-way set-associative BTB and a high-space-utilization RAS to achieve a hybrid direction prediction mechanism. Specifically, the dynamic collaboration between the TAGE and Local predictors improves direction prediction accuracy; a four-way set-associative BTB compresses storage space; and a RAS with a counter optimizes recursive call prediction. By dynamically selecting global and local history prediction strategies and combining a storage-optimized branch target prediction mechanism, this invention significantly improves branch prediction accuracy, optimizes storage efficiency, and reduces hardware resource overhead, achieving overall performance improvement. It solves the problems of poor history type adaptability and low storage efficiency in traditional technologies and is suitable for high-performance processor designs. Attached Figure Description

[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a structural diagram of the hybrid branch predictor provided by the present invention; Figure 2 This is the overall structure diagram of the TAGE provided by the present invention; Figure 3 This is a structural diagram of the Local predictor provided by the present invention; Figure 4 This is a schematic diagram illustrating the impact of BHR width on prediction accuracy provided by the present invention; Figure 5 This is a diagram of the BTB structure provided by the present invention; Figure 6 This is a diagram of the BTB structure of the four-way group interconnection structure provided by the present invention; Figure 7 This is a schematic diagram of the branch instruction types provided by the present invention; Figure 8 This is a schematic diagram of the four-way group-connected structure BTB update process provided by the present invention; Figure 9 This is a schematic diagram of a three-level nested subroutine call provided by the present invention; Figure 10 This is a schematic diagram of the prediction process of CALL and RET by RAS provided by the present invention; Figure 11 This is a schematic diagram of the RAS structure for increasing space utilization provided by the present invention; Figure 12 This is a schematic diagram of the test platform for the hybrid branch direction predictor provided by the present invention; Figure 13 This is a schematic diagram of the simulation waveform of the hybrid branch direction predictor provided by the present invention; Figure 14 This is a schematic diagram showing the comparison results between the hybrid direction predictor provided by this invention and other branch direction predictors; Figure 15 This is a schematic diagram of the simulation waveform of the branch target predictor function provided by the present invention. Detailed Implementation

[0011] To provide a clearer understanding of the technical features, objectives, and effects of the present invention, specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0012] Figure 1 A schematic diagram of the hybrid branch predictor of this embodiment is shown. In this embodiment, the hybrid branch predictor includes a branch direction predictor and a branch target predictor; the branch direction predictor is used to predict whether a branch instruction will result in a jump; the branch target predictor is used to predict the specific jump target address when the branch direction is predicted to be a jump.

[0013] In one exemplary embodiment, the branch direction predictor includes a global history prediction module, a local history prediction module, and a selection module; The global history prediction module is used to make predictions based on branch instructions that depend on the global branch history. The local history prediction module is used to provide supplementary predictions for branch instructions that rely on their own local history. The selection module is used to dynamically select the prediction results of the global historical prediction module or the local historical prediction module as the final direction.

[0014] In one exemplary embodiment, the global historical prediction module includes a base predictor, a first label prediction table counter, a second label prediction table counter, a third label prediction table counter, and a fourth label prediction table counter.

[0015] In one exemplary embodiment, the base predictor is composed of a saturation counter, and the tag prediction table counter consists of a tag for table entry matching, a saturation counter, and a usefulness counter indicating the usefulness of the table entry.

[0016] In one exemplary embodiment, the global history prediction module is specifically configured to: perform hash operations using global history data of different lengths and program counter data to generate indexes and labels for each label table; the prediction result of the label table is valid only when the label matches, and the long history matching result is used first; if all label tables are not matched, the prediction result of the basic predictor is used by default.

[0017] In one exemplary embodiment, the global history prediction module is further configured to: adjust the prediction confidence based on the actual jump result, determine whether the entry is useful, write to the global history register when a new entry is allocated, and write the current jump direction to the global history register to ensure that the prediction adapts to changes in branch behavior.

[0018] In one exemplary embodiment, the local history prediction module includes a local branch history table register and a local pattern history table counter. The local branch history table register is a branch history register used to store the local history of each branch instruction and record the past jump patterns of the branch instruction itself. The local mode history table counter consists of a two-bit saturation counter, used to store the branch jump mode.

[0019] In one exemplary embodiment, the local history prediction module is specifically configured as follows: using the lower 10 bits of the program counter to index the local branch history table register, the branch history register of the branch is read out; the branch history register is concatenated with bits 7-10 of the program counter to form a 10-bit index, the local mode history table counter is accessed, and the counter value is read out as the prediction result.

[0020] In an exemplary embodiment, the local history prediction module is further configured to: after branch execution, write the actual jump direction into the branch history register of the local branch history table register, update the low bits and discard the high bits; and adjust the corresponding counter of the local mode history table counter.

[0021] In one exemplary embodiment, the selection module consists of a three-bit saturation counter, which is used to dynamically select the prediction result of the global historical prediction module or the local historical prediction module as the final prediction result based on the branch historical bias using a multiplexer.

[0022] In one exemplary embodiment, the branch target predictor includes a four-way set-associative branch target cache and a return address stack with a counter; The branch target cache is used to provide address prediction for branch instructions with fixed target addresses (such as direct jumps and conditional jumps) and to store the jump target address; The return address stack is used to predict the target address of indirect jump instructions.

[0023] In some embodiments, the hybrid branch predictor described above can also be implemented in the following ways.

[0024] In this embodiment, the branch prediction unit is divided into a branch direction predictor and a branch target predictor, and its overall framework is as follows: Figure 1 As shown in the diagram, MUX stands for Multiplexer; PC stands for Program Counter; in the design of the branch direction predictor, a hybrid predictor capable of dynamically selecting different prediction methods is used, mainly including two predictors: TAGE (Tagged Geometric History Length) and Local, with Choice responsible for selecting the prediction method; in the design of the branch target predictor, two prediction modes are used for different types of branch instructions: a four-way set-associative BTB (Branch Target Buffer) and a space-efficient RAS. The final predicted address is determined by the branch direction predictor to be the address of the next instruction to be fetched sequentially. It is still the prediction result of the branch target predictor.

[0025] 1. Branch direction predictor based on hybrid prediction method (1) TAGE predictor based on global history The TAGE branch direction predictor mainly includes a basic prediction table. and M tagged prediction tables representing label matching During prediction, different prediction tables are accessed using branch history information of varying lengths to obtain the optimal prediction result. The TAGE structure designed in this embodiment is as follows: Figure 2 As shown in the figure, the TAGE designed in this embodiment has a basic predictor. And four tagged prediction tables, respectively , , as well as ,in The basic predictor, Bimodal, consists only of saturation counters. The TAGE prediction table consists of a tag for matching the entry, a saturation counter ctr, and a usefulness counter u that indicates the usefulness of the entry. The parameters of each prediction table in TAGE are shown in Table 1.

[0026] Table 1: Parameter Allocation for Each Prediction Table in TAGE

[0027] `hist` represents the global history of branch instructions, indicating the branching behavior of executed branch instructions within the global history. The history length used by each prediction table... The calculation formula is as follows, where MinHist and MaxHist represent the minimum and maximum history lengths, respectively, and (int) represents the rounding operation.

[0028]

[0029]

[0030]

[0031] To avoid multiple branch instructions indexing the same table entry, the index address of the tagged prediction table is hashed with the PC using historical information of geometric length. The pseudocode for calculating the index of the table item in the prediction table is shown in Table 2.

[0032]

[0033] In the pseudocode above, This is the ID of the current tagged prediction table. The record contains the historical length that the bank forecast table can utilize. For the depth of the bank forecast table, This indicates the preprocessing of the path history. This indicates the absolute value operation. This refers to the directional history after it has been folded. The specific calculation process is shown in Table 3, where Indicates path history, Indicates the length of the path history:

[0034] In the prediction process of TAGE, The prediction table uses only the saturation counter in the PC address index entry, utilizing its sign bit to generate the prediction direction. For T i The prediction table first determines the relationship between PC and... Does the tag value obtained by hashing the length of the hist match the length of the tag? In the prediction table If they are the same, it means that the prediction table has hit, and only the prediction result given when it hits is valid. The calculation of the tag only uses a POR operation between the PC and the collapsed global history. After calculating the hit status of all tagged prediction tables, the final prediction result is selected according to the length of the global history used by each prediction table. The prediction table with a longer history has a higher priority. Referring to the provider and alt-pred defined by the TAGE authors, they represent the prediction results of the longest and second longest tagged matching prediction tables, respectively. That is, the provider represents the prediction table that finally provides the prediction. For the sake of simplifying the design, this embodiment generally uses alt-pred. The prediction results. If If neither is hit, then both the provider and alt-pred are basic predictors. The prediction results. The TAGE predictor updates primarily target the contents of table entries and global historical information; the specific update operation depends on the actual prediction situation. For the update of u: an update is only performed when the prediction results provided by the provider and alt-pred differ. If the provider predicts correctly, it means using a longer global history for prediction is necessary, and the u value for the corresponding table entry in the provider is incremented by 1. If the provider predicts incorrectly, but alt-pred predicts correctly, it means that a longer global history is not necessary, and the u value for the corresponding table entry is decremented by 1. Simultaneously, the u bit is periodically reset to ensure the predictor always maintains relatively recent historical information. For the update of ctr: when... When none of them hit, If prediction results are provided, only update The ctr function increments by 1 if the actual result is a jump, and decrements by 1 if an error occurs; when Upon a hit, the system checks if the provider's prediction is correct. If correct and the actual jump result is a jump, the corresponding saturation counter increments by 1; otherwise, it decrements by 1. If the alt-pred prediction is incorrect, it updates according to the actual jump result. If a jump occurs, the ctr increments by 1; otherwise, it decrements by 1. For tag updates: this mainly occurs during new entry allocation. When the final prediction is incorrect and the provider is not the prediction table with the longest history, it attempts to allocate a new entry in a prediction table with a longer history than the provider. Only one entry is allocated at a time, with prediction tables closer to the provider having higher priority. During allocation, it primarily searches for prediction tables where the corresponding entry's u bit is 0, writes a new tag bit, and resets u to a strongly useful state. Simultaneously, it resets the ctr based on the actual jump result. If a jump occurs, it resets to a weak jump; otherwise, it resets to a weak no jump. If no entry with u is 0 exists, no new entry is allocated, and the corresponding u in all prediction tables with a history longer than the provider's is decremented by 1. For updating `hist`: the actual jump direction of the branch instruction is written to the low-order bits of the global history register `hist`, and the high-order bits are discarded. Compared to other branch direction predictors, TAGE has excellent prediction accuracy, mainly due to its use of partial tag matching, which reduces the occurrence of branch aliases. Furthermore, multiple tagged prediction tables use global history information of different lengths, enabling it to handle different branch environments. However, its prediction accuracy is slightly insufficient when dealing with branch instructions that are more related to local history.

[0035] (2) Based on local history predictor For the same branch instruction, different prediction methods may need to be dynamically selected based on the specific context during instruction execution. A simple TAGE predictor based on global history cannot adapt to complex branching environments. To address this issue, this embodiment designs a local history-based branch predictor, Local, that works in conjunction with TAGE. TAGE is used to predict branch instructions that rely more on global history, while Local is responsible for predicting branch instructions that rely on local history. Its structure is as follows: Figure 3As shown in the figure. This predictor effectively captures the local jump patterns of branch instructions by recording local history information, thus providing accurate prediction results even when the correlation of global history information is weak. Local, as a two-level predictor, mainly consists of two parts: the Local Branch History Register (LBHT) and the Local Pattern History Table (LPHT). The LBHT is composed of the Branch History Register (BHR), which records the past jump patterns of the branch instructions themselves, while the LPHT is composed of a two-bit saturation counter. In this study, this embodiment focuses on the impact of the BHR width on prediction accuracy. The BHR width determines the length of historical jump information that the branch history register can record. Theoretically, increasing the BHR width can capture longer historical jump patterns, thereby improving prediction accuracy. However, increasing the BHR width also means a significant increase in hardware storage overhead. Therefore, this embodiment, based on a trade-off between storage overhead and prediction performance, limits the research range of the BHR width to between 0 and 8. The experimental results are shown in the figure. Figure 4As shown, the prediction accuracy initially rises rapidly and then gradually levels off as the BHR width increases. When the BHR is 0, meaning local historical information is not used for prediction and the PC is used directly to index the LPHT, the prediction accuracy is the lowest, at only 88.3%. As the BHR width increases, the prediction accuracy gradually increases, stabilizing at 6 bits with a prediction accuracy of 92.8%. Further increasing the BHR width reduces the improvement in prediction accuracy, reaching its highest at 8 bits (93.3%), but the number of LPHT entries increases exponentially with the BHR width. Therefore, considering all factors, the BHR width in this embodiment is chosen to be 6 bits. Since this embodiment's Local predictor uses only one LPHT, there is a problem where different branch instructions with the same BHR index the same LPHT entry. To solve this problem, this embodiment uses bit concatenation, concatenating bits from different PC addresses with the BHR to form a new index address for accessing the LPHT. Therefore, the prediction process of the Local predictor is as follows: First, the lower 10 bits of the branch instruction's PC are used to generate the LBHT index address INDEX_BHT. This address is used to access the LBHT and read the value of the corresponding branch's BHR. Second, the 6-bit wide BHR result is concatenated with bits 7 to 10 of the PC address to form a 10-bit signal, which serves as the LPHT index address INDEX_PHT. Finally, the INDEX_PHT index is used to read the saturation counter value of the LPHT as the prediction result of the Local predictor. The Local update is divided into two parts. For the LBHT update, the actual jump result of the branch instruction is written to the lower bits of the corresponding BHR, and the higher bits are discarded. For the LPHT update, if the actual result of the branch instruction is a jump, the saturation counter in the corresponding entry in the LPHT is incremented by 1; otherwise, it is decremented by 1.

[0036] 2. Storage-optimized multi-target branch target predictor (1) Branch target cache with four-way group associative structure Based on the foregoing analysis, the jump target address for conditional jump instructions and direct jump instructions falls into only two categories: one, when the branch instruction does not jump, the jump target address is the current instruction's PC address plus the instruction's byte count; and two, when the branch instruction jumps, the jump target address is the current PC address plus an address offset. The address offset is an immediate value carried in the instruction; therefore, for direct jump branch instructions, the jump target address remains unchanged. A table can be used to record the jump target addresses of each direct jump instruction. The predicted jump target address can be obtained by looking up the table using the PC address the next time the instruction is executed. This table is called the BTB, and its structure is as follows: Figure 5As shown. The Branch Target Address (BTB) is essentially a cache that uses the PC address as its index. Due to the limited capacity of the BTB, only a portion of the PC address can be used as the index. Therefore, instructions with the same index address but different PC addresses may be indexed into the same entry, affecting the prediction accuracy of the BTB. Thus, in addition to recording the Branch Target Address (BTA), the BTB also stores a portion of the PC address as a tag to distinguish different instructions. The BTB also contains a Valid bit indicating whether the entry is valid. During prediction, the PC index is used to access the BTB, the prediction result in the entry is read, and then the Tag bit in the entry is compared with the PC address portion. If they match and the Valid part of the entry is 1, it indicates a hit, meaning the current predicted target jump address is valid. If there is a miss, the predicted jump target address is PC+4, which is the address of the next instruction to be fetched sequentially. During updates, if the BTB hits and the prediction is correct, the entry is not updated. If the prediction is incorrect, there are generally two scenarios: one is a direction prediction error, meaning the branch jumped in the previous iteration but not this one; this usually occurs in the last loop of a program, and the update simply sets Valid to 0, indicating that the branch does not jump. The other scenario is an incorrect jump target address, meaning different branch instructions index the same entry; in this case, the new target jump address is written to the corresponding entry in the BTB. If the BTB misses, and the actual result of the branch instruction is a jump, a new entry is created, and the jump target address of the branch instruction and a Tag consisting of bits from the PC address are written to the BTB, with Valid set to 1. If the actual result of the branch instruction is not a jump, the BTB entry is not updated; only branch instructions that actually result in a jump are recorded. This allows for the prediction of more jump branch instructions within the same memory space. In this single-target branch-oriented BTB structure, multiple branch instructions indexing the same entry can lead to frequent changes in the target address within the BTB. Therefore, this embodiment adopts a four-way group-associative BTB structure, which divides the entry corresponding to each branch instruction into multiple groups, where each group has 4 channels and can store the jump target addresses of 4 different branch instructions, as shown in the structure below. Figure 6 As shown. This embodiment adds a Least Recently Used (LRU) counter consisting of two bits to the traditional four-way group-associative BTB structure, as well as a 2-bit Type bit indicating the branch instruction type. The LRU counter counts the least recently used entries in the four-way group and evicts them when more than four branch instructions are indexed into the same group. Storing the complete jump target address in each entry would occupy a lot of storage space; therefore, the Type bit is used to indicate which of the RISC-V32I instruction sets (B, I, J) the branch instruction type belongs to. The branch instruction types are as follows:Figure 7 As shown, the `Type` parameter is stored as binary 01 for B-type branch instructions, binary 10 for I-type instructions, and binary 11 for J-type instructions. Depending on the instruction type, only the address offset represented by the immediate value is recorded in the BTB. For J-type instructions, the immediate value is the longest at 20 bits. To facilitate recording, each entry uses 20 bits of storage space for the immediate value. When an entry is hit, the instruction offset is first expanded according to the branch instruction type, and then the PC address is added to generate the predicted jump target address. This eliminates the need to store the entire jump target address, effectively saving BTB storage space. However, it adds an addition calculation, causing additional path delay, which is within an acceptable range compared to the more important area metrics. In this embodiment, the four-way set-associative BTB prediction first uses the lower 10 bits of the PC address as the set index and simultaneously accesses the corresponding entries of the four BTBs within the set. Then, the remaining part of the PC address is compared with the Tag parts of the four BTBs within the set. If one matches and the Valid bit in the matching path is 1, it indicates a hit. The immediate value in the matching path is expanded according to the branch instruction type and added to the current PC address to become the predicted jump target address. If none of the four Tag bits in the set match, or if a matching path exists but the Valid bit is 0, it indicates a BTB hit. In this case, the default branch instruction does not jump, and the jump target address is the PC address plus the instruction byte count. The update logic of the four-way set-associative BTB in this embodiment is as follows: Figure 8As shown in the diagram, Alloc represents creating a new entry. When the four-way set-associative structure BTB is not hit, if the actual result of the branch instruction is a jump, a new entry is created in the one with the largest LRU count (least recently used) among the four ways. If all four LRU counts are at their maximum (e.g., all LRU counters are set to maximum during initialization), a new entry is created by default in the first way, writing the branch instruction type, address offset, and tag. The Valid bit is also set to 1, indicating that the entry's result is valid. Simultaneously, the LRU counter of the newly created entry is initialized to the latest value. If the actual result of the branch instruction is no jump, the BTB is not updated. When a four-way set-associative BTB hits, if the actual result is a jump and the BTB prediction matches the actual jump target address, it indicates that the current branch instruction prediction is correct. The LRU counter of the predicted path is updated to the latest, indicating that the entry was recently used. If the predicted address does not match the actual jump target address, it indicates that multiple branches index to the same entry. In this case, a new entry is created in the path with the largest LRU among the four paths, and the LRU of that path is initialized to the latest. The LRU counters of the other paths are incremented by one, indicating that the other paths are used. If the actual result is no jump, the Valid value of the predicted path entry is set to zero, indicating that the branch does not jump, and the information of the other branch instructions in the entry remains unchanged. Only when the branch instruction jumps again is the Valid value reset to 1 during the update. The four-way set-associative BTB designed in this embodiment can provide very effective prediction for conditional jump instructions with fixed jump target addresses and direct jump instructions, significantly improving the processor's execution efficiency. However, accurate prediction is still not possible for indirect jump instructions whose target addresses frequently change (such as the RET subroutine return instruction emulated by JARL). This is because the target address of indirect jump instructions typically depends on runtime calculations and has significant uncertainty, making effective prediction difficult for BTB.

[0037] (2) High space utilization return address stack In the RISC-V instruction set, JAR / JARL instructions are typically used to call and return subroutines. The call is called a CALL instruction, and the return is called a RET instruction. When the destination register RD address of a JAR instruction is x1 or x5, it usually indicates that the JAR instruction is a CALL instruction; while for JARL instructions, the instruction type is determined by the linking of RD and RS1, as shown in Table 4.

[0038] Table 4: CALL and RET instructions in RISC-V (link indicates "x1 or x5")

[0039] The subroutine called each time is fixed, so the target address of the CALL instruction it represents is generally fixed and can be predicted using BTB. However, the location where a subroutine is called is not fixed. For example, the printf function in C may be called anywhere, which causes the return address of the RET instruction to frequently change, making it unsuitable for BTB prediction. During program execution, subroutines may call other subroutines, resulting in nested subroutine calls. This nested calling causes the return addresses of the RET instructions to form a stack structure, meaning the subroutine called later returns first. CALL and RET instructions usually appear in pairs, and the return address of the RET instruction always points to the address of the instruction following the corresponding CALL instruction. Figure 9 This diagram illustrates a nested subroutine call. It shows that when the main program executes the CALL0 instruction to call subroutine A, the return address of the corresponding RET0 instruction points to the address of the instruction following CALL0, MUL. Before the execution of RET0, CALL1 and CALL2 instructions appear, calling subroutines B and C respectively. The return addresses of their corresponding RET1 and RET2 instructions point to SUB and ADD instructions respectively. Only after subroutine C completes the RET2 instruction can it return to subroutine B to execute the RET1 instruction, then return to subroutine A to execute the RET0 instruction before returning to the main program.

[0040] Based on the characteristics of the RET instruction, this embodiment uses the Return Address Stack (RAS) for prediction. The RAS is a Last-In-First-Out (LIFO) memory, and its top address is controlled by the pointer top_ptr. The prediction process for CALL and RET instructions is as follows: Figure 10 As shown. Each time a CALL instruction is executed, top_ptr is incremented by one, and the address of the next instruction of the CALL instruction is pushed onto the stack top pointed to by top_ptr in the RAS. Simultaneously, a four-way set-associative (BTB) structure is used to predict the jump destination address of the CALL instruction. When a RET instruction is encountered, the address at the top of the stack is popped as the predicted jump target address, and the pointer top_ptr is decremented to point to the new stack top. In actual design, the RAS has limited capacity. When dealing with recursive function calls, the return address of the same CALL instruction may be pushed onto the RAS multiple times, thus occupying a large amount of RAS space, causing RAS overflow, and consequently leading to prediction errors. To address this situation where the RAS contains a large number of return addresses of the same CALL instruction, this embodiment uses an 8-bit counter to count the number of times the return address of the same CALL instruction is pushed onto the RAS. Its structure is as follows: Figure 11As shown. When executing a CALL instruction, the return address of the CALL instruction about to be pushed onto the RAS is compared with the return address of the previously pushed instruction. If they are the same, it indicates that it is the return address of the same CALL instruction, and only the counter value is incremented without moving the address pointer of the RAS. When executing a RET instruction, the value of the counter in the current entry is decremented by 1 each time the return address is read from the top of the RAS stack, until the counter reaches 0, at which point the address pointer points to the return address of the next CALL instruction. By adding a counter, the RAS can handle scenarios such as recursive calls more efficiently, reduce the waste of RAS space, and improve the accuracy of prediction, thus reducing the occupation of hardware resources while ensuring performance.

[0041] In the branch direction predictor, considering the different branch history information relied upon by different branches, a tagged geometric history length (TAGE) is used to predict instructions biased towards global history branches, while a local history predictor (Local) is used to predict instructions biased towards local history branches. The resulting hybrid predictor can effectively adapt to different branch situations. For branch instructions with multiple target addresses that the branch target cache (BTB) cannot predict, a four-way set-associative BTB structure is adopted, incorporating an Least Recently Used (LRU) counter and a Type (bit) component. By statistically analyzing instruction types and the age of entries, hardware resource consumption is effectively saved. To improve the accuracy of indirect instruction prediction, a Return Address Stack (RAS) is designed, using a counter to track the return addresses of the same instruction, increasing space utilization.

[0042] 3. Branch predictor function verification (1) Testing and analysis of hybrid branch direction predictor To verify whether the prediction accuracy of the branch direction predictor based on the hybrid prediction method designed in this embodiment can meet expectations, a test platform was built using Verilog language and simulated using Modelsim software. The excitation signals adopted the test benchmark provided by the 3rd Branch Predictor Competition CBP-3, which includes five types (CLIENT, INT, MM, SERVER, WS), totaling approximately 50 million microinstructions. The test platform structure is as follows: Figure 12As shown. First, the hybrid branch direction predictor is instantiated, generating the corresponding clock and reset signals. Then, the test stimuli are read into the instantiated model. Only one instruction is read at a time. Through the control of the internal state machine, the read address pointer is automatically incremented by 1 after each prediction-update operation. The prediction result is compared with the actual result. If they are the same, it means the prediction is correct and is sent to the statistics module for statistical analysis. At the same time, the actual result is also used to update the hybrid branch direction predictor. The statistics module is responsible for counting the total number of correctly predicted instructions, the total number of instructions, and the number of correctly predicted instructions per 100,000 instructions. Finally, these results are printed to the display. Before testing the prediction accuracy of the hybrid predictor, its functional correctness is verified through a test platform, as shown below. Figure 13 The diagram shows the simulation waveform of the hybrid predictor's prediction logic. As can be seen from the diagram, when the PC address is 004ef532, the indices of the four tagged prediction tables are 288, 188, 041, and 009 in hexadecimal representation, and the tags are 37, 38, 11f, and 0a6 in hexadecimal representation. The Bimodal index is the lower 12 bits of the PC address. eq0, eq1, eq2, and eq3 represent the hit status of the four tagged prediction tables. It can be seen that only tagge0 and tagged1 are hits, but tagged1 uses a longer hist, thus having higher priority. Therefore, the final prediction result for TAGE is provided by tagged1, and the prediction result `prediction_branch_prediction` does not jump, while the prediction result `local_result` of the Local predictor jumps. `Choice` consists of a three-bit saturation counter, responsible for selecting the final prediction result based on the historical bias of the branch using `update_select`. A 1 in the highest bit indicates TAGE is selected, and a 0 in the highest bit indicates Local is selected. At this point, the highest bit of `update_select` is 1, meaning that this branch instruction is more inclined to predict based on global history. Therefore, the final prediction result is provided by TAGE, indicating that this branch will not jump. The actual jump result `Actual_branch` also shows no jump, resulting in a correct prediction. To verify the advantages of the hybrid branch direction predictor designed in this embodiment, it was compared with several classic branch direction predictors on a test platform: Bimodal, Local, Gshare, and the original TAGE before optimization. The comparison metric was the number of prediction errors per thousand instructions (MPKI). The results are as follows: Figure 14 As shown in Table 5, the PHT depths of the three branch direction predictors—Bimodal, Local, and Gshare—are all 8192. The parameter allocation of the hybrid branch direction predictor designed in this embodiment is shown in Table 5. The original TAGE base predictor before optimization. The depth is 8192, and it is in the hybrid branch direction predictor. The total hardware consumption of Local and Choice is the same.

[0043] Table 5: Parameter Allocation of Hybrid Branch Direction Predictor

[0044] As can be seen, the hybrid predictor proposed in this embodiment has the lowest MPKI under different test stimuli. It exhibits significant advantages over single prediction methods based on local history (Local) and global history (Gshare). Compared to the Local predictor, the average MPKI decreases by 36.6, and the false prediction rate decreases by 56.7%; compared to the Gshare predictor, the average MPKI decreases by 33.5, and the false prediction rate decreases by 54.5%. Compared to Bimodal, which is simply composed of saturated counters, the MPKI decreases by 44.2, and the false prediction rate decreases by 61.2%. Furthermore, under the same hardware resource consumption, compared with the traditional TAGE, it achieves the most significant prediction performance under the INT03 test stimulus, with an MPKI decrease of 7.8, an overall average MPKI decrease of 4.7, and a false prediction rate decrease of 14.3%.

[0045] (2) Functional verification and testing of the branch target predictor The testing of branch target predictors differs from that of branch direction predictors. It requires simultaneously judging both the branch jump result and the predicted branch target address; only when both results are correct is the prediction considered correct. Therefore, the testing method used for branch direction predictors cannot be continued. Instead, the five-stage pipelined microprocessor based on the RV32IM instruction designed in this embodiment is selected as the test platform. During testing, the branch target predictor's function is integrated into the instruction fetch stage to predict the jump target address of branch instructions. During the execution stage, the microprocessor judges the actual jump result and target jump address of the branch instruction and compares them with the prediction result. To statistically analyze the performance of the branch target predictor, this embodiment adds two counters during the execution stage to count the number of branch instructions and the number of prediction errors, respectively, to obtain the prediction accuracy of the branch predictor. Table 6 compares the parameters of each entry in the four-way set-associative BTB structure designed in this embodiment with those in the traditional four-way set-associative BTB structure. Both have a depth of 1024. It can be seen that each entry in the four-way set-associative BTB structure designed in this embodiment reduces hardware resource consumption by 8 bits, resulting in an overall resource saving of 16.3%.

[0046] Table 6: BTB parameter allocation in this embodiment compared to the traditional four-way interconnect structure

[0047] The benchmark used was CoreMark, a widely used embedded processor performance testing program written in C. It was compiled into a binary file using the GCC 10.2.0 compiler with the -O3 optimization option and 1000 loop iterations. To load the test program into the processor's instruction memory, a Python script converted the binary file to a hexadecimal file, and finally, Verilog was used for the final output. The function is written to the instruction ROM of the five-stage pipeline. The simulator uses Modelsim to observe the behavior of the branch target predictor during actual operation and records the prediction results. For example... Figure 15 The simulation results for the branch target predictor are shown. It can be seen that at instruction address 000009c4, the instruction type `btb_inst_type` is binary 01, representing a type B branch instruction. Comparing the PC address and Tag bit shows that the instruction was predicted, with the predicted address `predicted_target` at 000009b4. In the second cycle, the instruction address changes to 000009b4, indicating successful prediction. The figure shows `branch_cnt` recording 4896 branch instructions executed during program execution, and `fault_cnt` recording 372 branch prediction failures. The formula for calculating the prediction accuracy of the branch target predictor is as follows:

[0048] The above testing methods verified the correctness of the branch target predictor function in this embodiment. Compared with the traditional four-way group-associative BTB structure, it saves 8 bits of storage space per entry and still achieves a prediction accuracy of 92.4% under the CoreMark benchmark, demonstrating high prediction capability.

[0049] The present invention also provides an application of the hybrid branch predictor as described above, applied to a microprocessor.

[0050] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A hybrid branch predictor comprising: The branch direction predictor and the branch target predictor are included; the branch direction predictor is used to predict whether a branch instruction jumps; the branch target predictor is used to predict a specific jump target address when the branch direction is predicted to jump.

2. The hybrid branch predictor of claim 1, wherein, The branch direction predictor includes a global history prediction module, a local history prediction module and a selection module; the global history prediction module is used to predict branch instructions depending on global branch history; the local history prediction module is used to provide supplementary prediction for branch instructions depending on their own local history; and the selection module is used to dynamically select the prediction result of the global history prediction module or the local history prediction module as the final direction.

3. The hybrid branch predictor of claim 2, wherein, The global history prediction module includes a basic predictor, a first tag prediction table counter, a second tag prediction table counter, a third tag prediction table counter and a fourth tag prediction table counter.

4. The hybrid branch predictor of claim 3, wherein, The basic predictor is composed of a saturation counter, and the tag prediction table counter is composed of a tag for table entry matching, a saturation counter and a usefulness counter representing the usefulness degree of the table entry.

5. The hybrid branch predictor of claim 2, wherein, The global history prediction module is specifically configured to perform hash operation on data of the global history and the program counter with different lengths to generate indexes and tags of each tag table; the prediction result of the tag table is valid only when the tags match, and the long history matching result is preferentially used; if all the tag tables miss, the prediction result of the basic predictor is used by default; and the prediction confidence is adjusted according to the actual jump result, it is judged whether the table entry is useful, a new table entry is allocated and written, and the current jump direction is written into the global history register, so as to ensure that the prediction adapts to the change of branch behavior.

6. The hybrid branch predictor of claim 2, wherein, The local history prediction module includes a local branch history table register and a local mode history table counter; the local branch history table register is a branch history register, used to store the local history of each branch instruction itself and record the past jump regularity of the branch instruction itself; and the local mode history table counter is composed of two saturation counters, used to store the jump mode of the branch.

7. The hybrid branch predictor of claim 2, wherein, The local history prediction module is specifically configured to index the local branch history table register by using the low 10 bits of the program counter, read out the branch history register of the branch, splice the branch history register and the 7-10 bits of the program counter into a 10-bit index, access the local mode history table counter, read out the counter value as the prediction result, write the actual jump direction into the branch history register of the local branch history table register after the branch is executed, update the low bits and discard the high bits, and adjust the corresponding counter of the local mode history table counter.

8. The hybrid branch predictor of claim 2, wherein, The selection module is composed of three saturation counters, used to dynamically select the prediction result of the global history prediction module or the local history prediction module as the final prediction result according to the branch history bias condition by using a multiplexer.

9. The hybrid branch predictor of claim 2, wherein, The branch target predictor comprises a branch target buffer of four-way set-associative structure and a return address stack with a counter; the branch target buffer is used to provide address prediction for branch instructions with fixed target addresses (such as direct jumps, conditional jumps) and store jump target addresses; the return address stack is used to predict target addresses of indirect jump instructions.

10. Use of a hybrid branch predictor as claimed in any of the claims 1-9, characterized in that, Applied to a microprocessor.

Citation Information

Cited By

  • RISC-V branch predictor closed loop verification method and system

    CN122018994A

  • Risc-v branch predictor closed loop verification method and system

    CN122018994B