Novel BTB structure for dual-branch prediction
By improving the parallel prediction of BTB structure and branch predictor, the problem that existing processors cannot predict multiple branches simultaneously in a single cycle is solved, and dual-branch parallel prediction is realized, which improves the processor's finger fetch bandwidth and throughput rate.
Patent Information
- Application Number
- CN202510776813.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing processors can only predict a single branch within a single cycle or two branches under limited conditions, which cannot meet the wide transmit processor's demand for finger fetch bandwidth, resulting in pipeline pauses and reduced throughput.
A new type of dual-branch prediction BTB structure is designed, using a multi-group-associated structure and multi-target address storage, and the jump results of two consecutive branches are predicted in parallel through a branch predictor, and multiple target addresses are output in the same period. The dual FetchBlock request is generated in combination with the multiple selector to meet the finger fetch bandwidth requirements of the wide transmit processor.
Parallel prediction of two branches in a single cycle is realized, the bandwidth utilization of the front-end finger fetch unit and the overall processor performance are improved, and pipeline delay and instruction starvation are reduced.
Smart Images

Figure CN120295672A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of BTB structures, and in particular to a novel BTB structure for dual-branch prediction. Background Art
[0002] BTB: Branch Target Buffer A cache used to store predicted branch instruction target addresses.
[0003] With the continuous evolution of processor microstructures, the CPU's launch width and pipeline depth continue to increase, placing higher requirements on the front-end instruction fetch bandwidth. In order to reduce pipeline stalls caused by branch jumps, the branch predictor usually needs to give the predicted position of the next instruction (PC) in a very short time. Most traditional branch prediction units only support one branch prediction in a single cycle, or can only predict two branches under very limited conditions. When the CPU needs to process multiple branches in the same cycle, if the existing technology cannot accurately and quickly give the target address after continuous branch jumps, it will reduce the efficiency of front-end instruction fetching, thereby affecting the overall throughput.
[0004] Therefore, how to design and optimize a front-end branch prediction structure that can handle two (or even more) branch jumps in a single cycle has become one of the key technologies to improve the performance of high-end processors. In response to this problem, the present invention proposes a new hardware structure and method based on BTB, which realizes parallel prediction of two conditional branches at the same time by improving the BTB data storage and multi-target selection method, so as to meet the demand of wide-issue processors for instruction fetch bandwidth.
[0005] In terms of front-end branch prediction technology, processors (CPUs) currently on the market can usually only predict one branch in one cycle, or can only predict two branches but with limitations in usage scenarios. As the processor's issue width and instruction fetch bandwidth continue to increase, the need to predict multiple branches in a single cycle has become increasingly prominent. However, conventional branch predictors and the corresponding branch target buffer (BTB) structures are difficult to unconditionally support the generation of predicted targets for two jump instructions in one cycle, resulting in the inability to fully utilize the bandwidth of the front-end instruction fetch unit, affecting the overall performance of the processor.
[0006] Defects of existing BTB technology: 1. Only single branches can be predicted or the prediction of multiple branches is limited: Most processors only provide one jump prediction in one cycle. If two consecutive conditional branches are encountered, additional cycles or logic are required for serial processing.
[0007] 2. It is impossible to truly implement dual-branch parallel prediction: Even if some architectures claim to support two branch predictions, there are many usage scenario restrictions in the specific implementation, and it is difficult to adapt to the more common two-hop consecutive branches.
[0008] 3. Front-end bandwidth waste: As the CPU issue width increases, if the front-end cannot provide enough instruction streams to the back-end at one time, it will cause pipeline stalls and reduce the processor throughput.
[0009] To address the above problems, the present invention provides a novel BTB design that can predict two consecutive conditional branches in the same cycle and output the corresponding target addresses, thereby significantly improving the bandwidth utilization of the front-end instruction fetch unit and the overall performance of the processor.
[0010] The present invention proposes a novel BTB structure with dual-branch prediction to solve this problem. Summary of the Invention
[0011] The object of the present invention is to unconditionally predict two jump conditional branches in one cycle and provide the corresponding target addresses. By improving the BTB data structure, it can output multiple potential target addresses in the same query, and combine with the branch predictor to simultaneously predict whether the two branches will jump, so as to quickly determine the next instruction fetch address, and realize the parallel prediction of "up to two conditional branches" in the hardware structure to fully match the instruction fetch bandwidth of the wide-issue processor and reduce the instruction starvation caused by insufficient branch prediction, thereby overcoming the problems in the above background technology.
[0012] Based on the above technical ideas, the technical solution adopted by the present invention is as follows: A novel BTB structure with dual-branch prediction, comprising the following steps: S1 Improve the BTB structure design step, which includes a multi-way set-associative structure link and a multi-target address storage link, and also includes branch information association. Each branch instruction associates the corresponding Tgt and tag fields according to its position in the instruction block; S2 BTB query and data reading step, which includes an input PC processing link, a Tag matching and hit judgment link, and a data reading link; S3 Parallel branch prediction step, which includes a branch predictor cooperation link and a prediction result output link; S4 Target address selection logic step, which includes a multiplexer configuration link and a dual FetchBlock generation link, and also includes bandwidth matching. By parallel fetching two Cachelines, it meets the demand of the wide-issue processor for high instruction fetch bandwidth; S5 Dynamic update and replacement policy step, which includes a BTB entry update link and a replacement policy implementation link; S6 Exception handling and recovery step, which includes a prediction error recovery link and a miss handling link.
[0013] For further limitation of the above technical solution, in the step of improving the BTB structure design in S1, the multi-way set-associative structure link in this step includes designing the BTB as a multi-way set-associative structure similar to Cache. Each group contains multiple channels. The high-order part of the program counter is used as the index value to determine the BTB group to be queried. Each BTB entry stores the high-order part of the PC as the Tag for matching the input PC during query.
[0014] For further limitation of the above technical solution, in the step of improving the BTB structure design in S1, the multi-target address storage link in this step includes that each BTB entry contains the following fields: Tgt0, Tgt1, Tgt2, which store the predicted target addresses of different branch combinations respectively; Tag0, Tag1, Tag2, which identify the branch types; LRU field, which records the recent usage information and supports the replacement algorithm.
[0015] For further limitation of the above technical solution, in the step of BTB query and data reading in S2, the input PC processing link in this step includes that the fetch unit inputs the current PC value into the BTB, extracts the high-order part of the PC as the index, and locates the target group in the BTB.
[0016] For further limitation of the above technical solution, in the step of BTB query and data reading in S2, the Tag matching and hit determination link in this step includes parallelly matching the Tag fields of all channels in the group with the high-order part of the input PC. If there is an entry with a Tag match, it is determined that the BTB hits; otherwise, it misses.
[0017] For further limitation of the above technical solution, in the step of BTB query and data reading in S2, the data reading link in this step includes that when hitting, read Tgt0, Tgt1, Tgt2 and the corresponding tag information from the matching BTB entry. When missing, use the default sequential address as the next fetch address.
[0018] For further limitation of the above technical solution, in the step of parallel branch prediction in S3, the collaborative work link of the branch predictor in this step includes that the branch predictor and the BTB work in the same clock cycle, independently predicting the jump results of two consecutive branches. Br0 corresponds to the first conditional branch in the instruction block, and Br1 corresponds to the second conditional branch. The prediction result is a binary combination, indicating the jump states of the two branches; the prediction result output link includes that the branch predictor outputs the prediction results of Br0 and Br1 to the target address selection module.
[0019] For further limitation of the above technical solution, in the step of target address selection logic in S4, the multiplexer configuration link in this step includes selecting the corresponding target address according to the combination of Br0 and Br1: No jump, select sequential address; Select Tgt0 as the next instruction fetch address; Select Tgt1 as NextPC; Select both Tgt1 and Tgt2 simultaneously to generate a dual FetchBlock request.
[0020] For further limitation of the above technical solution, in the S4 target address selection logic step, the dual FetchBlock generation link in this step includes that when both Br0 and Br1 are jumps (1, 1), the instruction fetch unit sends two requests to the instruction cache simultaneously: Request 1: Instruction block with Tgt1 as the target address; Request 2: Instruction block with Tgt2 as the target address.
[0021] For further limitation of the above technical solution, in the S5 dynamic update and replacement strategy step, the BTB entry update link in this step includes that when the actual execution result of the branch is inconsistent with the prediction, update the BTB entry: Add or replace entry: Select the channel to be replaced within the group according to the LRU policy; Write the new Tag, Tgt, and tag information to overwrite the old data.
[0022] Compared with the prior art, the beneficial effects of the present invention are: 1. Up to two jump target addresses can be generated within the same cycle, improving the branch prediction throughput rate.
[0023] 2. Adopt a multi-way BTB structure similar to Cache, with strong scalability, and it is easy to adjust the indexing method and the number of ways according to the processor requirements.
[0024] 3. It can effectively reduce the latency caused by traditional serial branch prediction, and improve the prediction accuracy and performance in high-concurrency branch scenarios. Brief Description of the Drawings
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0026] Figure 1 It is a flow schematic diagram of a novel dual-branch prediction BTB structure of the present invention. Detailed Embodiments
[0027] The following is combined with the attached Figure 1A further detailed description of the present invention is provided.
[0028] Embodiment 1: This embodiment provides a novel BTB structure with dual-branch prediction. As Figure 1 shown, it includes the following steps: S1 Step of improving the BTB structure design, which includes a multi-way set-associative structure link and a multi-target address storage link. This step also includes branch information association, where each branch instruction associates corresponding Tgt and tag fields according to its position in the instruction block. S2 BTB query and data reading step, which includes an input PC processing link, a Tag matching and hit judgment link, and a data reading link. S3 Parallel branch prediction step, which includes a branch predictor collaborative work link and a prediction result output link. S4 Target address selection logic step, which includes a multiplexer configuration link and a dual FetchBlock generation link. This step also includes bandwidth matching, where by fetching two Cachelines in parallel, the high fetch bandwidth requirement of the wide-issue processor is met. S5 Dynamic update and replacement strategy step, which includes a BTB entry update link and a replacement strategy implementation link. S6 Exception handling and recovery step, which includes a prediction error recovery link and a miss handling link.
[0029] S1 Step of improving the BTB structure design. In this step, the multi-way set-associative structure link includes designing the BTB as a multi-way set-associative structure similar to a Cache. Each set contains multiple channels, and the high-order bits of the program counter are used as the index value to determine the BTB set to be queried. Each BTB entry stores the high-order part of the PC as the Tag for matching the input PC during query.
[0030] S1 Step of improving the BTB structure design. In this step, the multi-target address storage link includes that each BTB entry contains the following fields: Tgt0, Tgt1, Tgt2, which store the predicted target addresses of different branch combinations respectively. Tag0, Tag1, Tag2, which identify the branch types. LRU field, which records the recent usage information to support the replacement algorithm.
[0031] S2 BTB query and data reading step. In this step, the input PC processing link includes that the fetch unit inputs the current PC value into the BTB, extracts the high-order bits of the PC as the index, and locates the target set in the BTB.
[0032] S2 BTB Query and Data Reading Step. In this step, the Tag matching and hit judgment section includes parallelly matching the Tag fields of all channels within the group with the high part of the input PC. If there is an entry with a Tag match, it is determined as a BTB hit; otherwise, it is a miss.
[0033] S2 BTB Query and Data Reading Step. In this step, the data reading section includes, when there is a hit, reading Tgt0, Tgt1, Tgt2 and the corresponding tag information from the matching BTB entry; when there is a miss, using the default sequential address (PC + instruction block length) as the next fetch address.
[0034] S3 Parallel Branch Prediction Step. In this step, the collaborative working section of the branch predictor includes the branch predictor and the BTB working in the same clock cycle, independently predicting the jump results of two consecutive branches. Br0 corresponds to the first conditional branch in the instruction block, and Br1 corresponds to the second conditional branch. The prediction result is a binary combination representing the jump states of the two branches; the prediction result output section includes the branch predictor outputting the prediction results of Br0 and Br1 to the target address selection module.
[0035] S4 Target Address Selection Logic Step. In this step, the multiplexer configuration section includes selecting the corresponding target address according to the combination of Br0 and Br1: No jump, select the sequential address; Select Tgt0 as the next fetch address; Select Tgt1 as NextPC; Select both Tgt1 and Tgt2 simultaneously to generate a dual FetchBlock request.
[0036] S4 Target Address Selection Logic Step. In this step, the dual FetchBlock generation section includes when both Br0 and Br1 are jumps (1, 1), the fetch unit sending two requests to the instruction cache simultaneously: Request 1: The instruction block with Tgt1 as the target address; Request 2: The instruction block with Tgt2 as the target address.
[0037] S5 Dynamic Update and Replacement Policy Step. In this step, the BTB entry update section includes when the actual execution result of the branch is inconsistent with the prediction, updating the BTB entry: Add or replace an entry: Select the channel within the group to be replaced according to the LRU policy; Write the new Tag, Tgt and tag information to overwrite the old data.
[0038] The S5 dynamic update and replacement strategy step, in which the replacement strategy implementation link includes updating the LRU counter each time the BTB is accessed to mark the recent usage. When replacement is needed, the entry with the largest LRU value in the group is selected for elimination.
[0039] The S6 exception handling and recovery step, in which the prediction error recovery includes that if the actual branch result does not match the prediction, the pipeline is flushed and instructions are fetched again from the correct address, and the BTB corrects the entry content according to the actual branch result; the miss handling link includes that when the BTB misses, instructions are fetched sequentially and branch instructions are monitored. After a new branch instruction is detected, its information is dynamically inserted into the BTB.
[0040] This method realizes zero additional latency through single-cycle dual-branch prediction and hardware parallelization; the dual FetchBlock mechanism makes full use of the ICache bandwidth to avoid pipeline starvation; the number of ways and index bit width of the BTB can be adjusted to adapt to different processor scales.
[0041] Embodiment 2: This embodiment provides a new BTB structure for dual-branch prediction, as Figure 1 shown, and also includes the structural improvement of the BTB: Cache-like organization method: The present invention designs the BTB as a structure similar to a cache (Cache), uses a part of the high bits of the PC (such as bit5 and above) as the index, and stores multiple branch information in each entry. Multiple channels (ways), such as 4-way, are set under each index for high-parallel storage and query.
[0042] Multi-target address storage: In each BTB entry, the storage includes: Tag: Used to match the input PC (or its high-bit part).
[0043] Src0, Src1, Src2: Record the source information corresponding to this branch, or additional information required for a specific branch type.
[0044] Tgt0, Tgt1, Tgt2: Correspond to the target addresses under different branch prediction combinations respectively.
[0045] Label 0, Label 1, Label 2: Can be used to distinguish the type of branch (conditional / unconditional / call / return, etc.) or other control information.
[0046] LRU: Record the information required for the replacement strategy (or other replacement algorithms).
[0047] Two-branch prediction adaptation: To adapt to the prediction of two consecutive conditional branches, three target addresses (tgt0, tgt1, tgt2) can be obtained after a single BTB lookup. The branch predictor simultaneously gives the taken / not-taken combinations of two branches (Br0 and Br1), and the next fetch address (Next PC) is selected from tgt0, tgt1, tgt2 through a simple multiplexer (MUX).
[0048] Branch prediction process: BTB query: The PC generated by the fetch unit is used to search the BTB through index indexing and tag comparison. If there is a hit, tgt0, tgt1, tgt2 are read. If there is a miss, the default sequential fetch address is used.
[0049] Branch predictor output: In the same cycle, the branch predictor gives the prediction results of Br0 and Br1 in parallel. For example, (Br0 = 1, Br1 = 0) means that the first branch jumps and the second branch does not jump.
[0050] Target address selection: (Br0 = 0, Br1 = 0) → It means no jump, continue with the next sequential address (Br0 = 0, Br1 = 1) → Select tgt0 (Br0 = 1, Br1 = 0) → Select tgt1 (Br0 = 1, Br1 = 1) → Select tgt1 and tgt2 Dual Fetch Block construction: Since two branches can be predicted in the same cycle, the present invention can simultaneously issue two possible fetch requests to meet the requirement of a wide-issue processor to fetch multiple Cachelines at a time, improving the front-end throughput.
[0051] The improvement scheme of this embodiment supports parallel prediction of two branches in a single cycle: By storing multi-target branch paths in the BTB and cooperating with the dual-branch prediction results, the next one or two fetch addresses are given at one time; The multi-way parallel query structure of the BTB: Using a multi-way organization similar to the Cache to store different branch paths; High fetch bandwidth matching for wide-issue CPUs: It can generate dual fetch requests in hardware, thus truly exerting the advantages of a multi-issue pipeline.
[0052] It can truly achieve dual-branch parallel prediction: for consecutive conditional branches, two jump target addresses can be output without additional logic or latency; the fetch bandwidth is significantly improved: the number of instructions fetched by the front end at one time is greatly increased, matching or exceeding the requirements of the wide-issue pipeline for instruction throughput; strong scalability: a BTB similar to the Cache structure is adopted, and the number of ways and the index bit width can be flexibly configured according to the processor scale; the hardware cost is controllable: it is extended on the basis of retaining the basic functions of the original BTB, and the impact on chip area and power consumption is relatively acceptable.
[0053] Embodiment 3: This embodiment provides a new BTB structure for dual-branch prediction, as Figure 1 shown, and an improvement scheme is also included for Embodiment 2: In terms of implementation, the quantity or logic of tgt0, tgt1, and tgt2 can be extended or streamlined. For example, it can be designed to only predict one conditional branch and one unconditional jump, or more target addresses can be added to support more complex branch scenarios.
[0054] By expanding the ways, the recorded branch paths are increased.
[0055] The improvement of the index algorithm can adopt a more complex hash algorithm to calculate the index.
[0056] If only single-branch jump prediction is required, this solution can be degraded for use, still having the advantages of high hit rate and flexible update mechanism.
[0057] The above content is a further detailed description of the present invention in combination with specific preferred implementation schemes, which is convenient for those skilled in the art of this technology to understand and apply the present invention. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions.
Claims
1. A novel BTB structure with dual-branch prediction, characterized in that, It includes the following steps: S1. Step of improving the BTB structure design, which includes a multi-way set-associative structure link and a multi-target address storage link. This step also includes branch information association. Each branch instruction associates the corresponding Tgt and tag fields according to its position in the instruction block; S2. BTB query and data reading step, which includes an input PC processing link, a Tag matching and hit judgment link, and a data reading link; S3. Parallel branch prediction step, which includes a branch predictor collaborative work link and a prediction result output link; S4. Target address selection logic step, which includes a multiplexer configuration link and a dual FetchBlock generation link. This step also includes bandwidth matching. By fetching two Cachelines in parallel, the requirement of the wide-issue processor for high fetch bandwidth is met; S5. Dynamic update and replacement policy step, which includes a BTB entry update link and a replacement policy implementation link; S6. Exception handling and recovery step, which includes a prediction error recovery link and a miss handling link.
2. A novel dual-branch prediction BTB structure according to claim 1, characterized in that, In the S1 step of improving the BTB structure design, the multi-way set-associative structure link in this step includes designing the BTB as a similar multi-way set-associative structure. Each group contains multiple channels. The high-order part of the program counter is used as the index value to determine the BTB group to be queried. Each BTB entry stores the high-order part of the PC as the Tag for matching the input PC during query.
3. The novel BTB structure with dual-branch prediction according to claim 2, characterized in that, In the S1 step of improving the BTB structure design, the multi-target address storage link in this step includes that each BTB entry contains the following fields: Tgt0, Tgt1, Tgt2, which store the predicted target addresses of different branch combinations respectively; Tag0, Tag1, Tag2, which identify the branch types; LRU field, which records the recent usage information and supports the replacement algorithm.
4. A novel BTB structure with dual-branch prediction according to claim 3, characterized in that, In the S2 BTB query and data reading step, the input PC processing link in this step includes that the fetch unit inputs the current PC value into the BTB, extracts the high-order part of the PC as the index, and locates the target group in the BTB.
5. A novel BTB structure with dual-branch prediction according to claim 4, characterized in that, In the S2 BTB query and data reading step, the Tag matching and hit judgment link in this step includes parallelly matching the Tag fields of all channels in the group with the high-order part of the input PC. If there is an entry with a Tag match, it is determined that the BTB hits; otherwise, it misses.
6. A novel BTB structure with dual-branch prediction according to claim 5, characterized in that, In the S2 BTB query and data reading step, the data reading link in this step includes that when hitting, read Tgt0, Tgt1, Tgt2 and the corresponding tag information from the matching BTB entry. When missing, use the default sequential address as the next fetch address.
7. A novel BTB structure with dual-branch prediction according to claim 6, characterized in that, In the S3 parallel branch prediction step, the branch predictor collaborative work link in this step includes that the branch predictor and the BTB work in the same clock cycle, independently predict the jump results of two consecutive branches. Br0 corresponds to the first conditional branch in the instruction block, and Br1 corresponds to the second conditional branch. The prediction result is a binary combination, indicating the jump states of the two branches. The prediction result output link includes that the branch predictor outputs the prediction results of Br0 and Br1 to the target address selection module.
8. A novel BTB structure with dual-branch prediction according to claim 7, characterized in that, The S4 target address selection logic step, in which the multiplexer configuration section includes selecting the corresponding target address according to the combination of Br0 and Br1: No jump, select the sequential address; Select Tgt0 as the next instruction fetch address; Select Tgt1 as NextPC; Select both Tgt1 and Tgt2 simultaneously to generate a dual FetchBlock request.
9. A novel BTB structure with dual-branch prediction according to claim 8, characterized in that, The S4 target address selection logic step, in which the dual FetchBlock generation section includes that when both Br0 and Br1 are jump address 1, the instruction fetch unit issues two requests to the instruction cache simultaneously: Request 1: An instruction block with Tgt1 as the target address; Request 2: An instruction block with Tgt2 as the target address.
10. A novel dual-branch prediction BTB structure according to claim 9, characterized in that, The S5 dynamic update and replacement policy step, in which the BTB entry update section includes updating the BTB entry when the actual branch execution result is inconsistent with the prediction: Add or replace an entry: Select the channel to be replaced within the group according to the LRU policy; Write the new Tag, Tgt and tag information to overwrite the old data.
Citation Information
Patent Citations
Prediction method and device for supporting simultaneous prediction of two unconditional branch instructions of continuous jump
CN116661872A
Multi-branch predictor and prediction method based on RISC-V architecture processor
CN118152012A
Branch target buffer compression
US20180060073A1
Dual branch format
US20220129278A1
Method and circuit for single cycle multiple branch history table access
US6347369B1