Branch target address optimization storage structure based on target address pool

By only storing indexes to Target PCs in the BRB of the CPU, using the Target Address Pool (TPP) solution, the problems of large storage overhead and long life cycle of BRB are solved, and storage resources are saved and energy efficiency is improved.

CN120407025AActive Publication Date: 2025-08-01BEIJING YIHUA CLOUD NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510907572.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

In modern CPU design, the branch record buffer module (BRB) stores a large number of Target PCs, resulting in increased chip area and power consumption, and a long life cycle, making it difficult to optimize the relationship between processor performance, power consumption and area (PPA).

Method used

Using the Target PC Pool (TPP) scheme, only indexes pointing to TP are stored in BRB, rather than a complete Target PC, reducing storage resource requirements by allocating and releasing Target PC information on demand.

Benefits of technology

Significantly saves storage area and power consumption, improves overall CPU energy efficiency, is compatible with existing branch prediction processes, adapts to multiple processor address widths, and realizes efficient Target PC management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407025A_ABST
    Figure CN120407025A_ABST
Patent Text Reader

Abstract

The invention provides a branch target address optimization storage structure based on a target address pool, and belongs to the technical field of address optimization storage structures, the branch target address optimization storage structure comprises the target address pool and a working process thereof, and further comprises branch record buffer transformation, and the working process of the target address pool comprises the following steps: S1, an allocation stage step, S2, an operation stage step, the step comprises a target address access link and a processing assembly line refreshing link; s3, a branch analysis stage step, wherein the step comprises an execution result receiving link, a prediction correct processing link and a prediction error processing link; and S4, updating and releasing a branch predictor, which comprises a correct prediction link, an error prediction link and a branch record buffer entry release link. Aiming at the problems of high storage overhead and long life cycle of a traditional address storage scheme, the method comprises the following steps of: storing an index of an address in a target address pool through a branch record buffer; therefore, the effect of reducing the storage space is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of address-optimized storage structures, and particularly to a branch target address-optimized storage structure based on a target address pool. Background Art

[0002] With the evolution of integrated circuit technology and the rapid growth of computing demands, modern processors (CPUs) are evolving towards higher performance and stronger parallelism. Whether in the server field or in embedded devices, the main frequency of the processor, the instruction issue width, and the instruction-level parallelism (ILP) are all continuously improving. To adapt to diverse application scenarios (such as artificial intelligence, big data, cloud computing, etc.), the requirements for the front-end instruction fetch bandwidth and branch prediction accuracy of the CPU are also getting higher and higher. At the same time, the constraints on chip power consumption and area (PPA) are becoming increasingly strict. How to minimize power consumption and reduce chip area while maintaining high performance has become a core challenge in contemporary processor microarchitecture design.

[0003] In the front-end unit (IFU) of a processor, in order to effectively track and predict updates of branch instructions, a large amount of branch information in the execution process often needs to be stored in a branch record buffer (BRB) or a similar module. Among them, the "target address (Target PC)" usually occupies a relatively large bit width (for example, 48 bits or 64 bits). Since modern CPUs often need to configure dozens to hundreds of entries in the BRB, storing only the Target PC will result in a considerable consumption of chip area and register file resources. Especially in the scenario of high-width-issue CPUs, the number of in-flight branches may be even more, and the storage overhead is further amplified. How to significantly compress the Target PC storage resources, reduce power consumption, and improve overall efficiency while ensuring high branch prediction performance and correctness is one of the difficult problems that need to be solved in the current CPU front-end design.

[0004] In modern CPU design, in order to improve processor performance, it is often necessary to set up more buffers in the front-end instruction fetch unit to record various information of in-flight branches. Taking the branch record buffer (BRB) as an example, in order to completely cover the entire life cycle of branch instructions from instruction fetch to retire, the BRB often needs a large number of entries (for example, 128), and the complete Target PC of the branch is recorded in each entry. If the PC width is large (for example, RV48 supports 48-bit addresses), then storing only these Target PCs requires 128 * 48 bits = 0. = 0.75KB.

[0005] However, there is a strong mutual restraint relationship among the performance, power consumption (Power), and area (Area) of the CPU (PPA). On the premise of ensuring high performance, it is particularly important to optimize power consumption and area to the extreme. If all the complete Target PCs of the in-flight branches are directly stored in the BRB, it will bring a large overhead. Therefore, the present invention proposes a Target PC Pool (TP) scheme, which significantly compresses the storage requirement of the BRB for the target address by only storing the index "pointing" to the TP in the BRB, and ensures the correctness and efficiency of branch prediction and update at the system level.

[0006] Disadvantages of the prior art: 1. Large storage overhead: When each BRB Entry stores the complete Target PC, the overall storage scale increases exponentially, resulting in an increase in chip area and power consumption.

[0007] 2. Long life cycle: The entire process from when a branch instruction is recognized to when it retires may be long, and the Target PC remains in the BRB during this period, failing to fully utilize the cache space.

[0008] 3. Difficult PPA optimization: The IFU itself is already a high area / power consumption hotspot, and with the addition of large-capacity Target PC storage, the overall energy efficiency ratio of the processor is further reduced.

[0009] The present invention proposes an optimized storage structure for branch target addresses based on a target address pool to solve the above problems. Summary of the Invention

[0010] The present invention designs a target address pool Target PC Pool (TPP) structure. When a branch instruction is written into the BRB, the complete Target PC information is placed in the TPP. Only a Pool Index needs to be recorded in the BRB. After branch parsing, if the prediction result is correct, the TPP can release the Target PC information of this entry earlier, without having to wait until all branch prediction unit components are updated before releasing. This scheme can design the TP capacity (the number of entries) to be smaller (for example, only 16 entries) to meet the life cycle requirements from branch allocation entry to parsing, thus significantly saving the storage resources that might otherwise need to be maintained in the BRB, thereby overcoming the problems in the above background technology.

[0011] Based on the above technical ideas, the technical solution adopted by the present invention is: An optimized storage structure for branch target addresses based on a target address pool, including a target address pool and its working process; The target address pool includes storing a complete 48-bit target address and 1-bit status bit, and its capacity is much smaller than the branch record buffer; It also includes the transformation of the branch record buffer, where the transformation of the branch record buffer includes storing only the index pointing to the TP instead of the complete target address; The working process of the target address pool includes the following steps: S1 Allocation stage step, which includes the branch instruction fetching link, the step of writing to the target address pool, the index generation link, and the branch record buffer update link; S2 Running stage step, which includes the target address access link and the processing pipeline flush link; S3 Branch resolution stage step, which includes the execution result receiving link, the correct prediction processing link, and the incorrect prediction processing link; S4 Branch predictor update and release step, which includes the correct prediction link, the incorrect prediction link, and the branch record buffer entry release link.

[0012] For further limitation of the above technical solution, in the S1 allocation stage step, in the branch instruction fetching link of this step, the branch instruction enters the front-end fetch unit, is recognized and allocated a branch record buffer entry. The step of writing to the target address pool includes writing the complete predicted target address of this branch to the idle entry of the target address pool and changing the status bit of this target address entry to 1.

[0013] For further limitation of the above technical solution, the index generation link includes the target address pool returning the index allocated to this entry. The branch record buffer update link includes storing only the index instead of the complete target address in the corresponding entry of the branch record buffer and recording other branch information.

[0014] For further limitation of the above technical solution, in the S2 running stage step, in the target address access link of this step, when the pipeline needs to read the branch target address, the branch record buffer looks up the target address pool through the index to obtain the complete address.

[0015] For further limitation of the above technical solution, the processing pipeline flush link includes if a pipeline flush occurs, changing the status bits of the entries of the branches to be flushed and subsequent branches in the target address pool to 0. The corresponding index in the branch record buffer is retained but no longer used.

[0016] For further limitation of the above technical solution, in the S3 branch resolution stage step, in the execution result receiving link of this step, the execution unit completes the branch calculation and returns the actual target address and prediction correctness to the front end; The correct prediction processing link includes locating the entry in the target address pool through the index in the branch record buffer, immediately releasing the TP entry, changing the status bit of the entry in the target address pool to 0, without waiting for subsequent operations, and retaining the branch record buffer entry until the branch retires.

[0017] For further limitation of the above technical solution, the prediction error handling process includes locating an entry in the target address pool through indexing, overwriting the error address, writing the correct target address calculated by the execution unit to the original entry position in the target address pool, keeping the status bit as 1, and the entry remains valid until the branch predictor is updated subsequently and then released.

[0018] For further limitation of the above technical solution, in step S4 of branch predictor update and release, the correct prediction process in this step includes that there is no need to update the branch predictor, the entry in the target address pool has been released in the branch resolution stage, and the status bit has been changed to 0. The incorrect prediction process includes updating the branch predictor. The branch record buffer entry reads the correct target address in the target address pool through indexing, updates the branch predictor in combination with the branch information, releases the entry in the target address pool, and after the branch predictor update is completed, changes the status bit of the entry in the target address pool to 0.

[0019] For further limitation of the above technical solution, the branch record buffer entry release process includes that after the branch instruction retires, the branch record buffer entry is released regardless of whether the prediction is correct or not.

[0020] For further limitation of the above technical solution, the address widths adapted by the target address pool include 16 / 32 / 48 / 64 bits. There is an address width bit in the entry of the target address pool. The address width bit is composed of two boolean characters. The content stored through the address width bit is 0 / 1 / 2 / 3 corresponding to 16 / 32 / 48 / 64-bit address widths, and the address content of different lengths is read subsequently by identifying the address width bit.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Separate storage concept: The branch instruction only records the index pointing to the Target PC Pool in the BRB, rather than saving the complete Target PC.

[0022] 2. On-demand allocation and early release of Target PC Pool (TPP): When the branch is resolved, if the branch prediction is correct, the entry can be released immediately without waiting for all subsequent update operations to end.

[0023] 3. Flexible and scalable: The size of the TPP can be set according to statistical analysis, which is usually much smaller than the number of entries in the BRB, maximizing the saving of area and power consumption.

[0024] 4. Compatibility with the existing branch prediction process: There is no need for fundamental changes to branch resolution, BPU update, pipeline flush, etc., and only index operations need to be added during allocation and release.

[0025] 5. Greatly save storage area: By splitting the storage cycle and space requirements of the Target PC, the TPP only needs to sufficiently cover the time window from branch allocation to parsing, and does not need to redundantly cover the retirement period of the branch.

[0026] 6. Higher energy efficiency: The read / write access volume to the Target PC in the IFU is reduced, the power consumption of the memory is improved, and the overall PPA of the CPU is enhanced.

[0027] 7. Little impact on other front-end structures: It not only ensures the normal logic of branch prediction but also does not damage fields such as branch identifiers and directions originally recorded in the BRB.

[0028] 8. Can be flexibly adapted to a variety of processors: Regardless of whether the address width is 32 bits, 48 bits, or 64 bits, similar savings effects can be achieved by adjusting the TP size and index bit width. Brief Description of the Drawings

[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0030] Figure 1 It is a schematic diagram of an optimized storage structure for branch target addresses based on a target address pool of the present invention. Detailed Embodiments

[0031] The following will further describe the present invention in detail Figure 1 in conjunction with the attached

[0032] Example 1: This example provides an optimized storage structure for branch target addresses based on a target address pool, as Figure 1 shown, including a target address pool and its working process; The target address pool includes storing a complete 48-bit target address and 1-bit status bit, and its capacity is much smaller than the branch record buffer; It also includes the transformation of the branch record buffer, which only stores the index pointing to the TP instead of the complete target address; The working process of the target address pool includes the following steps: Step S1 in the allocation stage, which includes the branch instruction fetching link, the writing to the target address pool link, the index generation link, and the branch record buffer update link; Step S2 in the running stage, which includes the target address access link and the processing pipeline refresh link; The steps of the S3 branch parsing stage, which include a received execution result link, a correct prediction processing link, and an incorrect prediction processing link; The steps of the S4 branch predictor update and release, which include a correct prediction link, an incorrect prediction link, and a branch record buffer entry release link.

[0033] In the steps of the S1 allocation stage, in the branch instruction fetch link of this step, the branch instruction enters the front-end fetch unit, is recognized and allocated a branch record buffer entry. In the write target address pool link, the complete predicted target address of this branch is written into the idle entry of the target address pool, and the status bit of this target address entry is changed to 1.

[0034] The generating index link includes the target address pool returning the index allocated to this entry. The updating branch record buffer link includes storing only the index, rather than the complete target address, in the corresponding entry of the branch record buffer, and recording other branch information.

[0035] In the steps of the S2 running stage, in the target address access link of this step, when the pipeline needs to read the branch target address, the branch record buffer looks up the target address pool through the index to obtain the complete address.

[0036] The processing pipeline flush link includes, if a pipeline flush occurs, changing the status bits of the entries of the branches to be flushed and subsequent branches in the target address pool to 0. The corresponding indexes in the branch record buffer are retained but no longer used.

[0037] In the steps of the S3 branch parsing stage, in the received execution result link of this step, the execution unit completes the branch calculation and returns the actual target address and prediction correctness to the front end; The correct prediction processing link includes locating the entry in the target address pool through the index in the branch record buffer, immediately releasing the TP entry, changing the status bit of the entry in the target address pool to 0, without waiting for subsequent operations, and retaining the branch record buffer entry until the branch retires.

[0038] The incorrect prediction processing link includes locating the entry in the target address pool through the index, overwriting the incorrect address, writing the correct target address calculated by the execution unit to the original entry position in the target address pool, keeping the status bit as 1, and the entry is still valid and will be released after the subsequent update of the branch predictor.

[0039] S4 Branch Predictor Update and Release Step. In this step, the correct prediction scenario includes not updating the branch predictor, the entry in the target address pool has been released during the branch resolution phase, and the status bit has been changed to 0. The incorrect prediction scenario includes updating the branch predictor. The branch record buffer entry reads the correct target address in the target address pool through indexing, and updates the branch predictor in combination with the branch information. Then it releases the entry in the target address pool. After the branch predictor update is completed, the status bit of the entry in the target address pool is changed to 0.

[0040] The branch record buffer entry release process includes that after the branch instruction retires, the branch record buffer entry is released regardless of whether the prediction is correct or not.

[0041] The address widths adapted by the target address pool include 16 / 32 / 48 / 64 bits. There is an address width bit in the entry of the target address pool. The address width bit is composed of two boolean characters. The content stored through the address width bit is 0 / 1 / 2 / 3 corresponding to 16 / 32 / 48 / 64-bit address widths. By identifying the address width bit, the address content of different lengths is read subsequently.

[0042] Embodiment 2: This embodiment provides a branch target address optimized storage structure based on a target address pool, as Figure 1 shown, and also includes a TPP structure: Target Address Pool (TPP) Structure: Each Target PC (TP) entry contains a 48-bit complete target PC and 1 bit status flag to mark whether the entry is valid.

[0043] It can be designed to have far fewer entries than BRB Entries (for example, TP only has 16 entries while BRB has 128 entries).

[0044] Target PC Pool (TPP) Workflow: Allocation: 1. After the branch instruction is fetched, while writing it into the branch record buffer, its corresponding complete predicted target address is written into TP.

[0045] 2. After writing to the allocated entry in TP, set the entry valid status bit to 1 and return the tp_idx corresponding to the entry. Assuming TP has 16 entries, then idx requires 4 bits.

[0046] 3. Write the tp_idx into the entry of the Branch Record Buffer (BRB) that records this branch instruction.

[0047] In-Flight: 1. Only the tp_idx is stored in the BRB, and the target address is always stored in the TPP, without occupying the register area of the BRB.

[0048] 2. If the pipeline is flushed, the target addresses stored in the TPP for the flushed branch and subsequent branch instructions will become invalid, and the status bit valid will be set to 0.

[0049] Resolve: When the branch instruction is resolved, the instruction execution unit will return the prediction result to the IFU.

[0050] 1. When the prediction is correct, the corresponding entry storing the target address can be immediately indexed through the tp_idx stored in the BRB, and its status bit is cleared to indicate that the current entry has been deallocated.

[0051] 2. If the prediction is incorrect, the correct target address from the execution unit is written to the corresponding entry in the TPP, overwriting the previously stored incorrect address.

[0052] BP Update: After the branch instruction is completed, the BRB needs to update the branch prediction unit by storing the content.

[0053] 1. Prediction correct: It means that the target branch stored in the branch prediction unit does not need to be updated.

[0054] 2. Prediction incorrect: In addition to reading the branch instruction information stored in itself, the BRB will also read the TP through the tp_idx to obtain the correct branch target address to finally complete the update of the BP.

[0055] Deallocation: 1. When the branch prediction is correct during branch resolution, the TP entry will be immediately deallocated.

[0056] 2. When the branch prediction is incorrect during branch resolution, the entry in the TPP will be deallocated after the branch predictor is updated.

[0057] Significant area savings: Only the TP index needs to be stored in the BRB, greatly reducing the width of the BRB entry.

[0058] The capacity of the TPP can also be much smaller than the number of BRB Entries because not all branches are in the pre-resolution state simultaneously; Take 128 Entry BRB and 48-bit address as an example: If there is no distinction, 128 × 48 bits = 0.75KB is required. If the TP only has 16 Entries and the Pool Index + prediction information is saved in the BRB (for example, the Pool Index requires 4 bits), nearly 0.6KB of storage can be saved. The specific calculation method is as follows: Original scheme: 128 × 48 bits = 6144 bits (≈0.75KB) TPP in the new scheme: 16 × 49 bits = 784 bits (≈0.1KB) BRB in the new scheme: 128 × 4 bits (Pool Index) = 512 bits (≈0.06KB) The total occupancy is approximately 0.16KB. Compared with 0.75KB, nearly 80% of the storage area is saved.

[0059] Embodiment 3: This embodiment provides a branch target address optimized storage structure based on a target address pool, as Figure 1 shown, and also includes a shared TP structure: Shared Target PC Pool: If multiple branch instructions may share the same target PC, redundant storage can be further reduced. However, additional logic is required to distinguish the reference count.

[0060] Variable word length: If some branch PCs only require a 32-bit address space, the storage bit width can be adaptively shortened, but an additional flag bit is required to indicate the address mode (32 / 64 bits, etc.).

[0061] The above content is a further detailed description of the present invention in combination with specific preferred implementation schemes, which is convenient for those skilled in the art of this technology to understand and apply the present invention. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions.

Claims

1. A branch target address optimized storage structure based on a target address pool, characterized in that, It includes a target address pool and its working process; The target address pool includes storing complete 48-bit target addresses and 1-bit status bits, and its capacity is smaller than that of the branch record buffer; It also includes the transformation of the branch record buffer, which includes storing only the index pointing to the TP instead of the complete target address; The working process of the target address pool includes the following steps: S1 Allocation stage steps, which include a branch instruction fetching link, a step of writing to the target address pool, an index generation link, and an update branch record buffer link; S2 Running stage steps, which include a target address access link and a processing pipeline flushing link; S3 Branch resolution stage steps, which include a receiving execution result link, a correct prediction processing link, and an incorrect prediction processing link; S4 Branch predictor update and release steps, which include a correct prediction link, an incorrect prediction link, and a branch record buffer entry release link.

2. The optimized storage structure for branch target addresses based on a target address pool according to claim 1, characterized in that, In the S1 allocation stage steps, in the branch instruction fetching link, the branch instruction enters the front-end fetch unit, is recognized and allocated a branch record buffer entry. The step of writing to the target address pool includes writing the complete predicted target address of the branch to the idle entry of the target address pool and changing the status bit of the target address entry to 1.

3. The optimized storage structure for branch target addresses based on a target address pool according to claim 2, wherein The index generation link includes the target address pool returning the index allocated to the entry. The update branch record buffer link includes storing only the index instead of the complete target address in the corresponding entry of the branch record buffer and recording other branch information.

4. A branch target address optimized storage structure based on a target address pool according to claim 3, characterized in that, In the S2 running stage steps, in the target address access link, when the pipeline needs to read the branch target address, the branch record buffer looks up the target address pool through the index to obtain the complete address.

5. The optimized storage structure for branch target addresses based on a target address pool according to claim 4, wherein, The processing pipeline flushing link includes if a pipeline flush occurs, changing the status bits of the entries of the flushed branch and subsequent branches in the target address pool to 0. The corresponding index in the branch record buffer is retained but no longer used.

6. The optimized storage structure for branch target addresses based on a target address pool according to claim 5, wherein, In the S3 branch resolution stage steps, in the receiving execution result link, the execution unit completes the branch calculation and returns the actual target address and prediction correctness to the front end; The correct prediction processing link includes locating the entry in the target address pool through the index in the branch record buffer, immediately releasing the TP entry, changing the status bit of the entry in the target address pool to 0, without waiting for subsequent operations. The branch record buffer entry is retained until the branch retires.

7. The optimized storage structure for branch target addresses based on a target address pool according to claim 6, wherein The incorrect prediction processing link includes locating the entry in the target address pool through the index, overwriting the incorrect address, writing the correct target address calculated by the execution unit to the original entry position in the target address pool, keeping the status bit as 1, the entry is still valid, and waiting to be released after updating the branch predictor subsequently.

8. An optimized storage structure for branch target addresses based on a target address pool according to claim 7, characterized in that, S4 Branch Predictor Update and Release Step. In the correct prediction part of this step, there is no need to update the branch predictor. The entry in the target address pool has been released during the branch resolution phase, and the status bit has been changed to 0. In the incorrect prediction part, the branch predictor is updated. The branch record buffer entry reads the correct target address in the target address pool through indexing, and updates the branch predictor in combination with the branch information. The entry in the target address pool is released. After the branch predictor update is completed, the status bit of the entry in the target address pool is changed to 0.

9. The optimized storage structure for branch target addresses based on a target address pool according to claim 8, characterized in that, The release part of the branch record buffer entry includes that after the branch instruction retires, the branch record buffer entry is released, regardless of whether the prediction is correct or not.

10. The optimized storage structure for branch target addresses based on a target address pool according to claim 9, wherein The address widths adapted by the target address pool include 16 / 32 / 48 / 64 bits. There is an address width bit in the entry of the target address pool. The address width bit is composed of two boolean characters. The content stored through the address width bit is 0 / 1 / 2 / 3 corresponding to the 16 / 32 / 48 / 64-bit address widths. By identifying the address width bit, the address content of different lengths is read subsequently.

Citation Information

Patent Citations

  • Branch target buffer addressing in a data processor

    CN102841777A

  • Branch prediction method, microprocessor thereof and data processing system

    CN113760371A

  • Instruction prediction method of multi-thread processor and related device

    CN114020441A

  • Restoring speculative history used for making speculative predictions for instructions processed in a processor employing control independence techniques

    US20220113976A1