Branch prediction processing system, method for reducing branch prediction power consumption, and storage medium
By splitting the branch target cache into multiple small-sized sub-modules and performing dynamic power consumption control, the problem of high power consumption of the branch target cache is solved, and power consumption optimization and performance guarantee of the processor system are achieved.
Patent Information
- Application Number
- CN202511227794.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-29
AI Technical Summary
Existing branch target cache architectures consume a lot of power, affecting the power consumption, performance, and area performance (PPA) of processor systems.
The branch target cache is split into multiple small-sized branch target cache sub-modules. During the path prediction process, only the corresponding marked branch target cache sub-module is activated in normal power consumption mode, while other sub-modules remain in low power consumption state. Dynamic power consumption control is performed through multi-level branch target cache modules and multiplexers.
Reduce the power consumption of invalid cache accesses, lower the overall circuit activity rate, and optimize the power consumption of the processor system while ensuring performance.
Smart Images

Figure CN120743357A_ABST
Abstract
Description
Technical Field
[0001] The present invention is applicable to the field of processor technology, and in particular relates to a branch prediction processing system, a method for reducing branch prediction power consumption, and a storage medium. Background Art
[0002] Branch prediction is a technology used in modern processors to improve execution efficiency. Conditional branches prevent the processor from determining the address of the next instruction until the branch instruction has completed execution, which can cause the processor to wait and reduce execution efficiency. Branch prediction technology predicts the outcome of a branch instruction, allowing the processor to fetch and execute subsequent instructions in advance. If the prediction is correct, the processor can continue execution seamlessly; if the prediction is incorrect, it must roll back to the correct path. The accuracy of branch prediction is crucial to processor performance.
[0003] To improve performance, the front-end fetch pipeline of a high-performance central processing unit (CPU) predicts the jump target address / direction based on branch history information when encountering a branch instruction, and then speculatively executes the branch instruction. After obtaining the branch jump address, it is passed to the instruction fetch module, which reads the instruction from the instruction cache and passes it to the instruction decode module, which then passes it to the instruction execution module (also known as the execution backend). If the branch prediction result is correct, processor performance will be improved. Conversely, if the branch prediction result is incorrect, the processor pipeline must clear the incorrectly executed instruction and roll back to the execution state before the branch. In this case, the pipeline will stall for several clock cycles. Therefore, reducing the processor stall time after a branch prediction failure will help improve performance.
[0004] The Branch Target Buffer (BTB) is a hardware component used in modern processors to optimize branch prediction. It is a key technology for improving branch prediction accuracy and reducing branch overhead. By storing and quickly retrieving branch target addresses, the BTB helps improve processor performance and efficiency. Its core function is to store the target address of each branch instruction, eliminating the need for the processor to recalculate the target address each time a branch jump is executed.
[0005] Typically, a branch target cache consists of a large, deep static random access memory (SRAM). According to existing branch target cache designs, every time a processor reads data from the branch target cache, it reads the entire SRAM. This requires that every portion of the SRAM be powered on, leading to high power consumption during branch prediction and a significant impact on the overall PPA (Power, Performance, Area) performance of the CPU system. Summary of the Invention
[0006] The present invention provides a branch prediction processing system, a method for reducing branch prediction power consumption, and a storage medium, aiming to solve the technical problem of high power consumption in existing branch target cache architectures.
[0007] To solve the above technical problems, in a first aspect, the present invention provides a branch prediction processing system, comprising a first-level multiplexer, a multi-level branch target cache module, a way prediction determination module, and a second-level multiplexer, wherein the multi-level branch target cache module includes a plurality of branch target cache sub-modules, each of the branch target cache sub-modules having an additional preset way prediction tag field, wherein: The first-level multiplexer is configured to obtain instruction address information and branch target cache way prediction information transmitted from the processor pipeline, and set the branch target cache submodule having the corresponding preset way prediction tag field to normal power consumption mode and enable read mode based on the branch target cache way prediction information and reread tag information, and then transmit the instruction address information to the branch target cache submodule, wherein the reread tag information is generated by the second-level multiplexer; the first-level multiplexer is further configured to transmit the table lookup result returned by the multi-level branch target cache module to the way prediction determination module; The multi-level branch target cache module is configured to perform a table lookup of a branch target address in the corresponding branch target cache submodule according to the instruction address information, obtain the table lookup result, and return the table lookup result to the first-level multiplexer; The way prediction determination module is used to compare the instruction address information with the table lookup result, obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmit the way prediction result to the second-stage multiplexer; The second-level multiplexer is used to output the way prediction result or the table lookup result to the first-level multiplexer or the processor pipeline based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, the table lookup result is output to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, the way prediction result and the re-read mark information are output to the first-level multiplexer and re-read.
[0008] Furthermore, in the multi-level branch target cache module, the number of the branch target cache submodules is defined as N1, N1 ≥ 2, and N1 is an integer, and the lowest N2 bits of the branch target address stored in each branch target cache submodule are used as a preset flag field for identifying address alias conflicts, N2 ≥ log2(N1), and N2 is an integer; The bit width of each branch target cache submodule is defined as N3, where N3>N2 and N3 is an integer. The (N3-N2) bits of the branch target address stored in each branch target cache submodule are used as the search address bit width. The cache depth of each branch target cache submodule is less than or equal to 2. (N3-N2) .
[0009] Furthermore, when the way prediction and determination module compares the instruction address information with the table lookup result, it compares the lowest N2 bits of the address corresponding to the instruction address information and the table lookup result.
[0010] Furthermore, the multi-level branch target cache module is further configured to: According to the highest (N3-N2) bit of the instruction address information, a table lookup is performed on the highest (N3-N2) bit of the branch target address in the corresponding branch target cache submodule, Each data item read out contains an alias identification field, which has N2 bits. The lowest N2 bits of the instruction address information are compared with the N2 bits of the alias identification field read out by the branch target cache submodule, wherein: Compare the lowest N2 bits of the instruction address information with the N2-bit alias identification field read from the branch target cache submodule: If the two are the same, the corresponding branch target address is output as the table lookup result; If the two are different, a data missing flag bit data is generated and output as the table lookup result.
[0011] Furthermore, the first-stage multiplexer is further configured to: setting one of the branch target cache submodules having the corresponding way prediction tag field to a normal power consumption mode and the remaining branch target cache submodules to a low power consumption mode according to the branch target cache way prediction information; Furthermore, the first-stage multiplexer is further configured to: When reading for the first time, the instruction address information is sent to the corresponding multi-level branch target cache module. If the instruction address information is not correctly read for the first time, the branch target address is simultaneously looked up in all the branch target cache submodules according to the instruction address information when reading again.
[0012] Furthermore, the first-stage multiplexer is further configured to: The data transmitted from the processor pipeline and the second-stage multiplexer are received simultaneously, and the data transmitted from the second-stage multiplexer is processed preferentially.
[0013] In a second aspect, the present invention further provides a method for reducing branch prediction power consumption, the method being implemented based on the branch prediction processing system as described above, the method for reducing branch prediction power consumption comprising the following steps: Acquire instruction address information and branch target cache way prediction information transmitted from the processor pipeline through the branch predictor, and transmit the instruction address information and the branch target cache way prediction information to the first-level multiplexer; Obtaining instruction address information and branch target cache way prediction information transmitted from the processor pipeline through the first-stage multiplexer, and setting the branch target cache submodule having the corresponding preset way prediction tag field to a read mode based on the branch target cache way prediction information and reread tag information, and then transmitting the instruction address information to the branch target cache submodule, wherein the reread tag information is generated by the second-stage multiplexer; performing a table lookup of a branch target address in the corresponding branch target cache submodule according to the instruction address information through a multi-level branch target cache module, obtaining the table lookup result, and returning the table lookup result to the first-level multiplexer; Transmitting the table lookup result returned by the multi-level branch target cache module to the way prediction and determination module through the first-level multiplexer; Comparing the instruction address information with the table lookup result through a way prediction determination module to obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmitting the way prediction result to the second-stage multiplexer; The way prediction result or the table lookup result is output to the first-level multiplexer or the processor pipeline through the second-level multiplexer based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, the table lookup result is output to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, the way prediction result and the re-read mark information are output to the first-level multiplexer and re-read.
[0014] In a third aspect, the present invention also provides a computer device comprising: a memory, a processor, and a program for reducing branch prediction power consumption stored in the memory and executable on the processor, wherein the processor implements the steps of the method for reducing branch prediction power consumption as described in any one of the above embodiments when executing the program for reducing branch prediction power consumption.
[0015] In a fourth aspect, the present invention also provides a storage medium, on which is stored a program for reducing branch prediction power consumption. When the program for reducing branch prediction power consumption is executed by a processor, the steps in the method for reducing branch prediction power consumption as described in any one of the above embodiments are implemented.
[0016] The beneficial effect achieved by the present invention is that a branch prediction processing system based on a multi-level branch target cache is proposed. The system splits the branch target cache into multiple small-sized branch target cache sub-modules. During the path prediction process, only the branch target cache sub-modules with corresponding marks are activated to normal power consumption mode and read mode, while the remaining sub-modules remain in a low power consumption state. This can reduce the power consumption of invalid cache accesses, reduce the overall circuit activity rate, and achieve power consumption optimization of the processor system. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The present invention will be described in detail below with reference to the accompanying drawings. The above and other aspects of the present invention will become clearer and easier to understand through the detailed description made with reference to the following drawings. In the accompanying drawings: Figure 1 It is a structural diagram of a branch prediction processing system provided by an embodiment of the present invention; Figure 2 This is a flowchart of the steps of a method for reducing branch prediction power consumption provided by an embodiment of the present invention; Figure 3 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0019] Please refer to Figure 1 , Figure 1 1 is a schematic diagram of the structure of a branch prediction processing system provided by an embodiment of the present invention. The branch prediction processing system 100 includes a first-level multiplexer 101, a multi-level branch target cache module 102, a way prediction determination module 103, and a second-level multiplexer 104. The multi-level branch target cache module 102 includes a plurality of branch target cache sub-modules 1021 ( Figure 1 3 as an example), each of the branch target cache submodules 1021 has an additional preset way prediction tag field, wherein: The first-level multiplexer 101 is configured to obtain instruction address information and branch target cache way prediction information transmitted from the processor pipeline, and set the branch target cache submodule 1021 having the corresponding preset way prediction tag field to normal power consumption mode and read mode based on the branch target cache way prediction information and reread tag information. Thereafter, the instruction address information is transmitted to the branch target cache submodule 1021. The reread tag information is generated by the second-level multiplexer 104. The first-level multiplexer 101 is further configured to transmit the table lookup result returned by the multi-level branch target cache module to the way prediction determination module 103. The multi-level branch target cache module 102 is configured to perform a table lookup of a branch target address in the corresponding branch target cache submodule 1021 according to the instruction address information, obtain the table lookup result, and return the table lookup result to the first-level multiplexer 101; The way prediction determination module 103 is used to compare the instruction address information with the table lookup result to obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmit the way prediction result to the second-stage multiplexer 104; The second-level multiplexer 104 is used to output the way prediction result or the table lookup result to the first-level multiplexer 101 or the processor pipeline based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, the table lookup result is output to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, the way prediction result and the re-read mark information are output to the first-level multiplexer 101 and re-read.
[0020] Specifically, in the multi-level branch target cache module 102, the number of the branch target cache submodules 1021 is defined as N1, where N1 ≥ 2 and N1 is an integer. The lowest N2 bits of the branch target address stored in each branch target cache submodule 1021 are used as a preset flag field for identifying address alias conflicts, where N2 ≥ log2(N1) and N2 is an integer. The bit width of each branch target cache submodule 1021 is defined as N3, where N3>N2 and N3 is an integer. The (N3-N2) bits of the branch target address stored in each branch target cache submodule 1021 are used as the search address bit width. The cache depth of each branch target cache submodule 1021 is less than or equal to 2. (N3-N2) .
[0021] When the way prediction and determination module 103 compares the instruction address information with the table lookup result, it compares the lowest N2 bits of the address corresponding to the instruction address information and the table lookup result.
[0022] Based on the design in the above embodiment, when performing a comparison, the way prediction determination module 103 first reads data based on the predicted preset way prediction flag field and verifies whether the prediction is correct. If the initial prediction fails, all branch target cache sub-modules 1021 are read in parallel to obtain data from the correct block. Because the comparison uses the lowest N2 bits of the address, and the lowest N2 bits of the instruction address determine which branch target cache sub-module 1021 the instruction's branch target should be stored in, this ensures that data is read from the correct branch target cache sub-module 1021 when the prediction is correct.
[0023] The multi-level branch target cache module 102 is further configured to: According to the highest (N3-N2) bits of the instruction address information, a table lookup is performed on the highest (N3-N2) bits of the branch target address in the corresponding branch target cache submodule 1021 to read out the read data including the N2-bit alias identification field, wherein: Compare the lowest N2 bits of the instruction address information with the N2-bit alias identification field read from the branch target cache submodule 1021: If the two are the same, the corresponding branch target address is output as the table lookup result; If the two are different, a data missing flag bit data is generated and output as the table lookup result.
[0024] In the embodiment of the present invention, the multi-level branch target cache module 102 has a two-level table lookup strategy, wherein, based on the design in the above embodiment, the branch target address is divided into two parts: High-order bits (N3-N2): used to quickly locate possible cache entries; The lower N2 bits are used for exact matching and resolving alias conflicts.
[0025] Specifically, the query method of the above table lookup strategy in actual implementation is similar to the cache query method, which first reads data from the multi-way set associative structure according to a partial address, and then compares to see if there is a hit.
[0026] The first-stage multiplexer 101 is further configured to: If the reread tag information is empty, one of the branch target cache submodules 1021 having the corresponding way prediction tag field is set to normal power consumption mode according to the branch target cache way prediction information, and the other branch target cache submodules 1021 are set to low power consumption mode.
[0027] Based on the re-read tag information, a dynamic power consumption optimization strategy of the multi-level branch target cache module 102 in an embodiment of the present invention can be implemented. In most cases, because the system's path prediction is correct, only the predicted branch target cache sub-module 1021 needs to be activated, and other blocks can enter a low power consumption mode to save energy.
[0028] Preferably, when a path prediction error occurs (a rare occurrence), all branch target cache submodules 1021 need to be read in parallel, and all blocks must enter normal power consumption mode to maintain speed. By dynamically controlling the power consumption mode of each branch target cache submodule, the system can minimize power consumption while maintaining performance.
[0029] In an embodiment of the present invention, the additional preset way prediction flag field is of arbitrary bit width and can be configured to the same initial value during system initialization, or can be manually defined as an arbitrary initial value. After each branch target cache access is completed, the access order is recorded in a lookup table. This lookup table is indexed by the instruction address, and the table entry contains the branch target cache submodule number of the next instruction; When a way prediction fails, the number of the branch target cache submodule 1021 of the next instruction is found from the lookup table, and the data is written into the preset way prediction tag field, thereby achieving dynamic update of the preset way prediction tag field.
[0030] The first-stage multiplexer 101 is further configured to: When reading for the first time, it is sent to the corresponding multi-level branch target cache module 102 according to the instruction address information; When the instruction is not correctly read for the first time and is read again, a branch target address table lookup is performed simultaneously in all the branch target cache sub-modules 1021 according to the instruction address information.
[0031] In this embodiment of the present invention, the reread flag information is set by the second-stage multiplexer 104 when a way prediction fails, indicating that a second read is required. As described in the above embodiment, when a way prediction error occurs, the second-stage multiplexer 104 returns a reread flag to enable parallel reading of all branch target cache submodules to implement data query.
[0032] The first-stage multiplexer 101 is further configured to: It simultaneously receives data transmitted from the processor pipeline and the second-level multiplexer 104, and gives priority to processing data transmitted from the second-level multiplexer 104. This design implements data flow priority control in the branch prediction processing system 100, which is a key mechanism to ensure the correctness and performance of the system. Specifically, the first-level multiplexer 101, as the data distribution hub of the branch prediction system, needs to process two types of requests: normal prediction requests from the processor pipeline, and error correction requests from the second-level multiplexer 104. Compared with normal prediction requests, error correction requests must be processed first. This is because branch prediction errors can cause pipeline flushing, resulting in serious performance losses. Rapidly processing error correction requests from the second-level multiplexer 104 can minimize this loss to ensure rapid recovery and reduce performance losses caused by branch prediction errors.
[0033] The beneficial effect achieved by the present invention is that a branch prediction processing system based on a multi-level branch target cache is proposed. The system splits the branch target cache into multiple small-sized branch target cache sub-modules. During the path prediction process, only the branch target cache sub-modules with corresponding marks are activated to normal power consumption mode and read mode, while the remaining sub-modules remain in a low power consumption state. This can reduce the power consumption of invalid cache accesses, reduce the overall circuit activity rate, and achieve power consumption optimization of the processor system.
[0034] The embodiment of the present invention also provides a method for reducing branch prediction power consumption, which is implemented based on the branch prediction processing system as described above. Figure 2 , Figure 2 1 is a flowchart of the steps of a method for reducing branch prediction power consumption provided by an embodiment of the present invention, wherein the method for reducing branch prediction power consumption comprises the following steps: S201, obtaining instruction address information and branch target cache way prediction information transmitted from a processor pipeline through a first-level multiplexer, and setting the branch target cache submodule having a corresponding preset way prediction flag field to a normal power consumption mode and a read mode based on the branch target cache way prediction information and reread flag information, and then transmitting the instruction address information to the branch target cache submodule, wherein the reread flag information is generated by the second-level multiplexer; S202: performing a table lookup of a branch target address in the corresponding branch target cache submodule according to the instruction address information through the multi-level branch target cache module to obtain the table lookup result, and returning the table lookup result to the first-level multiplexer; S203, transmitting the table lookup result returned by the multi-level branch target cache module to the way prediction and determination module through the first-level multiplexer; S204: Compare the instruction address information with the table lookup result by a way prediction determination module to obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmit the way prediction result to the second-stage multiplexer; S205. Output the way prediction result or the table lookup result to the first-level multiplexer or the processor pipeline through the second-level multiplexer based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, output the table lookup result to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, output the way prediction result and the re-read mark information to the first-level multiplexer and re-read.
[0035] The method for reducing branch prediction power consumption can implement corresponding functional steps based on the system for reducing branch prediction power consumption in the above embodiment, and can achieve the same technical effects. Please refer to the description in the above embodiment and will not be repeated here.
[0036] The embodiment of the present invention also provides a computer device, please refer to Figure 3 , Figure 3 3 is a structural diagram of a computer device provided in an embodiment of the present invention. The computer device 300 includes: a memory 302, a processor 301, and a program for reducing branch prediction power consumption stored in the memory 302 and executable on the processor 301.
[0037] The processor 301 calls the program for reducing branch prediction power consumption stored in the memory 302 and executes the steps of the method for reducing branch prediction power consumption provided by the embodiment of the present invention. Figure 3 , specifically including the following steps: S201, obtaining instruction address information and branch target cache way prediction information transmitted from a processor pipeline through a first-level multiplexer, and setting the branch target cache submodule having a corresponding preset way prediction flag field to a normal power consumption mode and a read mode based on the branch target cache way prediction information and reread flag information, and then transmitting the instruction address information to the branch target cache submodule, wherein the reread flag information is generated by the second-level multiplexer; S202: performing a table lookup of a branch target address in the corresponding branch target cache submodule according to the instruction address information through the multi-level branch target cache module to obtain the table lookup result, and returning the table lookup result to the first-level multiplexer; S203, transmitting the table lookup result returned by the multi-level branch target cache module to the way prediction and determination module through the first-level multiplexer; S204: Compare the instruction address information with the table lookup result by a way prediction determination module to obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmit the way prediction result to the second-stage multiplexer; S205. Output the way prediction result or the table lookup result to the first-level multiplexer or the processor pipeline through the second-level multiplexer based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, output the table lookup result to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, output the way prediction result and the re-read mark information to the first-level multiplexer and re-read.
[0038] The computer device 300 provided in the embodiment of the present invention can implement the steps in the method for reducing branch prediction power consumption in the above embodiment and can achieve the same technical effects. Please refer to the description in the above embodiment and will not be repeated here.
[0039] An embodiment of the present invention further provides a storage medium, on which is stored a program for reducing branch prediction power consumption. When the program for reducing branch prediction power consumption is executed by a processor, the various processes and steps in the method for reducing branch prediction power consumption provided in an embodiment of the present invention are implemented, and the same technical effects can be achieved. To avoid repetition, they will not be described here.
[0040] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware (such as a mobile phone, computer, server, air conditioner, or network equipment) using a program that reduces branch prediction power consumption. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0041] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0042] The embodiments of the present invention are described above in conjunction with the accompanying drawings. What is disclosed is only a preferred embodiment of the present invention. However, the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms and equivalent changes without departing from the scope of protection of the purpose of the present invention and the claims, which are all within the protection of the present invention.
Claims
1. A branch prediction processing system, characterized in that: The invention comprises a first-level multiplexer, a multi-level branch target cache module, a way prediction determination module, and a second-level multiplexer, wherein the multi-level branch target cache module comprises a plurality of branch target cache sub-modules, each of the branch target cache sub-modules having an additional preset way prediction tag field, wherein: The first-level multiplexer is configured to obtain instruction address information and branch target cache way prediction information transmitted from the processor pipeline, and set the branch target cache submodule having the corresponding preset way prediction tag field to normal power consumption mode and enable read mode based on the branch target cache way prediction information and reread tag information, and then transmit the instruction address information to the branch target cache submodule, wherein the reread tag information is generated by the second-level multiplexer; the first-level multiplexer is further configured to transmit the table lookup result returned by the multi-level branch target cache module to the way prediction determination module; The multi-level branch target cache module is configured to perform a table lookup of a branch target address in the corresponding branch target cache submodule according to the instruction address information, obtain the table lookup result, and return the table lookup result to the first-level multiplexer; The way prediction determination module is used to compare the instruction address information with the table lookup result, obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmit the way prediction result to the second-stage multiplexer; The second-level multiplexer is used to output the way prediction result or the table lookup result to the first-level multiplexer or the processor pipeline based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, the table lookup result is output to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, the way prediction result and the re-read mark information are output to the first-level multiplexer and re-read.
2. The branch prediction processing system according to claim 1, wherein: In the multi-level branch target cache module, the number of the branch target cache submodules is defined as N1, where N1 ≥ 2 and N1 is an integer, and the lowest N2 bits of the branch target address stored in each of the branch target cache submodules are used as a preset flag field for identifying address alias conflicts, where N2 ≥ log2(N1) and N2 is an integer; The bit width of each branch target cache submodule is defined as N3, where N3>N2 and N3 is an integer. The (N3-N2) bits of the branch target address stored in each branch target cache submodule are used as the search address bit width. The cache depth of each branch target cache submodule is less than or equal to 2. (N3-N2) .
3. The branch prediction processing system according to claim 2, wherein: When the way prediction and determination module compares the instruction address information with the table lookup result, it compares the lowest N2 bits of the address corresponding to the instruction address information and the table lookup result.
4. The branch prediction processing system according to claim 2, wherein: The multi-level branch target cache module is further configured to: According to the highest (N3-N2) bits of the instruction address information, a table lookup is performed on the highest (N3-N2) bits of the branch target address in the corresponding branch target cache submodule to read out read data including the N2-bit alias identification field, wherein: Compare the lowest N2 bits of the instruction address information with the N2-bit alias identification field read from the branch target cache submodule: If the two are the same, the corresponding branch target address is output as the table lookup result; If the two are different, a data missing flag bit data is generated and output as the table lookup result.
5. The branch prediction processing system according to claim 1, wherein: The first-stage multiplexer is further configured to: If the reread tag information is empty, one of the branch target cache submodules having the corresponding way prediction tag field is set to normal power consumption mode according to the branch target cache way prediction information, and the remaining branch target cache submodules are set to low power consumption mode.
6. The branch prediction processing system according to claim 5, wherein: The first-stage multiplexer is further configured to: When reading for the first time, the instruction address information is sent to the corresponding multi-level branch target cache module; When the instruction is not correctly read for the first time and is read again, a branch target address table lookup is performed simultaneously in all the branch target cache sub-modules according to the instruction address information.
7. The branch prediction processing system according to claim 1, wherein: The first-stage multiplexer is further configured to: The data transmitted from the processor pipeline and the second-stage multiplexer are received simultaneously, and the data transmitted from the second-stage multiplexer is processed preferentially.
8. A method for reducing branch prediction power consumption, characterized in that: The method is implemented based on the branch prediction processing system according to claims 1-7, and the method for reducing branch prediction power consumption comprises the following steps: Obtaining instruction address information and branch target cache way prediction information transmitted from the processor pipeline through the first-level multiplexer, and setting the branch target cache submodule having the corresponding preset way prediction flag field to a normal power consumption mode and a read mode based on the branch target cache way prediction information and the reread flag information, and then transmitting the instruction address information to the branch target cache submodule, wherein the reread flag information is generated by the second-level multiplexer; performing a table lookup of a branch target address in the corresponding branch target cache submodule according to the instruction address information through a multi-level branch target cache module, obtaining the table lookup result, and returning the table lookup result to the first-level multiplexer; Transmitting the table lookup result returned by the multi-level branch target cache module to the way prediction and determination module through the first-level multiplexer; Comparing the instruction address information with the table lookup result through a way prediction determination module to obtain a way prediction result as to whether the instruction address information and the table lookup result are equal, and transmitting the way prediction result to the second-stage multiplexer; The way prediction result or the table lookup result is output to the first-level multiplexer or the processor pipeline through the second-level multiplexer based on the way prediction result, wherein: if the way prediction result indicates that the instruction address information and the table lookup result are equal, the table lookup result is output to the processor pipeline; if the way prediction result indicates that the instruction address information and the table lookup result are not equal, the way prediction result and the re-read mark information are output to the first-level multiplexer and re-read.
9. A computer device, characterized in that: include: A memory, a processor, and a program for reducing branch prediction power consumption stored in the memory and executable on the processor, wherein the processor implements the steps of the method for reducing branch prediction power consumption as claimed in claim 8 when executing the program for reducing branch prediction power consumption.
10. A storage medium, characterized in that: The storage medium stores a program for reducing branch prediction power consumption, and when the program for reducing branch prediction power consumption is executed by the processor, the steps of the method for reducing branch prediction power consumption according to claim 8 are implemented.
Citation Information
Patent Citations
Low-power-consumption branch target buffer with two-stage prediction mechanism and design method
CN116302112A
Branch prediction method of multilevel fetch target buffer based on processor
CN117472446A
Dynamic set associative cache apparatus for processor and access method thereof
US20140344522A1
Filtered branch prediction structures of a processor
US20200065106A1