Branch predictor selection method and selection device, and storage medium
By dynamically selecting branch predictors, using local and global historical data and processor load states, the problems of low resource utilization efficiency and high power consumption of existing branch predictors in complex branches and high load situations are solved, achieving higher prediction accuracy and energy efficiency balance.
Patent Information
- Application Number
- CN202510278380.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Existing branch predictors are inefficient in resource utilization when processing complex branches, resulting in increased power consumption, especially under high load conditions, and cannot achieve optimal utilization of processor resources.
A branch predictor selection method is proposed. By obtaining recorded data in local history registers and global history registers, determining the instruction type of the target jump instruction and the load state of the processor, and dynamically selecting the appropriate branch predictor (such as TAGE or Bimodal) to improve prediction accuracy and reduce energy consumption.
This method improves the accuracy of branch prediction, has stronger adaptability, reduces energy consumption, and achieves the best balance between prediction accuracy and energy efficiency.
Smart Images

Figure CN120216036A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of predictors, and in particular, to a method for selecting a branch predictor, a device for selecting a branch predictor, and a computer-readable storage medium. Background Art
[0002] In modern processor design, a branch predictor is an important component for improving performance. It enables the processor to continue executing instructions on the predicted path before the branch condition is determined by predicting the result of the branch instruction in advance, avoiding pipeline stalls and maintaining execution efficiency. The processor uses pipeline technology to divide the instruction execution process into multiple stages, allowing it to process multiple instructions simultaneously. However, when encountering branch instructions such as conditional judgments or loop controls, the processor must wait for the calculation result of the branch condition, which causes pipeline stalls and significantly reduces performance. The branch predictor helps the processor quickly determine the execution path through direction prediction (whether the branch jumps) and target prediction (jump target address). It can "guess" the execution path of the program in advance, enabling the processor to continue executing instructions on the predicted path before the branch condition is determined. Through this mechanism, the processor can maintain the continuous operation of the pipeline and reduce the waiting time for branch results. With the increasing complexity of modern processor pipeline designs, branch prediction is crucial for improving processor performance.
[0003] The design of a branch predictor not only needs to consider the accuracy of prediction but also pay attention to power consumption optimization. The necessity of low-power design is to ensure that the processor does not consume excessive system resources when performing branch prediction tasks, thereby achieving efficient performance and a long service life in applications such as embedded processors. Although existing branch prediction technologies are diverse and have been widely used in many application scenarios, each of them has its own drawbacks.
[0004] For example, Bimodal is a simple local history predictor that makes predictions by recording the local history states of branch instructions. A 2-bit saturating counter is usually used to predict whether a branch jumps or not. This predictor is very effective for simple, fixed-pattern branch instructions (such as loop exit conditions). Although the standalone Bimodal branch predictor has low hardware overhead, in some high-performance applications, its simplicity may not meet the demand for high prediction accuracy. The Bimodal branch predictor usually uses a single history record to predict branch behavior and cannot fully capture complex patterns in the program. Secondly, since the Bimodal predictor only relies on the low bits of the branch address to index the table, it is prone to conflicts when faced with similar or overlapping branch addresses. Such conflicts cause different branch instructions to share the same prediction entry, further reducing prediction accuracy. Therefore, in scenarios requiring higher performance, designers may turn to more complex predictors, such as the TAGE predictor, to achieve better performance.
[0005] TAGE is an advanced, hierarchical global history branch predictor. It dynamically selects the optimal history length for prediction based on the history length and behavior pattern of branch instructions. The core idea of TAGE is to make predictions based on global histories of different lengths, and it has a high prediction accuracy for complex branch instructions that rely on long-term histories. However, for embedded systems and edge devices, there are also some issues to consider with the TAGE branch predictor. When using the TAGE predictor alone, although it can handle complex branch instructions, it often has low resource utilization efficiency when faced with simple, fixed-pattern branches. The overly complex algorithm and hardware architecture lead to wasted prediction resources in these cases, not only occupying a large amount of hardware resources but also increasing power consumption. Especially when the processor load is high, the power consumption problem of TAGE is particularly obvious.
[0006] Therefore, the above approach adopts the same prediction strategy for all branch instructions, lacking flexibility and prone to resource waste. Especially when dealing with simple, fixed-pattern branches, the excessive calculation of complex predictors not only fails to improve efficiency but instead increases power consumption, and the optimal utilization of processor resources cannot be achieved. Summary of the Invention
[0007] This application aims to solve at least one of the technical problems in the related art to some extent. To this end, the first object of this application is to propose a method for selecting a branch predictor, which obtains the recorded data of the target jump instruction in the local history register and the global history register based on the target jump instruction, determines the instruction type of the target jump instruction based on the recorded data, determines the load status of the processor based on the running parameters of the processor, determines the target predictor according to the instruction type and the load status, and outputs the branch prediction result based on the target predictor. Thus, the accuracy of the prediction can be improved, the adaptability can be enhanced, the energy consumption can be reduced, and the best balance between prediction accuracy and energy efficiency can be achieved.
[0008] The second object of this application is to propose a device for selecting a branch predictor.
[0009] The third object of this application is to propose a computer-readable storage medium.
[0010] To achieve the above object, an embodiment of the first aspect of this application proposes a method for selecting a branch predictor, and the method includes: obtaining the recorded data of the target jump instruction in the local history register and the global history register based on the target jump instruction; determining the instruction type of the target jump instruction based on the recorded data; determining the load status of the processor based on the running parameters of the processor; determining the target predictor according to the instruction type and the load status, and outputting the branch prediction result based on the target predictor.
[0011] According to the method for selecting a branch predictor of the embodiment of this application, the recorded data of the target jump instruction in the local history register and the global history register is obtained based on the target jump instruction, the instruction type of the target jump instruction is determined based on the recorded data, the load status of the processor is determined based on the running parameters of the processor, the target predictor is determined according to the instruction type and the load status, and the branch prediction result is output based on the target predictor. Thus, this method can improve the accuracy of the prediction, has stronger adaptability, reduces the energy consumption, and achieves the best balance between prediction accuracy and energy efficiency.
[0012] In addition, the method for selecting a branch predictor according to the above embodiment of this application may also have the following additional technical features:
[0013] According to an embodiment of the present application, determining a target predictor based on the instruction type and the load status includes: when the instruction type is a complex instruction type and the load status is a low load, determining the target predictor as a first predictor; when the instruction type is a simple instruction type, determining the target predictor as a second predictor; when the instruction type is a complex instruction type and the load status is a high load, determining the target predictor as the second predictor; wherein, the first predictor is a TAGE predictor, and the second predictor is a Bimodal predictor.
[0014] According to an embodiment of the present application, the operating parameters of the processor include at least one of an instruction issue rate, an execution speed, a cache hit rate, and a pipeline state. Determining the load status of the processor based on the operating parameters of the processor includes: when the instruction issue rate is greater than a preset issue rate threshold, or the execution speed is greater than a preset speed threshold, or the cache hit rate is less than a preset hit rate threshold, or the pipeline state is a blocked state, determining the load status as a high load; when the instruction issue rate is less than or equal to the preset issue rate threshold, the execution speed is less than or equal to the preset speed threshold, the cache hit rate is greater than or equal to the preset hit rate threshold, and the pipeline state is not a blocked state, determining the load status as a low load.
[0015] According to an embodiment of the present application, determining the instruction type of the target jump instruction based on the recorded data includes: when the recorded data of the target jump instruction in the local history register are all jump or non-jump instructions, or the recorded data of the target jump instruction in the global history register are regular, determining the instruction type as a simple instruction type; when the recorded data of the target jump instruction in the local history register simultaneously include jump and non-jump instructions, or the recorded data of the target jump instruction in the global history register are irregular, determining the instruction type as a complex instruction type.
[0016] According to an embodiment of the present application, the method further includes: determining a prediction accuracy rate based on the branch prediction result and the actual result of the target jump instruction; when the prediction accuracy rate corresponding to the target jump instruction is lower than a preset accuracy rate threshold, determining the target predictor corresponding to the next target jump instruction as the first predictor.
[0017] According to an embodiment of the present application, determining the prediction accuracy rate based on the branch prediction result and the actual result of the target jump instruction includes: when the branch prediction result is consistent with the actual result of the target jump instruction, controlling a first counter to count; when the branch prediction result is inconsistent with the actual result of the target jump instruction, controlling a second counter to start counting; and determining the prediction accuracy rate according to the ratio of the count value of the first counter to the sum of the count values of the first counter and the second counter.
[0018] To achieve the above object, an embodiment of the second aspect of the present application provides a selection device for a branch predictor. The device includes: a branch instruction classification module, configured to obtain record data of the target jump instruction in a local history register and a global history register based on the target jump instruction, and determine the instruction type of the target jump instruction based on the record data; a load monitoring module, configured to determine the load status of the processor based on the operating parameters of the processor; and a branch predictor module, configured to determine a target predictor according to the instruction type and the load status, and output a branch prediction result based on the target predictor.
[0019] In the selection device for a branch predictor according to an embodiment of the present application, the branch instruction classification module is configured to obtain record data of the target jump instruction in a local history register and a global history register based on the target jump instruction, and determine the instruction type of the target jump instruction based on the record data. The load monitoring module is configured to determine the load status of the processor based on the operating parameters of the processor. The branch predictor module is configured to determine a target predictor according to the instruction type and the load status, and output a branch prediction result based on the target predictor. Thus, the device can improve the prediction accuracy, has stronger adaptability, reduces energy consumption, and achieves the best balance between prediction accuracy and energy efficiency.
[0020] In addition, the selection device for a branch predictor according to the above embodiment of the present application may further have the following additional technical features:
[0021] According to an embodiment of the present application, the branch predictor module includes: an OR gate unit, a first input end of the OR gate unit is connected to an output end of the branch instruction classification module, and a second input end of the OR gate unit is connected to an output end of the load monitoring module; a NOT gate unit, an input end of the NOT gate unit is connected to an output end of the OR gate unit, and an output end of the NOT gate unit is connected to an enable end of a first predictor; a selection unit, a first input end of the selection unit is connected to an output end of the first predictor, and a second input end of the selection unit is connected to an output end of a second predictor. Wherein, when the instruction type is a complex instruction type and the load state is a low load, the enable end of the first predictor is enabled, and the first predictor serves as the target predictor; when the instruction type is a simple instruction type, or when the instruction type is a complex instruction type and the load state is a high load, the enable end of the second predictor is enabled, and the second predictor serves as the target predictor.
[0022] According to an embodiment of the present application, the branch predictor module further includes: an AND gate unit, a first input end of the AND gate unit is connected to an output end of the OR gate unit, a second input end of the AND gate unit is configured to receive a feedback result, and an output end of the AND gate unit is respectively connected to an enable end of the second predictor and an input end of the NOT gate unit. Wherein, the feedback result is a prediction accuracy rate of the branch prediction result. When the prediction accuracy rate is lower than a preset accuracy rate threshold, the enable end of the first predictor is enabled, and the first predictor serves as the target predictor.
[0023] To achieve the above object, an embodiment of the third aspect of the present application proposes a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the above-mentioned method for selecting a branch predictor is implemented.
[0024] According to the computer-readable storage medium of the embodiment of the present application, by implementing the above-mentioned method for selecting a branch predictor when executed, the prediction accuracy can be improved, the adaptability is stronger, the energy consumption is reduced, and the best balance between prediction accuracy and energy efficiency is achieved.
[0025] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Description of the Drawings
[0026] Figure 1 It is a flowchart of the method for selecting a branch predictor according to an embodiment of the present application;
[0027] Figure 2 It is a top-level structure diagram of a low-power hybrid branch predictor according to a specific example of the present application;
[0028] Figure 3 Internal structure diagram of the branch instruction classification module according to an embodiment of the present application;
[0029] Figure 4 Structure diagram of the load detection module according to an embodiment of the present application;
[0030] Figure 5 Block diagram of the selection device of the branch predictor according to an embodiment of the present application;
[0031] Figure 6 Internal structure diagram of the branch predictor module according to an embodiment of the present application. Detailed implementation manners
[0032] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals indicate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0033] In the related art, only the scheme using the Bimodal or TAGE branch predictor is used for prediction. For example, the Bimodal predictor makes predictions through a simple indexing method, which is suitable for fixed-pattern branches that frequently repeat, but is prone to a high misprediction rate when dealing with complex and diverse branches. The TAGE predictor can provide high prediction accuracy when dealing with complex branches due to its multi-level historical information storage and complex tag matching mechanism, but it also incurs a large energy consumption overhead.
[0034] This hybrid branch predictor combines two predictors, namely the Bimodal and the TAGE. When dealing with complex branches, the TAGE predictor can more accurately capture the behavioral patterns of complex branches by leveraging multi-level history information and a tagging mechanism, thus improving the prediction accuracy compared to a single Bimodal branch predictor. In addition, this hybrid predictor also has a dynamic feedback and adjustment mechanism that can analyze the performance of each branch predictor in real time and flexibly adjust the predictor selection according to the historical behavior of different branch instructions and the current load of the processor. This enables the system to further improve the accuracy of branch prediction when dealing with complex workloads. Therefore, compared with a single Bimodal predictor, the hybrid branch predictor not only improves the accuracy but also has stronger adaptability. When a simple fixed-pattern branch is identified, the Bimodal predictor is preferentially used. The Bimodal predictor has a simple structure and low power consumption, and can quickly give a prediction result, avoiding the high energy consumption of the complex calculations of the TAGE predictor. In addition, the hybrid predictor also incorporates a load monitoring module to detect the workload of the processor in real time. Under high load conditions, the system automatically selects the Bimodal predictor with lower energy consumption to operate, thus minimizing power consumption while maintaining prediction efficiency. This dynamic adjustment mechanism effectively solves the high energy consumption problem of TAGE in complex situations and achieves a balance between performance and energy consumption. Therefore, compared with a pure TAGE predictor, this hybrid branch predictor not only maintains high accuracy but also significantly optimizes energy consumption.
[0035] In addition, when the TAGE branch predictor deals with simple fixed-pattern branches, there is often a phenomenon of resource waste. When dealing with some fixed-pattern branches with simple behavioral rules and strong repeatability, this complex calculation process is unnecessary and will instead lead to excessive occupation of processor resources, increasing power consumption and latency, thus reducing the overall efficiency of the system. The present invention cleverly solves this problem by introducing a branch instruction classification mechanism. First, each branch instruction is analyzed to distinguish whether it is a simple fixed-pattern branch or a complex branch. For those fixed-pattern branches with simple behavioral patterns and low prediction difficulty, the Bimodal predictor is preferentially used for prediction. The Bimodal predictor has a simple structure, a direct prediction logic, and requires extremely few computing resources, and can quickly give a prediction result, which not only improves the prediction efficiency but also significantly reduces the resource consumption of the processor. In summary, the innovation of the present invention lies in comprehensively considering the instruction characteristics and the processor load status, and the intelligent branch prediction strategy improves the accuracy and energy efficiency of branch prediction, effectively enhancing the performance. These improvements not only enhance the flexibility and adaptability of the system but also provide important technical references for the design of future embedded systems and edge devices.
[0036] The method for selecting a branch predictor, the apparatus for selecting a branch predictor, and the computer-readable storage medium proposed in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0037] Figure 1 It is a flowchart of the method for selecting a branch predictor according to the embodiments of the present application.
[0038] As Figure 1 shown, the method for selecting a branch predictor according to the embodiments of the present application may include the following steps:
[0039] S1. Obtain the recorded data of the target jump instruction in the local history register and the global history register based on the target jump instruction.
[0040] S2. Determine the instruction type of the target jump instruction based on the recorded data.
[0041] S3. Determine the load status of the processor based on the operating parameters of the processor.
[0042] S4. Determine the target predictor according to the instruction type and the load status, and output a branch prediction result based on the target predictor.
[0043] Specifically, first, the recorded data of the target jump instruction in the local history register and the global history register can be obtained according to the target jump instruction. Among them, the target jump instruction refers to the branch instruction currently being processed, that is, the instruction in the program that needs to decide whether to jump to another address. The local history register (LHR, local history register) is a register that stores the execution results of the target jump instruction in the recent several times. It records whether the instruction jumps in the past several executions. The information of the local history is usually used to capture the short-term behavior pattern of the instruction. The global history register (GHR, global history register) is a register that stores the execution results of the target jump instruction in a longer time range. It contains the jump history of the instruction during the program execution and is used to capture the long-term behavior pattern of the instruction. Obtaining the recorded data involves extracting the data related to the target jump instruction from the LHR and GHR. These data will be used to analyze the historical behavior of the instruction in order to predict its behavior when it is executed in the future.
[0044] Suppose there is a branch instruction whose position in the program is fixed, and one wants to predict whether it will jump during the next execution. To do this, one can look at the instruction's past execution history: For example, a local history register might record the jump situations of this instruction in the past 4 executions, such as: jump, no jump, jump, jump. This information can help understand the short-term behavior trend of the instruction. A global history register might record the jump situations of this instruction in the past 100 executions, such as: 70 jumps and 30 no jumps. This information can help understand the long-term behavior pattern of the instruction. By analyzing these recorded data, the branch predictor can make more accurate predictions. For example, if the LHR shows that the instruction always jumps in the recent several executions, and the GHR shows that the instruction also tends to jump in a longer history, then the predictor might predict that this instruction will also jump during the next execution. Such a prediction can help the processor prepare the instructions at the jump target address in advance, thus reducing the time waiting for the branch result and improving the execution efficiency of the processor.
[0045] After determining the recorded data, the instruction type of the target jump instruction can be determined based on the recorded data. For example, by analyzing the recorded data in the LHR and GHR, determine whether the target jump instruction is a simple instruction type or a complex instruction type. For example, if the recorded data shows that the instruction behavior is consistent or regular, it is classified as a simple instruction type; if the recorded data shows inconsistent or irregular behavior, it is a complex instruction type. Suppose a jump instruction shows that it has jumped in the past 5 executions in the LHR, but shows 70 jumps and 30 no jumps in the past 100 executions in the GHR. In this case, the LHR shows consistency, but the GHR shows a certain degree of complexity. If the consistency of the LHR is sufficient to cover the complexity of the GHR, this instruction might be classified as a simple instruction type.
[0046] Moreover, the load status of the processor can also be determined based on the running parameters of the processor. For example, by monitoring the running parameters of the processor, such as the instruction issue rate, execution speed, cache hit rate, and pipeline status, to evaluate the current load status of the processor. If any of these parameters exceeds or is lower than a preset threshold, or the pipeline status is blocked, the processor might be in a high-load state. Or the load status of the processor can also be determined based on the number of tasks waiting to be executed in the task queue (or scheduling queue) of the processor. If the length of the task queue continues to increase or remains at a high level, this might indicate that the processor cannot process all tasks in a timely manner and is in a high-load state.
[0047] After determining the load status, the target predictor can be determined according to the instruction type and the load status, so as to output a branch prediction result based on the target predictor. That is to say, the most suitable branch predictor can be selected according to the instruction type and the processor load status. For example, if the instruction type is complex and the processor load is high, a predictor with high accuracy but possibly high power consumption may be selected. Then, the selected predictor will output a branch prediction result to guide whether the processor should jump. If the instruction type is simple, or the instruction type is complex and the processor load status is high, a low-power predictor may be selected, and the selected predictor will output a branch prediction result to guide whether the processor should jump. Thus, the branch predictor can intelligently select the most suitable prediction strategy according to the characteristics of the instruction and the current state of the processor to improve the efficiency and performance of the processor.
[0048] According to an embodiment of the present application, determining the target predictor according to the instruction type and the load status includes: when the instruction type is a complex instruction type and the load status is a low load, determining the target predictor as the first predictor; when the instruction type is a simple instruction type, determining the target predictor as the second predictor; when the instruction type is a complex instruction type and the load status is a high load, determining the target predictor as the second predictor; wherein, the first predictor is a TAGE predictor, and the second predictor is a Bimodal predictor.
[0049] Specifically, when determining the target predictor according to the instruction type and the load status, the current instruction type and the load status are judged. When the instruction type is a simple type, that is, those instructions with relatively fixed behavior patterns and easy to predict, regardless of whether the load status is a low load status or a high load status, the target predictor can be determined as the second predictor (Bimodal predictor). This is because the Bimodal predictor has a simple structure and low power consumption, and is accurate enough for simple fixed-mode branch instructions, while reducing the energy consumption of the processor.
[0050] When the instruction type is a complex instruction type and the load status is a low load, that is, those instructions with behavior patterns that are not easy to predict and require more complex historical information to improve the prediction accuracy, and the current load of the processor is low, the target predictor can be determined as the first predictor (TAGE predictor). This is because the TAGE predictor can handle complex branch behaviors and provide a higher prediction accuracy.
[0051] When the instruction type is a complex instruction type and the load state is a high load, that is, it is recognized that the current branch instruction belongs to the complex instruction type and the current load of the processor is high, the target predictor can be determined as the second predictor (Bimodal predictor). This is because in the high load state, the processor needs to respond quickly, and using the Bimodal predictor can reduce the power consumption of the processor while ensuring a certain prediction accuracy..
[0052] Thus, it is possible to intelligently select an appropriate branch predictor. This dynamic selection mechanism allows the processor to adaptively adjust under different working conditions to achieve the best balance between performance and power consumption. That is, this hybrid branch predictor can flexibly select the most suitable predictor according to the complexity of the instruction and the load state of the processor, thereby optimizing power consumption and improving system performance while ensuring prediction accuracy.
[0053] According to an embodiment of the present application, the operating parameters of the processor include at least one of the instruction issue rate, execution speed, cache hit rate, and pipeline state. Determining the load state of the processor based on the operating parameters of the processor includes: determining that the load state is a high load when the instruction issue rate is greater than a preset issue rate threshold, or the execution speed is greater than a preset speed threshold, or the cache hit rate is less than a preset hit rate threshold, or the pipeline state is a blocked state; determining that the load state is a low load when the instruction issue rate is less than or equal to the preset issue rate threshold, the execution speed is less than or equal to the preset speed threshold, the cache hit rate is greater than or equal to the preset hit rate threshold, and the pipeline state is not a blocked state. Among them, the preset issue rate threshold and the preset hit rate threshold can be determined according to the actual situation.
[0054] Specifically, the operating parameters of the processor include one or more of the instruction issue rate, execution speed, cache hit rate, and pipeline state. Among them, the instruction issue rate refers to the number of instructions issued by the processor per clock cycle. A high issue rate usually means that the processor is executing a large number of instructions. For example, the number of instructions issued per clock cycle can be tracked by a counter to monitor the instruction issue rate. The execution speed refers to the time from when an instruction enters the execution stage to its completion. A longer execution time may indicate that the processor is handling complex or resource-intensive tasks. For example, the time difference between when an instruction enters the execution stage and when it completes can be recorded by a timer to measure the execution speed. The cache hit rate refers to the ratio of the successful retrieval of the required data when the processor accesses the cache. A low cache hit rate may result in more memory accesses, thereby increasing the load on the processor. For example, during each memory access, the hit and miss situations of the cache can be monitored, and a counter integrated in the cache controller is used to record the number of these hits and misses to determine the cache hit rate based on the ratio of the number of hits to the sum of the number of hits and misses. The pipeline state refers to the operating state of the processor pipeline. The blocked state means that some stages in the pipeline are paused due to waiting for data or resources. For example, by inserting monitoring logic at each stage of the pipeline, the register state and control signals of the pipeline can be detected to identify pipeline stalls, blocks, or data dependencies.
[0055] When determining the load status of the processor based on its operating parameters, the results provided by the workload analyzer can be compared with a preset load threshold to determine whether the current load of the processor meets the criteria of "high load" or "low load". For example, if the instruction issue rate is greater than the preset issue rate threshold, such as the instruction issue rate exceeds 80%-90% of the maximum processable capacity designed for the processor, then the processor is considered to be in a high-load state. Or, if the execution speed is greater than the preset speed threshold, such as the average execution time of a certain instruction is significantly higher than its historical benchmark, then the processor is considered to be in a high-load state. Or, if the cache hit rate is less than the preset hit rate threshold, such as when the cache hit rate is lower than 70%-75%, it can be determined that the processor is in a high-load state, which means more memory accesses and potential latency. Or, when the pipeline state is in a blocked state, that is, some stages in the pipeline are paused due to waiting for data or resources, this also indicates that the processor is in a high-load state. For example, assume that a processor can issue 1000 instructions per second, and the preset issue rate threshold is 800 instructions / second. If the actual issue rate is 850 instructions / second, exceeding the preset issue rate threshold, then the processor is considered to be in a high-load state.
[0056] Conversely, if the instruction issue rate is less than or equal to a preset issue rate threshold, the execution speed is less than or equal to a preset speed threshold, the cache hit rate is greater than or equal to a preset hit rate threshold, and the pipeline state is not a blocked state, it can be determined that the load state is a low load. For example, if the instruction issue rate of a processor is 700 instructions per second (less than or equal to the preset issue rate threshold), the execution speed is within the normal range (less than or equal to the preset speed threshold), the cache hit rate is 85% (greater than or equal to the preset hit rate threshold), and the pipeline is not blocked, then this processor is considered to be in a low load state.
[0057] Thus, the processor can dynamically adjust the branch predictor according to the current workload to achieve the best balance between performance and energy consumption. This adaptive mechanism enables the processor to operate efficiently under different working conditions.
[0058] According to an embodiment of the present application, determining the instruction type of a target jump instruction based on recorded data includes: when the recorded data of the target jump instruction in the local history register are all jump or non-jump instructions, or when the recorded data of the target jump instruction in the global history register are regular, determining that the instruction type is a simple instruction type; when the recorded data of the target jump instruction in the local history register simultaneously include jump and non-jump instructions, or when the recorded data of the target jump instruction in the global history register are irregular, determining that the instruction type is a complex instruction type.
[0059] Specifically, when determining the instruction type of a target jump instruction based on the recorded data, the recorded data are judged. When the recorded data of the target jump instruction in the local history register are all jump or non-jump instructions, or when the recorded data of the target jump instruction in the global history register are regular, it can be determined that the instruction type is a simple instruction type. That is to say, if the historical data of the target jump instruction recorded in the local history register shows that this instruction always jumps or always does not jump in the past, it is considered to be a simple instruction type, which means that the behavior pattern of the instruction is consistent and easy to predict. In addition, when the data recorded in the global history register has recognizable regularity, for example, the jump behavior of the instruction follows a certain periodicity or sequence pattern, this indicates that although the behavior of the instruction may not be completely consistent, it has a certain predictability, so it is classified as a simple instruction type.
[0060] When the recorded data of the target jump instruction in the local history register has both jump and non-jump instructions, or when the recorded data of the target jump instruction in the global history register is irregular, the instruction type can be determined as a complex instruction type. That is to say, if the data recorded in the local history register indicates that the historical behavior of the target jump instruction changes between jump and non-jump, without an obvious consistent pattern, it is considered a complex instruction type, which indicates that the behavior pattern of the instruction is unstable and difficult to predict. In addition, if the data recorded in the global history register has no obvious regularity, that is, the jump behavior of the instruction seems to be random and has no recognizable pattern, it indicates that the behavior of the instruction is difficult to predict and is therefore classified as a complex instruction type.
[0061] Thus, when the recorded data in the local history register or the global history register shows that the instruction behavior is consistent or regular, the processor classifies the instruction as a simple instruction type. This classification helps to select a branch predictor suitable for simple and predictable behavior, such as a Bimodal predictor. When the recorded data in the local history register or the global history register shows that the instruction behavior is inconsistent or irregular, the processor classifies the instruction as a complex instruction type. This classification helps to select a branch predictor that can handle complex and unpredictable behavior, such as a TAGE predictor.
[0062] Suppose the record of a jump instruction in the local history register shows that it always jumps in the past 10 executions. In this case, even if the global history register shows a more complex pattern, the consistency of the local history register is sufficient to classify the instruction as a simple instruction type. Another example is that the record of a jump instruction in the global history register shows that it follows a specific periodic pattern, such as jumping every 5 executions. Then, even if the local history register shows a mixed jump and non-jump behavior, the regularity of the global history register is sufficient to classify the instruction as a simple instruction type. By this method, the processor can intelligently select the most suitable branch predictor according to the historical behavior of the instruction, thereby improving the prediction accuracy and the overall performance of the processor.
[0063] According to an embodiment of the present application, the method for selecting a branch predictor further includes: determining the prediction accuracy rate based on the branch prediction result and the actual result of the target jump instruction; and when the prediction accuracy rate corresponding to the target jump instruction is lower than a preset accuracy rate threshold, determining that the target predictor corresponding to the next target jump instruction is the first predictor. The preset accuracy rate threshold can be determined according to the actual situation.
[0064] Specifically, in the embodiments of the present application, to continuously improve the prediction performance, the performance of the branch predictor is monitored in real time through a dynamic feedback and adjustment mechanism, and the prediction accuracy is continuously optimized. Through this adaptive mechanism, continuous adjustment can be made according to the actual operating conditions to ensure that the branch prediction performance is always in the best state under different workloads and instruction types, thus improving the overall efficiency of the processor. That is, the prediction accuracy rate can be determined based on the branch prediction result and the actual result of the target jump instruction. That is to say, after each target jump instruction is executed, the processor compares the prediction result provided by the branch predictor with the actual jump result of the instruction (i.e., whether it actually jumps). If the prediction result is consistent with the actual result (for example, the predictor predicts a jump and it actually jumps, or the predictor predicts no jump and it actually does not jump), it is counted as an accurate prediction. If the prediction result is inconsistent with the actual result, it is counted as an incorrect prediction.
[0065] And the processor maintains an accuracy counter for each branch predictor to count the number of correct predictions and the total number of predictions of the predictor. That is, the accuracy rate can be determined by the following formula: accuracy rate = number of correct predictions / total number of predictions × 100%. In addition, a preset accuracy threshold is set for each predictor, and this threshold is determined based on system performance requirements and power consumption considerations. Regularly or after each prediction, it is checked whether the accuracy rate of the predictor corresponding to the target jump instruction is lower than the preset threshold. If the accuracy rate of the predictor corresponding to the target jump instruction is lower than the preset threshold, the target predictor corresponding to the next target jump instruction may be determined as the first predictor (for example, the TAGE predictor) because the first predictor usually has higher prediction accuracy. If the accuracy rate is within an acceptable range, the system may continue to use the current predictor or select a second predictor (for example, the Bimodal predictor) based on other factors (such as power consumption).
[0066] For example, assume that a branch predictor has 85 correct predictions in the last 100 predictions. Its accuracy rate is 85%. If the preset accuracy threshold is 80%, then the performance of this predictor is acceptable. If another branch predictor has only 70 correct predictions in the last 100 predictions, with an accuracy rate of 70%, which is lower than the preset 80% threshold. In this case, the system will adjust the predictor and assign the next prediction task to the first predictor. Thus, through this dynamic feedback and adjustment mechanism, the processor can optimize the selection of the branch predictor according to the actual prediction performance, so as to keep the branch prediction performance always in the best state under different workloads and instruction types, and improve the overall efficiency of the processor.
[0067] Further, according to an embodiment of the present application, determining the prediction accuracy rate based on the branch prediction result and the actual result of the target jump instruction includes: when the branch prediction result is consistent with the actual result of the target jump instruction, controlling the first counter to count; when the branch prediction result is inconsistent with the actual result of the target jump instruction, controlling the second counter to start counting; determining the prediction accuracy rate according to the ratio of the count value of the first counter to the sum of the count values of the first counter and the second counter.
[0068] Specifically, when determining the prediction accuracy rate based on the branch prediction result and the actual result of the target jump instruction, when the branch prediction result is consistent with the actual result of the target jump instruction, controlling the first counter to count, that is, the first counter is used to record the number of correct predictions. Whenever the branch prediction result is consistent with the actual result of the target jump instruction, this counter will increase. When the branch prediction result is inconsistent with the actual result of the target jump instruction, controlling the second counter to start counting, that is, the second counter is used to record the number of incorrect predictions. Whenever the branch prediction result is inconsistent with the actual result of the target jump instruction, this counter will increase. That is, when the branch predictor predicts that a jump instruction will jump and it actually jumps during execution, or predicts no jump and actually does not jump, in this case, the first counter counts. When the branch predictor predicts that a jump instruction will jump, but it does not jump during actual execution, or predicts no jump but actually jumps, in this case, the second counter starts counting.
[0069] The prediction accuracy rate is determined by comparing the count value of the first counter with the sum of the count values of the two counters, that is, prediction accuracy rate = count value of the first counter / (count value of the second counter + count value of the first counter) * 100%. The closer this ratio is to 1, the higher the prediction accuracy rate; the closer it is to 0, the lower the prediction accuracy rate. Suppose that within a certain period of time, the first counter records 90 correct predictions and the second counter records 10 incorrect predictions. Then, the prediction accuracy rate can be determined to be 90%. This means that within this period of time, the prediction accuracy rate of the branch predictor is 90%, which is a very high accuracy rate, indicating that the predictor has good performance.
[0070] Thus, the prediction accuracy rate can be used to evaluate the performance of the current branch predictor and as a basis for selecting or switching to a different predictor. If the accuracy rate is lower than the preset threshold, it may be selected to switch to another predictor to improve the overall prediction performance. Through this method, the processor can monitor and adjust the performance of the branch predictor in real time to ensure that the accuracy of branch prediction is optimized under different workloads and instruction types.
[0071] The following combines Figures 2 to 4 to describe the method of the present application.
[0072] As a specific example, as Figure 2 shown, first, the branch instruction classification module is responsible for receiving the PC (program counter) and the branch prediction result. Through the classification algorithm, it analyzes the historical behavior of the instruction to determine whether it is a simple fixed-mode branch or a complex branch. This module classifies by reading the LHR (local history register) and the GHR (global history register), and records the credibility of the prediction result in the classification feedback adjustment module for subsequent adjustment. At the same time, the load monitoring module can detect the workload of the processor in real time, including the instruction issue rate, execution speed, cache hit rate, and pipeline status. This module compares the monitoring result with the preset standard through a threshold comparator to judge the current load status and outputs a logic level to indicate whether the processor is in a high-load or low-load state. The predictor selection mechanism intelligently selects an appropriate branch predictor based on the outputs of the classification module and the load detection module. Finally, the dynamic feedback and adjustment mechanism records the accuracy of each predictor in real time, analyzes the performance of different instructions, and adjusts the use of the predictor or optimizes the classification algorithm according to the feedback result to ensure the highest prediction accuracy. This mechanism enables the entire system to adaptively adjust under different working conditions, thus achieving the best balance between performance and energy consumption.
[0073] As Figure 3 shown, Figure 3The presented branch instruction classification module is responsible for identifying and classifying the input branch instructions to optimize subsequent branch prediction. This module analyzes historical behavior data by receiving the PC and prediction results to determine whether the instruction is a simple fixed-pattern branch or a complex branch. Through a feedback adjustment mechanism, it can continuously optimize the classification results based on the success or failure of the prediction and finally transmit the classification information to the predictor selection mechanism. The design of this module aims to improve the accuracy of branch prediction, thereby enhancing the processor performance and energy efficiency. The branch instruction classification module can be divided into the following four sub-modules: 1 Input interface: Receives the PC of the instruction and the branch prediction result; 2 Classification algorithm module: According to the PC of the extracted instruction, reads the content stored in the corresponding local history register (LHR) and global history register (GHR), and conducts instruction classification analysis. If the behaviors recorded in the LHR are all jumps or non-jumps, or the behavior pattern recorded in the GHR, it returns that the instruction type is a simple fixed-pattern branch; otherwise, it returns that the instruction type is a complex branch; 3 Branch type feedback adjustment module: Based on the feedback from the classification algorithm module and the branch prediction result, records the credibility of the current instruction classification result. If the branch prediction fails, it records and adjusts the classification result to the opposite type in the next classification; 4 Output interface: Transmits the classification result to the predictor selection mechanism. If it is a simple fixed-pattern branch, it outputs a logic level 1; if it is a complex branch, it outputs a logic level 0.
[0074] Figure 4The presented workload monitoring module is responsible for evaluating the processor's workload in real time to dynamically adjust the selection of the branch predictor. This module determines whether the current workload is high or low by monitoring multiple metrics such as instruction issue rate, execution speed, cache hit rate, and pipeline status. Its output result will directly affect the predictor selection mechanism, so that a predictor with lower power consumption is preferentially used under high load to improve energy efficiency; while under low load, it switches back to a predictor with high accuracy. This design aims to achieve the best balance between performance and energy consumption and ensure the efficient operation of the processor under different working conditions. The workload monitoring module plays a crucial role in the hybrid branch predictor design and consists of the following key parts: 1 Input monitoring interface: The input to monitor the processor's status data includes the following aspects: The instruction issue rate is monitored by tracking the number of instructions issued in each clock cycle through a counter. The execution speed of the instruction is measured by recording the time difference between the instruction entering the execution stage and the completion stage through a timer. Each time a memory access occurs, the cache hit and miss situations are monitored, and the counters integrated in the cache controller are used to record the number of these hits and misses. Monitoring logic is inserted at each stage of the pipeline to detect the register status and control signals of the pipeline in order to identify pipeline stalls, blocks, or data dependencies. 2 Threshold comparator: Compares the result provided by the workload analyzer with a preset workload threshold to determine whether the current processor's workload reaches the criteria of "high load" or "low load". The following situations are considered that the processor is in a high load state, otherwise it is in a low load state: The instruction issue rate exceeds 80%-90% of the maximum processable capacity designed by the processor, and the average execution time of a certain instruction is significantly higher than its historical benchmark, and the cache hit rate is lower than 70%-75%, and in a certain clock cycle, multiple stages in the pipeline are in a blocked state. 3 Output control interface: According to the comparison result, if a certain metric in the threshold comparator reaches the high load threshold, it is determined that the processor is in a high load state, otherwise it is in a low load state. Output the processor's load condition and pass it to the predictor selection mechanism: If it is in a high load state, output a logic level 1; if it is in a low load state, output a logic level 0.
[0075] Thus, the predictor selection mechanism intelligently selects an appropriate branch predictor based on the output of the branch instruction classification module. This mechanism first identifies the instruction type to select the corresponding branch predictor. In addition, this mechanism also combines the output of the load monitoring module and automatically adjusts the predictor selection according to the load status of the processor, thereby achieving optimization of performance and energy consumption. For example, for simple fixed-mode branches, the Bimodal branch predictor is used. This predictor is based on a single historical information and makes predictions through direct indexing. It is suitable for frequently occurring simple branches, has low power consumption, and a simple design. For complex branches, the TAGE branch predictor is used. This predictor utilizes a multi-level history storage and tagging mechanism, can efficiently handle complex branch behaviors, and provides a higher accuracy rate.
[0076] In addition, after each branch instruction is executed, the actual branch result is compared with the output of the predictor. If the prediction result is consistent with the actual behavior (such as jump or no jump), it is counted as an accurate prediction; otherwise, it is counted as an incorrect prediction. There is a dedicated accuracy counter inside the system, which is used to count the number of correct predictions and the total number of predictions for each predictor. Through these counters, the accuracy rate of the predictor can be calculated and recorded in real time, the prediction performance of different instructions can be analyzed, and feedback can be provided. If the prediction accuracy rate of a certain instruction is lower than the set threshold, the system will adjust the type of predictor used to the TAGE predictor when the instruction arrives next time to ensure the highest prediction accuracy: if the prediction accuracy rate is lower than the set threshold, the output logic level is 0; otherwise, the output logic level is 1. In summary, this application comprehensively considers the instruction characteristics and the processor load status. The intelligent branch prediction strategy improves the branch prediction accuracy and energy efficiency, effectively improves the performance. These improvements not only enhance the flexibility and adaptability of the system, but also provide an important technical reference for the design of future embedded systems and edge devices.
[0077] In summary, according to the method for selecting a branch predictor in an embodiment of this application, based on the target jump instruction, the recorded data of the target jump instruction in the local history register and the global history register is obtained. Based on the recorded data, the instruction type of the target jump instruction is determined. Based on the operating parameters of the processor, the load status of the processor is determined. According to the instruction type and the load status, the target predictor is determined. Based on the target predictor, a branch prediction result is output. Thus, this method can improve the prediction accuracy, has a stronger adaptability, reduces energy consumption, and achieves the best balance between prediction accuracy and energy efficiency.
[0078] Corresponding to the above embodiment, this application also proposes a device for selecting a branch predictor.
[0079] As Figure 5 shown, the device 100 for selecting a branch predictor in an embodiment of this application includes: a branch instruction classification module 110, a load monitoring module 120, and a branch predictor module 130.
[0080] Among them, the branch instruction classification module 110 is used to obtain the recorded data of the target jump instruction in the local history register and the global history register based on the target jump instruction, and determine the instruction type of the target jump instruction based on the recorded data. The load monitoring module 120 is used to determine the load status of the processor based on the operating parameters of the processor. The branch predictor module 130 is used to determine the target predictor according to the instruction type and the load status, and output a branch prediction result based on the target predictor.
[0081] According to an embodiment of the present application, as Figure 6 shown, the branch predictor module 130 includes: an OR gate unit 131, a first input end of the OR gate unit 131 is connected to an output end of the branch instruction classification module 110, and a second input end of the OR gate unit 131 is connected to an output end of the load monitoring module 120; a NOT gate unit 134, an input end of the NOT gate unit 134 is connected to an output end of the OR gate unit 131, and an output end of the NOT gate unit 134 is connected to an enable end of a first predictor 135; a selection unit 136, a first input end of the selection unit 136 is connected to an output end of the first predictor 135, and a second input end of the selection unit 136 is connected to an output end of a second predictor 133, wherein, when the instruction type is a complex instruction type and the load status is a low load, the enable end of the first predictor 135 is enabled, and the first predictor 135 serves as the target predictor; when the instruction type is a simple instruction type, or when the instruction type is a complex instruction type and the load status is a high load, the enable end of the second predictor 133 is enabled, and the second predictor 133 serves as the target predictor.
[0082] Specifically, the OR gate unit 131 receives two inputs: one is the output of the branch instruction classification module 110, and the other is the output of the load monitoring module 120. That is, if any input is at a high level (logical 1), the output of the OR gate unit 131 will be at a high level, which means that if the instruction type is a simple instruction type, or if the processor is in a high load state, the output of the OR gate will be activated.
[0083] The input end of the NOT gate unit 134 is connected to the output end of the OR gate unit 131, the output end of the NOT gate unit 134 is connected to the enable end of the first predictor 135, and the function of the selection unit 136 is to select the output of the first predictor 135 or the second predictor 133 as the final branch prediction result according to the enable signal. The selection unit 136 has two input ends: the first input end is connected to the output end of the first predictor 135, and the second input end is connected to the output end of the second predictor 133. According to the enable states of the first predictor 135 and the second predictor 133, the selection unit 136 will select the output of one of the predictors as the output of the target predictor.
[0084] If the instruction type is a complex instruction type and the processor is in a low load state, the enable terminal of the first predictor 135 will be enabled, and the first predictor 135 will be used as the target predictor. If the instruction type is a simple instruction type, or the instruction type is complex and the processor is in a high load state, the enable terminal of the second predictor 133 will be enabled, and the second predictor 133 will be used as the target predictor.
[0085] For example, assume that the branch instruction classification module 110 outputs a high level, indicating that the instruction type is a simple instruction type. The output of the OR gate unit 131 will be high level, enabling the second predictor 133. The NOT gate unit 134 will output a low level, disabling the first predictor 135. The selection unit 136 will select the output of the second predictor 135. Assume that the branch instruction classification module 110 outputs a low level, but the load monitoring module 120 outputs a high level, indicating that the instruction type is a complex instruction type but the processor is in a high load state. The same logic applies, and the second predictor 135 will be enabled, disabling the first predictor 135. Assume that both modules output a low level, indicating that the instruction type is a complex instruction type and the processor is in a low load state. In this case, the output of the OR gate unit 131 will be low level, disabling the second predictor 133. The NOT gate unit 134 will output a high level, enabling the first predictor 135, and the selection unit 136 will select the output of the first predictor 135. Thus, this design allows the processor to flexibly select the most suitable branch predictor according to the current instruction type and load state to optimize performance and energy consumption.
[0086] According to an embodiment of the present application, as Figure 6 shown, the branch predictor module 130 further includes: an AND gate unit 137. The first input terminal of the AND gate unit 137 is connected to the output terminal of the OR gate unit 131. The second input terminal of the AND gate unit 137 is used to receive a feedback result, where the feedback result is the prediction accuracy rate of the branch prediction result. In the case where the prediction accuracy rate is lower than a preset accuracy rate threshold, the enable terminal of the first predictor 135 is enabled, and the first predictor 135 is used as the target predictor.
[0087] Specifically, the branch predictor module 130 may further include an AND gate unit 137. The AND gate unit 137 is a logic gate that outputs a high level only when both input terminals are at a high level. This unit is used to combine two conditions to determine whether to enable the second predictor 133 or adjust the selection of the predictor. Its first input terminal is connected to the output terminal of the OR gate unit 131. This output terminal is at a low level when the first predictor 135 is enabled and at a high level when the enable terminal of the first predictor 135 is disabled. The second input terminal is used to receive the feedback result, that is, the prediction accuracy rate of the branch prediction result. This feedback result is used to evaluate the performance of the current predictor. The output terminal of the AND gate unit 137 is connected to the enable terminal of the second predictor 133 and the input terminal of the NOT gate unit 134. This output determines whether the second predictor 133 is enabled and whether it is necessary to adjust the enable state of the first predictor 135 through the NOT gate unit 134. Additionally, the feedback result is based on the prediction accuracy rate of the branch prediction result. If the accuracy rate is lower than the preset accuracy rate threshold, it indicates that the performance of the current predictor is poor and it is necessary to switch to another predictor.
[0088] When the prediction accuracy rate is lower than the preset accuracy rate threshold, the second input terminal of the AND gate unit 137 will receive a low-level signal (indicating poor performance). If the output terminal of the OR gate unit 131 is at a high level, the AND gate unit 137 will output a low level and transmit this low-level signal to the input terminal of the NOT gate unit 134. After receiving the low level output by the AND gate unit 137, the NOT gate unit 134 will output a high level. This high-level signal will enable the enable terminal of the first predictor 135, thereby switching to the first predictor 135 as the target predictor. Thus, when the prediction is inaccurate, it can switch to the first predictor 135 as the target predictor to improve the prediction result.
[0089] It should be noted that for the details not disclosed in the branch predictor selection device of the embodiments of the present application, please refer to the details disclosed in the branch predictor selection method of the embodiments of the present application, and specific details will not be elaborated here.
[0090] According to the branch predictor selection device of the embodiments of the present application, the branch instruction classification module is used to obtain the recorded data of the target jump instruction in the local history register and the global history register based on the target jump instruction, and determine the instruction type of the target jump instruction based on the recorded data. The load monitoring module is used to determine the load status of the processor based on the operating parameters of the processor. The branch predictor module is used to determine the target predictor according to the instruction type and the load status, and output a branch prediction result based on the target predictor. Thus, the device can improve the prediction accuracy, has stronger adaptability, reduces energy consumption, and achieves the best balance between prediction accuracy and energy efficiency.
[0091] Corresponding to the above embodiments, the present application also provides a computer-readable storage medium.
[0092] The computer-readable storage medium of the embodiments of the present application stores a program thereon, and when the program is executed by a processor, the above-mentioned method for selecting a branch predictor is implemented.
[0093] According to the computer-readable storage medium of the embodiments of the present application, by executing the above-mentioned method for selecting a branch predictor, the prediction accuracy can be improved, the adaptability can be enhanced, the energy consumption can be reduced, and the best balance between prediction accuracy and energy efficiency can be achieved.
[0094] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0095] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0096] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0097] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of this application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0098] In this application, unless otherwise clearly specified and defined, terms such as "mounted", "connected", "coupled", "fixed", etc. shall be construed in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal communication of two components or the interaction relationship between two components, unless otherwise clearly defined. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0099] Although the embodiments of this application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting this application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A method for selecting a branch predictor, characterized in that: The method comprises: Acquire the record data of the target jump instruction in the local history register and the global history register based on the target jump instruction; Determining the instruction type of the target jump instruction based on the recorded data; determining a load state of the processor based on an operating parameter of the processor; A target predictor is determined according to the instruction type and the load status, and a branch prediction result is output based on the target predictor.
2. The method for selecting a branch predictor according to claim 1, wherein: Determining a target predictor according to the instruction type and the load state includes: In a case where the instruction type is a complex instruction type and the load state is a low load, determining the target predictor to be a first predictor; In the case where the instruction type is a simple instruction type, determining the target predictor to be a second predictor; When the instruction type is a complex instruction type and the load state is a high load, the target predictor is determined to be the second predictor; wherein the first predictor is a TAGE predictor and the second predictor is a Bimodal predictor.
3. The method for selecting a branch predictor according to claim 1 or 2, characterized in that: The operating parameters of the processor include at least one of an instruction emission rate, an execution speed, a cache hit rate, and a pipeline state. The load state of the processor is determined based on the operating parameters of the processor, including: When the instruction emission rate is greater than a preset emission rate threshold, or the execution speed is greater than a preset speed threshold, or the cache hit rate is less than a preset hit rate threshold, or the pipeline state is a blocking state, determining that the load state is a high load; When the instruction emission rate is less than or equal to the preset emission rate threshold, the execution speed is less than or equal to the preset speed threshold, the cache hit rate is greater than or equal to the preset hit rate threshold, and the pipeline state is not a blocking state, the load state is determined to be low load.
4. The method for selecting a branch predictor according to claim 1 or 2, characterized in that: Determining the instruction type of the target jump instruction based on the recorded data includes: In the case where the recorded data of the target jump instruction in the local history register are all jump or no-jump instructions, or the recorded data of the target jump instruction in the global history register has a regular pattern, determining that the instruction type is a simple instruction type; When the recorded data of the target jump instruction in the local history register contains both jump instructions and non-jump instructions, or the recorded data of the target jump instruction in the global history register is irregular, it is determined that the instruction type is a complex instruction type.
5. The method for selecting a branch predictor according to claim 2, wherein: The method further comprises: Determining a prediction accuracy rate based on the branch prediction result and an actual result of the target jump instruction; When the prediction accuracy rate corresponding to the target jump instruction is lower than a preset accuracy rate threshold, it is determined that the target predictor corresponding to the next target jump instruction is the first predictor.
6. The method for selecting a branch predictor according to claim 5, characterized in that: Determining a prediction accuracy rate based on the branch prediction result and an actual result of the target jump instruction includes: When the branch prediction result is consistent with the actual result of the target jump instruction, controlling the first counter to count; When the branch prediction result is inconsistent with the actual result of the target jump instruction, controlling the second counter to start counting; The prediction accuracy is determined according to a ratio of a count value of the first counter to a sum of count values of the first counter and a second counter.
7. A branch predictor selection device, characterized in that: The device comprises: A branch instruction classification module, used for acquiring the record data of the target jump instruction in the local history register and the global history register based on the target jump instruction, and determining the instruction type of the target jump instruction based on the record data; A load monitoring module, used to determine the load state of the processor based on the operating parameters of the processor; The branch predictor module is used to determine a target predictor according to the instruction type and the load state, and output a branch prediction result based on the target predictor.
8. The branch predictor selection device according to claim 7, characterized in that: The branch predictor module comprises: An OR gate unit, wherein a first input end of the OR gate unit is connected to an output end of the branch instruction classification module, and a second input end of the OR gate unit is connected to an output end of the load monitoring module; A NOT gate unit, wherein the input end of the NOT gate unit is connected to the output end of the OR gate unit, and the output end of the NOT gate unit is connected to the enable end of the first predictor; A selection unit, wherein a first input terminal of the selection unit is connected to an output terminal of the first predictor, and a second input terminal of the selection unit is connected to an output terminal of the second predictor, wherein: When the instruction type is a complex instruction type and the load state is a low load, the enable end of the first predictor is enabled, and the first predictor serves as the target predictor; when the instruction type is a simple instruction type, or the instruction type is a complex instruction type and the load state is a high load, the enable end of the second predictor is enabled, and the second predictor serves as the target predictor.
9. The branch predictor selection device according to claim 8, characterized in that: The branch predictor module also includes: An AND gate unit, wherein the first input end of the AND gate unit is connected to the output end of the OR gate unit, the second input end of the AND gate unit is used to receive a feedback result, and the output end of the AND gate unit is respectively connected to the enable end of the second predictor and the input end of the NOT gate unit, wherein the feedback result is the prediction accuracy of the branch prediction result, and when the prediction accuracy is lower than a preset accuracy threshold, the enable end of the first predictor is enabled, and the first predictor is used as a target predictor.
10. A computer-readable storage medium, characterized in that: A branch predictor selection program is stored thereon, and when the branch predictor selection program is executed by a processor, the branch predictor selection method according to any one of claims 1-6 is implemented.
Citation Information
Cited By
Instruction jump prediction method and device, computer equipment, readable storage medium and program product
CN120596151A
Recognition processing method and system for difficult-to-predict branches
CN120994366A