Single-Layer Perceptron Branch Prediction Method and System in Asynchronous Superscalar Processors

By introducing a single-layer perceptron branch prediction method into an asynchronous superscalar processor, the mechanism of weighted calculation and dynamic update of weights is used to solve the problem of insufficient instruction processing speed and prediction accuracy in complex branch modes, and efficient branch prediction and low-power processor performance are achieved.

CN119806650BActive Publication Date: 2025-07-01LANZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510286782.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-01
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing perceptron branch prediction algorithm needs to improve the instruction processing speed and prediction accuracy in complex branch modes, and the traditional synchronous superscalar processors are limited by the clock frequency increase in processing speed and energy consumption.

Method used

The single-layer perceptron branch prediction method in the asynchronous superscalar processor is adopted. The perceptron branch prediction module loads the historical data of the branch instructions before the instruction is executed, performs weighted calculations and dynamic updates of the weights to realize branch prediction and correction operations.

Benefits of technology

It effectively improves the instruction processing speed and prediction accuracy in complex branch mode, reduces the cost caused by misprediction, reduces power consumption, and adapts to the needs of complex instruction flow environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119806650B_ABST
    Figure CN119806650B_ABST
Patent Text Reader

Abstract

The present invention discloses a single-layer perceptron branch prediction method and system in an asynchronous superscalar processor. In the asynchronous superscalar processor architecture, after receiving a prediction request signal and a data packet sent by the instruction fetch module, the top-level control module of branch prediction sends prediction data to the perceptron branch prediction module. The perceptron branch prediction module loads the historical data of branch instructions and stores branch instruction information before instruction execution, performs weighted calculation using the branch instruction information, and determines whether the predicted branch instruction jumps according to the weighted result. Meanwhile, the historical record and weight of the branch instruction information are updated, and the prediction information is returned to the instruction fetch module for bidirectional instruction fetch. The correction information of various types of instructions in branch prediction is sent by the out module. The present invention is based on asynchronous circuit design. By introducing a dynamic branch prediction method based on perceptron and a bidirectional addressing mechanism, the instruction processing speed and prediction accuracy in complex and dynamic branch patterns are improved, and the power consumption is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of processors, and particularly relates to a single-layer perceptron branch prediction method and system in an asynchronous superscalar processor. Background Art

[0002] Superscalar processors have achieved a significant improvement in processor performance by executing multiple instructions in the same clock cycle. Through multiple pipeline designs, the superscalar architecture realizes instruction-level parallelism, that is, it improves throughput without increasing the processor's main frequency.

[0003] However, traditional synchronous superscalar processors are gradually restricted by the synchronous clock in terms of processing speed and energy consumption. The increase in clock frequency brings higher power consumption and heat generation, making it more difficult to further optimize processor performance. With the increasing demand for high performance and low power consumption, asynchronous superscalar processors have gradually attracted attention. Asynchronous processors do not rely on a global clock signal, but manage the instruction scheduling and parallel processing of execution units based on an event-driven manner. Different from the traditional synchronous architecture, the asynchronous architecture can flexibly schedule instructions, reduce the dependence on the clock signal, and thus reduce the impact of clock jitter and clock skew on performance.

[0004] In the design of asynchronous superscalar processors, efficient branch prediction is of great significance for performance improvement. Traditional branch prediction techniques mainly include static prediction and dynamic prediction. Static prediction relies on fixed rules or simple statistical information to judge branch behavior and does not make adjustments based on runtime data. For example, static prediction usually assumes that all backward jump branch instructions will be executed, while forward jump branches will not be executed. Static prediction is simple to implement, but has a low prediction accuracy when dealing with complex branch patterns and is difficult to meet the high-performance requirements of modern processors. Dynamic prediction records the branch behavior history at runtime and adjusts the prediction strategy according to the previous actual jump situations; the two-bit saturating counter is the most basic dynamic predictor. Each branch instruction corresponds to a two-bit counter, which makes predictions based on the previous branch behaviors. The counter accumulates or decreases when the prediction is correct, which is equivalent to forming a kind of "signal strength" to judge the future branch direction.

[0005] In recent years, in order to improve the prediction accuracy, the algorithm of perceptron branch prediction has been gradually introduced. This method can handle complex long-history branch patterns. The perceptron branch predictor makes predictions about future branch directions through a neural network model, combining branch history data, and updates the weights of the neural network with the actual result of each branch. It is applicable to complex instruction streams and has higher prediction accuracy. However, the existing algorithms of perceptron branch prediction need to improve the instruction processing speed and prediction accuracy in complex branch patterns. Summary of the Invention

[0006] In view of the problems existing in the above-mentioned background art, the object of the present invention is to provide a single-layer perceptron branch prediction method and system in an asynchronous superscalar processor, which reduces the influence of misprediction on the instruction stream under complex branch patterns, effectively improves the instruction processing speed and prediction accuracy, and reduces power consumption at the same time.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] The single-layer perceptron branch prediction method in an asynchronous superscalar processor includes the following steps:

[0009] In the asynchronous superscalar processor architecture, when the branch prediction top-level control module is in the prediction state, after sending a prediction request signal and the corresponding data packet through the instruction fetch module, the branch prediction top-level control module controls the data to be sent to the perceptron branch prediction module for prediction. The perceptron branch prediction module loads the historical data of the branch instruction before the instruction execution, stores the branch instruction information as several historical bits, and performs weighted calculation using the branch instruction information. If the predicted weighted result meets a specific threshold, it is predicted that the branch instruction jumps; otherwise, it is predicted that the branch instruction does not jump. At the same time, the perceptron branch prediction module dynamically updates the historical record and weight of the branch instruction information according to the change of the branch behavior, and returns the prediction information to the branch prediction top-level control module. Then, the branch prediction top-level control module sends the data to the instruction fetch module for two-way instruction fetch. When the instruction fetch module finishes receiving a round of instructions, the branch prediction top-level control module changes its state to the correction state;

[0010] When the branch prediction top-level control module is in the correction state, after sending a prediction correction request signal and the corresponding data packet through the out-of-order module, the branch prediction top-level control module sends the data to the perceptron branch prediction module. The perceptron branch prediction module performs a correction operation on the branch instruction, including modifying the historical record and weight of the branch instruction information. When the out-of-order module issues a correction end signal, the branch prediction top-level control module sends the base address for the next round of instruction fetch to the instruction fetch module. At the same time, the branch prediction top-level control module changes its state to the prediction state.

[0011] Further, the branch prediction top-level control module sends the data to the perceptron branch prediction module to perform branch prediction on the branch instruction by selecting the instruction fetch data branch prediction path. The branch prediction top-level control module sends the data to the perceptron branch prediction module to perform a correction operation on the branch instruction by selecting different out-of-order instruction correction slots. When the out-of-order end event and data occur, the branch prediction top-level control module ends the correction operation by selecting the out-of-order data correction end path and changes the correction state to the prediction state.

[0012] Further, a counter is set in the branch prediction top-level control module to control the conversion between the prediction state and the correction state.

[0013] Furthermore, the branch instruction information includes an instruction sequence number, an instruction type, a original PC value, a next instruction PC value, and a jump PC value.

[0014] The present invention further provides a system for implementing a single-layer perceptron branch prediction method in an asynchronous superscalar processor, including a top-level branch prediction control module, a perceptron branch prediction module, an instruction fetching module, and an outlier module, where:

[0015] The top-level branch prediction control module: is configured to receive a prediction request signal and a corresponding data packet sent by the instruction fetching module, or receive a prediction correction request signal and a corresponding data packet sent by the outlier module, and send data to the perceptron branch prediction module;

[0016] The perceptron branch prediction module: includes a prediction sub-module and a correction sub-module. The prediction sub-module includes a prediction data unpacking module, a B-type prediction module, a Jalr-type prediction module, a Call-type prediction module, and a Ret-type prediction module; the correction sub-module includes a correction data unpacking module, a B-type correction module, and a Jalr-type correction module;

[0017] The prediction data unpacking module is configured to find the first jump instruction from the prediction data packet received by the perceptron branch prediction module, and record the branch instruction information of the jump instruction. The prediction data unpacking module detects whether there is a jump instruction in the data packet according to the instruction type encoding; the data packet to be branch-predicted enters the corresponding prediction sub-module according to the branch instruction type code after being processed by the prediction data unpacking module, performs the prediction of the corresponding jump instruction, and outputs a prediction result;

[0018] The correction data unpacking module is configured to split the instructions in the prediction correction data packet received by the perceptron branch prediction module, and send them to the corresponding correction sub-module according to the branch instruction type code for branch instruction correction.

[0019] Furthermore, the B-type prediction module adopts a dynamic branch prediction model implemented by a single-layer perceptron, including a history record table sub-module, a weight table sub-module, and a prediction logic sub-module. The history record table sub-module internally sets a global history register and a counter, and the global history register records the prediction results of each B-type branch instruction; the weight table sub-module adopts a 256*72 SRAM, and a 72-bit weight value table is stored for each low 8 bits of the PC of a B-type instruction. The 72 bits are divided into 9 8-bit weights for predicting B-type jump instructions; the 72-bit weight value table is fetched from the weight table sub-module through the low 8 bits of the input B-type instruction PC, and is sent to the prediction logic sub-module together with the low 8 bits of the global history register in the history record table sub-module. The prediction logic sub-module divides the 72-bit weight value table into 9 8-bit weight registers, and sets a summation register for predicting B-type jump instructions.

[0020] Further, a two-dimensional historical record table of 16 * 37 bits is set inside the Jalr type prediction module, which is used to record the PC value and the number of times of the correct jump address of each Jalr type instruction.

[0021] Further, the Call type prediction module manages Call type instructions and the Ret type prediction module manages Ret type instructions in the way of simulating a stack. When a Call type instruction appears, the next instruction of the original instruction is pushed onto the stack, and the PC value of the original Call type jump instruction is output. When a Ret type instruction appears, the instruction popped from the top of the stack is output as the predicted jump instruction.

[0022] Compared with the disadvantages and deficiencies of the prior art, the present invention has the following beneficial effects:

[0023] (1) Improve the accuracy of branch prediction: By introducing a dynamic branch prediction method based on a perceptron, after each prediction is completed, the perceptron model updates the weight vector according to the actual jump situation of the branch instruction, so that the predictor can improve the instruction prediction accuracy under complex and dynamic branch patterns;

[0024] (2) Reduce the cost brought by misprediction: In the perceptron branch prediction design of the present invention, through the weighted calculation and weight adjustment mechanism, the number of mispredictions is effectively reduced, and the impact of mispredictions on the pipeline is reduced. It not only improves the throughput of the processor and the instruction processing speed under complex branch patterns, but also reduces the waste of computing resources caused by correction operations and re-fetching due to mispredictions;

[0025] (3) The single-layer perceptron branch prediction system in the asynchronous superscalar processor of the present invention is simpler in structure, lower in power consumption, better in performance, and can meet the requirements of complex instruction stream environments, enabling the processor to have significant advantages in high-performance and low-power application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a schematic diagram of the internal mechanism of the branch prediction top-level control module provided by an embodiment of the present invention;

[0027] Figure 2 It is a schematic diagram of the internal mechanism of the perceptron branch prediction module provided by an embodiment of the present invention;

[0028] Figure 3 It is a schematic diagram of the instruction encoding representation provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0029] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0030] Branch prediction plays an important role in improving the continuity of instruction execution in the processor architecture. Especially in the superscalar architecture, by predicting in advance the execution path of branch instructions, the pipeline stalls caused by branch instructions can be reduced. The present invention proposes a single-layer perceptron branch prediction method in an asynchronous superscalar processor, which determines whether the upcoming branch instruction will jump based on the branch behavior characteristics and historical information of the instruction, and accordingly determines the fetch address of the next instruction. The specific method is as follows:

[0031] In the asynchronous superscalar processor architecture, when the branch prediction top-level control module is in the prediction state, after sending a prediction request signal and the corresponding data packet through the instruction fetch module, the branch prediction top-level control module controls the data to be sent to the perceptron branch prediction module for prediction. To improve the prediction accuracy, the perceptron branch prediction module loads the historical data of branch instructions before instruction execution to identify and learn the rules of branch behavior. During this process, the perceptron branch prediction module stores the branch instruction information as several historical bits. The branch instruction information includes instruction sequence number, instruction type, original PC value, PC value of the next instruction, and jump PC value, and uses this branch instruction information for weighted calculation. If the predicted weighted result meets a specific threshold, it is predicted that the branch instruction will jump; otherwise, it is predicted that the branch instruction will not jump. At the same time, the perceptron branch prediction module dynamically updates the historical record and weight of the branch instruction information according to the change of branch behavior to adapt to the change of branch behavior. The perceptron branch prediction module returns the prediction information to the branch prediction top-level control module, and then the branch prediction top-level control module sends the data to the instruction fetch module for two-way instruction fetch. When the instruction fetch module finishes receiving a round of instructions, the branch prediction top-level control module changes its state to the correction state;

[0032] When the branch prediction top-level control module is in the correction state, after sending a prediction correction request signal and the corresponding data packet through the out-of-order module, the branch prediction top-level control module selects different out-of-order instruction correction slots to send data to the perceptron branch prediction module. The perceptron branch prediction module performs correction operations on the branch instructions, including modifying the historical record and weight of the branch instruction information. When the out-of-order module issues a correction end signal, the branch prediction top-level control module sends the base address for the next round of instruction fetch to the instruction fetch module to start the next round of instruction fetch operation. At the same time, the branch prediction top-level control module changes its state to the prediction state.

[0033] The system for implementing the single-layer perceptron branch prediction method in an asynchronous superscalar processor includes a top-level branch prediction control module, a perceptron branch prediction module, an instruction fetch module, and an out-of-order module.

[0034] The internal mechanism of the top-level branch prediction control module is as Figure 1 shown. The top-level branch prediction control module has two states: the prediction state and the correction state.

[0035] In the prediction state, the top-level branch prediction control module only interacts with the instruction fetch module. When the top-level branch prediction control module receives a data packet from the instruction fetch module, it sends the data to the perceptron branch prediction module for branch prediction through the selected instruction fetch data branch prediction path Fetch_rou, and sends the result to the instruction fetch module after the prediction ends. It judges whether an exception occurs in the instruction fetch module according to the highest bit of the received instruction fetch data packet. If an exception occurs, it does not send events and data to the perceptron branch prediction module, but bypasses the perceptron branch prediction module and directly sends a packet of all zeros to the instruction fetch module. The instruction fetch module is designed with 4 instruction slots with a depth of 16 for two-way instruction fetch. At initialization, the top-level branch prediction control module is in the prediction state, and the counter is initialized to 0. Since the asynchronous superscalar processor pipeline is empty after initialization, the time point for this module to switch to the correction state to process out-of-order events and data is after the instruction fetch module has received three rounds of 64 instructions. Therefore, a counter is set to control the transition between the prediction and correction states.

[0036] In the correction state, the top-level branch prediction control module obtains events and data from the out-of-order module and the exception handling module (only interacts when an exception occurs). After processing, it sends the next instruction fetch address to the instruction fetch module and simultaneously switches to the prediction state. Each out-of-order instruction correction slot (Correchtion_rou1~Correchtion_rou4 respectively) corresponds to 1 path interface. After switching to the correction state, it will receive events and data from at least one of the 4 out-of-order instruction correction slots in the out-of-order, as well as the events and data indicating the end of out-of-order. The strategy of four-instruction-slot correction is adopted, enabling the processor to accurately predict the jump direction of branch instructions within one cycle, thereby fetching more instructions from the instruction cache and improving the instruction throughput. When receiving the events and data of the out-of-order instruction correction slot, it is sent to the perceptron branch prediction module for correction through an asynchronous control chain, and at the same time, the correct address in the data brought by each out-of-order instruction correction slot is overwritten and recorded. Because the data given by out-of-order contains the base address for fetching instructions starting from this out-of-order instruction correction slot, and the order of out-of-order must be in the sequence of Correchtion_rou1~Correchtion_rou4, the finally recorded base address must be the base address for the next instruction fetch (except in exceptional cases). When the events and data indicating the end of out-of-order occur, they are processed according to Figure 1Perform the event chain route of the retirement termination_rou5 for the out-of-round data correction end path, convert the correction status to the prediction status, and select the path to control whether to send the base address of the record through an exception signal.

[0037] The internal mechanism of the perceptron branch prediction module is as Figure 2 shown, including a prediction sub-module and a correction sub-module. The prediction sub-module includes a prediction data unpacking module (Data_Unpacking), a B-type prediction module (B_type), a Jalr-type prediction module (Jalr_type), a Call-type prediction module (Call_type), and a Ret-type prediction module (Ret_type). The correction sub-module includes a correction data unpacking module (Data_Unpacking), a B-type correction module (B_correction), and a Jalr-type correction module (Jalr_correction). When the branch prediction top-level control module sends a prediction request signal and a data packet (Inst_Buffer), the perceptron branch prediction module will select to enter the prediction unit (Prediction_Unit) to perform the prediction process; when the branch prediction top-level control module sends a prediction correction request signal and a data packet (C_Data), the perceptron branch prediction module will select to enter the correction unit (Prediction_Unit).

[0038] The prediction data unpacking module is used to find the first jump instruction from the data packet to be branch-predicted received by the perceptron branch prediction module, and record the branch instruction information of the jump instruction, including the instruction sequence number (index), instruction type (o_type), original PC (o_pc) value, next instruction PC (o_next_pc) value, and jump PC (o_addr) value. The prediction data unpacking module detects whether there is a jump instruction in the data packet according to the instruction type encoding. The specific instruction encoding table is as Figure 3 shown. The detection method is to detect whether the 3-bit instruction encoding of each instruction is one of 000, 001, 010, 011, 100, 111. After the data packet to be branch-predicted is processed by the prediction data unpacking module, it enters the corresponding prediction sub-module according to the branch instruction type code to perform the prediction of the corresponding jump instruction and output the prediction result.

[0039] The B-type prediction module adopts a dynamic branch prediction model implemented by a single-layer perceptron, including a history record table sub-module, a weight table sub-module, and a prediction logic sub-module. Among them, a 20-bit global history register and a counter are set inside the history record table sub-module. The global history register records the prediction results of each B-type branch instruction; the weight table sub-module uses a 256*72 SRAM, and a 72-bit weight value table is stored for each low 8 bits of the PC of a B-type instruction. The 72 bits are divided into 9 8-bit weights for predicting B-type jump instructions; the 72-bit weight value table is fetched from the weight table sub-module through the low 8 bits of the input B-type instruction PC and sent to the prediction logic sub-module together with the low 8 bits of the global history register in the history record table sub-module. The prediction logic sub-module divides the 72-bit weight value table into 9 8-bit weight registers and sets a sum register for predicting B-type jump instructions.

[0040] The Jalr-type prediction module internally sets a 16*37-bit two-dimensional history record table for recording the PC value and the number of times of the correct jump address of each Jalr-type instruction (37 bits = 32-bit PC + 5-bit number of times).

[0041] The Call-type prediction module manages Call-type instructions and the Ret-type prediction module manages Ret-type instructions in a way that simulates a stack. Call-type instructions and Ret-type instructions normally appear in pairs. When a Call-type instruction appears, the next instruction of the original instruction is pushed onto the stack, and the PC value of the original Call-type jump instruction is output. When a Ret-type instruction appears, the instruction popped from the top of the stack is output as the predicted jump instruction.

[0042] The correction data unpacking module is used to split the instructions in the prediction correction data packet received by the perceptron branch prediction module and send them to the corresponding correction sub-module according to the branch instruction type code for correcting the branch instructions.

[0043] The B-type correction module first fetches the 72-bit weight value table and the low 8 bits of the global history register from the weight table and the history record table and sends them to the B-type correction module. Then, the 72-bit global history register value is divided into 9 8-bit registers. Traverse the history register, update the register values, store the updated weight table in the weight table sub-module. The history record table sub-module judges according to the B-type prediction result given by the outlier. If the B-type prediction is correct, the counter is decremented by 1. If the B-type prediction given by the outlier is incorrect, the global history register is shifted to the right by the number of bits of the counter value to restore to the value at the time of predicting the corresponding B-type instruction, and the counter is set to zero.

[0044] The Jalr - type correction module first traverses 16 entries in the historical weight table to find the entry with the smallest lower five - bit value, and records the corresponding frequency and entry number. Sixteen flag registers are set to indicate whether the corrected jump address exists in each entry of the historical record table.

[0045] In the single - layer perceptron branch prediction system of the asynchronous superscalar processor of the present invention, a two - phase single - rail self - timed event - driven circuit based on the bounded data - driven (BBD) handshake protocol - that is, a BBD - type asynchronous circuit is used. Its implementation mechanism is based on the Click controller family in asynchronous circuits, which can effectively achieve clockless control and optimize timing management within a local range. The BBD - type asynchronous circuit is a technology that decomposes the timing - dependence problem, dividing the circuit design structure into a data path and a control path to ensure stable data transmission and operation at different stages.

[0046] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A single-layer perceptron branch prediction method in an asynchronous superscalar processor, characterized in that: The steps include: In an asynchronous superscalar processor architecture, when a branch prediction top-level control module is in a prediction state, after sending a prediction request signal and a corresponding data packet through an instruction fetch module, the branch prediction top-level control module controls the data to be sent to a perceptron branch prediction module for prediction. The perceptron branch prediction module loads the historical data of branch instructions before the instruction is executed, stores the branch instruction information as a number of historical bits, and uses the branch instruction information for weighted calculation. If the predicted weighted result meets a specific threshold, the branch instruction is predicted to jump, otherwise the branch instruction is predicted not to jump. At the same time, the perceptron branch prediction module dynamically updates the historical record and weight of the branch instruction information according to the change of branch behavior, and returns the prediction information to the branch prediction top-level control module, which then sends data to the instruction fetch module for bidirectional instruction fetching. The instruction fetch module is designed with 4 16-depth instruction slots for bidirectional instruction fetching. When the instruction fetch module receives a full round of instructions, the branch prediction top-level control module switches to a correction state. When the branch prediction top-level control module is in the correction state, after sending the prediction correction request signal and the corresponding data packet through the outgoing module, the branch prediction top-level control module sends data to the perceptron branch prediction module, and the perceptron branch prediction module performs correction operations on the branch instruction, including modifying the historical records and weights of the branch instruction information. When the outgoing module sends a correction end signal, the branch prediction top-level control module sends the base address of the next round of instruction fetching to the instruction fetching module, and at the same time, the branch prediction top-level control module switches the state to the prediction state; The branch prediction top-level control module selects the instruction fetch data branch prediction path to send data to the perceptron branch prediction module to perform branch prediction on the branch instruction. The branch prediction top-level control module selects different out-of-order instruction correction slots to send data to the perceptron branch prediction module to perform correction operations on the branch instruction. After switching to the correction state, at least the event and data of one of the four out-of-order instruction correction slots and the out-of-order end event and data will be received from the out-of-order. When the out-of-order end event and data occur, the branch prediction top-level control module ends the correction operation by selecting the out-of-order data correction end path, and switches the correction state to the prediction state.

2. The single-layer perceptron branch prediction method in an asynchronous superscalar processor according to claim 1, characterized in that: A counter is set in the branch prediction top-level control module, and the conversion between the prediction state and the correction state is controlled by the counter.

3. The single-layer perceptron branch prediction method in an asynchronous superscalar processor according to claim 1, characterized in that: The branch instruction information includes instruction number, instruction type, original PC value, next instruction PC value, and jump PC value.

4. A single-layer perceptron branch prediction system in an asynchronous superscalar processor that implements the single-layer perceptron branch prediction method in an asynchronous superscalar processor according to any one of claims 1 to 3, characterized in that: It includes a branch prediction top-level control module, a perceptron branch prediction module, an instruction fetch module and an exit module, among which: Branch prediction top-level control module: used to receive the prediction request signal and the corresponding data packet sent by the instruction fetch module, or receive the prediction correction request signal and the corresponding data packet sent by the exit module, and send data to the perceptron branch prediction module; Perceptron branch prediction module: includes a prediction submodule and a correction submodule, wherein the prediction submodule includes a prediction data unpacking module, a B-type prediction module, a Jalr-type prediction module, a Call-type prediction module, and a Ret-type prediction module; the correction submodule includes a correction data unpacking module, a B-type correction module, and a Jalr-type correction module; The prediction data unpacking module is used to find the first jump instruction from the prediction data packet received by the perception machine branch prediction module, and record the branch instruction information of the jump instruction. The prediction data unpacking module detects whether the data packet contains the jump instruction according to the instruction type code; the data packet to be branch predicted enters the corresponding prediction submodule according to the branch instruction type code after being processed by the prediction data unpacking module, predicts the corresponding jump instruction and outputs the prediction result; The correction data unpacking module is used to split the instructions in the prediction correction data packet received by the perception machine branch prediction module, and send them to the corresponding correction submodule according to the branch instruction type code to correct the branch instruction.

5. The single-layer perceptron branch prediction system in an asynchronous superscalar processor according to claim 4, characterized in that: The type B prediction module adopts a dynamic branch prediction model implemented by a single-layer perceptron, including a history record table submodule, a weight table submodule and a prediction logic submodule. A global history register and a counter are set inside the history record table submodule, and the global history register records the prediction results of each type B branch instruction; the weight table submodule adopts a 256*72 SRAM, and the lower 8 bits of the PC of each type B instruction store a 72-bit weight value table, and the 72 bits are divided into 9 8-bit weights for predicting type B jump instructions; the 72-bit weight value table is taken out from the weight table submodule through the lower 8 bits of the PC of the input type B instruction, and is sent to the prediction logic submodule together with the lower 8 bits of the global history register in the history record table submodule, and the prediction logic submodule divides the 72-bit weight value table into 9 8-bit weight registers, and sets a summation register to predict type B jump instructions.

6. The single-layer perceptron branch prediction system in an asynchronous superscalar processor according to claim 4, characterized in that: A 16*37-bit two-dimensional history record table is set inside the Jalr type prediction module to record the PC value and number of correct jump addresses of each Jalr type instruction.

7. The single-layer perceptron branch prediction system in an asynchronous superscalar processor according to claim 4, characterized in that: The Call type prediction module manages the Call type instructions and the Ret type prediction module manages the Ret type instructions using a simulated stack. When a Call type instruction appears, the next instruction of the original instruction is pushed onto the stack and the value of the PC of the original Call type jump instruction is output. When a Ret type instruction appears, an instruction is popped from the top of the stack and output as a predicted jump instruction.

Citation Information

Patent Citations

  • Branch prediction method and branch predictor applied to processor

    CN115686639A

  • Branch prediction method, electronic equipment and storage medium

    CN117234594A