Branch predictor, operation method, processor, electronic device and storage medium

By setting a unique label offset value for each thread, the tag prediction table shared by multiple threads corresponds to different storage areas, the problem of branch prediction accuracy degradation in a multi-threaded environment is solved, and efficient branch prediction and multi-user control flow isolation is achieved.

CN120179293AActive Publication Date: 2025-06-20HYGON INFORMATION TECH CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202510299681.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-20
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

In a multi-threaded environment, existing branch predictors have reduced branch prediction accuracy due to capacity competition and interference between threads, and it is difficult to achieve isolation of control flow by multiple users, affecting the security and reliability of the system.

Method used

By setting a unique label offset value for each thread, the label prediction table shared by multiple threads corresponds to different storage areas, thereby alleviating capacity competition, reducing inter-thread interference, and optimizing the use of storage areas based on demand information by adjusting the label offset value.

Benefits of technology

It improves branch prediction accuracy, reduces interference between threads, realizes the isolation of control flow by multiple users, and enhances the security and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179293A_ABST
    Figure CN120179293A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a branch predictor for multiple threads, an operation method, a processor, electronic equipment and a storage medium. The branch predictor comprises a plurality of label prediction tables of different levels, a plurality of threads share the plurality of label prediction tables, and the operation method comprises the following steps: setting label offset values for the plurality of threads; and in response to an operation of a target thread in the plurality of threads on the plurality of label prediction tables, determining a storage area, corresponding to the target thread, of each label prediction table based on a target label deviation value corresponding to the target thread and respective prediction table identifiers of the plurality of label prediction tables. According to the operation method, the problem of capacity competition of different threads for the same branch predictor can be relieved, the interference between the threads is reduced, and the branch prediction precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a branch predictor for multiple threads, an operation method thereof, a processor, an electronic device, and a storage medium. Background Art

[0002] In a high-performance out-of-order execution processor, an accurate branch predictor plays a crucial role in maximizing the throughput of the processor. It can prospectively guess the branch decisions on the program execution path, thereby allowing the processor to start executing the predicted instruction stream before actually determining the branch condition. In this way, the branch predictor can effectively avoid pipeline stalls and interruptions in the instruction execution sequence caused by waiting for branch results, thus significantly improving the overall instruction-level parallelism and the utilization rate of processor resources, ensuring that it continuously operates at a high efficiency. Especially when facing a large amount of code that depends on branch logic, accurate branch prediction is a key factor in system performance optimization. Summary of the Invention

[0003] At least one embodiment of the present disclosure provides an operation method for a branch predictor for multiple threads. The branch predictor includes multiple tag prediction tables of different levels, and multiple threads share the multiple tag prediction tables. The operation method includes: setting tag offset values for the multiple threads; and in response to an operation of a target thread among the multiple threads on the multiple tag prediction tables, determining a storage area corresponding to each tag prediction table and the target thread based on the target tag offset value corresponding to the target thread and the prediction table identifiers of the multiple tag prediction tables.

[0004] For example, in the operation method provided by at least one embodiment of the present disclosure, the branch predictor includes N tag prediction tables and N storage areas, and each storage area stores 2m table entries respectively, where N and m are positive integers. Determining a storage area corresponding to each tag prediction table and the target thread based on the target tag offset value corresponding to the target thread and the prediction table identifiers of the multiple tag prediction tables includes: for a first tag prediction table among the N tag prediction tables, in response to (Tage_Offset + Table_id) ≤ N, determining that the area identifier of the storage area corresponding to the first tag prediction table and the target thread is (Tage_Offset + Table_id); or, in response to (Tage_Offset + Table_id) > N, determining that the area identifier of the storage area corresponding to the first tag prediction table and the target thread is (Tage_Offset + Table_id - N), where Table_id is the prediction table identifier of the first tag prediction table, and Tage_Offset is the target tag offset value.

[0005] For example, in the operation method provided by at least one embodiment of the present disclosure, the operations of the target thread on multiple tag prediction tables include query operations. After determining the storage area corresponding to each tag prediction table and the target thread, the operation method further includes: reading the entries stored in the storage area corresponding to each tag prediction table and the target thread to query whether the branch instruction to be predicted by the target thread hits N tag prediction tables.

[0006] For example, the operation method provided by at least one embodiment of the present disclosure further includes: in response to the branch instruction to be predicted by the target thread hitting the first entry in at least one tag prediction table, taking the tag prediction table with the highest rank among the at least one hit tag prediction tables as the target tag prediction table, and reading the prediction result of the first entry located in the storage area corresponding to the target tag prediction table and the target thread to obtain the target prediction result.

[0007] For example, the operation method provided by at least one embodiment of the present disclosure further includes: in response to the target prediction result being incorrect, updating the first entry in the at least one hit tag prediction table based on the actual result of the branch instruction to be predicted, and applying for and allocating a new entry in the tag prediction table at a level higher than the target tag prediction table to store the instruction information and the actual result corresponding to the branch instruction to be predicted.

[0008] For example, the operation method provided by at least one embodiment of the present disclosure further includes: in response to the target prediction result being incorrect and the target tag prediction table having the highest rank among the multiple tag prediction tables, updating the first entry in the at least one hit tag prediction table based on the actual result of the branch instruction to be predicted.

[0009] For example, the operation method provided by at least one embodiment of the present disclosure further includes: recording the demand information of multiple threads for each tag prediction table, and adjusting the tag offset values of the multiple threads based on the demand information.

[0010] For example, in the operation method provided by at least one embodiment of the present disclosure, the demand information includes the number of times each tag prediction table is applied for and allocated a new entry. Adjusting the tag offset values of the multiple threads based on the demand information includes: in response to the number of times at least one tag prediction table is applied for and allocated a new entry exceeding a preset threshold within a preset time, adjusting the tag offset values of the multiple threads.

[0011] For example, the operation method provided by at least one embodiment of the present disclosure further includes: dividing the multiple threads into at least two groups of threads, where each group of threads shares the same tag offset value.

[0012] At least one embodiment of the present disclosure further provides a branch predictor for multiple threads. The branch predictor includes: multiple tag prediction tables of different levels, a shift register, and a branch prediction control logic module. The multiple threads share the multiple tag prediction tables. The shift register is configured to store tag offset values set for the multiple threads. The branch prediction control logic module is configured to, in response to an operation of a target thread among the multiple threads on the multiple tag prediction tables, determine a storage area corresponding to each tag prediction table and the target thread based on the target tag offset value corresponding to the target thread and the prediction table identifiers of the multiple tag prediction tables respectively.

[0013] For example, the branch predictor provided by at least one embodiment of the present disclosure further includes a hardware control logic module. The hardware control logic module is configured to record demand information of the multiple threads for each tag prediction table, and adjust the tag offset values stored in the shift register based on the demand information.

[0014] For example, in the branch predictor provided by at least one embodiment of the present disclosure, the tag offset values in the shift register are configured to be externally modifiable.

[0015] At least one embodiment of the present disclosure further provides a processor, which includes the branch predictor provided by at least one embodiment of the present disclosure, wherein the processor is configured to support multiple threads.

[0016] At least one embodiment of the present disclosure further provides an electronic device, which includes the processor provided by at least one embodiment of the present disclosure.

[0017] At least one embodiment of the present disclosure further provides an electronic device, which includes at least one storage unit and at least one processing unit. The at least one storage unit is configured to store computer-readable instructions; the at least one processing unit is configured to execute the computer-readable instructions stored in the storage unit to implement the operation method for the branch predictor provided by at least one embodiment of the present disclosure.

[0018] At least one embodiment of the present disclosure further provides a computer-readable storage medium. Computer-readable instructions are stored in the computer-readable storage medium, and when a processor executes the computer-readable instructions, the operation method for the branch predictor provided by at least one embodiment of the present disclosure is implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings of the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.

[0020] Figure 1A FIG. shows a schematic diagram of a pipeline of a processor core;

[0021] Figure 1B Shows a schematic diagram of a TAGE branch predictor;

[0022] Figure 2A And Figure 2B Shows a schematic diagram of two threads sharing a tag prediction table;

[0023] Figure 3 Shows a flowchart of an operation method of a branch predictor provided by at least one embodiment of the present disclosure;

[0024] Figure 4 Shows a schematic diagram of a thread querying a tag prediction table provided by at least one embodiment of the present disclosure;

[0025] Figure 5A Shows a schematic diagram of adjusting a tag offset value provided by at least one embodiment of the present disclosure;

[0026] Figure 5B Shows a schematic diagram of a storage area used by a thread after adjusting a tag offset value provided by at least one embodiment of the present disclosure;

[0027] Figure 6 Shows a schematic structural diagram of a branch predictor provided by at least one embodiment of the present disclosure;

[0028] Figure 7 Shows a schematic diagram of adjusting a tag offset value in a shift register provided by at least one embodiment of the present disclosure;

[0029] Figure 8 Shows a schematic block diagram of a processor provided by at least one embodiment of the present disclosure; and

[0030] Figure 9 Shows a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure. Detailed implementation manners

[0031] In order to make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0032] Unless otherwise defined, technical terms or scientific terms used in this disclosure shall have the ordinary meanings as understood by those of ordinary skill in the art to which this disclosure pertains. The terms "first", "second" and similar terms used in this disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "comprising" or "including" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. Terms such as "upper", "lower", "left" and "right" are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0033] The following illustrates this disclosure through several specific embodiments. To keep the following description of the embodiments of this disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of an embodiment of this disclosure appears in more than one drawing, the component is denoted by the same or similar reference numeral in each drawing.

[0034] Flowcharts are used in this disclosure to illustrate the operations performed by the systems according to the embodiments of this application. It should be understood that the operations before or below do not necessarily have to be executed precisely in order. Instead, various steps can be processed in reverse order or simultaneously as needed. Also, other operations can be added to these processes, or one or several steps can be removed from these processes.

[0035] Figure 1A A schematic diagram of a pipeline of a processor core is shown, and the dashed arrows in the figure represent the redirected instruction flow.

[0036] As Figure 1AAs shown, the processor core (such as a CPU core) of a single-core processor or a multi-core processor improves the instruction-level parallelism through pipeline technology. Inside the processor core, there are multiple pipeline stages. For example, after the program counter from various sources is fed into the pipeline and the next program counter (PC) is selected through a multiplexer (Mux), the instruction corresponding to this program counter has to go through branch prediction, instruction fetch, instruction decoding, instruction dispatch and renaming, instruction execution, instruction completion, etc. Waiting queues are set between each pipeline stage as needed, and these queues are usually first-in-first-out (FIFO) queues. For example, after the branch prediction unit, there is a branch prediction (BP) FIFO queue to store the branch prediction results; after the instruction fetch unit, there is an instruction cache (IC) FIFO to cache the fetched instructions; after the instruction decoding unit, there is a decoding (DE) FIFO to cache the decoded instructions; after the instruction dispatch and renaming unit, there is a retirement (RT) FIFO to cache the instructions waiting for confirmation of completion after execution. At the same time, the pipeline of the processor core also includes an instruction queue to cache the instructions waiting for the instruction execution unit to execute after instruction dispatch and renaming.

[0037] To support a high operating frequency, each pipeline stage may further contain multiple pipeline levels (operation cycles), and each pipeline level performs limited operations to further improve the performance of the processor core. Although each pipeline level performs limited operations, in this way, each clock cycle can be made the shortest, and the performance of the CPU core is improved by increasing the operating frequency of the CPU. Each pipeline level can also further improve the performance of the processor core by accommodating more instructions (i.e., superscalar technology). Superscalar refers to a method of parallelly executing multiple instructions in one cycle. A processor with increased instruction-level parallelism that can process multiple instructions in one cycle is called a superscalar processor. A superscalar processor adds additional resources on the basis of a normal scalar processor to create multiple pipelines, and each pipeline executes the instructions assigned to it to achieve parallelization.

[0038] Branch prediction is an important part of a high-performance, multi-pipeline-stage processor core. By predicting the execution path of conditional branch instructions in a program, the processor can start prefetching and executing subsequent instructions before the actual result of the branch is determined, thereby reducing the latency caused by waiting for the branch result. If a multi-pipeline-stage CPU core does not have branch prediction, it has to wait until the execution of each branch instruction ends to know which instruction to jump to, which will cause multiple pipeline stages from the front end to the execution to idle, resulting in a large performance loss. As Figure 1AAs shown, branch prediction is at the very front end of the processor core pipeline, and continuously predicts the start and end addresses of the next instruction based on the result of the previous branch prediction. When a later pipeline stage (such as after instruction decoding or after instruction execution) discovers a branch prediction error, all instructions in the pipeline that are younger than the mispredicted branch instruction will be flushed, that is, a pipeline flush is performed, and then the branch predictor continues to predict the instruction stream from that point and fill it into the pipeline.

[0039] Branch prediction includes the following three steps: identifying whether the instruction obtained in the current instruction fetch stage is a branch instruction; if it is a branch instruction, determining whether the branch instruction is a conditional branch instruction; if it is a conditional branch instruction, determining its jump target address. During the execution of a program, the target address of a conditional branch instruction usually remains fixed, that is, the target address is determined at program compilation.

[0040] The Global History Register (GHR) is a register with a variable bit width, used to record the execution results of all branch instructions in the recent period (also known as the "global branch history"). Each time a new branch instruction is executed, the global history register is updated to reflect the latest branch execution result. By using the information stored in the global history register, the execution results of other branch instructions can be used to assist in prediction, and this prediction method is called global history-based branch prediction.

[0041] For example, the Tagged Geometric History Length (TAGE) branch predictor is a typical global history-based branch predictor. The TAGE branch predictor is designed by combining the advantages of the Partial Pattern Matching (PPM) technique and the Optimal Geometric History Length (OGHL / OGEHL) technique. Specifically, the TAGE branch predictor includes a base prediction table and multiple tag prediction tables (hereinafter also referred to as "TAGE tables"). For different tag prediction tables, different lengths of branch history sequences in the global branch history can be used for index calculation. Thus, the TAGE branch predictor can simultaneously perform branch prediction on a certain branch instruction according to different lengths of branch history sequences, and evaluate the accuracy rate of the branch instruction under each branch history sequence, and select the one with the highest historical accuracy rate as the judgment criterion for the final branch prediction.

[0042] For example, Figure 1B shows a schematic diagram of a TAGE branch predictor.

[0043] As shown Figure 1B in the figure, the TAGE branch predictor includes a base prediction table T0 and tag prediction tables T1 to T4. The base prediction table (Base Predictor) T0 is directly indexed using the address of the branch instruction to be predicted (the program counter value, hereinafter referred to as the PC value), and is used to provide a basic prediction result when none of the tag prediction tables T1 to T4 can hit. The tag prediction tables at different levels use branch history sequences of different lengths in the Global Branch History (GBH) to index table entries, thereby allowing each level of prediction table to capture historical patterns of different lengths. For example, the lower-level tag prediction table (e.g., T1) corresponds to a shorter branch history sequence, and the higher-level tag prediction table (e.g., T4) corresponds to a longer branch history sequence. Specifically, each tag prediction table is indexed using the hash value obtained by performing a first hash operation (i.e., Figure 1B Hash1 in the figure) on the PC value and the branch history sequence of the corresponding length. The branch predictor also includes a global branch history register ( Figure 1B not shown in the figure), which is used to record the global branch history.

[0044] As the level of the tag prediction table increases, the length of the corresponding branch history sequence increases in a geometric progression. For example, the calculation formula can be , where L(i) represents the length of the branch history sequence corresponding to the i-th level tag prediction table Ti, is the scale factor of the geometric progression, and L(1) is the length of the branch history sequence corresponding to the tag prediction table T1. The branch history sequence corresponding to the length L(i) can be expressed as ghist[0:L(i)].

[0045] The base prediction table and each tag prediction table each include a certain number of entries. As Figure 1B shown in the figure, each entry in the base prediction table T0 can include a 2-bit saturation counter. Each entry in each tag prediction table (also referred to as the "TAGE table") includes a tag value tag, a prediction value pred, and a valid value u.

[0046] For example, as Figure 1B shown in the figure, the tag value tag is obtained by performing a second hash operation (i.e., Figure 1BThe hash value obtained by Hash2) in is used to confirm whether the indexed table entry corresponds to the current branch instruction to be predicted (i.e., whether the table entry is hit). Here, the first hash operation Hash1 for generating the index is different from the second hash operation Hash2 for generating the tag value. The predicted value pred is the target address of the current branch instruction to be predicted, that is, the prediction result provided by the corresponding table entry. The valid value u represents the reliability of the corresponding table entry. The larger the valid value u, the more reliable the corresponding table entry. When performing a table entry replacement operation on the tag prediction table, the table entry with a smaller valid value u can be preferentially replaced.

[0047] For example, the process of performing branch prediction using the TAGE branch predictor shown in Figure 1B is as follows.

[0048] For the tag prediction table Ti (i = 1, 2, 3, or 4), the PC value and ghist[0:L(i)] are subjected to the first hash operation (Hash1) to obtain the index index_i, and the PC value and ghist[0:L(i)] are subjected to the second hash operation (Hash2) to obtain the tag value tag_i. The tag value tag1 of the table entry (such as table entry 1) of the tag prediction table Ti is read according to the index index_i, and it is judged whether table entry 1 is hit by comparing the tag value tag_i and the tag value tag1 (this step corresponds to Figure 1B the "=?" in ). Specifically, if the tag value tag_i is equal to the tag value tag recorded in table entry 1, it is judged that table entry 1 of the tag prediction table Ti is hit, and the predicted value pred recorded in table entry 1 of the tag prediction table Ti is provided as the prediction result obtained by performing branch prediction based on the tag prediction table Ti. When no table entries in multiple tag prediction tables T1~T4 are hit, the prediction result provided by the base prediction table T0 is used as the target prediction result. When table entries in multiple tag prediction tables are hit, the predicted value pred of the hit table entry in the tag prediction table at the highest level is selected as the target prediction result. If the target prediction result is different from the actual result (i.e., the branch prediction is incorrect), the actual result is used to update the hit table entries in multiple tag prediction tables. Moreover, when the branch prediction is incorrect, if the tag prediction table (such as T2) that gives the target prediction result is not the tag prediction table at the highest level, a new table entry needs to be applied for allocation from a higher-level tag prediction table (such as T3).

[0049] Since the number of branches with different branch correlation distances in different programs is different, during the running of the program, the program will gradually apply for table entries from the tag prediction tables at higher levels in the TAGE branch predictor until the branch correlation distance that the tag prediction table can cover is greater than or equal to the branch correlation distance required by the branch instructions of the program.

[0050] To make more efficient use of processor resources and improve the system's concurrent execution ability, Simultaneous Multi-Threading (SMT) technology has been applied to some processors. In a multi-threaded processor (such as an SMT processor), the branch predictor is shared by multiple threads (hardware threads) in the processor at different timesharing. Different threads commonly use, for example, the TAGE branch predictor for branch prediction. During this process, each thread competes for the capacity of each tag prediction table, and there is interference between threads, resulting in a decrease in branch prediction accuracy.

[0051] For example, Figure 2A and Figure 2B shows a schematic diagram of two threads sharing a tag prediction table.

[0052] In this example, "gray" represents the table entries occupied by thread 0 at the current moment, "black" represents the table entries occupied by thread 1 at the current moment, and "white" represents the table entries that are idle at the current moment.

[0053] As Figure 2A shown, in the SMT2 scenario, both thread 0 and thread 1 start applying for new table entries from the tag prediction table T1. Thread 0 and thread 1 are very likely to compete for the same table entry in the tag prediction table T1. For example, when thread 1 applies for a new table entry from the tag prediction table T1, it first generates an index corresponding to the branch instruction of thread 1 through the first hash operation (for example, denoted as "index1") to locate the position of the new table entry in the tag prediction table T1. If the generated index index1 is the same as the index of table entry 2 in the tag prediction table T1, and table entry 2 is currently occupied by thread 0, then thread 1 will directly replace the table entry 2 occupied by thread 0 with the new table entry corresponding to thread 1, instead of querying for an idle table entry in the tag prediction table T1 for replacement.

[0054] In addition, the demand of the same program for each tag prediction table is approximately normally distributed. That is, compared with the lower-level tag prediction tables (such as tag prediction table T1) and the higher-level tag prediction tables (such as tag prediction table T4), the program has a greater demand for the intermediate-level tag prediction tables (such as tag prediction tables T2 and T3). Therefore, due to the obvious differences in the demand of different programs for each tag prediction table, when multiple threads run simultaneously, it will further exacerbate the imbalance in the demand of the program for each tag prediction table, making the capacity competition of multiple threads for the intermediate-level tag prediction tables more intense, and ultimately resulting in a decrease in branch prediction accuracy.

[0055] For example, as Figure 2BAs shown, since Thread 0 and Thread 1 have a high demand for the tag prediction tables T2 and T3, and all the entries in the tag prediction tables T2 and T3 are occupied by Thread 0 and Thread 1, there are no idle entries. Therefore, Thread 0 and Thread 1 will frequently access and replace the entries in the tag prediction tables T2 and T3. Due to the frequent replacement of the entries in the tag prediction tables T2 and T3, the retention time of the high-value entries in the tag prediction tables T2 and T3 will be reduced and the branch history information recorded during the branch prediction period will be damaged. In addition, the capacity competition between Thread 0 and Thread 1 for the tag prediction tables T2 and T3 will be intensified, ultimately resulting in a decrease in the branch prediction accuracy.

[0056] In addition, as the number of threads increases, the capacity competition among multiple threads for the intermediate-level tag prediction tables will become more intense. The shared TAGE resource cannot achieve multi-user control flow isolation, and the attacker process can infer or manipulate the control flow of the victim process during speculative execution, thereby leaking the memory content of the victim process.

[0057] At least one embodiment of the present disclosure provides an operation method for a branch predictor for multiple threads. The branch predictor includes multiple tag prediction tables of different levels, and the multiple threads share the multiple tag prediction tables. The operation method includes: setting tag offset values for the multiple threads; and in response to an operation of a target thread among the multiple threads on the multiple tag prediction tables, determining a storage area corresponding to each tag prediction table and the target thread based on the target tag offset value corresponding to the target thread and the prediction table identifiers of the multiple tag prediction tables respectively.

[0058] In the above operation method of the branch predictor, by setting tag offset values for each thread, the storage areas corresponding to the same tag prediction table when used by different threads are different, thereby alleviating the capacity competition of different threads for the same branch predictor, reducing the interference between threads, and thus improving the branch prediction accuracy. At the same time, the operation method realizes the isolation of the control flow of multiple users by breaking the shared mode of multiple threads / programs for the branch predictor, and enhances the security of the multi-user system.

[0059] The operation method of the branch predictor provided by at least one embodiment of the present disclosure is applicable to the simultaneous multi-threading (SMT) application scenario. For example, the branch predictor can be the above-mentioned TAGE branch predictor. The branch predictor includes multiple tag prediction tables of different levels, and the multiple tag prediction tables are shared by different threads (hardware threads) of the multi-threaded processor for branch prediction. The multiple tag prediction tables respectively include multiple entries, and each entry includes a tag value tag, a prediction value pred, and a valid value u. Specifically, reference can be made to the above description of Figure 1B which will not be elaborated here.

[0060] The present disclosure is described below through several specific embodiments. In order to keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of known functions and known components may be omitted. When any component of the embodiments of the present disclosure appears in more than one figure, the component is represented by the same or similar reference numeral in each figure.

[0061] Figure 3 A schematic flow chart of an operation method of a branch predictor provided by at least one embodiment of the present disclosure is shown.

[0062] For example, Figure 3 As shown, the operation method of the branch predictor provided by at least one embodiment of the present disclosure includes steps S310 to S320.

[0063] Step S310: setting label offset values ​​for multiple threads.

[0064] Step S320: In response to the target thread among the multiple threads operating on the multiple label prediction tables, based on the target label offset value corresponding to the target thread and the prediction table identifiers of the multiple label prediction tables, determine the storage area corresponding to each label prediction table and the target thread.

[0065] For example, for step S310, different label offset values ​​may be set for different threads. For example, the label offset value may be stored in a shift register in a binary data format. Taking the SMT4 scenario as an example, for example, the label offset value of thread 0 may be set to 00, the label offset value of thread 1 may be set to 01, the label offset value of thread 2 may be set to 10, and the label offset value of thread 3 may be set to 11.

[0066] For example, for step S310, multiple threads may be divided into at least two groups of threads, each group of threads sharing the same offset value. Taking the SMT4 scenario as an example, threads 0 and 1 may be divided into a first group, and threads 2 and 3 may be divided into a second group, wherein threads 0 and 1 share the same label offset value (e.g., 00), and threads 2 and 3 share another label offset value (e.g., 10). The division method here is only for illustration, and the present disclosure does not limit this.

[0067] The embodiments of the present disclosure do not limit the setting method of the label offset value, and the label offset value of each thread is not fixed and can be adjusted and changed according to actual conditions, thereby more efficiently alleviating the capacity competition problem between different threads, and further improving the branch prediction accuracy of the branch pre-predictor for multiple threads.

[0068] For example, for step S320, the target thread is the thread that performs branch prediction operations based on multiple tag prediction tables among multiple threads. The operations of the target thread on the multiple tag prediction tables include operations such as querying, replacing, creating new table entries, and updating.

[0069] For example, each tag prediction table is set with a corresponding prediction table identifier (or prediction table number). For example, Figure 1B the prediction table identifiers of the 4 tag prediction tables T1 to T4 in [reference] are 1, 2, 3, and 4 respectively. The prediction table identifier can also represent the level of the corresponding tag prediction table. For example, the larger the value of the prediction table identifier, the higher the level of the tag prediction table corresponding to the prediction table identifier.

[0070] For example, the storage area determined in step S320 (also referred to as the "physical storage area" or "physical area") is configured to store the table entries for the target thread in the corresponding tag prediction table. That is, when multiple threads perform this operation on the same tag prediction table, the table entries for different threads (here, "different threads" refer to threads with different tag offset values) in the same tag prediction table are mapped (stored) to these different storage areas.

[0071] In at least one embodiment of the present disclosure, the branch predictor includes N tag prediction tables and N storage areas, and each storage area stores 2 m table entries, where N and m are positive integers. Each storage area is set with a corresponding area identifier.

[0072] For example, the above step S320 may include: for the first tag prediction table among the N tag prediction tables, in response to (Tage_Offset + Table_id) ≤ N, determining that the area identifier of the storage area corresponding to the first tag prediction table and the target thread is (Tage_Offset + Table_id); or, in response to (Tage_Offset + Table_id) > N, determining that the area identifier of the storage area corresponding to the first tag prediction table and the target thread is (Tage_Offset + Table_id - N), where Table_id is the prediction table identifier of the first tag prediction table, and Tage_Offset is the target tag offset value.

[0073] Here, the first tag prediction table is any one of the N tag prediction tables.

[0074] In at least one embodiment of the present disclosure, the number of tag prediction tables in the branch predictor is the same as the number of storage areas. The number of tag prediction tables can be set according to the actual situation, generally 4 to 8, and the present disclosure does not limit this.

[0075] For example, Figure 4A schematic diagram of a thread query tag prediction table provided by at least one embodiment of the present disclosure is shown.

[0076] As Figure 4 shown, in the SMT4 scenario, the tag offset values of threads 0 to 3 are set to 00, 01, 10, and 11 in binary (i.e., 0, 1, 2, and 3 in decimal data format, and the following uses a similar expression). The branch predictor includes four tag prediction tables T1 to T4 and four storage areas 1 to 4.

[0077] When the target thread is thread 0, in response to thread 0 (i = 0) querying the tag prediction table T1, the tag offset value of thread 0 is 00 (i.e., 0). Since Tage_Offset + Table_id = 0 + 1 = 1 ≤ 4, the area identifier of the storage area corresponding to thread 0 in tag prediction table T1 is 1 (as Figure 4 shown by the solid arrow with i = 0 in Figure 4 ), and the entry in tag prediction table T1 for thread 0 is mapped to storage area 1. And so on, the entry in tag prediction table T1 for thread 1 is mapped to storage area 2 (as Figure 4 shown by the solid arrow with i = 1 in Figure 4 ), the entry in tag prediction table T1 for thread 2 is mapped to storage area 3 (as

[0078] shown by the solid arrow with i = 2 in Figure 4 ), and the entry in tag prediction table T1 for thread 3 is mapped to storage area 4 (as Figure 4 shown by the solid arrow with i = 3 in Figure 4 ). Figure 4 ).

[0079] In the same way as determining the storage area corresponding to each tag prediction table for the target thread (step S320) described above, when the target threads are thread 0, thread 1, thread 2, and thread 3 in sequence, Figure 4 the mapping relationships between the tag prediction tables T1 to T4 in

[0080] and the storage areas 1 to 4 are shown in Table 1.

[0081]

[0082] In Figure 4 the example shown, before setting the tag offset values for multiple threads (i.e., threads 0 to 3), the entries in the tag prediction tables are mapped to the storage areas whose region identifiers are equal to the prediction table identifiers of the tag prediction tables. For example, the entries in the tag prediction table T1 are all mapped to the storage area 1, the entries in the tag prediction table T2 are all mapped to the storage area 2, the entries in the tag prediction table T3 are all mapped to the storage area 3, and the entries in the tag prediction table T4 are all mapped to the storage area 4. Since the demands of the program for each tag prediction table are not evenly distributed, for example, following a normal distribution, threads 0 to 3 all have a relatively large demand for the intermediate-level tag prediction tables (such as tag prediction tables T2 and T3). That is, the entries in the tag prediction tables T2 and T3 for threads 0 to 3 are centrally mapped (stored) to the storage areas 2 and 3. Since the number of entries that each storage area can store is limited, threads 0 to 3 will perform relatively frequent replacement operations on the entries stored in the storage areas 2 and 3. That is, compared with the competition of threads 0 to 3 for the storage areas 1 and 4, the competition of threads 0 to 3 for the storage areas 2 and 3 is more intense.

[0083] In Figure 4 the example shown, after setting the tag offset values for multiple threads (i.e., threads 0 to 3), referring to Table 1 above, the entries in the tag prediction table T2 for threads 0, 1, 2, and 3 are respectively mapped to the storage areas 2, 3, 4, and 1, and the entries in the tag prediction table T3 for threads 0, 1, 2, and 3 are respectively mapped to the storage areas 3, 4, 1, and 2. Although threads 0 to 3 still have a relatively large demand for the tag prediction tables T2 and T3, through the set tag offset values, the entries in the tag prediction tables T2 and T3 for threads 0 to 3 are stored dispersedly in the storage areas 1, 2, 3, and 4, rather than being centrally stored in the storage areas 2 and 3, thereby effectively alleviating the competition between different threads for the storage areas, improving the overall resource utilization rate of the storage areas, and improving the branch prediction accuracy.

[0084] For example, the operation method provided by at least one embodiment of the present disclosure further includes: recording the demand information of multiple threads for each tag prediction table, and adjusting the tag offset values of the multiple threads based on the demand information.

[0085] For example, the demand information is used to characterize the capacity competition situation of multiple threads for each tag prediction table. The demand information may include information in multiple dimensions. For example, the number of times each tag prediction table is applied to allocate new table entries, the number of times each tag prediction table provides a target prediction result, the number of times the target prediction result provided by each tag prediction table is incorrect, etc. The present disclosure places no restrictions on this.

[0086] For example, when the demand information includes the number of times each tag prediction table is applied to allocate new table entries, adjusting the tag offset values of multiple threads based on the demand information includes: in response to the number of times at least one tag prediction table is applied to allocate new table entries exceeding a preset threshold within a preset time, adjusting the tag offset values of the multiple threads.

[0087] For example, Figure 5A FIG. shows a schematic diagram of adjusting the tag offset value provided by at least one embodiment of the present disclosure.

[0088] As Figure 5A shown, a hardware control logic module 530 may be set in the branch predictor. The hardware control logic module 530 may record the demand information of multiple threads for each tag prediction table. For example, the number of times thread 0 and thread 1 apply to allocate new table entries to each tag prediction table (i.e., Figure 5A TageTable1_AllocNum_Thread0, TageTable2_AllocNum_Thread0, TageTable3_AllocNum_Thread0, etc. in ). Then, the hardware control logic module 530 adjusts the tag offset values of thread 0 and thread 1 based on the demand information, thereby realizing the dynamic adjustment of the tag offset values, so that the branch predictor can adapt to the demands of the program for each tag prediction table.

[0089] It should be noted that when adjusting the tag offset values of multiple threads, the tag offset value of each thread may be adjusted, or only the tag offset values of some threads may be adjusted, and the tag offset values of the remaining threads remain unchanged. For example, the specific adjustment method may be set in advance (such as burning) in the hardware control logic module 530, so that when the demand information meets certain conditions, the hardware control logic module 530 can automatically adjust the tag offset values of the threads.

[0090] For example, as Figure 5AAs shown, in the STM2 scenario, the initial tag offset values of Thread 0 and Thread 1 are both set to 00. Then, the entries in tag prediction table T1 are all mapped to storage area 1, the entries in tag prediction table T2 are all mapped to storage area 2, the entries in tag prediction table T3 are all mapped to storage area 3, and the entries in tag prediction table T4 are all mapped to storage area 4. If the hardware control logic module 530 detects that the sum of the number of times each thread (i.e., Thread 0 and Thread 1) applies to tag prediction table T2 for allocating new entries (i.e., TageTable2_AllocNum_Thread0 + TageTable3_AllocNum_Thread1) exceeds a preset threshold (e.g., 500 times) within a preset time, and the sum of the number of times each thread (i.e., Thread 0 and Thread 1) applies to tag prediction table T3 for allocating new entries (i.e., TageTable3_AllocNum_Thread0 + TageTable3_AllocNum_Thread1) also exceeds the preset threshold (e.g., 500 times) within the preset time, it is determined that Thread 0 and Thread 1 have a greater demand for tag prediction tables T2 and T3.

[0091] As Figure 5A shown, the entries for Thread 0 and Thread 1 almost occupy all of storage area 2 and storage area 3, and there are still free storage locations in storage area 1 and 4 (i.e., Figure 5A the white entries in it), which indicates that both Thread 0 and Thread 1 use storage areas 2 and 3 more and storage areas 1 and 4 less, that is, Thread 0 and Thread 1 compete fiercely for storage areas 2 and 3. At this time, the tag offset value of Thread 1 can be adjusted from 00 to 10, so that the storage areas corresponding to the entries for Thread 1 in tag prediction tables T1~T4 are changed.

[0092] For example, after adjusting the tag offset value of Thread 1 from 00 to 10, as Figure 5A shown by the dashed arrow in it (or see Table 2 below), the entry for Thread 1 in tag prediction table T1 is mapped to storage area 3, the entry for Thread 1 in tag prediction table T2 is mapped to storage area 4, the entry for Thread 1 in tag prediction table T3 is mapped to storage area 1, and the entry for Thread 1 in tag prediction table T4 is mapped to storage area 2.

[0093] Table 2: Mapping relationship between tag prediction tables and storage areas

[0094]

[0095] For example, Figure 5B shows a schematic diagram of the storage areas used by threads after adjusting the tag offset value provided by at least one embodiment of the present disclosure.

[0096] As Figure 5B shown, by adjusting the tag offset value of Thread 1 from 00 to 10, the competition between Thread 0 and Thread 1 for storage areas 2 and 3 can be alleviated, enabling Thread 0 and Thread 1 to use storage areas 1 to 4 more evenly. The specific method for determining the area identifier can refer to the example Figure 4 shown above, which will not be elaborated here. It should be noted that Figure 5B shown, the usage of storage areas 1 to 4 by Thread 0 and Thread 1 after adjusting the tag offset value is only illustrative, and the actual usage may be different from Figure 5B this.

[0097] For example, after adjusting the tag offset value of Thread 1 from 00 to 10, the requirements of Thread 0 and Thread 1 for tag prediction tables T1 to T4 may change again. For example, if the sum of the number of times Thread 0 and Thread 1 apply for allocating new table entries to tag prediction table T1 (i.e., TageTable1_AllocNum_Thread0 + TageTable1_AllocNum_Thread1) exceeds a preset threshold within a preset time, and the sum of the number of times Thread 0 and Thread 1 apply for allocating new table entries to tag prediction table T3 (i.e., TageTable3_AllocNum_Thread0 + TageTable3_AllocNum_Thread1) also exceeds the preset threshold within the preset time, it is determined that Thread 0 and Thread 1 have a greater demand for tag prediction tables T1 and T3. If the current tag offset values (the tag offset value of Thread 0 is 00, and the tag offset value of Thread 0 is 10) are continued to be used, referring to Table 2, both Thread 0 and Thread 1 use storage areas 1 and 3 more, and use storage areas 2 and 4 less, resulting in intense competition between Thread 0 and Thread 1 for storage areas 1 and 3. Therefore, it is necessary to adjust the tag offset values of Thread 0 and Thread 1 again. For example, the tag offset value of Thread 1 can be adjusted from 10 to 01, so that the storage areas corresponding to the table entries for Thread 1 in tag prediction tables T1 to T4 change. For details, please refer to Table 3 below.

[0098] As shown in Table 3, after adjusting the tag offset value of Thread 1 to 01, the table entries for Thread 1 in tag prediction table T1 are mapped to storage area 4, the table entries for Thread 1 in tag prediction table T2 are mapped to storage area 1, the table entries for Thread 1 in tag prediction table T3 are mapped to storage area 2, and the table entries for Thread 1 in tag prediction table T4 are mapped to storage area 3. The adjusted tag offset values enable Thread 0 and Thread 1 to use storage areas 1 to 4 more evenly, thus alleviating the competition between different threads for storage areas.

[0099] Table 3: Mapping relationship between tag prediction tables and storage areas

[0100]

[0101] For example, the operations of the target thread on multiple tag prediction tables include query operations. After determining the storage areas corresponding to each tag prediction table and the target thread, the operation method provided by at least one embodiment of the present disclosure further includes: reading the entries stored in the storage areas corresponding to each tag prediction table and the target thread to query whether the branch instruction to be predicted by the target thread hits N tag prediction tables. The specific way to query whether it hits can refer to the example of Figure 1B above, which will not be elaborated here.

[0102] For example, the operation method provided by at least one embodiment of the present disclosure further includes: in response to the branch instruction to be predicted by the target thread hitting the first entry in at least one tag prediction table, taking the tag prediction table with the highest level in the at least one hit tag prediction table as the target tag prediction table, and reading the prediction result of the first entry located in the storage area corresponding to the target tag prediction table and the target thread to obtain the target prediction result.

[0103] For example, the operation method provided by at least one embodiment of the present disclosure further includes: in response to the target prediction result being incorrect, updating the first entry in the at least one hit tag prediction table based on the actual result of the branch instruction to be predicted, and applying for allocation of a new entry in a tag prediction table at a level higher than the target tag prediction table to store the instruction information and the actual result corresponding to the branch instruction to be predicted.

[0104] For example, the operation method provided by at least one embodiment of the present disclosure further includes: in response to the target prediction result being incorrect and the target tag prediction table having the highest level among multiple tag prediction tables, updating the first entry in the at least one hit tag prediction table based on the actual result of the branch instruction to be predicted.

[0105] For example, referring to Figure 4 , if it is determined that the branch instruction to be predicted by the target thread (e.g., thread 2) hits the first entry in tag prediction tables T2, T3, and T4, then take the tag prediction table T4 with the highest level as the target tag prediction table, read the prediction result of the first entry in the storage area 2 corresponding to tag prediction table T4 and thread 2, and the prediction result of the first entry in storage area 2 is the target prediction result obtained by thread 2 performing the branch prediction operation. In response to the target prediction result being incorrect, that is, the target prediction result is different from the actual result of the branch instruction to be predicted by thread 2, then update the first entry in the hit tag prediction tables T2, T3, and T4 based on this actual result. Since the target tag prediction table (i.e., tag prediction table T4) has the highest level among all tag prediction tables T1 to T4, there is no need to apply for a new entry in a tag prediction table at a level higher than tag prediction table T4.

[0106] IfFigure 4 The branch predictor shown also includes a tag prediction table T5 at a higher level than the tag prediction table T4. In this case, a new entry needs to be applied for in the tag prediction table T5, and a new entry is created or an old entry is replaced with a new entry in the storage area corresponding to thread 2 in the tag prediction table T5 to store the instruction information and the actual result corresponding to the branch instruction to be predicted for thread 2.

[0107] Therefore, in the operation method of the branch predictor provided by at least one embodiment of the present disclosure, by setting a tag offset value for each thread to divide the storage areas corresponding to the same tag prediction table for different threads, the capacity competition problem of different threads for the same branch predictor can be alleviated, the interference between threads can be reduced, and thus the branch prediction accuracy can be improved.

[0108] Figure 6 FIG. shows a schematic structural diagram of a branch predictor provided by at least one embodiment of the present disclosure. The branch predictor is applicable to multiple threads.

[0109] As Figure 6 shown, the branch predictor 600 includes a plurality of tag prediction tables T1 to TN at different levels, a shift register 610, and a branch prediction control logic module 620.

[0110] Multiple threads share the plurality of tag prediction tables T1 to TN. The shift register 610 is configured to store the tag offset values set for the multiple threads. The branch prediction control logic module 620 is configured to, in response to an operation of a target thread among the multiple threads on the plurality of tag prediction tables, determine the storage area corresponding to each tag prediction table and the target thread based on the target tag offset value corresponding to the target thread and the prediction table identifiers of the plurality of tag prediction tables respectively.

[0111] For example, in at least one embodiment of the present disclosure, the branch predictor 600 further includes a hardware control logic module, which is configured to record the demand information of the multiple threads for each tag prediction table and adjust the tag offset values stored in the shift register 610 based on the demand information. For specific details, please refer to the above description of Figure 5A which will not be elaborated here.

[0112] For example, in at least one embodiment of the present disclosure, the demand information includes the number of times each tag prediction table is applied for and allocated a new entry. The hardware control logic module is further configured to adjust the tag offset values of the multiple threads in response to the number of times at least one tag prediction table is applied for and allocated a new entry exceeding a preset threshold within a preset time.

[0113] For example, software can also be used to record and analyze the demand information of each tag prediction table by multiple threads, and then the hardware control logic module adjusts the tag offset value stored in the shift register 610 according to the analysis result. For another example, the hardware control logic module may further include a storage component to store the adjustment method of the tag offset value for multiple threads.

[0114] For example, in at least one embodiment of the present disclosure, the tag offset value in the shift register 610 is configured to be modifiable externally. For example, a software interface is set in the branch predictor 600 so that the user can dynamically adjust the tag offset value in the shift register 610 through the software interface.

[0115] For example, Figure 7 FIG. shows a schematic diagram of adjusting the tag offset value in the shift register provided by at least one embodiment of the present disclosure.

[0116] As Figure 7 As shown, in a multi-user scenario, the software (or software scheduler) is responsible for managing and allocating system resources (such as CPU time) to different tasks or users. For example, the software manages a pool including multiple users such as User 1 to User 4 (also referred to as "user pool"). The software presets different tag offset values Tage_Offset for each user. For example, the Tage_Offset corresponding to User 1 is 00, the Tage_Offset corresponding to User 2 is 01, the Tage_Offset corresponding to User 3 is 10, and the Tage_Offset corresponding to User 4 is 11. Through the scheduling management of the software, multiple users share (or time-share) the same branch predictor by using the corresponding threads (also referred to as "physical threads") in the processor at different times. Here, the processor can be a multi-core SMT processor, that is, the processor supports multiple threads and at least includes processor core 0 and processor core 1.

[0117] For example, the software schedules User 1 to thread 0 of processor core 0 in the first operation cycle, and schedules User 2 to thread 0 of processor core 0 in the second operation cycle, then User 1 and User 2 time-share the same branch predictor.

[0118] For example, when the software schedules a certain user to any physical thread of any processor core, the software sets the offset register for this thread to the tag offset value preset for this user. As Figure 7As shown, when the software schedules User 2 and User 3 to Thread 0 and Thread 1 of Processor Core 0 respectively, the software will adjust the label offset value of Thread 0 of Processor Core 0 stored in Shift Register 0 to the Tage_Offset corresponding to User 2 (i.e., 01), and adjust the label offset value of Thread 1 of Processor Core 0 to the Tage_Offset corresponding to User 3 (i.e., 10). When User 2 and User 3 use the shared branch predictor, the storage areas corresponding to User 2 and User 3 in the same label prediction table are different, thus ensuring the isolation between User 2 and User 3 when using the shared branch predictor.

[0119] Therefore, in a multi-user scenario, by setting different label offset values for different users and adjusting the label offset values corresponding to each thread in the shift register based on the software scheduling, it can ensure the isolation between different users when using the shared branch predictor, thereby achieving the independence and isolation of the control flow between multiple users and significantly improving the overall security and reliability of the multi-user system.

[0120] For example, in at least one embodiment of the present disclosure, the branch predictor 600 includes N label prediction tables and N storage areas, and each storage area stores 2 m table entries, where N and m are positive integers. The branch prediction control logic module 620 is further configured to: for the first label prediction table among the N label prediction tables, in response to (Tage_Offset + Table_id) ≤ N, determine that the area identifier of the storage area corresponding to the first label prediction table and the target thread is (Tage_Offset + Table_id); or, in response to (Tage_Offset + Table_id) > N, determine that the area identifier of the storage area corresponding to the first label prediction table and the target thread is (Tage_Offset + Table_id - N), where Table_id is the prediction table identifier of the first label prediction table, and Tage_Offset is the target label offset value.

[0121] For example, in at least one embodiment of the present disclosure, the operations of the target thread on multiple label prediction tables include query operations, and the branch predictor 600 further includes a query module. After determining the storage area corresponding to each label prediction table and the target thread, the query module is configured to read the table entries stored in the storage area corresponding to each label prediction table and the target thread to query whether the branch instruction to be predicted by the target thread hits the N label prediction tables.

[0122] For example, in at least one embodiment of the present disclosure, the branch predictor 600 further includes a prediction module. The prediction module is configured to: in response to a to-be-predicted branch instruction of a target thread hitting a first entry in at least one tag prediction table, use the tag prediction table with the highest rank in the at least one hit tag prediction table as the target tag prediction table, and read the prediction result of the first entry in the storage area corresponding to the target tag prediction table and the target thread, so as to obtain a target prediction result.

[0123] For example, in at least one embodiment of the present disclosure, the branch predictor 600 further includes a judgment module. The judgment module is configured to judge whether the target prediction result is correct.

[0124] For example, in at least one embodiment of the present disclosure, the judgment module is further configured to: in response to the target prediction result being incorrect, update the first entry in the at least one hit tag prediction table based on the actual branch result of the to-be-predicted branch instruction, and apply for and allocate a new entry in a tag prediction table at a level higher than the target tag prediction table to store the instruction information and the actual branch result corresponding to the to-be-predicted branch instruction.

[0125] For example, in at least one embodiment of the present disclosure, the judgment module is further configured to: in response to the target prediction result being incorrect and the target tag prediction table having the highest rank, update the first entry in the at least one hit tag prediction table based on the actual branch result of the to-be-predicted branch instruction.

[0126] For example, in at least one embodiment of the present disclosure, the branch predictor 600 further includes a thread partitioning module, and the thread partitioning module is configured to partition a plurality of threads into at least two groups of threads, where each group of threads shares the same tag offset value.

[0127] For example, the above modules (such as a branch prediction control logic module, a hardware control logic module, a thread partitioning module) or units can be implemented by hardware, firmware, software, or any combination thereof.

[0128] Figure 8 Fig. shows a schematic block diagram of a processor provided by at least one embodiment of the present disclosure.

[0129] For example, as Figure 8 shown, the processor 800 provided by at least one embodiment of the present disclosure includes a branch predictor 801. The processor 800 is, for example, an SMT processor. The maximum number of threads supported by the SMT processor can be, for example, 2, 4, 8, etc. The SMT processor can be a single-core or multi-core processor. For example, the branch predictor 801 can be the branch predictor provided by any of the above embodiments.

[0130] For example, in addition to the branch predictor 801, the processor 800 may further include an instruction fetch unit, a decoding unit, an allocation unit, a fixed-point execution unit, a floating-point execution unit, a memory access unit, etc.; the processor includes multiple pipeline stages, and the instruction corresponding to the program counter enters the instruction fetch unit after the branch prediction by the branch predictor 801 to enter the pipeline of the processor for processing. The specific operations are not elaborated here.

[0131] At least one embodiment of the present disclosure further provides an electronic device, which includes at least one storage unit and at least one processing unit. The at least one storage unit is configured to store computer-readable instructions. The at least one processing unit is configured to execute the computer-readable instructions stored in the storage unit to implement the operation method for the branch predictor provided in any embodiment of the present disclosure.

[0132] At least one embodiment of the present disclosure further provides an electronic device, which includes the processor of any of the above embodiments.

[0133] Figure 9 A schematic block diagram of the electronic device provided by at least one embodiment of the present disclosure is shown.

[0134] The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The shown electronic device 900 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0135] For example, as Figure 9 shown, in some examples, the electronic device 900 includes a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, and the processing device may include the processor of any of the above embodiments, and may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the computer system are also stored. The processor 901, the ROM 902, and the RAM 903 are connected to each other through the bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.

[0136] For example, the following components can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, such as, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; a communication device 909 which can also include, for example, a network interface card such as a LAN card, a modem, etc. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data and perform communication processing via a network such as the Internet. The driver 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the driver 910 as needed so that a computer program read from it can be installed into the storage device 908 as needed. Although Figure 9 the electronic device 900 including various devices is shown, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices can be implemented or included.

[0137] For example, the electronic device 900 can further include a peripheral interface (not shown in the figure), etc. The peripheral interface can be various types of interfaces, such as a USB interface, a Lightning interface, etc. The communication device 909 can communicate with the network and other devices wirelessly. The network can be, for example, the Internet, an intranet, and / or a wireless network such as a cellular phone network, a wireless local area network (LAN), and / or a metropolitan area network (MAN). The wireless communication can use any one of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), WiMAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.

[0138] For example, the electronic device 900 can be any device such as a mobile phone, a tablet computer, a laptop computer, an e-book, a game console, a television, a digital photo frame, a navigator, a server, etc., or can also be any combination of a data processing device and hardware. The embodiments of the present disclosure are not limited thereto.

[0139] At least one embodiment of the present disclosure further provides a computer-readable storage medium storing computer-readable instructions, which, when executed by a processor, implement the operation method for a branch predictor provided in any embodiment of the present disclosure.

[0140] It should be noted that the computer-readable medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0141] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device.

[0142] Although the present disclosure has been described in detail above using general descriptions and specific embodiments, based on the embodiments of the present disclosure, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present disclosure fall within the scope of protection required by the present disclosure.

[0143] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.

[0144] For the present disclosure, the following points also need to be noted:

[0145] (1) The drawings of the embodiments of the present disclosure only relate to the structures involved in the embodiments of the present disclosure, and other structures may refer to the general design.

[0146] (2) For clarity, in the drawings used to describe the embodiments of the present disclosure, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn to actual scale.

[0147] (3) Without conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other to obtain new embodiments.

[0148] The above are only the specific implementation manners of the present disclosure, but the protection scope of the present disclosure is not limited thereto. The protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A method for operating a branch predictor for multiple threads, wherein: The branch predictor includes a plurality of label prediction tables of different levels, the plurality of threads share the plurality of label prediction tables, and the operation method includes: Setting label offset values ​​for the plurality of threads; In response to the operation of the multiple label prediction tables by a target thread among the multiple threads, a storage area corresponding to each label prediction table and the target thread is determined based on a target label offset value corresponding to the target thread and prediction table identifiers of the multiple label prediction tables.

2. The operating method according to claim 1, wherein: The branch predictor includes N label prediction tables and N storage areas, each of which stores 2 m table entries, N and m are positive integers, The determining, based on the target label offset value corresponding to the target thread and the prediction table identifiers of the plurality of label prediction tables, a storage area corresponding to each label prediction table and the target thread, comprises: For a first tag prediction table among the N tag prediction tables, in response to (Tage_Offset + Table_id) ≤ N, determining that the region identifier of the storage region corresponding to the first tag prediction table and the target thread is (Tage_Offset + Table_id); or, in response to (Tage_Offset + Table_id)>N, determining that the region identifier of the storage region corresponding to the first tag prediction table and the target thread is (Tage_Offset + Table_id-N), Among them, Table_id is the prediction table identifier of the first tag prediction table, and Tage_Offset is the target tag offset value.

3. The operating method according to claim 2, wherein: The target thread's operation on the multiple tag prediction tables includes a query operation, After determining the storage area corresponding to each tag prediction table and the target thread, the operation method further includes: The table entries stored in the storage area corresponding to each of the label prediction tables and the target thread are read to query whether the to-be-predicted branch instruction of the target thread hits the N label prediction tables.

4. The operating method according to claim 3, further comprising: In response to the predicted branch instruction of the target thread hitting the first table entry in at least one label prediction table, the label prediction table with the highest level among the at least one label prediction table that is hit is used as the target label prediction table, and the prediction result of the first table entry located in the storage area corresponding to the target label prediction table and the target thread is read to obtain the target prediction result.

5. The operating method according to claim 4, further comprising: In response to the target prediction result being wrong, the first table entry in the at least one label prediction table that is hit is updated based on the actual result of the branch instruction to be predicted, and a new table entry is requested to be allocated in a label prediction table at a higher level than the target label prediction table to store the instruction information corresponding to the branch instruction to be predicted and the actual result.

6. The operating method according to claim 4, further comprising: In response to the target prediction result being wrong and the target tag prediction table having the highest rank among the multiple tag prediction tables, a first table entry in the at least one tag prediction table that is hit is updated based on an actual result of the to-be-predicted branch instruction.

7. The operating method according to any one of claims 1 to 6, further comprising: The requirement information of the multiple threads for each label prediction table is recorded, and the label offset values ​​of the multiple threads are adjusted based on the requirement information.

8. The operating method according to claim 7, wherein: The demand information includes the number of times each label prediction table is applied to allocate new table entries, and adjusting the label offset values ​​of the multiple threads based on the demand information includes: In response to the number of times that at least one label prediction table is applied to allocate the new entry exceeding a preset threshold within a preset time, the label offset values ​​of the multiple threads are adjusted.

9. The method according to any one of claims 1 to 6, further comprising: The plurality of threads are divided into at least two groups of threads, wherein each group of threads shares a same tag offset value.

10. A branch predictor for multiple threads, comprising: a plurality of label prediction tables of different levels, wherein the plurality of threads share the plurality of label prediction tables; a shift register configured to store tag offset values ​​set for the plurality of threads; The branch prediction control logic module is configured to respond to the operation of the target thread among the multiple threads on the multiple label prediction tables, and determine the storage area corresponding to each label prediction table and the target thread based on the target label offset value corresponding to the target thread and the prediction table identifiers of the multiple label prediction tables.

11. The branch predictor according to claim 10, further comprising: The hardware control logic module is configured to record the demand information of the multiple threads for each tag prediction table, and adjust the tag offset value stored in the shift register based on the demand information.

12. The branch predictor according to claim 10, wherein: The tag offset value in the shift register is configured to be modified externally.

13. A processor comprising a branch predictor according to any one of claims 10 to 12, wherein: The processor is configured to support multiple threads.

14. An electronic device comprising the processor according to claim 13.

15. An electronic device, comprising: at least one storage unit configured to store computer-readable instructions; as well as At least one processing unit is configured to execute the computer-readable instructions stored in the storage unit to implement the operating method according to any one of claims 1-9.

16. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-readable instructions. When the processor executes the computer-readable instructions, the operating method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Method and apparatus for controlling branch prediction logic

    CN103218209A

  • Branch predictor design of simultaneous thread processor

    CN103593166A

  • Branch prediction method, device, medium and equipment

    CN112596792A

  • Instruction prediction method of multi-thread processor and related device

    CN114020441A

  • Scheduling method and device for multiple threads and processor

    CN117055961A