An instruction scheduling method, apparatus, system, product, and medium
Through the delay cache mechanism, the instruction history delay information is recorded, combined with data dependence and hardware resource information, the problem of insufficient dynamic delay prediction in the out-of-order scheduling strategy is solved, and more accurate instruction scheduling and processor performance improvement is achieved.
Patent Information
- Application Number
- CN202510600780.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The existing out-of-order scheduling strategies cannot accurately predict dynamic changes in instruction launch time, especially in scenarios such as memory access latency and cache misses, resulting in waste of resources and inability to fully utilize processor performance.
By introducing a delay cache mechanism, the historical delay information of the instruction is recorded, used for prediction of dynamic delays, combined with data dependencies and hardware resource availability information, the execution order of the scheduling instruction, especially for instructions with dynamic delays, the historical delay information is used to accurately predict to determine the optimal transmission time.
Improve the accuracy and efficiency of instruction scheduling, reduce pipeline pauses, better utilize processor resources, and give full play to processor performance.
Smart Images

Figure CN120104191B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of processor design, and particularly to an instruction scheduling method, apparatus, system, product and medium. Background Art
[0002] In modern processor design, the instruction scheduling efficiency directly affects the performance of the processor. Especially in the design of high-performance processors, the scheduling mechanism directly determines the pipeline utilization rate and system throughput. Currently, the mainstream design solution is to adopt an out-of-order execution mechanism, and dynamically schedule instructions through an out-of-order scheduling strategy to fill the pipeline gaps, thereby significantly improving the execution efficiency.
[0003] However, most common out-of-order scheduling strategies construct the scheduling order based on static latency or dependencies, without considering the dynamic changes in instruction issue times. Especially for scenarios with uncertain latency times such as memory access latency and cache misses, the current scheduling strategies are difficult to accurately predict the optimal issue time, resulting in resource waste, pipeline stalls, and the inability to fully utilize the processor performance.
[0004] Therefore, there is an urgent need for a method for instruction scheduling in the art to solve the problem that the current out-of-order scheduling strategy cannot accurately schedule the dynamic changes in instruction issue times, and the scheduling effect needs to be further improved. Summary of the Invention
[0005] The purpose of the present invention is to provide an instruction scheduling method, apparatus, system, product and medium to further improve the execution efficiency of processor instructions and reduce pipeline stalls.
[0006] To solve the above technical problems, the present invention provides an instruction scheduling method, including:
[0007] Obtain the data dependencies and instruction execution latencies of each instruction.
[0008] Determine the predicted issue time of each instruction according to the data dependencies and the instruction execution latencies.
[0009] Obtain the hardware resource availability information, and schedule the execution order of each instruction based on the predicted issue time, the data dependencies, and the hardware resource availability information.
[0010] Wherein, the types of the instruction execution latencies include dynamic latency and static latency; when the instruction execution latency is dynamic latency, obtaining the instruction execution latency includes:
[0011] Obtain the historical latency information corresponding to the current instruction from the latency cache, and predict the dynamic latency through the historical latency information.
[0012] The delay cache is used to record the actual past delay time of each of the instructions as the historical delay information.
[0013] In a possible embodiment, obtaining the instruction execution delay includes:
[0014] Determine whether the instruction type of the instruction is a load instruction or a non-load instruction.
[0015] If it is a load instruction, query the status of the data accessed by the instruction in the cache.
[0016] If the status is cache hit, obtain the static delay corresponding to the instruction, and use the static delay as the instruction execution delay.
[0017] If the status is cache miss, query the corresponding historical delay information in the delay cache, predict the dynamic delay according to the historical delay information, and use the dynamic delay as the instruction execution delay.
[0018] If it is a non-load instruction, obtain the static delay corresponding to the instruction, and use the static delay as the instruction execution delay.
[0019] In a possible embodiment, determining the predicted issue time of each of the instructions according to the data dependence relationship and the instruction execution delay includes:
[0020] Determine the data dependence chain between each of the instructions according to the data dependence relationship, and use the first-level instruction in each data dependence chain as the producer instruction, and the remaining instructions as the consumer instructions.
[0021] For the producer instruction, the predicted issue time is the current time.
[0022] For the consumer instruction, the predicted issue time is the maximum value of the instruction completion times corresponding to the previous-level instructions of the current consumer instruction.
[0023] Wherein, the previous-level instruction is all the instructions located at the previous level of the current consumer instruction in the data dependence chain; the instruction completion time is determined by the predicted issue time and the instruction execution delay.
[0024] In a possible embodiment, for the producer instruction, after determining the predicted issue time, the method further includes:
[0025] Judge whether the instruction execution delay corresponding to the producer instruction is the static delay or the dynamic delay.
[0026] If it is the dynamic delay, the instruction completion time is determined according to the predicted emission time, the dynamic delay, and the stability weighting factor.
[0027] In a possible embodiment, determining the instruction completion time according to the predicted emission time, the dynamic delay, and the stability weighting factor includes:
[0028] Determine a margin stable delay according to the product of the stability weighting factor and the dynamic delay.
[0029] Determine the instruction completion time according to the sum of the margin stable delay and the predicted emission time.
[0030] Wherein, the stability weighting factor is greater than or equal to 1, and the smaller the delay fluctuation corresponding to the instruction, the closer the value of the stability weighting factor is to 1.
[0031] In a possible embodiment, the historical delay information includes a plurality of the actual delay times;
[0032] Then, before determining the margin stable delay, the method further includes:
[0033] Determine the average value and the range of each of the actual delay times.
[0034] Determine the stability weighting factor according to the ratio of the range to the average value.
[0035] In a possible embodiment, the method further includes:
[0036] When the instruction is emitted, start counting by a time counter.
[0037] When the instruction is completed, stop the counting of the time counter, obtain the current count value as the actual delay time of the instruction, and write it into the delay cache.
[0038] In a possible embodiment, at most N actual delay times for the same instruction are reserved in the delay cache, where N is any integer greater than 1.
[0039] When a new actual delay time of the instruction for which N actual delay times are reserved in the delay cache is written, replace the earliest written actual delay time corresponding to the instruction in the delay cache.
[0040] In a possible embodiment, when writing the actual delay time into the delay cache, the method further includes:
[0041] Judge whether the delay cache is full.
[0042] If so, based on the least recently used algorithm, delete the earliest written one among all the actual latency times corresponding to the least frequently used instruction.
[0043] Write the actual latency time to the idle position and establish a mapping relationship with the corresponding instruction.
[0044] In a possible embodiment, scheduling the execution order of each instruction based on the predicted issue time, the data dependency, and the hardware resource availability information includes:
[0045] Adjust the order of each instruction in the priority queue according to the sequence of the predicted issue time. Among them, the instruction with an earlier predicted issue time has a higher order in the priority queue.
[0046] For the instruction with the highest order in the priority queue:
[0047] Judge whether there is a data dependency for the instruction according to the data dependency.
[0048] If there is, lower the order of the instruction in the priority queue.
[0049] If not, judge whether the hardware resources corresponding to the instruction are available according to the hardware resource availability information.
[0050] If not available, lower the order of the instruction in the priority queue.
[0051] If available, issue the instruction.
[0052] In a possible embodiment, before judging whether there is a data dependency for the instruction according to the data dependency, it further includes:
[0053] Judge whether the instruction type of the instruction is a load instruction.
[0054] If so, judge whether the actual execution time of the load instruction at the current moment reaches the corresponding instruction execution latency.
[0055] If not, wait until the corresponding instruction execution latency is reached.
[0056] If reached, go to the step of judging whether there is a data dependency for the instruction according to the data dependency.
[0057] To solve the above technical problems, the present invention further provides an instruction scheduling device, including:
[0058] An acquisition module, configured to acquire the data dependency and the instruction execution latency of each instruction.
[0059] A prediction module, configured to determine the predicted issue time of each of the instructions according to the data dependency relationship and the instruction execution latency.
[0060] A scheduling module, configured to obtain hardware resource availability information, and schedule the execution order of each of the instructions based on the predicted issue time, the data dependency relationship, and the hardware resource availability information.
[0061] Wherein, the types of the instruction execution latency include dynamic latency and static latency, the obtaining module includes a dynamic latency obtaining unit, and the dynamic latency obtaining unit is configured to obtain the instruction execution latency of the type of dynamic latency.
[0062] The dynamic latency obtaining unit is further configured to: obtain historical latency information corresponding to the current instruction from a latency cache, and predict the dynamic latency through the historical latency information.
[0063] Wherein, the latency cache is configured to record the actual latency time of each of the instructions in the past as the historical latency information.
[0064] To solve the above technical problems, the present invention further provides an instruction scheduling system, including:
[0065] A memory, configured to store a computer program.
[0066] A processor, configured to implement the steps of the instruction scheduling method as described above when executing the computer program.
[0067] To solve the above technical problems, the present invention further provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the instruction scheduling method as described above are implemented.
[0068] To solve the above technical problems, the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the instruction scheduling method as described above are implemented.
[0069] An instruction scheduling method provided by the present invention provides a dynamic delay prediction mechanism for the dynamic delay in the instruction execution delay where the specific delay time cannot be determined in advance. Specifically, the historical delay information corresponding to each instruction is stored in a delay cache. When the instruction execution delay corresponding to an instruction is a dynamic delay, the dynamic delay is predicted through the historical delay information stored in the delay cache, so as to obtain an accurate dynamic delay for subsequent instruction scheduling. Based on the above dynamic delay prediction mechanism, this method incorporates the dynamic delay whose delay time cannot be determined in advance into the consideration scope of instruction scheduling. When performing instruction scheduling based on this method, for dynamic change scenarios in instruction issue times such as memory access delay and cache miss, the optimal issue timing of the instruction can also be accurately predicted, thereby further improving the effect of instruction scheduling on processor performance improvement.
[0070] The instruction scheduling device, system, computer program product, and computer-readable storage medium provided by the present invention correspond to the above method and have the same effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] In order to more clearly illustrate the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0072] Figure 1 It is a flowchart of an instruction scheduling method provided by the present invention.
[0073] Figure 2 It is a scheduling flowchart of a load instruction provided by the present invention.
[0074] Figure 3 It is a scheduling flowchart of a non-load instruction provided by the present invention.
[0075] Figure 4 It is a control flowchart of a delay cache provided by the present invention.
[0076] Figure 5 It is a structural diagram of an instruction scheduling architecture provided by the present invention.
[0077] Figure 6 It is a structural diagram of an instruction scheduling system provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0078] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.
[0079] The core of the present invention is to provide an instruction scheduling method, device, system, product and medium.
[0080] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0081] In modern processor design, out-of-order execution mechanism is a mainstream design solution. The out-of-order execution mechanism can disrupt the execution order of instructions in the pipeline, dynamically schedule instructions to fill pipeline gaps, thereby significantly improving execution efficiency. Specifically, in the instruction scheduling of a processor, the issue time of an instruction depends on three major factors: data dependency, hardware resource availability, and instruction execution latency.
[0082] Among them, data dependency refers to the sequential dependency formed between instructions due to the read-write relationship of operands. Only when the dependent instruction is executed and the corresponding operand becomes valid can the instruction depending on that operand be issued for execution.
[0083] Hardware resource availability refers to whether the hardware unit (such as an execution unit) required for instruction execution is idle in the current cycle. If the resource is unavailable, the instruction cannot be issued for execution even if the data dependency is satisfied.
[0084] Instruction execution latency refers to the time required for an instruction to complete from issuance (usually expressed in terms of the number of cycles), which is affected by factors such as the type of instruction and whether the cache is hit. Among them, instruction execution latency can be further divided into static latency and dynamic latency. Static latency is a fixed latency that can be determined at the processor design stage, such as the latency of a load instruction when the cache is hit. Dynamic latency refers to the execution time of an instruction being affected by changes in the runtime cache state or memory hierarchy, resulting in the execution latency being unable to be determined at the design or compilation stage, and generally needs to be determined according to the actual situation at runtime.
[0085] Based on the above, since the accurate dynamic latency cannot be obtained during the processor design phase and cannot be obtained before the instruction execution is completed in actual applications. Therefore, the current out-of-order execution mechanism mostly constructs the scheduling order based on static latency or dependencies, without considering the dynamic changes in the instruction issue time. Especially in scenarios such as memory access latency and cache misses, the current scheduling strategy is difficult to accurately predict the optimal issue time and perform corresponding instruction scheduling, resulting in resource waste, pipeline stalls, and the inability to fully utilize the processor performance.
[0086] To solve the above problems, the present invention provides an instruction scheduling method, as Figure 1 shown, including:
[0087] S1: Obtain the data dependencies and instruction execution latencies of each instruction.
[0088] S2: Determine the predicted issue time of each instruction according to the data dependencies and instruction execution latencies.
[0089] S3: Obtain the hardware resource availability information, and schedule the execution order of each instruction based on the predicted issue time, data dependencies, and hardware resource availability information.
[0090] Among them, the types of instruction execution latencies include dynamic latency and static latency. When the instruction execution latency is dynamic latency, obtaining the instruction execution latency includes:
[0091] Obtain the historical latency information corresponding to the current instruction from the latency cache, and predict the dynamic latency through the historical latency information; the latency cache is used to record the actual latency time of each instruction in the past as the historical latency information.
[0092] Step S1 is the key step of this method. By accurately predicting the dynamic latency that cannot be accurately determined during the design or compilation phase, this part of the instructions is incorporated into the instruction scheduling scope, realizing a larger range, more comprehensive, and more accurate instruction scheduling. Thereby better improving the execution efficiency of the processor pipeline and reducing the pipeline stall time.
[0093] For step S1, the key point of this method lies in how to obtain the dynamic latency that cannot be accurately determined during the design or compilation phase. A latency cache mechanism is introduced in this method. The latency cache is a dedicated part of the cache used to store information related to latency. Specifically, the latency cache in this method is used to store the historical latency information of instructions. The historical latency information corresponding to a certain instruction may include one or more actual latency times when the instruction was executed in the past, which is used to characterize the latency situation of the instruction in history. It should be noted that although this method does not limit which specific instructions are recorded with historical latency information in the latency cache. However, from the above description of instruction execution latency, static latency is a fixed latency that can be determined during the design phase and does not require recording historical latency information and prediction through the latency cache. Therefore, the instruction objects for which historical latency information is recorded in the latency cache in this method refer to some instructions whose instruction execution latency is dynamic latency, so as to reduce the requirement for the capacity of the latency cache and improve the scheduling efficiency.
[0094] Furthermore, this embodiment does not limit how to predict the dynamic latency based on the historical latency information recorded in the latency cache. In a relatively simple implementation, only one past actual latency time is included in the historical latency information of each instruction. Then, in this instruction scheduling, this actual latency time can be directly used as the predicted dynamic latency for this instruction scheduling. It is not difficult to understand that other processing can also be done on this actual latency time to achieve more accurate dynamic latency prediction. Also, when the historical latency information includes multiple actual latency times, other schemes for predicting dynamic latency can also be adopted, such as taking the average of multiple actual latency times, weighted summation, etc. This embodiment does not limit any of these.
[0095] For steps S2 and S3, from the description in the above embodiment part, it can be seen that the key to achieving accurate and efficient instruction scheduling lies in accurately predicting the optimal issue time of each instruction. And the optimal issue time of an instruction is related to three factors: data dependence relationship, hardware resource availability information, and instruction execution latency. Among them, the main factors determining whether an instruction can be issued are whether the operands on which the instruction depends are satisfied and whether the hardware units required for instruction execution are idle. Whether the above two conditions are satisfied can be known by real-time query. However, whether the data dependence is satisfied can be predicted through the data dependence relationship and instruction execution latency. According to the data dependence relationship, it can be determined which instructions each operand on which an instruction depends corresponds to. These instructions are the instructions on which the current instruction depends, and these instructions must be at the previous level of the current instruction on the data dependence chain (a relationship chain determined based on the data dependence relationship, and the execution of any level of instruction in the relationship chain directly depends on the operands of the previous level instruction), so they are called previous-level instructions.
[0096] Based on data dependency and instruction execution delay, the expected completion time of all previous instructions of any instruction can be predicted, that is, the expected time for the instruction to meet data dependency. Therefore, step S2 in this method first determines the expected time for each instruction to meet data dependency through data dependency and instruction execution delay, that is, predicts the emission time, and provides a preliminary parameter basis for the scheduling of each instruction. Based on the predicted emission time, each instruction can be preliminarily scheduled before actually meeting data dependency and hardware resource availability, which is conducive to improving the accuracy and efficiency of subsequent actual instruction scheduling. In addition, it is too rough to schedule instructions based only on whether data dependency is met and whether hardware resources are available. Both of them only have two situations of satisfaction and non-satisfaction, and the reference features brought are too few for realizing the scheduling of massive instructions in the processor. The expected completion time determined in step S2 of this method is a specific time value, which can contain more data features, which is conducive to more accurate scheduling of massive instructions in the processor.
[0097] However, it should be noted that the data dependency satisfaction time represented by the predicted launch time is only an expected time after all. In actual applications, it cannot be guaranteed that the instruction will satisfy the data dependency when the predicted launch time is reached. Moreover, the predicted launch time cannot predict when the instruction will meet the availability of hardware resources. Therefore, when actually performing instruction scheduling in step S3, the predicted launch time is only one of the reference factors. It is also necessary to judge whether the instruction truly satisfies the data dependency based on the data dependency relationship, and to judge whether the instruction truly satisfies the hardware resource availability based on the hardware resource availability information. Because the availability of the hardware resources required for instruction execution changes in real time, the hardware resource availability information obtained in advance is useless. This is why this method does not obtain the hardware resource availability information until step S3.
[0098] However, it should be noted that, as can be seen from the above description of step S1, the focus of this method is to accurately predict the dynamic delay that was originally not considered in the instruction scheduling, so as to accurately predict the best time to issue these instructions and achieve more comprehensive instruction scheduling. That is, this effect can be achieved based on step S1 provided above in this method.
[0099] That is, the specific implementation of steps S2 and S3 in this embodiment is only a possible implementation scheme, which provides more comprehensive and rich data support for instruction scheduling by predicting the emission time, which is conducive to improving the efficiency and accuracy of instruction scheduling. In practical applications, other currently mature instruction scheduling schemes can also be adopted. As long as the dynamic delay prediction method provided in step S1 of this method is adopted, better instruction scheduling effect can be achieved.
[0100] As can be seen from the above, an instruction scheduling method provided by the present application records the actual delay time in history for some instructions with a dynamic delay in the execution delay in the delay cache by introducing a delay cache mechanism. Furthermore, when these instructions participate in instruction scheduling subsequently, the dynamic delay can be predicted according to the historical delay information recorded in the delay cache, and the accurate prediction of the optimal emission timing of these instructions can be achieved, so that these instructions can be incorporated into the instruction scheduling scope. The instruction scheduling implemented based on this method can cover a larger range of instructions in the processor pipeline processing process. In particular, for scenarios where the emission time of instructions such as memory access delay and cache miss changes dynamically, the accurate prediction of the optimal emission timing can also be achieved, which can make better use of processor resources, reduce pipeline stalls, and thus give full play to the performance of the processor more fully.
[0101] On the other hand, in the above embodiment, it mainly illustrates how to obtain the dynamic delay among the two types of instruction execution delays. Then, for how to determine whether the instruction execution delay corresponding to an instruction is a dynamic delay or a static delay, and how to obtain the static delay, this embodiment provides a possible implementation solution. The above step S1 specifically further includes:
[0102] S11: Determine whether the instruction type of the instruction is a load instruction or a non-load instruction.
[0103] S12: If it is a load instruction, query the status of the data accessed by the instruction in the cache.
[0104] S13: If the status is cache hit, obtain the static delay corresponding to the instruction, and use the static delay as the instruction execution delay.
[0105] S14: If the status is cache miss, query the corresponding historical delay information in the delay cache, predict the dynamic delay according to the historical delay information, and use the dynamic delay as the instruction execution delay.
[0106] S15: If it is a non-load instruction, obtain the static delay corresponding to the instruction, and use the static delay as the instruction execution delay.
[0107] First of all, it should be noted that the above loading instructions refer to the instructions that need to read data from the cache or main memory, and the latency uncertainty is relatively high, and there is a possibility that the corresponding instruction execution latency is a dynamic latency. Non-loading instructions include arithmetic, logical, and control operation instructions, and the instruction latency is relatively stable, and the corresponding instruction execution latency generally belongs to static latency. Therefore, in this embodiment, it can be determined whether the type of the instruction belongs to the loading instruction in the decoding stage; if it is a non-loading instruction, it can be directly considered that the instruction execution latency of the instruction is a static latency, and the corresponding fixed latency time determined in advance can be obtained; if it is a loading instruction, there is a possibility that the instruction execution latency is a dynamic latency. At this time, in this embodiment, the latency type is further determined by querying the state of the data accessed by the instruction in the cache; if the cache hits, it means that the latency type of the instruction is a static latency, and the fixed latency time determined in advance can be directly obtained; if the cache misses, it means that the latency type of the instruction is a dynamic latency, and the corresponding historical latency information needs to be queried from the latency cache, and a specific dynamic latency is predicted based on the historical latency information as the instruction execution latency of the instruction.
[0108] It should also be noted that this embodiment mainly gives further possible implementation schemes for the part of obtaining the instruction execution latency in step S1, and there are still no restrictions on the part of obtaining the data dependence relationship and the hardware resource availability information in step S1. The above steps S11 to S15 do not represent all of the steps S1.
[0109] As can be seen from the above method flow, in this embodiment, by determining the instruction type, all instructions are divided into two categories: loading instructions and non-loading instructions, and it is initially distinguished whether there is a possibility that the corresponding instruction execution latency is a dynamic latency. For loading instructions, there is a possibility of dynamic latency, so it is further determined whether it is a dynamic latency by judging the cache state of the data accessed by the instruction; if the cache hits, it is a static latency; if the cache misses, it is a dynamic latency, and it is necessary to implement the prediction of the specific latency time based on the latency cache mechanism. This embodiment provides a two-layer query scheme based on the instruction type and the cache state to determine whether the latency type corresponding to the instruction is a dynamic latency or a static latency, and obtain the specific latency time. And the two-layer query scheme provided by this embodiment uses the instruction type that can be determined in the decoding stage to reduce the number of steps of querying the cache state in the follow-up. For non-loading instructions with relatively stable and fixed latency, it can be directly determined that the instruction execution latency is a static latency without querying the cache state, reducing the additional query time-consuming to improve the overall efficiency.
[0110] On the other hand, as can be seen from the above description, step S2 of the present method predicts the time when each instruction is expected to satisfy data dependencies based on data dependencies and instruction execution latency, that is, obtains the predicted issue time. Based on the predicted issue time, preliminary scheduling can be performed before the instruction actually satisfies data dependencies and hardware resource availability, and it can also provide data support for instruction scheduling from other aspects, which is conducive to improving the accuracy and efficiency of instruction scheduling.
[0111] Therefore, for how to specifically determine the predicted issue time in the above step S2, this embodiment also provides a possible implementation scheme. The above step S2 specifically further includes:
[0112] S21: Determine the data dependency chains among the instructions according to the data dependencies, and take the head instructions in each data dependency chain as producer instructions, and the remaining instructions as consumer instructions.
[0113] S22: For the producer instructions, the predicted issue time is the current time.
[0114] S23: For the consumer instructions, the predicted issue time is the maximum value among the instruction completion times corresponding to the previous instructions of the current consumer instruction.
[0115] Among them, the previous instructions are all the instructions at the previous level of the current consumer instruction in the data dependency chain, and the instruction completion time is determined by the predicted issue time and the instruction execution latency.
[0116] In this embodiment, based on the position of the instruction in the data dependency chain, the instructions are divided into two different types: producer instructions and consumer instructions. Among them, the above producer instructions refer to the instructions that generate operands for other instructions to use, usually load write instructions or write-back instructions. Consumer instructions refer to the instructions that depend on the results of producer instructions as inputs, and their issue times are restricted by the completion times of producer instructions.
[0117] It should be noted that the dependence of the above consumer instructions on producer instructions is not necessarily a direct dependence, and it can be an indirect dependence with other consumer instructions in between. Exemplarily, assume there is a simple data dependency chain like "Instruction 1 → Instruction 2 → Instruction 3". Among them, the execution of Instruction 3 depends on the operands of Instruction 2, the execution of Instruction 2 depends on the operands of Instruction 1, and the execution of Instruction 1 does not depend on the operands of any other instructions. Then Instruction 1 is a producer instruction, and both Instruction 2 and Instruction 3 are consumer instructions, and Instruction 1 is the previous instruction of Instruction 2, and Instruction 2 is the previous instruction of Instruction 3. It should also be noted that the above example is only one possible situation. In fact, the previous instructions of an instruction are not necessarily unique, and there may be multiple previous instructions at the same time.
[0118] Therefore, as in the above example, step S21 can establish a data dependence chain based on the data dependence relationship. It is not difficult to understand that the data dependence chain is established for a single instruction or a small number of instructions. When the data dependence chains corresponding to all instructions are determined, they can be further combined into a data dependence network. However, considering that the prediction of the issue time is based on a single instruction as a unit, only the data dependence chain corresponding to this instruction needs to be determined. And to determine the data dependence chain for predicting the issue time of this instruction, it only needs to be determined up to the end of this instruction.
[0119] After that, for step S22, it is to determine the predicted issue time for the producer instruction defined above. It is not difficult to understand that since the producer instruction does not depend on the operands of any other instructions, it satisfies the data dependence at any time. Therefore, step S22 takes the earliest time that can be obtained currently, that is, the current time, as the predicted issue time of the producer instruction.
[0120] For step S23, it is to determine the predicted issue time for the consumer instruction defined above. As can be seen from the above, when the consumer instruction satisfies the data dependence depends on whether all its preceding instructions have been executed. Since the instruction completion time (moment) can simply be considered as the instruction issue time (moment) + the instruction execution delay. Therefore, the predicted issue time of each instruction can be determined by the following formulas:
[0121] (1);
[0122] In the formula, TComplete(p) represents the instruction completion time of the producer instruction p; TIssue(p) represents the instruction issue time of the producer instruction p (not equivalent to the predicted issue time); TDelay(p) represents the instruction execution delay of the producer instruction p.
[0123] It should be noted that there is no corresponding relationship between the producer instruction and the consumer instruction defined in this embodiment and the load instruction and the non-load instruction in the above embodiment. The producer instruction can be a load instruction or a non-load instruction, and the corresponding instruction execution delay can be a static delay or a dynamic delay, which can be obtained through the instruction execution delay acquisition scheme provided in the above embodiment, and this embodiment does not limit this.
[0124] (2);
[0125] In the formula, TPredicted(c) represents the predicted issue time of the consumer instruction c; TComplete(i) represents the instruction completion time of the preceding instruction i of the consumer instruction c; N represents the total number of all preceding instructions i corresponding to the consumer instruction c. It is not difficult to understand that the preceding instruction of the consumer instruction may be a producer instruction or other consumer instructions.
[0126] Based on this embodiment, the predicted issue time of each instruction can be determined based on the data dependence relationship, thereby providing data support for subsequent instruction scheduling from another aspect, which is beneficial to improving the efficiency and accuracy of instruction scheduling.
[0127] Furthermore, as can be seen from the above embodiment, when the instruction execution delay corresponding to the producer instruction is a dynamic delay, it is necessary to obtain the corresponding historical delay information from the delay cache to predict the dynamic delay. Considering the large volatility of the dynamic delay, in order to more accurately predict the predicted issue time of the instruction, this embodiment also provides a further implementation scheme. After step S22 of the above method, it further includes:
[0128] S24: Determine whether the instruction execution delay corresponding to the producer instruction is a static delay or a dynamic delay;
[0129] S25: If it is a dynamic delay, determine the instruction completion time according to the predicted issue time, the dynamic delay, and the stability weighting factor.
[0130] That is, in this embodiment, the predicted dynamic delay is further corrected by the stability weighting factor. The stability weighting factor is used to characterize the degree of delay fluctuation of the instruction during historical execution. Based on the stability weighting factor, if the delay fluctuation of the instruction is large, the dynamic delay can be appropriately increased to increase the issue waiting time of the instruction, thereby reducing the risk of prediction failure.
[0131] Furthermore, for the value of the stability weighting factor and its specific way of adjusting the dynamic delay, this embodiment provides a possible implementation scheme. The above step S25 specifically further includes:
[0132] S251: Determine the margin stable delay according to the product of the stability weighting factor and the dynamic delay.
[0133] S252: Determine the instruction completion time according to the sum of the margin stable delay and the predicted issue time.
[0134] Among them, the stability weighting factor is greater than or equal to 1, and the smaller the delay fluctuation of the corresponding instruction, the closer the value of the stability weighting factor is to 1.
[0135] That is, the above formula (1) becomes the following formula (3):
[0136] (3);
[0137] In the formula, β(p)·TDelay(p) represents the margin stable delay; β(p) represents the stability weighting factor of the producer instruction p.
[0138] In this embodiment, the stability weighting factor is used to reflect the degree of fluctuation in the latency of instructions during historical execution. When the latency fluctuation is relatively stable, the credibility of the predicted dynamic latency based on the historical latency information recorded in the latency cache is relatively high. Therefore, the value of the stability weighting factor is closer to 1, meaning that there is less need to increase the dynamic latency. However, when the latency fluctuation is large, it indicates that the credibility of the predicted dynamic latency is low. At this time, it is necessary to appropriately increase the predicted dynamic latency as the instruction execution latency during instruction scheduling, that is, the value of the stability weighting factor is appropriately increased to reduce the risk of dynamic latency prediction failure by increasing the issue waiting time of the instruction.
[0139] Furthermore, regarding how to determine the value of the stability weighting factor in the above embodiment, the above embodiment has already explained that the value of the stability weighting factor is related to the magnitude of the historical latency fluctuation of the instruction. And the historical latency fluctuation of the instruction can be reflected by the historical latency information recorded in the latency cache. Therefore, the stability weighting factor can be determined through the historical latency information.
[0140] However, this embodiment also provides a possible scheme for determining the value of the stability weighting factor:
[0141] The historical latency information includes multiple actual latency times.
[0142] Then, before step S251, the method further includes:
[0143] S250: Determine the average value and range of each actual latency time, and determine the stability weighting factor according to the ratio of the range to the average value.
[0144] Specifically, the above step S250 can be expressed by the following formula:
[0145] (4);
[0146] In the formula, γ represents a preset sensitivity coefficient, which characterizes the sensitivity of the stability weighting factor to latency fluctuations and is used to control the influence degree of latency fluctuations on the stability weighting factor. represents the range of the actual latency time, represents the average value of the actual latency time, reflects the relative fluctuation degree of the historical latency. If the fluctuation is larger, it indicates that cache misses are more frequent. At this time, the stability weighting factor is increased.
[0147] In the above formula (4) scheme, a linear relationship between β(p) and is established through the sensitivity coefficient. In another possible implementation scheme, the stable weighting factor can also be determined directly by using the interval mapping method, as shown in Table 1 below.
[0148] Table 1 Stability Weighting Factor Mapping Relationship Table
[0149]
[0150] It is not difficult to see from Table 1 above that in this embodiment, several relatively reasonable β(p) setting values that vary in gradient are directly preset based on experience or tests. A mapping relationship is established between each β(p) setting value and a range. Thus, after falls into a certain range, the corresponding unique β(p) setting value can be directly determined through the mapping relationship.
[0151] It is not difficult to see that both of the above two solutions can achieve the purpose of determining the stability weighting factor according to the ratio of the range to the average value as described in step S250, and in practical applications, a suitable implementation solution can be freely selected according to needs.
[0152] On the other hand, the above embodiments give several possible implementation solutions for how to specifically implement step S2, and this embodiment also gives a more likely implementation solution for the implementation of step S3. In the above embodiments, it is explained that step S3 realizes the scheduling of instructions based on three major factors: predicted emission time, data dependency, and hardware resource availability information. Among them, whether the instruction satisfies data dependency and whether it satisfies hardware resource availability directly determine whether the instruction can be emitted. The predicted emission time represents the time when the instruction is expected to satisfy data dependency, and more provides a scheduling reference in the time dimension. Based on this, this embodiment provides a possible implementation solution. The specific instruction scheduling in the above step S3 includes:
[0153] S31: Adjust the order of each instruction in the priority queue according to the sequence of the predicted emission time.
[0154] Among them, the earlier the predicted emission time of the instruction, the higher its order in the priority queue.
[0155] For the instruction with the highest order in the priority queue:
[0156] S32: Judge whether there is data dependency for the instruction according to the data dependency relationship; if there is, go to step S33; if not, go to step S34.
[0157] S33: Lower the order of the instruction in the priority queue.
[0158] S34: Judge whether the hardware resources corresponding to the instruction are available according to the hardware resource availability information; if not, go to step S35; if available, go to step S36.
[0159] S35: Lower the order of the instruction in the priority queue.
[0160] S36: Emit the instruction.
[0161] It should be noted that although both step S33 and step S35 are to lower the order of the instruction in the priority queue (i.e., the priority of emission), the specific adjustment schemes are different. In step S33, since the instruction does not satisfy data dependence and needs to lower the priority, its order should be adjusted after other instructions that satisfy data dependence. Similarly, in step S35, since it does not satisfy the availability of hardware resources and needs to lower the priority, its order should be adjusted after other instructions that satisfy the availability of hardware resources. When both data dependence and the availability of hardware resources are satisfied, the instruction is allowed to be emitted, and the instructions are emitted in sequence according to the order in the priority queue for execution.
[0162] This embodiment provides a specific instruction scheduling strategy, which preliminarily schedules the priority of instructions based on the predicted emission time. Then, in the order in the priority queue, it is sequentially determined whether each instruction satisfies the emission conditions (data dependence and the availability of hardware resources); if any condition is not satisfied, its order is adjusted after other instructions that satisfy this condition; if all conditions are satisfied, the instructions can be emitted sequentially according to the order in the priority queue, thereby realizing the effective scheduling of instructions.
[0163] Furthermore, as can be seen from the above embodiment for step S1, instructions can be divided into two types: load instructions and non-load instructions. Among them, those with a dynamic delay in instruction execution belong to load instructions (but load instructions are not necessarily with dynamic delay). And load instructions refer to instructions that need to read data from the cache or main memory. Therefore, for the emission of load instructions, it is necessary to additionally consider whether the data has been read. If the data has not been read, the instruction cannot be emitted, that is, there is no need for the scheduling and emission judgment as in steps S31 to S36 above.
[0164] Based on this, for load instructions, this embodiment also provides an adaptive precise check scheme. Before the above step S32, this method further includes:
[0165] S37: Determine whether the instruction type of the instruction is a load instruction.
[0166] S38: If it is, determine whether the actual execution time of the load instruction at the current moment reaches the corresponding instruction execution delay; if not, wait until the corresponding instruction execution delay is reached; if it reaches, go to step S32.
[0167] However, as can be seen from the above embodiment for step S1, the judgment on whether the instruction is a load instruction has already been carried out as early as the stage of obtaining the instruction execution delay. Therefore, for load instructions, there is an instruction scheduling process as Figure 2 shown, including the whole process of obtaining the instruction execution delay, instruction scheduling, and instruction emission. From Figure 2It is not difficult to see that the load instruction can still perform preliminary instruction scheduling based on step S31. However, before making a subsequent determination on whether to allow emission and updating the priority, it is necessary to ensure that the instruction has waited until the corresponding predicted emission time is reached, and then perform precise checks as carried out in steps S32 to S36.
[0168] Similarly, for non-load instructions that do not need to read data from the cache or main memory, there is no need to wait until after the predicted emission time, and precise checks can be started immediately. The entire process of instruction scheduling for non-load instructions is as Figure 3 shown.
[0169] In summary, the entire process of implementing instruction scheduling by this method is obtained. In this embodiment, the instructions are divided into load instructions and non-load instructions and different instruction scheduling processes are provided. For load instructions that need to read data from the cache or main memory, it is first determined whether the load instruction has reached the predicted emission time before performing precise checks; if it has reached, resources are consumed to perform precise checks; if it has not reached, it waits until it reaches, and during this period, resources are not wasted to perform precise checks, so as to further improve the efficiency of instruction scheduling.
[0170] On the other hand, as can be seen from the above, one of the keys to improving the efficiency and accuracy of instruction scheduling by this method is the accurate prediction of dynamic latency. And the prediction of dynamic latency depends on the implementation of a latency cache. The latency cache records the actual execution time of instructions during the historical execution process as historical latency information for use in the prediction of dynamic latency during subsequent instruction scheduling.
[0171] Therefore, for how the latency cache implements the recording of historical latency information, this embodiment provides a possible implementation scheme. The above method further includes:
[0172] S41: When an instruction is emitted, start counting by a time counter.
[0173] S42: When the instruction is completed, stop the counting of the time counter, obtain the current count value as the actual latency time of the instruction, and write it into the latency cache.
[0174] That is to say, in this embodiment, after an instruction is emitted, the time counter starts counting until the instruction is executed and then stops counting, so as to obtain the actual latency time of the instruction and record it in the latency cache. Then, when it is necessary to predict the dynamic latency of the instruction during subsequent instruction scheduling, the dynamic latency can be accurately predicted based on the actual latency time recorded in the latency cache.
[0175] In addition, this embodiment also provides a possible implementation scheme for the structure of the latency cache:
[0176] The latency cache adopts a direct mapped memory structure. The data structure is shown in Table 2, which records the instruction ID, latency period (i.e., the actual latency time represented by the number of cycles), and valid bit.
[0177] Table 2 Latency Cache Data Structure
[0178]
[0179] Among them, the instruction ID is used to uniquely identify an instruction (load instruction), usually using the program counter (PC) value. The execution time counter (i.e., the time counter used to count the instruction execution time in steps S41 and S42) records the execution latency period of the instruction in the most recent N times (N = 5 is used as an example in Table 2), that is, the actual latency time. The valid bit is used to indicate the validity of this record, which is set to 1 only when a cache miss occurs for the instruction, and no latency value is recorded when a hit occurs.
[0180] Furthermore, as can be seen from the above embodiments, when instructions are divided into two categories: load instructions and non-load instructions, generally only load instructions may have dynamic latency. And it is not easy to determine whether the instruction execution latency of an instruction is static latency or dynamic latency during the decoding stage, but it can be determined whether the instruction is a load instruction or a non-load instruction. And for load instructions, it can be judged whether it corresponds to dynamic latency by querying whether its access to the data cache hits.
[0181] Therefore, as Figure 4 shown, the instruction objects targeted by the latency cache control flow provided in this embodiment only include load instructions. For load instructions, first query whether its cache hits. If it hits, it means that the load instruction corresponds to static latency, and there is no need to provide historical latency information, nor to record historical latency information. If the cache misses, the latency cache needs to provide historical latency information. Further, there are two cases: one is that the latency cache hits, which means that the historical latency time of this instruction exists in the latency cache and can be read out; the other is that the latency cache misses, which means that there is no historical latency information of this instruction in the latency cache, so historical latency information needs to be recorded in the latency cache.
[0182] On the other hand, the above embodiments also illustrate that for each instruction in the latency cache, one or more actual latency times can be recorded as historical latency information. In practical applications, to improve the accuracy of dynamic latency prediction, and a determination scheme of a stability weighting factor is also required with multiple actual latency times in some of the above embodiments, so this embodiment provides a better implementation scheme:
[0183] The latency cache retains multiple actual latency times for the same instruction,
[0184] Furthermore, since the latency cache is a separate area partitioned from the cache, if its capacity is too large, it will affect the normal function of the processor. When the latency cache is full, it is unable to store new historical latency times for subsequent instruction scheduling. To solve this problem, this embodiment also provides a possible implementation:
[0185] In the latency cache, at most N actual latency times are retained for the same instruction, where N is any integer greater than 1. And when a new actual latency time is written for an instruction that already has N actual latency times retained in the latency cache, the earliest written actual latency time corresponding to that instruction in the latency cache is replaced.
[0186] That is, in this embodiment, at most N (such as 5) actual latency times can be retained for the same instruction. When N actual latency times are reached, and a new actual latency time is written, the earliest written actual latency time is automatically overwritten. On the one hand, the amount of data of the historical latency information for each instruction is controlled, and on the other hand, the automatic update of the historical latency information is achieved. The historical latency information recorded in the latency cache can always reflect the actual latency situation of the latency fluctuations of the instruction in the recent period, which is beneficial to improving the accuracy of dynamic latency prediction.
[0187] Furthermore, in addition to the automatic update scheme when the historical latency information of a single instruction reaches the specified quantity, this embodiment also provides a writing scheme for the latency cache. When writing the actual latency time into the latency cache, the above method further includes:
[0188] S51: Determine whether the latency cache is full; if so, go to step S52; if not, go to step S53.
[0189] S52: Based on the least recently used algorithm, delete the earliest written one among all the actual latency times corresponding to the least frequently used instruction.
[0190] S53: Write the actual latency time to the idle position and establish a mapping relationship with the corresponding instruction.
[0191] The specific process is as follows Figure 4As shown, it can be seen from the above embodiments that a delay cache miss indicates that there is no historical delay information of the current instruction in the delay cache that needs to be delayed, and it is necessary to start timing and record the actual delay time. At this time, based on step S51, it is necessary to first determine whether the delay cache is full; if it is full, then corresponding to step S52, the least recently used (LRU) algorithm is used to determine the least frequently used instruction as the replacement target, and since what is written is an actual delay information, so the earliest written one among all the actual delay times corresponding to the least frequently used instruction is used as the replacement object for replacement writing to ensure that the capacity of the delay cache is not exceeded to achieve data update. Further, in combination with the delay cache mapping structure provided by the corresponding embodiments in Table 2 above, when it is necessary to record the actual delay time, record the ID of the instruction, raise the valid bit, and start counting through the execution time counter. Until the instruction is submitted, the execution time counter stops counting and records the current count value.
[0192] As can be seen from the above, this embodiment provides a method for replacement and update of a delay cache when it is full of data, ensuring that accurate prediction of dynamic delay can still be guaranteed on the premise that the capacity requirement for the delay cache is not too large.
[0193] On the other hand, for the hardware architecture implemented by an instruction scheduling method provided in the above embodiments, this embodiment also provides a possible implementation scheme. As Figure 5 shown, an instruction scheduling architecture 10 includes: a decoding module 11, a dependency record table 12, a delay cache 13, an issue time prediction module 14, a scheduling module 15, and a priority queue 16.
[0194] Among them, the decoding module is an original module in the processor, which is used to decode instructions and can obtain information such as instruction types and data dependency relationships. The data dependency relationships are recorded in the dependency record table, and when the instruction corresponds to a dynamic delay and there is no historical delay information record in the delay cache, the delay cache records the historical delay information of the instruction.
[0195] The dependency record table is used to manage the dependency relationships of physical registers in an out-of-order execution environment, maintain the ready state of registers, and further manage the data dependencies between instructions to facilitate the subsequent issue time prediction module to predict the issue time. A possible data structure of the dependency record table is shown in Table 3 below, which records the source register, the mapped physical register, and the ready flag bit.
[0196] Table 3 Data Structure of Dependency Record Table
[0197]
[0198] Specifically, the ready flag is used to indicate whether the data in the physical register is valid. If the ready flag is 1, it means that the value in the source register has been calculated and the instruction can be directly used. If the ready flag is 0, it means that the value in the physical register depends on the calculation result of the previous instruction, forming a read after write (RAW) dependency. The subsequent instruction must wait for the previous instruction to complete before it can be executed, which is a true dependency. In actual execution, if a load instruction waits for the main memory to return data due to a cache miss, the corresponding target register will be in an "unready" state for a long time, thus affecting the scheduling of consumer instructions.
[0199] The emission time prediction module therein is the functional module for implementing the functions of steps S1 and S2 above. By obtaining the data dependency relationship from the dependency record table and predicting the dynamic delay through the delay cache, the predicted emission time of each instruction is determined.
[0200] The scheduling module corresponds to step S3 above. After determining the predicted emission time, the initial scheduling of each instruction is performed through the predicted emission time. The priority queue is a queue for managing the execution order of each instruction. Based on the position of the instruction in the priority queue, the management and scheduling of the instruction emission time and execution order are realized.
[0201] Moreover, to better illustrate an instruction scheduling method provided by the present method, this embodiment also gives a specific example in combination with a specific implementation scenario:
[0202] Assume that the processor is a single-issue processor with 13 instructions, some of which have data dependencies and include multiple load instructions. Assume that the L1 cache hits with a delay of 1 cycle, the L2 cache hits with a delay of 2 cycles, and the cache miss has a delay of 8 cycles. The emission time prediction adopts equations 1 to 4. If the load instruction misses, the correction value in the delay cache is used for delay adjustment. The instructions and instruction descriptions are shown in Table 4. Assume that instruction 1 hits the L1 cache, instruction 3 hits the L2 cache, and instruction 5 misses the cache.
[0203] Table 4 Priority queue and instruction description before adjustment
[0204]
[0205] As can be seen from the above, instruction 1 hits the L1 cache with a static delay of 1 cycle; instruction 3 hits the L2 cache with a static delay of 2 cycles; instruction 5 misses the cache. Assume that the historical delays in the delay cache are four actual delay times of 8, 9, 7, and 8 (expressed in cycle numbers), and the average value is , the fluctuation range , according to formula 4, take , and get , used to correct the execution delay to 8 * 1.25 = 10 cycles. That is, the dynamic delay of the predicted instruction 5 after correction is 10 cycles.
[0206] Furthermore, schedule the instructions shown in Table 4 based on the above method to obtain the optimized Table 5.
[0207] Table 5 Adjusted Priority Queue and Instruction Description
[0208]
[0209] In the above emission timing, the efficiency of optimizing instruction scheduling by this method can be seen. First, the load instruction is predicted and corrected through the delay cache in the scheduling stage. Taking instruction 5 (LOAD R6) as an example, due to cache miss, its delay is no longer directly set to a constant, but combined with historical execution records, and the actual delay of 10 cycles is obtained through weighted factor correction. The processor predicts its completion time based on this, avoiding resource waiting and execution failure caused by the premature occurrence of the consumer instruction R7.
[0210] Secondly, in scenarios where there are data dependencies among instruction 2 (ADD R2), instruction 4 (MUL R5), instruction 7 (XOR R8), etc., dynamically deduce their available emission times through Equations 1 and 2 to ensure that consumer instructions do not enter the execution stage prematurely when the operands are not ready.
[0211] Thirdly, for non-load instructions without direct dependencies, such as instruction 11 (AND R11) and instruction 12 (NOR R12), the scheduling module automatically identifies whether they can be emitted in the current cycle and flexibly inserts them during the waiting for dependent data resources or resource conflicts, effectively filling the pipeline gaps and improving the single-cycle emission efficiency.
[0212] In addition, in the second half of the instruction scheduling, instruction R13 (NAND R13) and instruction R10 (STORE R10) respectively depend on multiple upstream operands. When the system calculates their predicted emission times, by integrating the completion times of multiple producers, it ensures that the final emission time meets the requirements of the latest dependency readiness, realizes fine-grained scheduling control, and avoids resource preemption.
[0213] Meanwhile, taking instruction 6 (SUB R7) as an example, its predicted issue time is the 12th cycle. This time is derived by the system based on the completion times of producer instructions R6 and R5. Among them, the R6 cache misses, and the delay is weighted and corrected to 10 cycles. After being issued in the 2nd cycle, it should be completed in the 12th cycle; R5 is completed in the 6th cycle and is ready earlier. When the system runs to the 12th cycle, the scheduling module determines that P6 can be issued according to the prediction result. However, to avoid scheduling failures caused by delay fluctuations or resource occupancy, the scheduling module performs an accurate check. Only when R6 and R5 are ready and the execution unit is idle and available, the instruction is issued. If any condition is not met, the instruction remains in the scheduling queue, and the priority is updated to wait for rechecking in the next cycle.
[0214] It can be seen that through the accurate prediction of the instruction issue time, this method optimizes and improves the instruction scheduling in the scenario of dynamic delay changes. While improving the instruction scheduling efficiency and reducing the idle cycles, it further improves the accuracy of the instruction issue time, and has certain energy efficiency optimization and system expansion functions, providing an effective optimization method for the design of modern high-performance processors.
[0215] In the above embodiment, an instruction scheduling method is described in detail. The present invention also provides an embodiment corresponding to an instruction scheduling device. This embodiment provides an instruction scheduling device, including:
[0216] An acquisition module, configured to acquire the data dependency relationship and instruction execution delay of each instruction;
[0217] A prediction module, configured to determine the predicted issue time of each instruction according to the data dependency relationship and instruction execution delay;
[0218] A scheduling module, configured to acquire the hardware resource availability information, and schedule the execution order of each instruction based on the predicted issue time, data dependency relationship, and hardware resource availability information;
[0219] Among them, the types of instruction execution delay include dynamic delay and static delay. The acquisition module includes a dynamic delay acquisition unit, and the dynamic delay acquisition unit is configured to acquire the instruction execution delay of the type of dynamic delay;
[0220] The dynamic delay acquisition unit is further configured to: acquire the historical delay information corresponding to the current instruction from the delay cache, and predict the dynamic delay through the historical delay information;
[0221] Among them, the delay cache is used to record the actual delay time of each instruction in the past as historical delay information.
[0222] Since the embodiments of the device part correspond to the embodiments of the method part, for the embodiments of the device part, please refer to the description of the embodiments of the method part, and will not be elaborated here for the time being.
[0223] In addition to the embodiments of an instruction scheduling method provided in the above embodiments, the present invention also provides an embodiment corresponding to a computer program product. A computer program product includes computer programs / instructions which, when executed by a processor, can implement the steps of the instruction scheduling method described in any of the above embodiments.
[0224] Since the embodiments of the computer program product part correspond to the embodiments of the method part, for the embodiments of the computer program product part, please refer to the description of the embodiments of the method part, which will not be elaborated here for the time being.
[0225] Figure 6 The following is a structural diagram of an instruction scheduling system provided in another embodiment of the present invention. As Figure 6 shown, an instruction scheduling system includes: a memory 20 for storing computer programs;
[0226] a processor 21 for implementing the steps of an instruction scheduling method as described in the above embodiment when executing the computer program.
[0227] The instruction scheduling system provided in this embodiment may include, but is not limited to, a mobile terminal, a personal computer, a workstation, etc.
[0228] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0229] The memory 20 may include one or more computer-readable storage media, which may be non-transitory. The memory 20 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, the related steps of an instruction scheduling method disclosed in any of the foregoing embodiments can be implemented. In addition, the resources stored in the memory 20 may also include an operating system 202, data 203, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, an instruction scheduling method, etc.
[0230] In some embodiments, an instruction scheduling system may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.
[0231] Those skilled in the art can understand that Figure 6 the structure shown in does not constitute a limitation on an instruction scheduling system, and may include more or fewer components than shown in the figure.
[0232] An instruction scheduling system provided by an embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, the following method can be implemented: an instruction scheduling method.
[0233] Finally, the present invention also provides an embodiment corresponding to a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps recorded in the foregoing method embodiment are implemented.
[0234] It can be understood that if the method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and executes all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0235] The above has introduced in detail an instruction scheduling method, apparatus, system, product, and medium provided by the present invention. The various embodiments in the specification are described in a progressive manner, and the key point of each embodiment is the difference from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and reference can be made to the description of the method part for related parts. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
[0236] It should also be noted that in this specification, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or device including the said element.
Claims
1. An instruction scheduling method, characterized in that, including: obtaining the data dependency relationship and instruction execution latency of each instruction; determining the predicted issue time of each instruction according to the data dependency relationship and the instruction execution latency; obtaining the hardware resource availability information, and scheduling the execution order of each instruction based on the predicted issue time, the data dependency relationship and the hardware resource availability information; wherein, the types of the instruction execution latency include dynamic latency and static latency; when the instruction execution latency is dynamic latency, obtaining the instruction execution latency includes: obtaining the historical latency information corresponding to the current instruction from the latency cache, and predicting the dynamic latency through the historical latency information; the latency cache is used to record the actual latency time of each instruction in the past as the historical latency information; obtaining the instruction execution latency includes: determining whether the instruction type of the instruction is a load instruction or a non-load instruction; if it is a load instruction, querying the status of the data accessed by the instruction in the cache; if the status is cache hit, obtaining the static latency corresponding to the instruction, and using the static latency as the instruction execution latency; if the status is cache miss, querying the corresponding historical latency information in the latency cache, predicting the dynamic latency according to the historical latency information, and using the dynamic latency as the instruction execution latency; if it is a non-load instruction, obtaining the static latency corresponding to the instruction, and using the static latency as the instruction execution latency.
2. The instruction scheduling method according to claim 1, wherein Determining the predicted issue time of each instruction according to the data dependency relationship and the instruction execution latency includes: determining the data dependency chain between each instruction according to the data dependency relationship, and taking the first-level instruction in each data dependency chain as the producer instruction, and the remaining instructions as the consumer instructions; for the producer instruction, the predicted issue time is the current time; for the consumer instruction, the predicted issue time is the maximum value of the instruction completion times corresponding to the previous-level instructions of the current consumer instruction; wherein, the previous-level instructions are all the instructions located in the previous level of the current consumer instruction in the data dependency chain; the instruction completion time is determined by the predicted issue time and the instruction execution latency.
3. The instruction scheduling method according to claim 2, wherein For the producer instruction, after determining the predicted issue time, the method further includes: judging whether the instruction execution latency corresponding to the producer instruction is static latency or dynamic latency; if it is the dynamic latency, determining the instruction completion time according to the predicted issue time, the dynamic latency and the stability weighting factor.
4. The instruction scheduling method according to claim 3, wherein Determining the instruction completion time according to the predicted issue time, the dynamic latency and the stability weighting factor includes: determining the margin stable latency according to the product of the stability weighting factor and the dynamic latency; determining the instruction completion time according to the sum of the margin stable latency and the predicted issue time; wherein, the stability weighting factor is greater than or equal to 1, and the smaller the latency fluctuation corresponding to the instruction, the closer the value of the stability weighting factor is to 1.
5. The instruction scheduling method according to claim 4, wherein The historical latency information includes a plurality of the actual latency times; Before determining the margin stable delay, the method further includes: Determining the average value and the range of each of the actual delay times; Determining the stability weighting factor according to the ratio of the range to the average value.
6. The instruction scheduling method according to claim 1, wherein The method further includes: When the instruction is issued, starting to count by a time counter; When the instruction is completed, stopping the counting of the time counter, obtaining the current count value as the actual delay time of the instruction, and writing it into the delay buffer.
7. The instruction scheduling method according to claim 6, wherein In the delay buffer, at most N actual delay times for the same instruction are retained, where N is any integer greater than 1; When a new actual delay time is written for an instruction for which N actual delay times are retained in the delay buffer, replacing the earliest-written one of the actual delay times corresponding to the instruction in the delay buffer.
8. The instruction scheduling method according to claim 6, wherein When writing the actual delay time into the delay buffer, the method further includes: Judging whether the delay buffer is full; If so, based on the least recently used algorithm, deleting the earliest-written one of all the actual delay times corresponding to the least frequently used instruction; Writing the actual delay time into the idle position and establishing a mapping relationship with the corresponding instruction.
9. The instruction scheduling method according to claim 1, wherein Scheduling the execution order of each instruction based on the predicted issue time, the data dependence relationship, and the hardware resource availability information includes: Adjusting the order of each instruction in the priority queue according to the sequence of the predicted issue times; wherein, the earlier the predicted issue time of an instruction, the higher the order of the instruction in the priority queue; For the instruction with the highest order in the priority queue: Judging whether there is a data dependence for the instruction according to the data dependence relationship; If there is, lowering the order of the instruction in the priority queue; If not, judging whether the hardware resources corresponding to the instruction are available according to the hardware resource availability information; If not available, lowering the order of the instruction in the priority queue; If available, issuing the instruction.
10. The instruction scheduling method according to claim 9, wherein Before judging whether there is a data dependence for the instruction according to the data dependence relationship, it further includes: Judging whether the instruction type of the instruction is a load instruction; If so, judging whether the actual execution time of the load instruction at the current moment reaches the corresponding instruction execution delay; If not, waiting until the corresponding instruction execution delay is reached; If reached, proceeding to the step of judging whether there is a data dependence for the instruction according to the data dependence relationship.
11. An instruction scheduling device, characterized in that, Including: An acquisition module, configured to acquire the data dependency relationship and instruction execution latency of each instruction; wherein, acquiring the instruction execution latency includes: determining whether the instruction type of the instruction is a load instruction or a non-load instruction; if it is a load instruction, querying the status of the data accessed by the instruction in the cache; if the status is cache hit, acquiring the static latency corresponding to the instruction, and using the static latency as the instruction execution latency; if the status is cache miss, querying the corresponding historical latency information in the latency cache, predicting the dynamic latency according to the historical latency information, and using the dynamic latency as the instruction execution latency; if it is a non-load instruction, acquiring the static latency corresponding to the instruction, and using the static latency as the instruction execution latency; A prediction module, configured to determine the predicted issue time of each instruction according to the data dependency relationship and the instruction execution latency; A scheduling module, configured to acquire the hardware resource availability information, and schedule the execution order of each instruction based on the predicted issue time, the data dependency relationship, and the hardware resource availability information; Wherein, the types of the instruction execution latency include dynamic latency and static latency, and the acquisition module includes a dynamic latency acquisition unit, and the dynamic latency acquisition unit is configured to acquire the instruction execution latency of the type of dynamic latency; The dynamic latency acquisition unit is further configured to: acquire the historical latency information corresponding to the current instruction from the latency cache, and predict the dynamic latency through the historical latency information; Wherein, the latency cache is configured to record the actual latency time of each instruction in the past as the historical latency information.
12. An instruction scheduling system, characterized in that, Including: A memory, configured to store a computer program; A processor, configured to implement the steps of the instruction scheduling method according to any one of claims 1 to 10 when executing the computer program.
13. A computer program product, comprising a computer program / instructions, characterized in that, The computer program / instruction, when executed by the processor, implements the steps of the instruction scheduling method according to any one of claims 1 to 10.
14. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the instruction scheduling method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Method and system for predicting Load instruction execution delay
CN111221579A
Instruction scheduling method and device based on timeliness priority, computer equipment and medium
CN117742795A